Quantile regression

Quantile regression is a type of regression analysis used in statistics and econometrics. Whereas the method of least squares estimates the conditional mean of the response variable across values of the predictor variables, quantile regression estimates the conditional median of the response variable. Quantile regression is an extension of linear regression used when the conditions of linear regression are not met.

Advantages and applications

One advantage of quantile regression relative to ordinary least squares regression is that the quantile regression estimates are more robust against outliers in the response measurements. However, the main attraction of quantile regression goes beyond this and is advantageous when conditional quantile functions are of interest. Different measures of central tendency and statistical dispersion can be useful to obtain a more comprehensive analysis of the relationship between variables.
In ecology, quantile regression has been proposed and used as a way to discover more useful predictive relationships between variables in cases where there is no relationship or only a weak relationship between the means of such variables. The need for and success of quantile regression in ecology has been attributed to the complexity of interactions between different factors leading to data with unequal variation of one variable for different ranges of another variable.
Another application of quantile regression is in the areas of growth charts, where percentile curves are commonly used to screen for abnormal growth.

Mathematics

The mathematical forms arising from quantile regression are distinct from those arising in the method of least squares. The method of least squares leads to a consideration of problems in an inner product space, involving projection onto subspaces, and thus the problem of minimizing the squared errors can be reduced to a problem in numerical linear algebra. Quantile regression does not have this structure, and instead leads to problems in linear programming that can be solved by the simplex method.

History

The idea of estimating a median regression slope, a major theorem about minimizing sum of the absolute deviances and a geometrical algorithm for constructing median regression was proposed in 1760 by Ruđer Josip Bošković, a Jesuit Catholic priest from Dubrovnik. He was interested in the ellipticity of the earth, building on Isaac Newton's suggestion that its rotation could cause it to bulge at the equator with a corresponding flattening at the poles. He finally produced the first geometric procedure for determining the equator of a rotating planet from three observations of a surface feature. More importantly for quantile regression, he was able to develop the first evidence of the least absolute criterion and preceded the least squares introduced by Legendre in 1805 by fifty years.
Other thinkers began building upon Bošković's idea such as Pierre-Simon Laplace, who developed the so-called "methode de situation." This led to Francis Edgeworth's plural median - a geometric approach to median regression - and is recognized as the precursor of the simplex method. The works of Bošković, Laplace, and Edgeworth were recognized as a prelude to Roger Koenker's contributions to quantile regression.
Median regression computations for larger data sets are quite tedious compared to the least squares method, for which reason it has historically generated a lack of popularity among statisticians, until the widespread adoption of computers in the latter part of the 20th century.

Quantiles

Let be a real valued random variable with cumulative distribution function. The th quantile of Y is given by
where
Define the loss function as, where is an indicator function.
A specific quantile can be found by minimizing the expected loss of with respect to :
This can be shown by setting the derivative of the expected loss function to 0 and letting be the solution of
This equation reduces to
and then to
Hence is th quantile of the random variable Y.

Example

Let be a discrete random variable that takes values 1,2,..,9 with equal probabilities. The task is to find the median of Y, and hence the value is chosen. The expected loss,, is
Since is a constant, it can be taken out of the expected loss function. Then, at u=3,
Suppose that u is increased by 1 unit. Then the expected loss will be changed by on changing u to 4. If, u=5, the expected loss is
and any change in u will increase the expected loss. Thus u=5 is the median. The Table below shows the expected loss for different values of u.

u	1	2	3	4	5	6	7	8	9
Expected loss	36	29	24	21	20	21	24	29	36

Intuition

Consider and let q be an initial guess for. The expected loss evaluated at q is
In order to minimize the expected loss, we move the value of q a little bit to see whether the expected loss will rise or fall.
Suppose we increase q by 1 unit. Then the change of expected loss would be
The first term of the equation is and second term of the equation is. Therefore, the change of expected loss function is negative if and only if, that is if and only if q is smaller than the median. Similarly, if we reduce q by 1 unit, the change of expected loss function is negative if and only if q is larger than the median.
In order to minimize the expected loss function, we would increase L if q is smaller than the median, until q reaches the median. The idea behind the minimization is to count the number of points that are larger or smaller than q and then move q to a point where q is larger than % of the points.

Sample quantile

The sample quantile can be obtained by solving the following minimization problem

Conditional quantile and quantile regression

Suppose the th conditional quantile function is. Given the distribution function of, can be obtained by solving
Solving the sample analog gives the estimator of.

Computation

The minimization problem can be reformulated as a linear programming problem
where
Simplex methods or interior point methods can be applied to solve the linear programming problem.

Asymptotic properties

For, under some regularity conditions, is asymptotically normal:
where
Direct estimation of the asymptotic variance-covariance matrix is not always satisfactory. Inference for quantile regression parameters can be made with the regression rank-score tests or with the bootstrap methods.

Equivariance

See invariant estimator for background on invariance or see equivariance.

Scale equivariance

For any and

Shift equivariance

For any and

Equivariance to reparameterization of design

Let be any nonsingular matrix and

Invariance to monotone transformations

If is a nondecreasing function on 'R, the following invariance property applies:
Example :
If and, then. The mean regression does not have the same property since

Bayesian methods for quantile regression

Because quantile regression does not normally assume a parametric likelihood for the conditional distributions of Y|X, the Bayesian methods work with a working likelihood. A convenient choice is the asymmetric Laplacian likelihood, because the mode of the resulting posterior under a flat prior is the usual quantile regression estimates. The posterior inference, however, must be interpreted with care. Yang, Wang and He provided a posterior variance adjustment for valid inference. In addition,
Yang and He showed that one can have asymptotically valid posterior inference if the working likelihood is chosen to be the empirical likelihood.

Machine learning methods for quantile regression

Beyond simple linear regression, there are several machine learning methods that can be extended to quantile regression. A switch from the squared error to the tilted absolute value loss function allows gradient descent based learning algorithms to learn a specified quantile instead of the mean. It means that we can apply all neural network and deep learning algorithms to quantile regression. Tree-based learning algorithms are also available for quantile regression.

Censored quantile regression

If the response variable is subject to censoring, the conditional mean is not identifiable without additional distributional assumptions, but the conditional quantile is often identifiable. For recent work on censored quantile regression, see: Portnoy
and Wang and Wang
Example :
Let and. Then. This is the censored quantile regression model: estimated values can be obtained without making any distributional assumptions, but at the cost of computational difficulty, some of which can be avoided by using a simple three step censored quantile regression procedure as an approximation.
For random censoring on the response variables, the censored quantile regression of Portnoy provides consistent estimates of all identifiable quantile functions based on reweighting each censored point appropriately.

Implementations

Numerous statistical software packages include implementations of quantile regression:

Matlab function quantreg
Eviews, since version 6.
gretl has the quantreg command.
R offers several packages that implement quantile regression, most notably quantreg by Roger Koenker, but also gbm, quantregForest, qrnn and qgam
Python, via Scikit-garden and statsmodels
SAS through proc quantreg and proc quantselect.
Stata, via the qreg command.
Vowpal Wabbit, via --loss_function quantile.
Statsmodels package for Python, via QuantReg
Mathematica package QuantileRegression.m hosted at the MathematicaForPrediction project at GitHub.

Popular movies

The Hunger Games (film) - 2012 American dystopian action thriller science fiction-adventure film directed by Gary Ross and based on Suzanne Collins’s 2008 novel of the same name. It is the first insta...
untitled Captain Marvel sequel - part of Marvel Cinematic Universe....
Killers of the Flower Moon (film project) - Killers of the Flower Moon - film project in United States of America. It was presented as drama, detective fiction, thriller. The film project starred Leonardo Dicaprio, Robert De Niro. Director of...
Five Nights at Freddy's (film) - Five Nights at Freddy's - film published in 2017 in United States of America. Scenarist of the film - Scott Cawthon....

Popular books

Book of Revelation - The Book of Revelation is the final book of the New Testament, and consequently is also the final book of the Christian Bible. Its title is derived from the first word of the Koine Greek text: apok...
Book of Genesis - account of the creation of the world, the early history of humanity, Israel's ancestors and the origins...
Gospel of Matthew - The Gospel According to Matthew is the first book of the New Testament and one of the three synoptic gospels. It tells how Israel's Messiah, rejected and executed in Israel, pronounces judgement on ...
Michelin Guide - Michelin Guides are a series of guide books published by the French tyre company Michelin for more than a century. The term normally refers to the annually published Michelin Red Guide , the oldest...
Psalms - The Book of Psalms , commonly referred to simply as Psalms , the Psalter or "the Psalms", is the first book of the Ketuvim , the third section of the Hebrew Bible, and thus a book of th...
Ecclesiastes - Ecclesiastes is one of 24 books of the Tanakh , where it is classified as one of the Ketuvim . Originally written c. 450–200 BCE, it is also among the canonical Wisdom literature of the Old Tes...
The 48 Laws of Power - non-fiction book by American author Robert Greene. The book...

Popular television series

The Crown (TV series) - historical drama web television series about the reign of Queen Elizabeth II, created and principally written by Peter Morgan, and produced by Left Bank Pictures and Sony Pictures Tel...
Friends - American sitcom television series, created by David Crane and Marta Kauffman, which aired on NBC from September 22, 1994, to May 6, 2004, lasting ten seasons. With an ensemble cast sta...
Young Sheldon - spin-off prequel to The Big Bang Theory and begins with the character Sheldon...
Modern Family - American television mockumentary family sitcom created by Christopher Lloyd and Steven Levitan for the American Broadcasting Company. It ran for eleven seasons, from September 23...
Loki (TV series) - upcoming American web television miniseries created for Disney+ by Michael Waldron, based on the Marvel Comics character of the same name. It is set in the Marvel Cinematic Universe, shar...
Game of Thrones - American fantasy drama television series created by David Benioff and D. B. Weiss for HBO. It...
Shameless (American TV series) - American comedy-drama television series developed by John Wells which debuted on Showtime on January 9, 2011. It...