Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2016;111(513):355-376.
doi: 10.1080/01621459.2015.1008363. Epub 2016 May 5.

Variable Selection with Prior Information for Generalized Linear Models via the Prior LASSO Method

Affiliations

Variable Selection with Prior Information for Generalized Linear Models via the Prior LASSO Method

Yuan Jiang et al. J Am Stat Assoc. 2016.

Abstract

LASSO is a popular statistical tool often used in conjunction with generalized linear models that can simultaneously select variables and estimate parameters. When there are many variables of interest, as in current biological and biomedical studies, the power of LASSO can be limited. Fortunately, so much biological and biomedical data have been collected and they may contain useful information about the importance of certain variables. This paper proposes an extension of LASSO, namely, prior LASSO (pLASSO), to incorporate that prior information into penalized generalized linear models. The goal is achieved by adding in the LASSO criterion function an additional measure of the discrepancy between the prior information and the model. For linear regression, the whole solution path of the pLASSO estimator can be found with a procedure similar to the Least Angle Regression (LARS). Asymptotic theories and simulation results show that pLASSO provides significant improvement over LASSO when the prior information is relatively accurate. When the prior information is less reliable, pLASSO shows great robustness to the misspecification. We illustrate the application of pLASSO using a real data set from a genome-wide association study.

Keywords: Asymptotic efficiency; Oracle inequalities; Solution path; Weak oracle property.

PubMed Disclaimer

Figures

Figure 1
Figure 1
An Example of pLASSO Solution Path. This path is illustrated using the diabetes data in Efron et al. (2004), with η on the horizontal axis and coefficient estimates on the vertical axis. There are 10 predictors in the diabetes data, each of which is standardized to have a unit norm. We choose a prior set [Image: see text] = {X5, X6, X7, X8}. λ is fixed at 316.07 with an initial set [Image: see text] = {3, 4, 9}. Each vertical line in the plot represents a change of the active set, with the inclusion/deletion of a variable noted at the top of the plot, and the corresponding η value noted at the bottom. The last vertical line is hypothetical to show the limit of coefficients when η → ∞.
Figure 2
Figure 2
Simulation results of unweighted penalization methods for linear regression. Rows 1–3 correspond to prior sets S1p-S3p in [Image: see text] respectively, and rows 4–6 correspond to prior sets S4p-S6p in [Image: see text] respectively. Columns 1–4 correspond to #CNZ, #INZ, Bias and MSR, respectively.
Figure 3
Figure 3
Simulation results of unweighted penalization methods for linear regression. Rows 1–3 correspond to prior sets S7p-S9p in [Image: see text] respectively, and rows 4–6 correspond to prior sets S10p-S12p in [Image: see text] respectively. Columns 1–4 correspond to #CNZ, #INZ, Bias and MSR, respectively.
Figure 4
Figure 4
Simulation results of unweighted penalization methods for logistic regression. Rows 1–3 correspond to prior sets S1p-S3p in [Image: see text] respectively, and rows 4–6 correspond to prior sets S4p-S6p in [Image: see text] respectively. Columns 1–5 correspond to #CNZ, #INZ, Bias, RME and MR, respectively.
Figure 5
Figure 5
Simulation results of unweighted penalization methods for logistic regression. Rows 1–3 correspond to prior sets S7p-S9p in [Image: see text] respectively, and rows 4–6 correspond to prior sets S10p-S12p in [Image: see text] respectively. Columns 1–5 correspond to #CNZ, #INZ, Bias, RME and MR, respectively.
Figure 6
Figure 6
Optimal selection of the tuning parameter η by pLASSO in linear regression (upper panel) and logistic regression (lower panel).
Figure 7
Figure 7
Simulation results of weighted penalization methods for linear regression. Rows 1–3 correspond to prior sets S1p-S3p in [Image: see text] respectively, and rows 4–6 correspond to prior sets S4p-S6p in [Image: see text] respectively. Columns 1–4 correspond to #CNZ, #INZ, Bias and MSR, respectively.
Figure 8
Figure 8
Simulation results of weighted penalization methods for linear regression. Rows 1–3 correspond to prior sets S7p-S9p in [Image: see text] respectively, and rows 4–6 correspond to prior sets S10p-S12p in [Image: see text] respectively. Columns 1–4 correspond to #CNZ, #INZ, Bias and MSR, respectively.
Figure 9
Figure 9
Simulation results of weighted penalization methods for logistic regression. Rows 1–3 correspond to prior sets S1p-S3p in [Image: see text] respectively, and rows 4–6 correspond to prior sets S4p-S6p in [Image: see text] respectively. Columns 1–5 correspond to #CNZ, #INZ, Bias, RME and MR, respectively.
Figure 10
Figure 10
Simulation results of weighted penalization methods for logistic regression. Rows 1–3 correspond to prior sets S7p-S9p in [Image: see text] respectively, and rows 4–6 correspond to prior sets S10p-S12p in [Image: see text] respectively. Columns 1–5 correspond to #CNZ, #INZ, Bias, RME and MR, respectively.

References

    1. Bach F. Self-concordant analysis for logistic regression. Electronic Journal of Statistics. 2010;4:384–414.
    1. Baum A, Akula N, Cabanero M, Cardona I, Corona W, Klemens B, Schulze T, Cichon S, Rietschel M, Nöthen M, et al. A genome-wide association study implicates diacylglycerol kinase eta (DGKH) and several other genes in the etiology of bipolar disorder. Molecular Psychiatry. 2007;13:197–207. - PMC - PubMed
    1. Baum A, Hamshere M, Green E, Cichon S, Rietschel M, Noethen M, Craddock N, McMahon F. Meta-analysis of two genome-wide association studies of bipolar disorder reveals important points of agreement. Molecular Psychiatry. 2008;13:466–467. - PMC - PubMed
    1. Bickel PJ, Ritov Y, Tsybakov AB. Simultaneous analysis of Lasso and Dantzig selector. The Annals of Statistics. 2009;37:1705–1732.
    1. Breiman L, Friedman J, Olshen R, Stone C. Classification and Regression Trees. Wadsworth International Group; 1984.

Publication types