AI Agents & Econometric Coding

Regression diagnostics · Other

Variance Inflation from Collinear Regressors

The variance inflation factor measures how strongly other regressors explain one regressor.

Isolating one regressor's independent variation

Regress $x_j$ on the remaining regressors and record $R_j^2$. The variance inflation factor is (O'Brien, 2007):

$$ VIF_j=\frac{1}{1-R_j^2}. $$

It compares the coefficient variance with its value under orthogonality, holding the residual variance and scale fixed.

The nonlinear increase

Auxiliary $R_j^2$$VIF_j$Standard-error multiplier $\sqrt{VIF_j}$
0.0011.00
0.5021.41
0.8052.24
0.95204.47

The increase accelerates as the auxiliary fit approaches one.

Assumptions and interpretation

VIF diagnoses the design matrix. It does not test omitted-variable bias or model validity. High collinearity can be a direct consequence of the research design, interactions, or polynomial terms.

Software differs on whether it includes the intercept and how it treats transformed terms. Exact aliasing makes the design matrix singular and the VIF undefined.

Geometrically, $x_j$ splits into a projection on the other regressors and an orthogonal remainder. The smaller remainder identifies its coefficient.

Interpretation and reporting

A VIF of $20$ means collinearity multiplies the coefficient variance by twenty under the comparison conditions. The corresponding standard-error multiplier is about $4.47$. It does not mean the coefficient itself is biased.

Report the auxiliary $R_j^2$, VIF, and regressor definition. Fixed cutoffs should not trigger automatic variable removal. Polynomial terms and interactions can have high VIF values because their construction creates dependence that the model requires.

Collinearity limits precision by leaving little independent variation in $x_j$. It does not diagnose omitted variables, reverse causality, or heteroskedasticity. Exact linear dependence is a separate problem because it prevents unique coefficient estimation.

Reproducible implementation

Use the regression's exact estimation sample and design transformations. For each nonintercept regressor, fit an auxiliary regression on all remaining columns with the intended intercept.

Calculate $1/(1-R_j^2)$ and compare it with the software output. Confirm that rescaling a regressor does not change its VIF. Detect aliased columns before interpreting extremely large values.

Inspect condition indices or singular values when dependence involves several columns at once. Center continuous variables before forming polynomials when that improves numerical interpretation. Preserve hierarchy when interactions are substantively required. Save the full column definitions so every auxiliary regression can be reconstructed (Belsley et al., 1980).

Further reading

O'Brien explains why fixed VIF thresholds have limited meaning (O'Brien, 2007). Belsley, Kuh, and Welsch discuss collinearity diagnostics based on the full design matrix (Belsley et al., 1980).

Sources and further reading

  1. Robert M. O'Brien. 2007. “A Caution Regarding Rules of Thumb for Variance Inflation Factors.” Quality and Quantity 41(5): 673--690. doi:10.1007/s11135-006-9018-6.
  2. David A. Belsley, Edwin Kuh, Roy E. Welsch. 1980. Regression Diagnostics: Identifying Influential Data and Sources of Collinearity. Wiley. doi:10.1002/0471725153.

Related reading

About this benchmark task

Status
In the benchmark
Identifier
vif
Family
Other
Software
Stata, R, Python
Source
Benchmark task set

Task statement

For the linear model y on x1, x2, x3, x4, and x5, compute the VIF for each regressor. Report rows vif_x1, vif_x2, vif_x3, vif_x4, and vif_x5.

Required output

vif_x1, vif_x2, vif_x3, vif_x4, vif_x5

Open every recorded run for this task.