The magnitude and direction of collider bias for binary variables

From MaRDI portal
Publication:2192289



Abstract: Suppose we are interested in the effect of variable X on variable Y. If X and Y both influence, or are associated with variables that influence, a common outcome, called a collider, then conditioning on the collider (or on a variable influenced by the collider -- its "child") induces a spurious association between X and Y, which is known as collider bias. Characterizing the magnitude and direction of collider bias is crucial for understanding the implications of selection bias and for adjudicating decisions about whether to control for variables that are known to be associated with both exposure and outcome but could be either confounders or colliders. Considering a class of situations where all variables are binary, and where X and Y either are, or are respectively influenced by, two marginally independent causes of a collider, we derive collider bias that results from (i) conditioning on specific levels of, or (ii) linear regression adjustment for, the collider (or its child). We also derive simple conditions that determine the sign of such bias.


Let \(Y\) be some dependent variable in a statistical model and let \(X\) be some explanatory vector. If \(Z\) is some variable influenced by both \(Y\) and \(X\), then the bias that occurs in estimating the effect of \(X\) by including \(Z\) in some model \(Y = X + Z\) is called collider bias. That is, collider bias occurs when the effect of a treatment variable \(X\) is overestimated or underestimated when we attempt to control for some variable \(Z\) which is actually the cause of both \(X\) and \(Y\). Collider bias is common in blind application of statistical models. It's easy to come up with ad-hoc numerical examples and hypothetical experimental situations where collider bias can occur in either direction. The present paper by Nguyen, Dafoe, and Ogburn take this to the systematic level and explicitly derive formulas for collider bias in the case of two binary variables (contingency tables) that are marginally independent with conditioning on collider variable. They give explicit formulas for bias on the covariance, risk difference, and in some cases the odds ratio for various causal relations among the variables. I believe this to be an important contribution to the statistical literature. It gives a quantitative force against bias in linear modelling, and it can be used to estimate the bias that can occur in erroneous models.











This page was built for publication: The magnitude and direction of collider bias for binary variables

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2192289)