Danish Epidemiological Society Conference
2026-09-22
Nevertheless, health inequalities are socially produced, and therefore, also potentially avoidable. However, effective political interventions require a scientific understanding of the causal mechanisms generating the strong and persistent correlations between social conditions and health outcomes.
Eikemo and Øversveen (2019)
What changed?
Explaining growing or shrinking gaps
Which ‘components’?
Bridge between descriptive and causal
What if?
Techniques often involve ‘counterfactual’ scenarios
Source: Den Nationale Sundhedsprofil 2010 (Tabel 4.1.1) and 2025 (Tabel 3.1.3), Statens Institut for Folkesundhed / Sundhedsstyrelsen.
If standardization alters a difference between two rates, it should be possible … to break it up into components attributable to the various factors for which the data were standardized.
Seminal paper by Kitagawa (1955)
may give a distorted view of the overall level of health inequality due to social class differentials as it does not take into account the changing distribution of social class within the whole population.
changes in the distribution of social class across the population are an important contributor to recent reductions in mortality
Heller, McElduff, and Edwards (2002)
Group rate: \(R_{i} = y_{i}/n_{i}\)
Group share: \(P_{i} = n_{i}/\sum_{i} n_{i}\)
Crude rate: \(\frac{\sum_{i=1}^{I} y_{i}}{\sum_{i=1}^{I} n_{i}} = \sum_{i=1}^{I} P_{i}R_{i}\)
Standardized rate: \(\frac{\sum_{i=1}^{I} y_{i}}{\sum_{i=1}^{I} s_{i}} = \sum_{i=1}^{I} S_{i}R_{i}\)
where \(S_{i}\) is the share of group i in a standard population (\(\sum_{i} S_{i} = 1\))
| Stratum | Pop | Count | Rate |
|---|---|---|---|
| \(1\) | \(n_{i}\) | \(y_{i}\) | \(r_{i}\) |
| \(\vdots\) | \(\vdots\) | \(\vdots\) | \(\vdots\) |
| \(I\) | \(N\) | \(Y\) | \(R\) |
| Group A | Group B | |||
|---|---|---|---|---|
| Age | Share | Rate | Share | Rate |
| Young | 30% | 10 | 50% | 20 |
| Old | 70% | 40 | 50% | 30 |
| Crude | 31 | 25 | ||
| Standardized to A | 31 | 27 | ||
What would the \(A-B\) difference be if \(B\) had \(A\)’s age distribution?
…if the difference between two standardized rates is subtracted from the corresponding difference in crude rates…the result is a weighted average of differences between the composition of the two groups.
But what are the weights?
Crude difference = \(\sum P_{Ai}R_{Ai} - \sum P_{Bi}R_{Bi}\)
→ Reflects differences in both rates and composition
\(A\)-Standardized difference = \(\sum P_{Ai}R_{Ai} - \sum P_{Ai}R_{Bi}\)
→ Reflects only differences in rates (\(P_{Ai}\) is constant)
Crude - Standardized = \(\sum R_{Bi}(P_{Ai}-P_{Bi})\)
→ Differences in composition, weighted by \(B\)s rates!
\[\begin{aligned} \underbrace{\text{Rate}_A - \text{Rate}_B}_{\text{crude difference}} &= \overbrace{\sum_i \left(\frac{R_{Ai}+R_{Bi}}{2}\right)(P_{Ai}-P_{Bi})}^{\text{composition effect}} \\ &+ \underbrace{\sum_i \left(\frac{P_{Ai}+P_{Bi}}{2}\right)(R_{Ai}-R_{Bi})}_{\text{rate effect}} \end{aligned}\]
\(R_{Ai}\) is the rate for group A in stratum i.
\(P_{Ai}\) is the i-stratum-specific weight in group A.
| Group A | Group B | |||
|---|---|---|---|---|
| Age | Share | Rate | Share | Rate |
| Young | 30% | 10 | 50% | 20 |
| Old | 70% | 40 | 50% | 30 |
| Composition effect | |||
|---|---|---|---|
| Group | \(\frac{R_{Ai}+R_{Bi}}{2}\) | \(P_{Ai}-P_{Bi}\) | Contribution |
| Young | 15 | −0.20 | −3.0 |
| Old | 35 | +0.20 | +7.0 |
| Sum | +4.0 | ||
| Rate effect | |||
|---|---|---|---|
| Group | \(\frac{P_{Ai}+P_{Bi}}{2}\) | \(R_{Ai}-R_{Bi}\) | Contribution |
| Young | 0.40 | −10 | −4.0 |
| Old | 0.60 | +10 | +6.0 |
| Sum | +2.0 | ||
Total = +6.0 = crude gap (31 − 25) →67% due to composition
Rate effect (+2) = standardized gap (avg reference) (28 − 26)
| Reference | Composition | Rate | Total |
|---|---|---|---|
| A's rates | 6.0 | 0.0 | 6.0 |
| B's rates | 2.0 | 4.0 | 6.0 |
| Average (Kitagawa) | 4.0 | 2.0 | 6.0 |
Crude gap = 31 − 25 = 6 in every row; only the split changes
Single references differ an interaction, \(\sum_i (P_{Ai}-P_{Bi})(R_{Ai}-R_{Bi})\) = 4.0, which the average splits evenly
What accounts for differences in neonatal mortality rates across European countries?
GA-specific rates or distribution?
Implications for interventions?
Sartorius et al. (2024)
Both DK and AT similar and ~0.3 deaths per 1000 higher than “Top 3”
But completely different contributions to the difference
Sartorius et al. (2024)
Sartorius et al. (2024)
Sartorius et al. (2024)
Implications for interventions
DK: improving care at early GA
AT: shifting the distribution of GA later
When to use
Comparing rates across groups or over time when their composition differs
Extensions
More than two groups or factors; three-part decomposition
Caveats
Not causal; results depend on the strata used and on the reference weights
Our findings show that unexplained factors associated to immigrant status determine to a great extent disparities in the probability of using hospital, specialist and emergency services of immigrants relative to Spaniards, while individual characteristics, in particular self-reported health and chronic conditions, are much more important in explaining the differences in the probability of using general practitioner services between immigrants and Spaniards
Jiménez-Rubio and Hernández-Quevedo (2010)
Studies by R. Oaxaca (1973) and Blinder (1973) applied regression-based decomposition methods to analyze the wage gap between men and women and between whites and blacks in the USA.
Focused on how much of wage gap was ‘explained’ by differences in observable characteristics
Recent attempts to reconcile O-B with Kitagawa (equivalent for binary outcome and categorical covariates)
R. L. Oaxaca and Sierminska (2025)
Decomposition methods are based on regression analyses, and thus all of the usual caveats about good specification apply.
If regressions are purely descriptive, they reveal the associations that characterize the health inequality. Then inequality is explained in a statistical sense but implications for policies to reduce inequality are limited.
O’Donnell et al. (2008)
OLS regression line always passes through \((\bar{x}, \bar{y})\) because residuals sum to zero.
The gap decomposes into endowments (different \(\bar{x}\), same coefficients) and coefficients (same \(\bar{x}\), different slopes/intercepts).
Equally valid to use the exposed group’s slope
Same total gap, but different split
What is the average difference in body mass index (BMI) between those with low vs. high education?
How much is due to differing determinants of BMI (age, gender, smoking, drinking, income)?
Any residual difference is due to education differences in the associations between those risk factors and BMI — i.e., the coefficients differ.
European Social Survey Round 11, Italy (n = 1743)
Body mass index as outcome (kg/m²): weight / height²
Overall difference by education: High ed (ISCED 5–7) vs Low ed (ISCED 1–4)
Potential determinants (the \(X\)s):
Source: Author’s calculations
Differences in age, smoking, etc. could explain part of the BMI gap
Source: Author’s calculations
Differences in returns (i.e., coefficients) form the ‘unexplained’ component
Overall gap = 1.61 kg/m². Endowments: 38%; Coefficients: 62%
| Endowments | Coefficients | |||
|---|---|---|---|---|
| Est | SE | Est | SE | |
| Total | 0.608 | 0.108 | 1.030 | 0.259 |
| Age (years) | 0.415 | 0.074 | 0.865 | 0.691 |
| Female | 0.080 | 0.051 | 0.466 | 0.181 |
| Current smoker | 0.004 | 0.008 | -0.145 | 0.115 |
| Binge drinking (monthly+) | 0.001 | 0.007 | 0.003 | 0.104 |
| Married | 0.012 | 0.016 | -0.037 | 0.208 |
| Income (decile) | 0.095 | 0.071 | 0.133 | 0.425 |
| Intercept | — | — | -0.256 | 0.837 |
Note: Low ed betas used as reference
| Endowments | ||
|---|---|---|
| Est | SE | |
| Total | 0.608 | 0.108 |
| Age (years) | 0.415 | 0.074 |
| Female | 0.080 | 0.051 |
| Current smoker | 0.004 | 0.008 |
| Binge drinking (monthly+) | 0.001 | 0.007 |
| Married | 0.012 | 0.016 |
| Income (decile) | 0.095 | 0.071 |
If the low-educated had the same covariate means as the high-educated (keeping the low-educated group’s own coefficients), their BMI would be 0.608 \(kg/m^2\) lower (38% of the gap).
Most of this is due to older age among the low-educated, which predicts higher BMI.
If the high-educated group’s covariates had the same relationship with BMI as they do in the low-educated group, their BMI would be 1.030 higher, accounting for 62% of the gap.
Smoking is negative, since smoking predicts higher BMI among the high-educated and is more common among high educated.
| Coefficients | ||
|---|---|---|
| Est | SE | |
| Total | 1.030 | 0.259 |
| Age (years) | 0.865 | 0.691 |
| Female | 0.466 | 0.181 |
| Current smoker | -0.145 | 0.115 |
| Binge drinking (monthly+) | 0.003 | 0.104 |
| Married | -0.037 | 0.208 |
| Income (decile) | 0.133 | 0.425 |
| Intercept | -0.256 | 0.837 |
Variables where groups differ and that predict BMI contribute most
Income and age are the biggest drivers
Endowments account for 36–44% of it; interaction is negligible
When to use
Explaining average gaps with potential determinants
Extensions
Nonlinear outcomes (binary, counts); pooled reference coefficients
Caveats
Not causal; specification errors; reference matters for categorical covariates
Explicitly accounts for changing composition
Research Question:
How much of inequality is due to other factors that are differential by \(x\) (income) and also affect \(y\) (health)?
\[RCI= \frac{2}{n\mu} \sum_{i=1}^{n}y_{i}R_{i}-1\]
where
Decomposition:
Develop a model for predicting \(y\) using several determinants, then plug back into the \(RCI\) equation
Estimates how much of the overall inequality in \(y\) is due to the association between income and other factors that predict health
Kakwani, Wagstaff, and Doorslaer (1997)
\(RCI\) is a function of health \((y_{i})\) and socioeconomic rank \((R_{i})\), i.e. \[RCI= \frac{2}{n\mu} \sum_{i=1}^{n}{\color{red}{y_{i}}}R_{i}-1\]
Then write a regression expressing health \((y_{i})\) as a function of several \(k_{i}\) determinants (e.g., age, gender, urban/rural status): \[\color{red}{y_{i}}=\alpha + \sum{\beta_{x}x_{k_{i}}}+\epsilon_{i}\]
Wagstaff, Doorslaer, and Watanabe (2003)
Now re-express \(RCI\) as:
\[RCI=\sum{(\beta_{k}\bar{x}_{k}/\mu)RCI_{k}}+gRCI_{e}/\mu\]
Where
Determinant impact depends on:
the strength of the relationship between each factor and income \(\color{blue}{RCI_{k}}\)
the strength of the relationship between each factor and health, and its prevalence in the population (elasticity) \(\color{red}{\beta_{k}\bar{x}_{k}/\mu}\)
\[y_{i}=\alpha + \sum{\beta_{x}x_{k_{i}}}+\epsilon_{i}\]
Calculate means of \(y\) \((\mu)\) and all \(x_{k}\) determinants
Calculate (\(RCI\)) for health and for each determinant \((RCI_{k})\), i.e., use each \(x_{k}\) as the “outcome” and estimate a \(RCI\) for age, education, etc.
\[(\beta_{k}\bar{x}_{k}/\mu)RCI_{k}\]
\[[(\beta_{k}\bar{x}_{k}/\mu)RCI_{k}]/RCI\]
| Unique | Missing Pct. | Mean | SD | Min | Median | Max | Histogram | |
|---|---|---|---|---|---|---|---|---|
| Current smoker | 2 | 0 | 0.1 | 0.2 | 0.0 | 0.0 | 1.0 | |
| Income decile | 10 | 0 | 6.6 | 2.6 | 1.0 | 7.0 | 10.0 | |
| Age (years) | 76 | 0 | 54.4 | 18.6 | 15.0 | 56.0 | 90.0 | |
| Low education (ISCED 1–4) | 2 | 0 | 0.4 | 0.5 | 0.0 | 0.0 | 1.0 | |
| Female | 2 | 0 | 0.5 | 0.5 | 0.0 | 0.0 | 1.0 | |
| Binge drinking (monthly+) | 2 | 0 | 0.5 | 0.5 | 0.0 | 0.0 | 1.0 | |
| Obese (BMI ≥ 30) | 2 | 0 | 0.1 | 0.4 | 0.0 | 0.0 | 1.0 | |
| Married | 2 | 0 | 0.5 | 0.5 | 0.0 | 0.0 | 1.0 | |
| Survey weight | 132 | 0 | 1.0 | 0.5 | 0.3 | 0.8 | 4.0 |
Source: Author’s calculations
Overall CI = -0.20
Smoking more concentrated among the poor
| Variable | β | OR | 95% CI | p |
|---|---|---|---|---|
| Age (years) | -0.001 | 0.999 | (0.99, 1.01) | 0.906 |
| Female | 0.141 | 1.152 | (0.73, 1.83) | 0.547 |
| Low education (ISCED 1–4) | 0.359 | 1.432 | (0.90, 2.27) | 0.130 |
| Married | -0.989 | 0.372 | (0.21, 0.65) | <0.001 |
| Binge drinking (monthly+) | 0.38 | 1.463 | (0.92, 2.32) | 0.107 |
| Obese (BMI ≥ 30) | 0.495 | 1.64 | (0.90, 2.99) | 0.106 |
\[RCI=\sum{ \underbrace{(\beta_{k}\bar{x}_{k}/\mu)}_{\text{elasticity}} RCI_{k}}+gRCI_{e}/\mu\]
Elasticity for low education is: (0.024 * 0.452 / .073) = 0.148
Interpretation: a 1% increase in low education increases smoking by 14.8% (not percentage points!).
What about the RCI for low education?
This is Step (2): Calculate the mean of \(y\) \((\mu)\) and of each of the \(x_{k}\) determinants
Step 3: \(RCI_{k}\)
Note the y-axis is cumulative share of low education
The poorest 50% account for roughly 60% of the share of low educated.
This is Step (3): Calculate the \(CI\) for each of the potential determinants \(x_{k}\)
Estimation for a specific factor: Low education
\[RCI=\sum{ \underbrace{(\beta_{k}\bar{x}_{k}/\mu)}_{\text{elasticity}} RCI_{k}}+gRCI_{e}/\mu\]
The elasticity of smoking (prior slide) = 0.148
Now we have \(RCI_{k}\) for low education = -0.265
The contribution of low education:
\[\text{Elasticity}\times RCI_{ed} = 0.148 * -0.265 = -.039\]
Thus low education accounts for -.039/ -0.203 = 19% of the overall \(RCI\)
RCI = -0.203
(-) Binge drinking: ↑ smoking but more common at higher incomes
(+) Married: ↓ smoking but more common at higher incomes
| Contribution | ||||
|---|---|---|---|---|
| Variable | Elasticity | CI_k | Absolute | % |
| Logistic regression; marginal effects used as elasticity weights. | ||||
| Age | -0.030 | 0.001 | -0.000 | 0.0 |
| Binge drinker | 0.152 | 0.103 | 0.016 | -7.7 |
| Female | 0.060 | -0.031 | -0.002 | 0.9 |
| Obese | 0.068 | -0.013 | -0.001 | 0.4 |
| Low education | 0.144 | -0.265 | -0.038 | 18.8 |
| Married | -0.353 | 0.284 | -0.100 | 49.4 |
| Residual | — | — | -0.077 | 38.1 |
Each contribution has sampling variability:
\[\underbrace{\frac{\hat{\beta}_k \bar{x}_k}{\hat{\mu}}}_{\text{elasticity}} \times \underbrace{\widehat{RCI}_k}_{\text{inequality in }x_k}\]
Can bootstrap the decomp \(B\) times to get empirical CIs.
When to use
Summarizing with RCI and asking which determinants drive it
Extensions
Nonlinear models; Oaxaca-type decomposition of differences and temporal changes
Caveats
‘Explained’ share depends on well-specified model; predictive, not causal
Prior decompositions (e.g., Kitagawa, O-B) ‘descriptive’
Causal mediation challenges
Challenges with non-manipulable exposures
How much would gaps change if we intervened on a target?
Jackson and VanderWeele (2018)
Reconciling the descriptive framework of KBO with causal inference.
Danish-origin adolescents die in accidents more often than immigrant-origin peers (IRR 1.7 vs first-generation)
Alcohol is a plausible pathway: 16% of fatal accidents; Danish-origin youth have 4x the odds of an alcohol-related death
But immigrant-origin families are poorer, so adjusting for income widens the gap (IRR 1.7 → 2.6)
What if Danish-origin youth drank like their immigrant-origin peers?
Steps for pursuing interventional decomposition
1 Define the gap and intervention
Consider feasibility, implementation
2 Draw the DAG
Which covariates to adjust for (allowability)
3 Defend the assumptions
Exchangeability; positivity; consistency
4 Estimate the model
Outcome model, weighting, doubly robust
5 Compare factual and counterfactual
Consider effect scale, uncertainty
we estimate the expected 1-year mortality among the low-income patients in a hypothetical world where they initiated medication as often as high-income patients…
Møller et al. (2026)
IDIE-exposed is the expected reduction in 1-year mortality among low-income heart failure patients in a hypothetical world where they initiated medication with the same probability as observed among similar high-income patients, rather than their observed probability.
Møller et al. (2026)
When to use
Counterfactual inequality under a hypothetical intervention on a mediator
Extensions
Stochastic interventions; weighting framework; nonparametric estimation
Caveats
Causal assumptions; well-defined intervention; choices about ‘allowable’ covariates
Various decomposition techniques exist that may be useful for analyzing social inequalities in health
Moves beyond measuring to explanations and interventions
Used responsibly, they can help to provide key evidence on why health inequalities exist and change over time.
Social Epidemiologic Aims
Descriptive
Existence of social differences in health
Etiologic
Causes of social differences in health
Interventions
Policies and programs to address causes