diff --git a/chapters/04-effect-modification.qmd b/chapters/04-effect-modification.qmd index aa38d50a..7c13d85c 100644 --- a/chapters/04-effect-modification.qmd +++ b/chapters/04-effect-modification.qmd @@ -518,8 +518,19 @@ The explanation (leaving aside random variability) is **qualitative effect modif The lesson: always specify the population, or subset, to which an effect measure corresponds. +::: {.callout-note title="Technical Point 4.3: Relation between marginal and conditional causal risk ratios"} +Knowing the stratum-specific causal risk ratios is not enough to tell +whether the marginal causal risk ratio lies above or below 1. +The marginal ratio averages the stratum-specific ratios, +with weights that depend on how large each stratum is +and on how common the outcome would be in it without treatment. +Solving for the weights turns the question into an inequality about those two ingredients +[@hernan2020causal, p. 58]. +@prp-marginal-rr-condition below works this out for Table 4.3. +::: + ::: {#prp-marginal-rr-condition} -## When is the marginal risk ratio below 1? (Technical Point 4.3) +## When is the marginal risk ratio below 1? Let $\RR_l \eqdef \Pr[Z^{a=1} = 1 \mid L = l] / \Pr[Z^{a=0} = 1 \mid L = l]$, $p_l \eqdef \Pr[Z^{a=0} = 1 \mid L = l]$, and $\pi_l \eqdef \Pr[L = l]$. With $\RR_0 = 2$ and $\RR_1 = 0.5$, the marginal risk ratio is below 1 if and only if @@ -539,14 +550,15 @@ $$ \end{aligned} $$ -Substituting $\RR_0 = 2$, $\RR_1 = 0.5$, and $w(1) = 1 - w(0)$: +Substituting $\RR_0 = 2$, $\RR_1 = 0.5$, and $w(1) = 1 - w(0)$, +and using $\Pr[Z^{a=0} = 1] = p_0 \pi_0 + p_1 \pi_1$: $$ \begin{aligned} 2\, w(0) + 0.5\,(1 - w(0)) < 1 &\iff 1.5\, w(0) < 0.5 \\ &\iff w(0) < 1/3 \\ -&\iff 3 p_0 \pi_0 < p_0 \pi_0 + p_1 \pi_1 && \text{(definition of } w(0)\text{; } \Pr[Z^{a=0} = 1] = p_0 \pi_0 + p_1 \pi_1\text{)} \\ +&\iff 3 p_0 \pi_0 < p_0 \pi_0 + p_1 \pi_1 && \text{(definition of } w(0)\text{)} \\ &\iff 2 p_0 \pi_0 < p_1 \pi_1 \\ &\iff \frac{p_1}{p_0} > \frac{2 \pi_0}{\pi_1}. \end{aligned} diff --git a/chapters/15-outcome-regression-propensity-scores.qmd b/chapters/15-outcome-regression-propensity-scores.qmd index eba7e75b..f917adea 100644 --- a/chapters/15-outcome-regression-propensity-scores.qmd +++ b/chapters/15-outcome-regression-propensity-scores.qmd @@ -348,13 +348,23 @@ Two caveats from the book [@hernan2020causal, pp. 201-202]: ::: ::: -::: {.notes} +--- + +::: {.callout-note title="Technical Point 15.1: Balancing Scores and Prognostic Scores"} +The propensity score is the simplest **balancing score** (@def-balancing-score): +once we fix its value, the treated and the untreated have the same distribution of $L$. Rosenbaum and Rubin proved that exchangeability and positivity given $L$ imply exchangeability and positivity given any balancing score [@rosenbaum1983central, Theorem 3; @hernan2020causal, p. 202]. +So if adjusting for $L$ suffices, adjusting for $b(L)$, and in particular for $\pi(L)$, suffices too. + Graphically, $\pi(L)$ sits between $L$ and $A$, with a deterministic arrow $L \to \pi(L)$ (book Figure 15.2); conditioning on $\pi(L)$ blocks the backdoor paths from $A$ to $Y$ through $L$ ($A \leftarrow \pi(L) \leftarrow L \to Y$), so adjusting for it suffices. -Methods based on prognostic scores need stronger assumptions and do not extend readily to time-varying treatments; +A **prognostic score** $s(L)$ is the outcome-side counterpart: +given $s(L)$, $L$ carries no further information about $Y^{a=0}$, +just as, given $b(L)$, it carries none about $A$ (@def-balancing-score). +Adjustment methods exist for both kinds of score, +but those based on prognostic scores need stronger assumptions and do not extend readily to time-varying treatments; see Hansen (2008) and Abadie et al. (2013), as cited in @hernan2020causal [p. 202]. ::: diff --git a/chapters/16-instrumental-variable-estimation.qmd b/chapters/16-instrumental-variable-estimation.qmd index 05b9dc57..1a189756 100644 --- a/chapters/16-instrumental-variable-estimation.qmd +++ b/chapters/16-instrumental-variable-estimation.qmd @@ -301,10 +301,19 @@ Versions 3 and 4 are special cases: they say "$e(U)$ is constant" and "$t(U)$ is --- -::: {.proof} -## Proof (Technical Point 16.5) +::: {.callout-note title="Technical Point 16.5: Proof of the General Homogeneity Condition"} +The proof of @thm-general-homogeneity rewrites the denominator and the numerator of the usual IV estimand +as averages over the unmeasured confounders $U$: +the denominator becomes $\E{t(U)}$ and the numerator $\E{e(U)\, t(U)}$. +Zero covariance then lets the numerator factor as $\E{e(U)}\, \E{t(U)}$, +and $\E{e(U)}$ is the average causal effect +[@hernan2020causal, p. 217]. +The proof below fills in each step. +::: +::: {.proof} In Figure 16.1, $Y^a \ind (A, Z) \mid U$ for each $a$ and $U \ind Z$. +Sums over $u$ become integrals when $U$ is continuous. **Step 1: the denominator is $\E{t(U)}$.** Because $U \ind Z$, $f(u \mid z) = f(u)$, so @@ -324,13 +333,17 @@ For each $z$ and $u$, $$ \begin{aligned} \E{A (Y^{a=1} - Y^{a=0}) \mid Z = z, U = u} -&= \Pr(A = 1 \mid Z = z, U = u)\, \E{Y^{a=1} - Y^{a=0} \mid A = 1, Z = z, U = u} \\ -&= \Pr(A = 1 \mid Z = z, U = u)\, \E{Y^{a=1} - Y^{a=0} \mid U = u} && (Y^a \ind (A, Z) \mid U) \\ +&= \Pr(A = 1 \mid Z = z, U = u) \\ +&\qquad \times \E{Y^{a=1} - Y^{a=0} \mid A = 1, Z = z, U = u} \\ +&= \Pr(A = 1 \mid Z = z, U = u)\, \E{Y^{a=1} - Y^{a=0} \mid U = u} \\ &= \Pr(A = 1 \mid Z = z, U = u)\, e(u), \\ -\E{Y^{a=0} \mid Z = z, U = u} &= \E{Y^{a=0} \mid U = u}. && (Y^{a=0} \ind Z \mid U) +\E{Y^{a=0} \mid Z = z, U = u} &= \E{Y^{a=0} \mid U = u}. \end{aligned} $$ +The second equality uses $Y^a \ind (A, Z) \mid U$, +and the last line uses $Y^{a=0} \ind Z \mid U$. + Averaging over $U$ (with $f(u \mid z) = f(u)$): $$ @@ -514,9 +527,21 @@ Greenland calls compliers "cooperative" and defiers "non-cooperative" to avoid c --- -::: {.proof} -## Proof (Technical Point 16.7) +::: {.callout-note title="Technical Point 16.7: Monotonicity and the Effect in the Compliers"} +Imbens and Angrist (1994) showed that, for a dichotomous causal instrument $Z$ +and no defiers (monotonicity, condition (iv)), +the usual IV estimand is the average causal effect in the compliers (@thm-late); +a closely related argument for binary outcomes is due to Baker and Lindeman (1994) +[@hernan2020causal, p. 224]. +The proof below assumes a causal instrument (Figure 16.1). + +For a surrogate instrument $Z$, Hernán and Robins (2006b) showed that the usual IV estimator still identifies the effect in the compliers (defined by $U_Z$), +but only if $Z$ is independent of $A$ and $Y$ given $U_Z$ and $U_Z$ is binary, and that independence is rarely plausible unless $U_Z$ is continuous [@hernan2020causal, p. 224]. +So when an observational study has at best a surrogate for $U_Z$, +interpreting the IV estimand as a complier effect is often doubtful. +::: +::: {.proof} Write the intention-to-treat effect as a weighted average over the strata (AT: always-takers, NT: never-takers, C: compliers, D: defiers): $$ @@ -547,12 +572,6 @@ $$ $$ ::: -::: {.notes} -The result is due to Imbens and Angrist (1994); -Baker and Lindeman (1994) had a related proof for a binary outcome [@hernan2020causal, p. 224]. -For a surrogate instrument $Z$, Hernán and Robins (2006b) showed that the usual IV estimator still identifies the effect in the compliers (defined by $U_Z$), -but only if $Z$ is independent of $A$ and $Y$ given $U_Z$ and $U_Z$ is binary, and that independence is rarely plausible unless $U_Z$ is continuous [@hernan2020causal, p. 224]. -::: --- diff --git a/chapters/21-g-methods-time-varying.qmd b/chapters/21-g-methods-time-varying.qmd index d1ed44b3..1de77f29 100644 --- a/chapters/21-g-methods-time-varying.qmd +++ b/chapters/21-g-methods-time-varying.qmd @@ -578,6 +578,30 @@ so the plug-in estimate stays in $[0, 1]$; the estimator in the main text is an instance of this construction. ::: +::: {.callout-note title="Technical Point 21.5: A Plug-In K+2 Robust Estimator"} +With a binary $Y$, the sample average of $\hat{U}_{TR}$ from Technical Point 21.4 can fall outside $[0, 1]$. +The ICE estimator, the sample average of $\hat{b}_0(L_0)$, cannot, +as long as both outcome regressions are logistic: +it is a **plug-in** estimator. +To make it triply robust as well, +fit the outcome models so that both correction terms of $\hat{U}_{TR}$ have sample mean exactly 0; +the average of $\hat{U}_{TR}$ then coincides with the average of $\hat{b}_0(L_0)$. + +- Fit the logistic model for $\hat{b}_1(L_0, L_1)$ by maximum likelihood among individuals with $A_0 = A_1 = 1$, + adding the single covariate $A_0 A_1 / (\hat{\pi}_0 \hat{\pi}_1)$ with its own coefficient. + The score equation for that coefficient is + $\sum_i \frac{A_{0i} A_{1i}}{\hat{\pi}_{0i} \hat{\pi}_{1i}} \paren{Y_i - \hat{b}_1(L_{0i}, L_{1i})} = 0$, + which is the first correction term + (individuals outside the fitting sample contribute 0 because $A_0 A_1 = 0$). +- Fit the logistic model for $\hat{b}_0(L_0)$ among individuals with $A_0 = 1$, + with response $\hat{b}_1(L_0, L_1)$ and the extra covariate $A_0 / \hat{\pi}_0$; + the same argument makes the second correction term sum to 0. + +The doubly robust estimator of @alg-dr-time-varying, which uses the time-varying weight as a covariate, +is an instance of this plug-in estimator +[@hernan2020causal, p. 289]. +::: + ::: {.notes} ::: {.callout-warning title="Fact-Check: A Term in the Book's First Display of the Estimator"} @@ -742,8 +766,19 @@ Sequential exchangeability then says $H_k(\psi)$ is independent of $A_k$ given t --- +::: {.callout-note title="Fine Point 21.3: G-Estimation with a Saturated Structural Nested Model"} +With a saturated SNMM, g-estimation needs no search: +at the true $\psi$, sequential exchangeability forces the mean of $H_k(\psi)$ +to agree between those treated and those untreated at time $k$ +within every stratum of the past. +Each such equality is one linear equation in the parameters, +and solving them from the last time point backwards recovers $\psi$ +[@hernan2020causal, p. 293]. +@exm-g-estimation-saturated below carries this out for @tbl-seq-rand. +::: + ::: {#exm-g-estimation-saturated} -## G-Estimation in the Sequentially Randomized Experiment (Fine Point 21.3) +## G-Estimation in the Sequentially Randomized Experiment **Time 1.** Within each stratum of $(A_0, L_1)$, the mean of $H_1(\psi)$ must not depend on $A_1$. @@ -888,6 +923,53 @@ but for an SNMM linear in $\beta$ the estimating equation has a closed-form solu ::: ::: +::: {.callout-note title="Technical Point 21.8: A Closed Form Estimator for Linear Structural Nested Mean Models"} +Suppose the blip is linear in $\beta$, +$\gamma_k(\bar{a}_{k-1}, \bar{l}_k, \beta) = \tp{\beta} R_k$, +where $R_k = r_k(\bar{L}_k, \bar{A}_{k-1})$ is a vector of known functions, +and the treatment model is $\logit \Pr[A_k = 1 \mid \bar{L}_k, \bar{A}_{k-1}] = \tp{\alpha} W_k$. +Then $H_k(\beta) = Y - \tp{\beta} S_k$ with $S_k \eqdef \sum_{j=k}^{K} A_j R_j$, +so the g-estimating equation + +$$\sum_{i=1}^{N} \sum_{k=0}^{K} X_{i,k}(\hat{\alpha})\, Q_{i,k}\, \paren{Y_i - \tp{S_{i,k}} \beta} = 0, +\qquad X_{i,k}(\hat{\alpha}) \eqdef A_{i,k} - \expit\paren{\tp{\hat{\alpha}} W_{i,k}},$$ + +is linear in $\beta$ and, when the matrix below is invertible, has the explicit solution + +$$\hb = \paren{\sum_{i=1}^{N} \sum_{k=0}^{K} X_{i,k}(\hat{\alpha})\, Q_{i,k}\, \tp{S_{i,k}}}^{-1} +\sum_{i=1}^{N} \sum_{k=0}^{K} Y_i\, X_{i,k}(\hat{\alpha})\, Q_{i,k}.$$ + +Here $Q_{i,k} = q_k(\bar{L}_{i,k}, \bar{A}_{i,k-1})$ has the dimension of $\beta$; +its choice changes the efficiency of $\hb$ but not its consistency +(Robins 1994 gives the optimal choice). + +A **multiply robust** version adds a working model $\tp{\varsigma} D_k$, with $D_k = d_k(\bar{L}_k, \bar{A}_{k-1})$, +for $\E{H_k(\beta) \mid \bar{L}_k, \bar{A}_{k-1}} = \E{Y^{\bar{A}_{k-1}, \underline{0}_k} \mid \bar{L}_k, \bar{A}_{k-1}}$, +and solves jointly for $(\tilde{\beta}, \tilde{\varsigma})$ + +$$\sum_{i,k} X_{i,k}(\hat{\alpha})\, Q_{i,k}\, \paren{H_{i,k}(\beta) - \tp{\varsigma} D_{i,k}} = 0, +\qquad +\sum_{i,k} D_{i,k}\, \paren{H_{i,k}(\beta) - \tp{\varsigma} D_{i,k}} = 0,$$ + +which are again linear, so the solution is in closed form. +$\tilde{\beta}$ is consistent and asymptotically normal if, at each $k$, +either the working outcome model or the treatment model is correct, +which makes it $2^{K+1}$ multiply robust +[@hernan2020causal, p. 296]. +::: + +::: {.notes} + +::: {.callout-warning title="Fact-Check: The Book's Closed-Form Display"} +*Fact-check note*: the book's display [@hernan2020causal, p. 296] multiplies by $A_{i,k}$ outside the sum +and defines $S_{i,k}$ as a sum of $R_{i,j}$ without the treatment indicators $A_{i,j}$. +Since $H_k(\beta) = Y - \sum_{j \ge k} A_j \gamma_j$ (@def-candidate-counterfactuals-snmm), +substituting the linear blip gives $\tp{\beta} \sum_{j \ge k} A_j R_j$, +which is the form used above; +the two agree when $K = 0$ but not in general. +::: +::: + ### From $\hb$ to Counterfactual Means ::: {#rem-counterfactual-means-after-g-estimation} @@ -1120,7 +1202,7 @@ which we can never know for certain in an observational study. --- ::: {#exm-front-door-big-g} -## The Front Door Formula from the Big G-Formula (Technical Point 21.11) +## The Front Door Formula from the Big G-Formula Under Figure 7.14 ($A \to M \to Y$, with $U$ a common cause of $A$ and $Y$), the big g-formula for $\Pr[Y^a = y]$ is @@ -1147,17 +1229,69 @@ This proof does not require that counterfactuals $Y^m$ exist [@hernan2020causal, p. 301]. ::: -::: {.notes} - -::: {#rem-front-door-other-proofs} -## Other proofs of the front door formula +--- -Technical Point 21.11 also gives a coupling argument, -and Technical Point 21.12 a proof based on a SWIG property (Shpitser et al. 2022): -if a fixed treatment node $a_m$ is d-separated from $B^a$ given $C^a$ on the SWIG, -then $\Pr[B^a = b \mid C^a = c]$ does not depend on $a_m$ -[@hernan2020causal, pp. 301-302]. +::: {.callout-note title="Technical Point 21.11: A Big G-Formula Proof of the Front Door Formula"} +Technical Point 7.4 derived the front door formula using counterfactuals $Y^m$. +Starting instead from the big g-formula, +the conditional independencies of Figure 7.14 alone reduce it to observed data +(@exm-front-door-big-g), +so the front door formula holds even if no well-defined $Y^m$ exists. + +A second route, also free of $Y^m$, is a **coupling** argument. +Even if everyone agrees that $Y^m$ is not well defined, +any law of the observed data that factorizes according to Figure 7.14 +can be generated by some FFRCISTG model (Technical Point 6.2) +that is "as detailed as the data" in the sense of Robins and Richardson (2010), +and such a model formally includes a variable $Y^m$. +Within that model, Technical Point 7.4 shows that the two formulas coincide. +A factorizing law for which they differed +could therefore not be generated by any such model, +so no such distribution exists +[@hernan2020causal, p. 301]. ::: + +--- + +::: {.callout-note title="Technical Point 21.12: A Front Door Formula Proof Using d-Separation of Treatment Nodes on SWIGs"} +**A SWIG property** (Shpitser et al. 2022). +Let $G(a)$ be the SWIG for strategy $a$, +assume only treatment counterfactuals are well defined, +and let $B^a$ and $C^a$ be disjoint sets of random (not fixed) observed nodes of $G(a)$, +such as $Y^a$, $M^a$, or the natural value of treatment $A$. +If, in $G(a)$, $B^a$ and the fixed node $a_m$ are d-separated given $C^a$, +then $\Pr[B^a = b \mid C^a = c]$ does not depend on $a_m$. +This compares distributions from different single worlds; +it is not a cross-world statement. + +**Front door example.** +In the SWIG for Figure 7.14, let $C^a = (M^a, A)$ and $B^a = Y^a$. +The only path from the fixed node $a$ to $Y^a$ passes through the non-collider $M^a$, which is conditioned on, +so $\E{Y^a \mid M^a, A} = \E{Y^{a'} \mid M^{a'}, A}$ for all $a, a'$. + +**Proof of the front door formula.** +Following Technical Point 7.4, it remains to show $\E{Y^a \mid M^a = m} = \sum_{a'} \E{Y \mid M = m, A = a'} \Pr[A = a']$: + +$$ +\begin{aligned} +\E{Y^a \mid M^a = m} +&= \sum_{a'} \E{Y^a \mid M^a = m, A = a'} \Pr[A = a' \mid M^a = m] \\ +&= \sum_{a'} \E{Y^a \mid M^a = m, A = a'} \Pr[A = a'] \\ +&= \sum_{a'} \E{Y^{a'} \mid M^{a'} = m, A = a'} \Pr[A = a'] \\ +&= \sum_{a'} \E{Y \mid M = m, A = a'} \Pr[A = a']. +\end{aligned} +$$ + +The four equalities use, in turn, +the law of total expectation, +the d-separation of $M^a$ from $A$ in the SWIG, +the SWIG property just shown, +and consistency. + +So $\E{Y^a \mid M^a}$ is the same for every $a$, +yet it generally differs from $\E{Y \mid M} = \sum_{a'} \E{Y \mid M, A = a'} \Pr[A = a' \mid M]$, +because the factual $M = M^A$, unlike $M^a$, is associated with $A$ +[@hernan2020causal, p. 302]. ::: ## Summary diff --git a/chapters/23-causal-mediation.qmd b/chapters/23-causal-mediation.qmd index ffc7e08f..7313006a 100644 --- a/chapters/23-causal-mediation.qmd +++ b/chapters/23-causal-mediation.qmd @@ -183,6 +183,20 @@ The book compresses the last two lines into one step --- +::: {.callout-note title="Technical Point 23.1: Proof of the mediation formula"} +The proof of @thm-mediation-formula makes three moves. +It averages over the value the mediator would take without treatment; +it drops the conditioning on that value by cross-world independence; +and it swaps the remaining counterfactual quantities for observed ones +by exchangeability and consistency. +Only the middle move is controversial: +the cross-world independence $Y^{a=1, m} \ind M^{a=0}$ holds if Figure 23.1 is read as an NPSEM-IE, +but not if it is read as an FFRCISTG model +[@hernan2020causal, p. 324]. +::: + +--- + ### Why the mediation formula is under attack - The mediation formula identifies a cross-world quantity @@ -305,11 +319,26 @@ that the g-formula be a function of the observed data distribution, which the last line shows it is. ::: -::: {.notes} +--- + +::: {.callout-note title="Technical Point 23.2: When the mediation formula is the g-formula"} +The separable components $N$ and $O$ are exchangeable (Figure 23.3), +so the g-formula would identify $\E{Y^{n=0, o=1}}$, +except that nobody in the trial has $(N = 0, O = 1)$, so positivity fails. +Positivity, however, is sufficient rather than necessary: +given exchangeability and consistency, +the g-formula identifies the mean whenever it can be written as a function of the observed data. +The deterministic arrows $A \to N$ and $A \to O$, +together with assumptions (i) and (ii) of no direct effect, +make that function exactly the mediation formula (@thm-mediation-g-formula). + The book calls this derivation "somewhat heuristic" because determinism between $A$, $N$, and $O$ creates null sets; Robins et al. (2022) give a rigorous proof -using the SWIG Markov property of Technical Point 21.12 [@hernan2020causal, p. 326]. +using the SWIG Markov property of Technical Point 21.12. +In the same way, under the expanded diagram of Figure 23.7 +the g-formula equals the front door formula (Technical Point 23.3) +[@hernan2020causal, p. 326]. ::: ---