Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 15 additions & 3 deletions chapters/04-effect-modification.qmd
Original file line number Diff line number Diff line change
Expand Up @@ -518,8 +518,19 @@ The explanation (leaving aside random variability) is **qualitative effect modif

The lesson: always specify the population, or subset, to which an effect measure corresponds.

::: {.callout-note title="Technical Point 4.3: Relation between marginal and conditional causal risk ratios"}
Knowing the stratum-specific causal risk ratios is not enough to tell
whether the marginal causal risk ratio lies above or below 1.
The marginal ratio averages the stratum-specific ratios,
with weights that depend on how large each stratum is
and on how common the outcome would be in it without treatment.
Solving for the weights turns the question into an inequality about those two ingredients
[@hernan2020causal, p. 58].
@prp-marginal-rr-condition below works this out for Table 4.3.
:::

::: {#prp-marginal-rr-condition}
## When is the marginal risk ratio below 1? (Technical Point 4.3)
## When is the marginal risk ratio below 1?

Let $\RR_l \eqdef \Pr[Z^{a=1} = 1 \mid L = l] / \Pr[Z^{a=0} = 1 \mid L = l]$, $p_l \eqdef \Pr[Z^{a=0} = 1 \mid L = l]$, and $\pi_l \eqdef \Pr[L = l]$.
With $\RR_0 = 2$ and $\RR_1 = 0.5$, the marginal risk ratio is below 1 if and only if
Expand All @@ -539,14 +550,15 @@ $$
\end{aligned}
$$

Substituting $\RR_0 = 2$, $\RR_1 = 0.5$, and $w(1) = 1 - w(0)$:
Substituting $\RR_0 = 2$, $\RR_1 = 0.5$, and $w(1) = 1 - w(0)$,
and using $\Pr[Z^{a=0} = 1] = p_0 \pi_0 + p_1 \pi_1$:

$$
\begin{aligned}
2\, w(0) + 0.5\,(1 - w(0)) < 1
&\iff 1.5\, w(0) < 0.5 \\
&\iff w(0) < 1/3 \\
&\iff 3 p_0 \pi_0 < p_0 \pi_0 + p_1 \pi_1 && \text{(definition of } w(0)\text{; } \Pr[Z^{a=0} = 1] = p_0 \pi_0 + p_1 \pi_1\text{)} \\
&\iff 3 p_0 \pi_0 < p_0 \pi_0 + p_1 \pi_1 && \text{(definition of } w(0)\text{)} \\
&\iff 2 p_0 \pi_0 < p_1 \pi_1 \\
&\iff \frac{p_1}{p_0} > \frac{2 \pi_0}{\pi_1}.
\end{aligned}
Expand Down
14 changes: 12 additions & 2 deletions chapters/15-outcome-regression-propensity-scores.qmd
Original file line number Diff line number Diff line change
Expand Up @@ -348,13 +348,23 @@ Two caveats from the book [@hernan2020causal, pp. 201-202]:
:::
:::

::: {.notes}
---

::: {.callout-note title="Technical Point 15.1: Balancing Scores and Prognostic Scores"}
The propensity score is the simplest **balancing score** (@def-balancing-score):
once we fix its value, the treated and the untreated have the same distribution of $L$.
Rosenbaum and Rubin proved that exchangeability and positivity given $L$ imply exchangeability and positivity given any balancing score
[@rosenbaum1983central, Theorem 3; @hernan2020causal, p. 202].
So if adjusting for $L$ suffices, adjusting for $b(L)$, and in particular for $\pi(L)$, suffices too.

Graphically, $\pi(L)$ sits between $L$ and $A$, with a deterministic arrow $L \to \pi(L)$ (book Figure 15.2);
conditioning on $\pi(L)$ blocks the backdoor paths from $A$ to $Y$ through $L$ ($A \leftarrow \pi(L) \leftarrow L \to Y$), so adjusting for it suffices.

Methods based on prognostic scores need stronger assumptions and do not extend readily to time-varying treatments;
A **prognostic score** $s(L)$ is the outcome-side counterpart:
given $s(L)$, $L$ carries no further information about $Y^{a=0}$,
just as, given $b(L)$, it carries none about $A$ (@def-balancing-score).
Adjustment methods exist for both kinds of score,
but those based on prognostic scores need stronger assumptions and do not extend readily to time-varying treatments;
see Hansen (2008) and Abadie et al. (2013), as cited in @hernan2020causal [p. 202].
:::

Expand Down
45 changes: 32 additions & 13 deletions chapters/16-instrumental-variable-estimation.qmd
Original file line number Diff line number Diff line change
Expand Up @@ -301,10 +301,19 @@ Versions 3 and 4 are special cases: they say "$e(U)$ is constant" and "$t(U)$ is

---

::: {.proof}
## Proof (Technical Point 16.5)
::: {.callout-note title="Technical Point 16.5: Proof of the General Homogeneity Condition"}
The proof of @thm-general-homogeneity rewrites the denominator and the numerator of the usual IV estimand
as averages over the unmeasured confounders $U$:
the denominator becomes $\E{t(U)}$ and the numerator $\E{e(U)\, t(U)}$.
Zero covariance then lets the numerator factor as $\E{e(U)}\, \E{t(U)}$,
and $\E{e(U)}$ is the average causal effect
[@hernan2020causal, p. 217].
The proof below fills in each step.
:::

::: {.proof}
In Figure 16.1, $Y^a \ind (A, Z) \mid U$ for each $a$ and $U \ind Z$.
Sums over $u$ become integrals when $U$ is continuous.

**Step 1: the denominator is $\E{t(U)}$.**
Because $U \ind Z$, $f(u \mid z) = f(u)$, so
Expand All @@ -324,13 +333,17 @@ For each $z$ and $u$,
$$
\begin{aligned}
\E{A (Y^{a=1} - Y^{a=0}) \mid Z = z, U = u}
&= \Pr(A = 1 \mid Z = z, U = u)\, \E{Y^{a=1} - Y^{a=0} \mid A = 1, Z = z, U = u} \\
&= \Pr(A = 1 \mid Z = z, U = u)\, \E{Y^{a=1} - Y^{a=0} \mid U = u} && (Y^a \ind (A, Z) \mid U) \\
&= \Pr(A = 1 \mid Z = z, U = u) \\
&\qquad \times \E{Y^{a=1} - Y^{a=0} \mid A = 1, Z = z, U = u} \\
&= \Pr(A = 1 \mid Z = z, U = u)\, \E{Y^{a=1} - Y^{a=0} \mid U = u} \\
&= \Pr(A = 1 \mid Z = z, U = u)\, e(u), \\
\E{Y^{a=0} \mid Z = z, U = u} &= \E{Y^{a=0} \mid U = u}. && (Y^{a=0} \ind Z \mid U)
\E{Y^{a=0} \mid Z = z, U = u} &= \E{Y^{a=0} \mid U = u}.
\end{aligned}
$$

The second equality uses $Y^a \ind (A, Z) \mid U$,
and the last line uses $Y^{a=0} \ind Z \mid U$.

Averaging over $U$ (with $f(u \mid z) = f(u)$):

$$
Expand Down Expand Up @@ -514,9 +527,21 @@ Greenland calls compliers "cooperative" and defiers "non-cooperative" to avoid c

---

::: {.proof}
## Proof (Technical Point 16.7)
::: {.callout-note title="Technical Point 16.7: Monotonicity and the Effect in the Compliers"}
Imbens and Angrist (1994) showed that, for a dichotomous causal instrument $Z$
and no defiers (monotonicity, condition (iv)),
the usual IV estimand is the average causal effect in the compliers (@thm-late);
a closely related argument for binary outcomes is due to Baker and Lindeman (1994)
[@hernan2020causal, p. 224].
The proof below assumes a causal instrument (Figure 16.1).

For a surrogate instrument $Z$, Hern谩n and Robins (2006b) showed that the usual IV estimator still identifies the effect in the compliers (defined by $U_Z$),
but only if $Z$ is independent of $A$ and $Y$ given $U_Z$ and $U_Z$ is binary, and that independence is rarely plausible unless $U_Z$ is continuous [@hernan2020causal, p. 224].
So when an observational study has at best a surrogate for $U_Z$,
interpreting the IV estimand as a complier effect is often doubtful.
:::

::: {.proof}
Write the intention-to-treat effect as a weighted average over the strata (AT: always-takers, NT: never-takers, C: compliers, D: defiers):

$$
Expand Down Expand Up @@ -547,12 +572,6 @@ $$
$$
:::

::: {.notes}
The result is due to Imbens and Angrist (1994);
Baker and Lindeman (1994) had a related proof for a binary outcome [@hernan2020causal, p. 224].
For a surrogate instrument $Z$, Hern谩n and Robins (2006b) showed that the usual IV estimator still identifies the effect in the compliers (defined by $U_Z$),
but only if $Z$ is independent of $A$ and $Y$ given $U_Z$ and $U_Z$ is binary, and that independence is rarely plausible unless $U_Z$ is continuous [@hernan2020causal, p. 224].
:::

---

Expand Down
156 changes: 145 additions & 11 deletions chapters/21-g-methods-time-varying.qmd
Original file line number Diff line number Diff line change
Expand Up @@ -578,6 +578,30 @@ so the plug-in estimate stays in $[0, 1]$;
the estimator in the main text is an instance of this construction.
:::

::: {.callout-note title="Technical Point 21.5: A Plug-In K+2 Robust Estimator"}
With a binary $Y$, the sample average of $\hat{U}_{TR}$ from Technical Point 21.4 can fall outside $[0, 1]$.
The ICE estimator, the sample average of $\hat{b}_0(L_0)$, cannot,
as long as both outcome regressions are logistic:
it is a **plug-in** estimator.
To make it triply robust as well,
fit the outcome models so that both correction terms of $\hat{U}_{TR}$ have sample mean exactly 0;
the average of $\hat{U}_{TR}$ then coincides with the average of $\hat{b}_0(L_0)$.

- Fit the logistic model for $\hat{b}_1(L_0, L_1)$ by maximum likelihood among individuals with $A_0 = A_1 = 1$,
adding the single covariate $A_0 A_1 / (\hat{\pi}_0 \hat{\pi}_1)$ with its own coefficient.
The score equation for that coefficient is
$\sum_i \frac{A_{0i} A_{1i}}{\hat{\pi}_{0i} \hat{\pi}_{1i}} \paren{Y_i - \hat{b}_1(L_{0i}, L_{1i})} = 0$,
which is the first correction term
(individuals outside the fitting sample contribute 0 because $A_0 A_1 = 0$).
- Fit the logistic model for $\hat{b}_0(L_0)$ among individuals with $A_0 = 1$,
with response $\hat{b}_1(L_0, L_1)$ and the extra covariate $A_0 / \hat{\pi}_0$;
the same argument makes the second correction term sum to 0.

The doubly robust estimator of @alg-dr-time-varying, which uses the time-varying weight as a covariate,
is an instance of this plug-in estimator
[@hernan2020causal, p. 289].
:::

::: {.notes}

::: {.callout-warning title="Fact-Check: A Term in the Book's First Display of the Estimator"}
Expand Down Expand Up @@ -742,8 +766,19 @@ Sequential exchangeability then says $H_k(\psi)$ is independent of $A_k$ given t

---

::: {.callout-note title="Fine Point 21.3: G-Estimation with a Saturated Structural Nested Model"}
With a saturated SNMM, g-estimation needs no search:
at the true $\psi$, sequential exchangeability forces the mean of $H_k(\psi)$
to agree between those treated and those untreated at time $k$
within every stratum of the past.
Each such equality is one linear equation in the parameters,
and solving them from the last time point backwards recovers $\psi$
[@hernan2020causal, p. 293].
@exm-g-estimation-saturated below carries this out for @tbl-seq-rand.
:::

::: {#exm-g-estimation-saturated}
## G-Estimation in the Sequentially Randomized Experiment (Fine Point 21.3)
## G-Estimation in the Sequentially Randomized Experiment

**Time 1.**
Within each stratum of $(A_0, L_1)$, the mean of $H_1(\psi)$ must not depend on $A_1$.
Expand Down Expand Up @@ -888,6 +923,53 @@ but for an SNMM linear in $\beta$ the estimating equation has a closed-form solu
:::
:::

::: {.callout-note title="Technical Point 21.8: A Closed Form Estimator for Linear Structural Nested Mean Models"}
Suppose the blip is linear in $\beta$,
$\gamma_k(\bar{a}_{k-1}, \bar{l}_k, \beta) = \tp{\beta} R_k$,
where $R_k = r_k(\bar{L}_k, \bar{A}_{k-1})$ is a vector of known functions,
and the treatment model is $\logit \Pr[A_k = 1 \mid \bar{L}_k, \bar{A}_{k-1}] = \tp{\alpha} W_k$.
Then $H_k(\beta) = Y - \tp{\beta} S_k$ with $S_k \eqdef \sum_{j=k}^{K} A_j R_j$,
so the g-estimating equation

$$\sum_{i=1}^{N} \sum_{k=0}^{K} X_{i,k}(\hat{\alpha})\, Q_{i,k}\, \paren{Y_i - \tp{S_{i,k}} \beta} = 0,
\qquad X_{i,k}(\hat{\alpha}) \eqdef A_{i,k} - \expit\paren{\tp{\hat{\alpha}} W_{i,k}},$$

is linear in $\beta$ and, when the matrix below is invertible, has the explicit solution

$$\hb = \paren{\sum_{i=1}^{N} \sum_{k=0}^{K} X_{i,k}(\hat{\alpha})\, Q_{i,k}\, \tp{S_{i,k}}}^{-1}
\sum_{i=1}^{N} \sum_{k=0}^{K} Y_i\, X_{i,k}(\hat{\alpha})\, Q_{i,k}.$$

Here $Q_{i,k} = q_k(\bar{L}_{i,k}, \bar{A}_{i,k-1})$ has the dimension of $\beta$;
its choice changes the efficiency of $\hb$ but not its consistency
(Robins 1994 gives the optimal choice).

A **multiply robust** version adds a working model $\tp{\varsigma} D_k$, with $D_k = d_k(\bar{L}_k, \bar{A}_{k-1})$,
for $\E{H_k(\beta) \mid \bar{L}_k, \bar{A}_{k-1}} = \E{Y^{\bar{A}_{k-1}, \underline{0}_k} \mid \bar{L}_k, \bar{A}_{k-1}}$,
and solves jointly for $(\tilde{\beta}, \tilde{\varsigma})$

$$\sum_{i,k} X_{i,k}(\hat{\alpha})\, Q_{i,k}\, \paren{H_{i,k}(\beta) - \tp{\varsigma} D_{i,k}} = 0,
\qquad
\sum_{i,k} D_{i,k}\, \paren{H_{i,k}(\beta) - \tp{\varsigma} D_{i,k}} = 0,$$

which are again linear, so the solution is in closed form.
$\tilde{\beta}$ is consistent and asymptotically normal if, at each $k$,
either the working outcome model or the treatment model is correct,
which makes it $2^{K+1}$ multiply robust
[@hernan2020causal, p. 296].
:::

::: {.notes}

::: {.callout-warning title="Fact-Check: The Book's Closed-Form Display"}
*Fact-check note*: the book's display [@hernan2020causal, p. 296] multiplies by $A_{i,k}$ outside the sum
and defines $S_{i,k}$ as a sum of $R_{i,j}$ without the treatment indicators $A_{i,j}$.
Since $H_k(\beta) = Y - \sum_{j \ge k} A_j \gamma_j$ (@def-candidate-counterfactuals-snmm),
substituting the linear blip gives $\tp{\beta} \sum_{j \ge k} A_j R_j$,
which is the form used above;
the two agree when $K = 0$ but not in general.
:::
:::

### From $\hb$ to Counterfactual Means

::: {#rem-counterfactual-means-after-g-estimation}
Expand Down Expand Up @@ -1120,7 +1202,7 @@ which we can never know for certain in an observational study.
---

::: {#exm-front-door-big-g}
## The Front Door Formula from the Big G-Formula (Technical Point 21.11)
## The Front Door Formula from the Big G-Formula

Under Figure 7.14 ($A \to M \to Y$, with $U$ a common cause of $A$ and $Y$),
the big g-formula for $\Pr[Y^a = y]$ is
Expand All @@ -1147,17 +1229,69 @@ This proof does not require that counterfactuals $Y^m$ exist
[@hernan2020causal, p. 301].
:::

::: {.notes}

::: {#rem-front-door-other-proofs}
## Other proofs of the front door formula
---

Technical Point 21.11 also gives a coupling argument,
and Technical Point 21.12 a proof based on a SWIG property (Shpitser et al. 2022):
if a fixed treatment node $a_m$ is d-separated from $B^a$ given $C^a$ on the SWIG,
then $\Pr[B^a = b \mid C^a = c]$ does not depend on $a_m$
[@hernan2020causal, pp. 301-302].
::: {.callout-note title="Technical Point 21.11: A Big G-Formula Proof of the Front Door Formula"}
Technical Point 7.4 derived the front door formula using counterfactuals $Y^m$.
Starting instead from the big g-formula,
the conditional independencies of Figure 7.14 alone reduce it to observed data
(@exm-front-door-big-g),
so the front door formula holds even if no well-defined $Y^m$ exists.

A second route, also free of $Y^m$, is a **coupling** argument.
Even if everyone agrees that $Y^m$ is not well defined,
any law of the observed data that factorizes according to Figure 7.14
can be generated by some FFRCISTG model (Technical Point 6.2)
that is "as detailed as the data" in the sense of Robins and Richardson (2010),
and such a model formally includes a variable $Y^m$.
Within that model, Technical Point 7.4 shows that the two formulas coincide.
A factorizing law for which they differed
could therefore not be generated by any such model,
so no such distribution exists
[@hernan2020causal, p. 301].
:::

---

::: {.callout-note title="Technical Point 21.12: A Front Door Formula Proof Using d-Separation of Treatment Nodes on SWIGs"}
**A SWIG property** (Shpitser et al. 2022).
Let $G(a)$ be the SWIG for strategy $a$,
assume only treatment counterfactuals are well defined,
and let $B^a$ and $C^a$ be disjoint sets of random (not fixed) observed nodes of $G(a)$,
such as $Y^a$, $M^a$, or the natural value of treatment $A$.
If, in $G(a)$, $B^a$ and the fixed node $a_m$ are d-separated given $C^a$,
then $\Pr[B^a = b \mid C^a = c]$ does not depend on $a_m$.
This compares distributions from different single worlds;
it is not a cross-world statement.

**Front door example.**
In the SWIG for Figure 7.14, let $C^a = (M^a, A)$ and $B^a = Y^a$.
The only path from the fixed node $a$ to $Y^a$ passes through the non-collider $M^a$, which is conditioned on,
so $\E{Y^a \mid M^a, A} = \E{Y^{a'} \mid M^{a'}, A}$ for all $a, a'$.

**Proof of the front door formula.**
Following Technical Point 7.4, it remains to show $\E{Y^a \mid M^a = m} = \sum_{a'} \E{Y \mid M = m, A = a'} \Pr[A = a']$:

$$
\begin{aligned}
\E{Y^a \mid M^a = m}
&= \sum_{a'} \E{Y^a \mid M^a = m, A = a'} \Pr[A = a' \mid M^a = m] \\
&= \sum_{a'} \E{Y^a \mid M^a = m, A = a'} \Pr[A = a'] \\
&= \sum_{a'} \E{Y^{a'} \mid M^{a'} = m, A = a'} \Pr[A = a'] \\
&= \sum_{a'} \E{Y \mid M = m, A = a'} \Pr[A = a'].
\end{aligned}
$$

The four equalities use, in turn,
the law of total expectation,
the d-separation of $M^a$ from $A$ in the SWIG,
the SWIG property just shown,
and consistency.

So $\E{Y^a \mid M^a}$ is the same for every $a$,
yet it generally differs from $\E{Y \mid M} = \sum_{a'} \E{Y \mid M, A = a'} \Pr[A = a' \mid M]$,
because the factual $M = M^A$, unlike $M^a$, is associated with $A$
[@hernan2020causal, p. 302].
:::

## Summary
Expand Down
33 changes: 31 additions & 2 deletions chapters/23-causal-mediation.qmd
Original file line number Diff line number Diff line change
Expand Up @@ -183,6 +183,20 @@ The book compresses the last two lines into one step

---

::: {.callout-note title="Technical Point 23.1: Proof of the mediation formula"}
The proof of @thm-mediation-formula makes three moves.
It averages over the value the mediator would take without treatment;
it drops the conditioning on that value by cross-world independence;
and it swaps the remaining counterfactual quantities for observed ones
by exchangeability and consistency.
Only the middle move is controversial:
the cross-world independence $Y^{a=1, m} \ind M^{a=0}$ holds if Figure 23.1 is read as an NPSEM-IE,
but not if it is read as an FFRCISTG model
[@hernan2020causal, p. 324].
:::

---

### Why the mediation formula is under attack

- The mediation formula identifies a cross-world quantity
Expand Down Expand Up @@ -305,11 +319,26 @@ that the g-formula be a function of the observed data distribution,
which the last line shows it is.
:::

::: {.notes}
---

::: {.callout-note title="Technical Point 23.2: When the mediation formula is the g-formula"}
The separable components $N$ and $O$ are exchangeable (Figure 23.3),
so the g-formula would identify $\E{Y^{n=0, o=1}}$,
except that nobody in the trial has $(N = 0, O = 1)$, so positivity fails.
Positivity, however, is sufficient rather than necessary:
given exchangeability and consistency,
the g-formula identifies the mean whenever it can be written as a function of the observed data.
The deterministic arrows $A \to N$ and $A \to O$,
together with assumptions (i) and (ii) of no direct effect,
make that function exactly the mediation formula (@thm-mediation-g-formula).

The book calls this derivation "somewhat heuristic"
because determinism between $A$, $N$, and $O$ creates null sets;
Robins et al. (2022) give a rigorous proof
using the SWIG Markov property of Technical Point 21.12 [@hernan2020causal, p. 326].
using the SWIG Markov property of Technical Point 21.12.
In the same way, under the expanded diagram of Figure 23.7
the g-formula equals the front door formula (Technical Point 23.3)
[@hernan2020causal, p. 326].
:::

---
Expand Down
Loading