# Shift-Share Instrumental Variables
```{r}
#| include: false
library(tidyverse)
library(fixest)
library(ggplot2)
library(knitr)
```
Shift-share, or Bartik, instruments answer one specific problem, so it is
clearest to write the regression down first.
We want the effect of a local economic shock on a local outcome. For region
$\ell$,
$$
Y_\ell \;=\; \beta X_\ell + Z_\ell'\gamma + u_\ell ,
$$ {#eq-shift-share-iv-structural}
where $X_\ell$ is the shock the region actually experienced --- local import
competition per worker, in the China-shock literature --- and $Y_\ell$ is the
outcome, say the change in the local manufacturing employment share. The
parameter we want is $\beta$.
The endogenous regressor is $X_\ell$. Regions do not receive their shocks at
random. A region whose import competition rises sharply may be a region whose
industries were already losing ground, and whatever caused that decline sits in
$u_\ell$ and moves $Y_\ell$ on its own. Then $\text{Cov}(X_\ell, u_\ell) \ne 0$,
and OLS mixes the effect of the shock with the local conditions that produced
it. We cannot sign that bias in advance; it depends on whether those local
conditions push the outcome up or down.
The shift-share instrument replaces the realised local shock with a predicted
one, assembled from two pieces that are meant to have nothing to do with
$u_\ell$: each region's *baseline* industry shares, fixed before the outcome
period, and *national* or foreign industry-level shocks, common across regions.
A region with a large baseline share in an industry is mechanically more exposed
to a shock hitting that industry, and that exposure is predicted rather than
realised.
Building the instrument is the easy part. The difficult part is saying which of
the two pieces carries the exogeneity assumption --- the shares, the shocks, or
both --- because the two answers lead to different identification arguments,
different diagnostics, and different standard errors. That is what the rest of
the chapter is about.
## The construction
For each location $\ell$ and industry $k$, let $s_{\ell k}$ be the share of
local employment in industry $k$ at baseline, and let $g_k$ be a national
growth rate (or trade shock) for that industry. The shift-share instrument is
$$
B_\ell \;=\; \sum_{k=1}^{K} s_{\ell k}\, g_k.
$$ {#eq-shift-share-iv-1}
This is "shift-share" because it combines a local share with an
industry-level shift. @bartik-1991 used local industry mix to predict local
employment growth. @autor-dorn-hanson-2013 used local industry shares
interacted with industry-level Chinese import growth in other high-income
countries.
The first-stage regression is then
$$
\Delta L_\ell \;=\; \pi_0 + \pi_1 B_\ell + X_\ell'\gamma + \epsilon_\ell,
$$ {#eq-shift-share-iv-2}
with $B_\ell$ instrumenting for the observed local shock in a 2SLS for the
outcome of interest.
## Two identification views
There are two ways to justify a shift-share instrument. They put the
exogeneity assumption on different objects.
### Share view (GPSS 2020)
Treat the shares $s_{\ell k}$ as the source of identifying variation, with the
shocks $g_k$ acting as weights. The instrument is valid if shares are
exogenous to the outcome (conditional on controls).
@goldsmith-pinkham-sorkin-swift-2020 (henceforth GPSS) prove that the
shift-share IV is numerically equivalent to a GMM combination of $K$
just-identified IVs, one per industry share:
$$
\hat\beta^{SS} \;=\; \sum_{k=1}^{K} \hat\alpha_k\, \hat\beta_k,
$$ {#eq-shift-share-iv-3}
where $\hat\beta_k$ is the just-identified IV using share $s_{\cdot k}$ alone
and $\hat\alpha_k$ is the **Rotemberg weight**. This gives an important
diagnostic: which industry shares are driving the estimate?
### Shock view (BHJ 2022)
Treat the shocks $g_k$ as the source of identifying variation. Shares are
exposure weights. Under this view, shocks should be uncorrelated with the
second-stage unobservables, conditional on industry-level controls.
@borusyak-hull-jaravel-2022 (henceforth BHJ) show how to rewrite the
regression at the shock level.
The two views are not mutually exclusive. In an application, we should be
clear which one is being defended and report diagnostics for both.
## Inference
Standard OLS or cluster-robust standard errors on the regional regression
**underestimate uncertainty** because the same industry shocks appear in many
regions, inducing cross-region correlation that region-level clustering does
not capture. Two corrections are now standard:
- **Cluster on shocks (BHJ)**: equivalent to running the regression at the
industry-shock level. Implemented via shock-level reweighting.
- **AKM SE** [@adao-kolesar-morales-2019]: explicit formula for SE accounting
for shock-level correlation. Available in the `ShiftShareSE` R package.
## Simulation: build intuition
The aim is to see the shift-share IV recover a known effect that OLS misses,
and then to break the identifying assumption on purpose and see which
diagnostic catches it.
We use 500 regions and 20 industries. Each region has a vector of industry
shares $s_{\ell k}$ summing to one, concentrated so that a few industries
dominate any given region. Each industry draws a shock $g_k \sim N(0,1)$,
independently of the shares and of everything at the region level. The
instrument is $B_\ell = \sum_k s_{\ell k} g_k$.
There is one region-level confounder $u_\ell \sim N(0,1)$. The observed local
shock and the outcome are
$$
X_\ell = B_\ell + 0.3u_\ell + \varepsilon_\ell, \qquad
Y_\ell = 0.5 X_\ell + u_\ell + \eta_\ell,
$$ {#eq-shift-share-iv-5}
with $\varepsilon$ and $\eta$ independent noise of standard deviation about 0.3
and 0.5. The true effect is $\beta = 0.5$.
Because $u$ enters both $X$ and $Y$, OLS is biased, and we can say by how much:
$\text{Cov}(X,u)/\text{Var}(X) = 0.3/0.315 = 0.95$, so OLS should return about
$0.5 + 0.95 = 1.45$.
The data are generated once and stored as CSVs so that this book and its Julia
companion produce identical numbers.
```{r}
#| label: sim-data
#| cache: true
n_region <- 500
n_industry <- 20
beta_true <- 0.5
df <- read_csv("data/shift_share_sim.csv", show_col_types = FALSE)
shares <- as.matrix(read_csv("data/shift_share_shares.csv", show_col_types = FALSE))
# This simulation is constructed to satisfy the Borusyak et al. (2022)
# shock-exogeneity assumption: the industry-level shocks are drawn
# independently of the region-level confounder `u` and of the shares
# themselves. The shares are allowed to be endogenous (correlated with
# region characteristics) -- it is specifically the shocks that must be
# "as good as randomly assigned" for this identification argument, which is
# a different (and in this simulation, the maintained) assumption from the
# Goldsmith-Pinkham et al. (2020) shares-exogeneity view discussed above.
shocks <- read_csv("data/shift_share_shocks.csv", show_col_types = FALSE)$shock
u <- df$u
head(df)
```
A naive OLS suffers from the confounder `u`:
```{r}
#| label: naive-ols
#| cache: true
ols <- feols(Y ~ X, data = df)
ivss <- feols(Y ~ 1 | X ~ B, data = df)
etable(ols, ivss, headers = c("OLS", "Shift-share IV"),
digits = 3, digits.stats = 3, fitstat = ~ . + ivf)
```
OLS gives 1.45 with a standard error of 0.075, matching the 1.45 predicted
above and nearly three times the true effect of `r beta_true`. The shift-share
IV gives 0.531 with a standard error of 0.131, within a quarter of a standard
error of the truth. The first-stage $F$ is 370.5, so the instrument is strong.
Note the price again: the IV standard error is 0.131 against 0.075 for OLS, and
the second-stage $R^2$ is 0.025 against 0.428. A biased estimator can fit well.
### Rotemberg weights
The GPSS decomposition says the shift-share IV is a weighted average of
$K$ just-identified IVs (one per industry share). Compute the weights:
```{r}
#| label: rotemberg
#| cache: true
# Rotemberg weight for industry k (using centered moments to match the
# demeaned 2SLS regression):
# alpha_k = g_k * Cov(s_{.k}, X) / sum_j g_j * Cov(s_{.j}, X)
# Each industry's just-identified IV estimate uses s_{.k} as instrument:
# beta_k = Cov(s_{.k}, Y) / Cov(s_{.k}, X)
# GPSS prove the shift-share IV estimate equals sum_k alpha_k * beta_k.
cov_sk_X <- sapply(1:n_industry, function(k) cov(shares[, k], df$X))
cov_sk_Y <- sapply(1:n_industry, function(k) cov(shares[, k], df$Y))
denom <- sum(shocks * cov_sk_X)
alpha <- shocks * cov_sk_X / denom
beta_k <- cov_sk_Y / cov_sk_X
rotemberg <- tibble(industry = 1:n_industry,
shock = shocks,
weight = alpha,
beta_k = beta_k)
kable(rotemberg, digits = 3,
caption = paste0("Rotemberg decomposition. Sum of alpha*beta_k = ",
round(sum(alpha * beta_k), 3),
" (= shift-share IV estimate of ",
round(coef(ivss)["fit_X"], 3), ")."))
```
The weighted sum reproduces the 2SLS estimate exactly: $\sum_k \alpha_k
\hat\beta_k = 0.531$, which is the shift-share IV coefficient. That is the GPSS
equivalence, and it holds numerically, not approximately.
The weights are concentrated. Industry 5 carries 0.387 and industry 16 carries
0.204, so those two supply nearly 60% of the identifying variation, and the
next three (industries 20, 6 and 9) bring the total past 80%. Under the share
view, the exogeneity of those few industries' shares is the assumption the
estimate rests on. Notice that the two dominant industries are also the two
with the largest shocks, 2.66 and 2.27: $\alpha_k$ is proportional to $g_k
\text{Cov}(s_{\cdot k}, X)$, so a big shock buys influence.
The $\hat\beta_k$ column is more surprising and worth reading carefully.
Industry 18 gives $-39.5$, industry 4 gives $-10.2$, industry 15 gives $-3.2$.
These are not near the true 0.5, and it would be easy to read them as evidence
of severe treatment-effect heterogeneity. They are not. Each $\hat\beta_k =
\text{Cov}(s_{\cdot k}, Y)/\text{Cov}(s_{\cdot k}, X)$ is a just-identified IV
using one share as the instrument, and when $\text{Cov}(s_{\cdot k}, X)$ is
near zero that ratio explodes. It is a weak-instrument artefact, one industry
at a time.
The decomposition is immune to it by construction. Since
$$
\alpha_k \hat\beta_k
= \frac{g_k \text{Cov}(s_{\cdot k}, X)}{\sum_j g_j \text{Cov}(s_{\cdot j}, X)}
\cdot \frac{\text{Cov}(s_{\cdot k}, Y)}{\text{Cov}(s_{\cdot k}, X)}
= \frac{g_k \text{Cov}(s_{\cdot k}, Y)}{\sum_j g_j \text{Cov}(s_{\cdot j}, X)},
$$ {#eq-shift-share-iv-6}
the troublesome $\text{Cov}(s_{\cdot k}, X)$ cancels. Every industry with an
exploding $\hat\beta_k$ has a weight near zero for exactly the same reason its
$\hat\beta_k$ blew up. Industry 18 carries a weight of 0.000 and industry 4 a
weight of 0.002.
Read the two columns together, then. Among the industries that carry real
weight, the estimates are 0.469, 0.746, 0.554, 1.025, 0.853 and 0.604 --- close
enough to 0.5 to be sampling noise, which is what we should see in a simulation
with a homogeneous effect. In real data, if industries with *large* weights
give very different estimates, the shift-share IV is averaging over genuine
heterogeneity and that should be reported.
The plot below makes the pairing visible: each industry's just-identified IV
estimate on the vertical axis, with point size proportional to its Rotemberg
weight. The outlying estimates are the smallest points.
```{r}
#| label: rotemberg-plot
#| cache: true
#| fig-height: 4
#| fig-width: 7
ggplot(rotemberg, aes(x = factor(industry), y = beta_k, size = abs(weight))) +
geom_hline(yintercept = beta_true, linetype = "dashed", color = "red") +
geom_point(alpha = 0.7) +
scale_size_continuous(name = "|Rotemberg weight|") +
labs(x = "Industry", y = "Just-identified IV estimate",
title = "Industry-level IV estimates and weights",
subtitle = paste("Red line = true effect =", beta_true)) +
theme_minimal()
```
### What goes wrong when shares are endogenous
Now break the share-exogeneity assumption on purpose. We introduce a second
region-level confounder $v_\ell$, add $0.15 v_\ell$ to industry 1's share
(renormalising the rows to sum to one), and let $v$ enter the outcome with
coefficient 0.6:
$$
Y_\ell = 0.5 X_\ell + u_\ell + 0.6 v_\ell + \eta_\ell .
$$ {#eq-shift-share-iv-7}
Everything else is unchanged. Industry 1's share is now correlated with the
second-stage error, so the instrument built from it is invalid, and we ask
which diagnostic notices.
```{r}
#| label: bad-shares
#| cache: true
# Re-do shares so industry 1's share covaries with a new confounder v
v <- read_csv("data/shift_share_bad_v.csv", show_col_types = FALSE)$v
shares_bad <- shares
shares_bad[, 1] <- pmax(0.01, shares[, 1] + 0.15 * v)
shares_bad <- shares_bad / rowSums(shares_bad) # renormalise to sum to 1
bad_noise <- read_csv("data/shift_share_bad_noise.csv", show_col_types = FALSE)
B_bad <- as.numeric(shares_bad %*% shocks)
X_bad <- B_bad + 0.3 * u + bad_noise$noise_x
Y_bad <- beta_true * X_bad + u + 0.6 * v + bad_noise$noise_y
df_bad <- tibble(X = X_bad, Y = Y_bad, B = B_bad)
ivss_bad <- feols(Y ~ 1 | X ~ B, data = df_bad)
cat("True beta:", beta_true, "\n")
cat("Shift-share IV with bad share 1:", round(coef(ivss_bad)["fit_X"], 3), "\n")
# Rotemberg weights recomputed under the bad shares
cov_sk_X_bad <- sapply(1:n_industry, function(k) cov(shares_bad[, k], X_bad))
alpha_bad <- shocks * cov_sk_X_bad / sum(shocks * cov_sk_X_bad)
cat("Rotemberg weight of industry 1:", round(alpha_bad[1], 3),
"-- rank", rank(-abs(alpha_bad))[1], "of", n_industry, "\n")
# Leave-one-industry-out: rebuild the instrument without industry 1
B_loo <- as.numeric(shares_bad[, -1] %*% shocks[-1])
ivss_loo <- feols(Y ~ 1 | X ~ B, data = tibble(X = X_bad, Y = Y_bad, B = B_loo))
cat("Shift-share IV without industry 1:", round(coef(ivss_loo)["fit_X"], 3), "\n")
```
The estimate moves from 0.531 to 0.798 against a true 0.5, so one endogenous
share out of twenty is enough to inflate the estimate by 60%. Note which
diagnostic catches it. Rotemberg weights
measure *influence* --- whose exogeneity the estimate leans on --- not
endogeneity: the offending industry carries a weight of only 0.11, third
largest of twenty, yet it alone moves the estimate from 0.5 to 0.8. The
leave-one-industry-out check is what isolates it: rebuilding the instrument
without industry 1 restores the estimate to 0.48.
## Empirical: the China shock (Autor, Dorn, Hanson 2013)
The standard shift-share application is the China-shock study of
@autor-dorn-hanson-2013 (henceforth ADH). Commuting zones are exposed
differently to Chinese import competition because their baseline industry
mixes differ. The instrument is
$$
\text{IV}_\ell \;=\; \sum_k \frac{L_{\ell k,\,1990}}{L_{\ell,\,1990}}\,
\frac{\Delta M^{other}_{k}}{L_{k,\,1990}},
$$ {#eq-shift-share-iv-4}
where the shocks $\Delta M^{other}_k$ are growth in Chinese imports into
other high-income countries. This leave-one-out construction removes
US-specific demand from the shock. In the actual ADH design the employment
shares are lagged one census behind the outcome period (e.g. 1980 employment
for the 1990--2000 stack), precisely to mitigate share endogeneity; we date
everything to 1990 here to keep the notation light.
```{r}
#| label: adh-skeleton
#| eval: false
#| echo: true
# David Dorn distributes the replication data at
# https://www.ddorn.net/data.htm — files needed:
# workfile_china.dta (commuting zone data, 1990-2007)
# industry_shares.csv (czone-industry shares)
# industry_imports.csv (industry-level imports)
library(haven)
library(fixest)
cz <- read_dta("workfile_china.dta")
fit <- feols(d_sh_empl_mfg ~ 1 | d_tradeusch_pw ~ d_tradeotch_pw_lag,
data = cz,
cluster = ~ statefip)
summary(fit)
```
Modern best-practice extensions of the basic ADH regression. These blocks are
schematic --- they do not run here, so check argument names against the
current package documentation before adapting them:
```{r}
#| label: adh-modern
#| eval: false
#| echo: true
# 1. Rotemberg weights via the bartik.weight package (Goldsmith-Pinkham)
# devtools::install_github("paulgp/bartik-weight")
library(bartik.weight)
rw <- bw(cz, master = master_spec, y = "d_sh_empl_mfg",
x = "d_tradeusch_pw", weight = "timepwt48",
G = G_growth, Z = Z_shares)
# Plot the Rotemberg weights to see which industries drive the estimate
plot(rw)
# 2. Adão-Kolesár-Morales standard errors
# devtools::install_github("kolesarm/ShiftShareSE")
library(ShiftShareSE)
ivreg_ss(d_sh_empl_mfg ~ d_tradeusch_pw + controls,
X = "d_tradeotch_pw_lag",
data = cz, W = shock_weights, region_cvar = "czone",
method = "akm0")
# 3. BHJ shock-level inference: collapse to shock (industry-period) level
# devtools::install_github("borusyak/shift-share")
library(ssaggregate)
shocks_data <- ssaggregate(data = cz, vars = c("d_sh_empl_mfg",
"d_tradeusch_pw"),
shock = "d_tradeotch_pw_lag", weights = "timepwt48",
l = "czone", n = "industry", t = "year",
s = "share")
# Now regress at shock level — interpretation: shock-level IV
feols(d_sh_empl_mfg ~ 1 | d_tradeusch_pw ~ shock, data = shocks_data)
```
For a new shift-share paper, I would expect Rotemberg weights, AKM standard
errors, and a BHJ shock-level version or an explanation for why it is not
appropriate.
## Summary
- Shift-share IV combines local shares and industry shocks.
- The GPSS view puts exogeneity on shares; the BHJ view puts it on shocks.
- Rotemberg weights show which industry shares drive the 2SLS estimate.
- Region-level clustering is usually too optimistic. Use AKM standard errors
or shock-level inference when possible.
- In R, `fixest`, `bartik.weight`, `ShiftShareSE`, and `ssaggregate` cover
most of the workflow.
For a longer treatment --- more on the GPSS decomposition, the Rotemberg weights,
and the inference options --- see
[The Bartik instrument](https://xiangao.github.io/blog_book/bartik-instrument.html)
in *Topics on Econometrics and Causal Inference*.