Job Training and Earnings
This vignette summarizes the National Supported Work (NSW) job-training example. NSW was a social experiment from the 1970s that targeted disadvantaged workers. Treated participants were offered a supported work placement for a limited period, and later earnings were measured from administrative records. The data are mirrored by RDatasets. The full Quarto source is in vignettes/Job_Training_NSW.qmd.
Causal Question
What is the effect of job-training participation on real earnings in 1978?
The treatment is treat, the outcome is re78, and the baseline covariates are age, education, race/ethnicity indicators, marital status, degree status, and pre-treatment real earnings in 1974 and 1975.
Data and Research Idea
The research idea is the classic program-evaluation problem: can an employment and training program raise later earnings? NSW is useful because it connects an experimental benchmark to the observational problem economists often face. If assignment is randomized, a simple comparison can estimate the effect. If participation is analyzed as if it came from a non-random labor market or administrative process, baseline earnings, education, age, race, marital status, and degree status become part of the identifying story.
The RDatasets file used here has 445 observations and 11 variables. The analysis below keeps the variables needed for the treatment, outcome, and baseline adjustment set.
Experimental Graph
If the NSW sample is treated as a randomized experiment, the graph is:
A -> Yusing CausalGraphs
experimental_graph = make_graph(
vertices = [:treat, :re78],
di_edges = [(:treat, :re78)],
)
identify(experimental_graph, :treat, :re78).strategy:a_fixabledraw_graph(experimental_graph; direction="LR")markov_pillow(experimental_graph, :treat; treatment=:treat)Symbol[]No baseline adjustment is required under this graph.
Observational Selection Graph
For an observational analysis, the graph allows measured baseline variables to affect both program participation and later earnings:
X -> A -> Y
X -----> Yconceptual_graph = make_graph(
vertices = [:X, :A, :Y],
di_edges = [(:X, :A), (:X, :Y), (:A, :Y)],
)
draw_graph(conceptual_graph; direction="LR")vertices = [
:age, :educ, :black, :hisp, :marr, :nodegree,
:re74, :re75, :treat, :re78,
]
baseline = setdiff(vertices, [:treat, :re78])
edges = vcat(
[(w, :treat) for w in baseline],
[(w, :re78) for w in baseline],
[(:treat, :re78)],
)
observational_graph = make_graph(vertices=vertices, di_edges=edges)
identify(observational_graph, :treat, :re78).strategy:a_fixabledraw_graph(observational_graph; direction="LR")markov_pillow(observational_graph, :treat; treatment=:treat)8-element Vector{Symbol}:
:age
:educ
:black
:hisp
:marr
:nodegree
:re74
:re75Under this graph, the effect is identified by adjusting for the measured baseline covariates.
Unmeasured Selection
If unmeasured motivation, caseworker discretion, health, networks, or local labor-market opportunity confound training and later earnings, encode that as a bidirected edge:
X -> A -> Y
X -----> Y
A <-> Yselection_graph = make_graph(
vertices = vertices,
di_edges = edges,
bi_edges = [(:treat, :re78)],
)
identify(selection_graph, :treat, :re78).strategy:not_identifieddraw_graph(selection_graph; direction="LR")With that additional assumption, CausalGraphs.jl reports the effect as not identified by the implemented criteria.
Estimation
The NSW data can be loaded directly from RDatasets. The estimate under the experimental graph uses only treat and re78, while the observational graph also adjusts for the baseline covariates through the Markov pillow of treat.
using DataFrames, DelimitedFiles, Downloads, Statistics
function read_rdatasets_csv(url, cols)
local_file = Downloads.download(url; timeout=120)
x, header = readdlm(local_file, ',', Any, '\n'; header=true)
raw = DataFrame(x, Symbol.(vec(header)))
DataFrame((c => Float64.(raw[!, c]) for c in cols)...)
end
url = "https://vincentarelbundock.github.io/Rdatasets/csv/causaldata/nsw_mixtape.csv"
cols = [
:treat, :age, :educ, :black, :hisp, :marr, :nodegree,
:re74, :re75, :re78,
]
data = read_rdatasets_csv(url, cols)
(n = nrow(data),
treated = Int(sum(data.treat .== 1)),
controls = Int(sum(data.treat .== 0)))(n = 445, treated = 185, controls = 260)Under the experimental graph, no baseline adjustment is made.
raw_difference = mean(data.re78[data.treat .== 1]) -
mean(data.re78[data.treat .== 0])
experimental_res = estimate_causal(
a = [1, 0],
data = select(data, [:treat, :re78]),
graph = experimental_graph,
treatment = :treat,
outcome = :re78,
)
r = x -> round(x, sigdigits=4)
(raw_difference = r(raw_difference),
TMLE_ACE = r(experimental_res[:TMLE].ACE),
lower_ci = r(experimental_res[:TMLE].lower_ci),
upper_ci = r(experimental_res[:TMLE].upper_ci))(raw_difference = 1794.0, TMLE_ACE = 1794.0, lower_ci = 482.5, upper_ci = 3106.0)Under the observational selection-on-observables graph, the estimator adjusts for the measured baseline variables.
observational_res = estimate_causal(
a = [1, 0],
data = data,
graph = observational_graph,
treatment = :treat,
outcome = :re78,
)
(TMLE_ACE = r(observational_res[:TMLE].ACE),
lower_ci = r(observational_res[:TMLE].lower_ci),
upper_ci = r(observational_res[:TMLE].upper_ci),
Onestep_ACE = r(observational_res[:Onestep].ACE),
Gcomp_ACE = r(observational_res[:Gcomp].ACE),
IPW_ACE = r(observational_res[:IPW].ACE))(TMLE_ACE = 1638.0, lower_ci = 321.8, upper_ci = 2953.0, Onestep_ACE = 1638.0, Gcomp_ACE = 1676.0, IPW_ACE = 1613.0)The key point is not the small numerical difference in this sample. It is that each estimate is tied to a different graph. Under the unmeasured-selection graph above, the effect is not identified, so estimate_causal should not be run.
Takeaway
The NSW example is useful because it makes the identifying assumptions visible:
- randomized assignment implies no adjustment;
- measured selection implies adjustment for baseline covariates;
- unmeasured selection makes the effect non-identified in this ADMG.