Job Training and Earnings

This vignette summarizes the National Supported Work (NSW) job-training example. NSW was a social experiment from the 1970s that targeted disadvantaged workers. Treated participants were offered a supported work placement for a limited period, and later earnings were measured from administrative records. The data are mirrored by RDatasets. The full Quarto source is in vignettes/Job_Training_NSW.qmd.

Causal Question

What is the effect of job-training participation on real earnings in 1978?

The treatment is treat, the outcome is re78, and the baseline covariates are age, education, race/ethnicity indicators, marital status, degree status, and pre-treatment real earnings in 1974 and 1975.

Data and Research Idea

The research idea is the classic program-evaluation problem: can an employment and training program raise later earnings? NSW is useful because it connects an experimental benchmark to the observational problem economists often face. If assignment is randomized, a simple comparison can estimate the effect. If participation is analyzed as if it came from a non-random labor market or administrative process, baseline earnings, education, age, race, marital status, and degree status become part of the identifying story.

The RDatasets file used here has 445 observations and 11 variables. The analysis below keeps the variables needed for the treatment, outcome, and baseline adjustment set.

Experimental Graph

If the NSW sample is treated as a randomized experiment, the graph is:

A -> Y
using CausalGraphs

experimental_graph = make_graph(
    vertices = [:treat, :re78],
    di_edges = [(:treat, :re78)],
)

identify(experimental_graph, :treat, :re78).strategy
:a_fixable
draw_graph(experimental_graph; direction="LR")
%3 treat treat re78 re78 treat->re78
markov_pillow(experimental_graph, :treat; treatment=:treat)
Symbol[]

No baseline adjustment is required under this graph.

Observational Selection Graph

For an observational analysis, the graph allows measured baseline variables to affect both program participation and later earnings:

X -> A -> Y
X -----> Y
conceptual_graph = make_graph(
    vertices = [:X, :A, :Y],
    di_edges = [(:X, :A), (:X, :Y), (:A, :Y)],
)

draw_graph(conceptual_graph; direction="LR")
%3 X X A A X->A Y Y X->Y A->Y
vertices = [
    :age, :educ, :black, :hisp, :marr, :nodegree,
    :re74, :re75, :treat, :re78,
]

baseline = setdiff(vertices, [:treat, :re78])
edges = vcat(
    [(w, :treat) for w in baseline],
    [(w, :re78) for w in baseline],
    [(:treat, :re78)],
)

observational_graph = make_graph(vertices=vertices, di_edges=edges)
identify(observational_graph, :treat, :re78).strategy
:a_fixable
draw_graph(observational_graph; direction="LR")
%3 age age treat treat age->treat re78 re78 age->re78 educ educ educ->treat educ->re78 black black black->treat black->re78 hisp hisp hisp->treat hisp->re78 marr marr marr->treat marr->re78 nodegree nodegree nodegree->treat nodegree->re78 re74 re74 re74->treat re74->re78 re75 re75 re75->treat re75->re78 treat->re78
markov_pillow(observational_graph, :treat; treatment=:treat)
8-element Vector{Symbol}:
 :age
 :educ
 :black
 :hisp
 :marr
 :nodegree
 :re74
 :re75

Under this graph, the effect is identified by adjusting for the measured baseline covariates.

Unmeasured Selection

If unmeasured motivation, caseworker discretion, health, networks, or local labor-market opportunity confound training and later earnings, encode that as a bidirected edge:

X -> A -> Y
X -----> Y
A <-> Y
selection_graph = make_graph(
    vertices = vertices,
    di_edges = edges,
    bi_edges = [(:treat, :re78)],
)

identify(selection_graph, :treat, :re78).strategy
:not_identified
draw_graph(selection_graph; direction="LR")
%3 age age treat treat age->treat re78 re78 age->re78 educ educ educ->treat educ->re78 black black black->treat black->re78 hisp hisp hisp->treat hisp->re78 marr marr marr->treat marr->re78 nodegree nodegree nodegree->treat nodegree->re78 re74 re74 re74->treat re74->re78 re75 re75 re75->treat re75->re78 treat->re78 treat->re78

With that additional assumption, CausalGraphs.jl reports the effect as not identified by the implemented criteria.

Estimation

The NSW data can be loaded directly from RDatasets. The estimate under the experimental graph uses only treat and re78, while the observational graph also adjusts for the baseline covariates through the Markov pillow of treat.

using DataFrames, DelimitedFiles, Downloads, Statistics

function read_rdatasets_csv(url, cols)
    local_file = Downloads.download(url; timeout=120)
    x, header = readdlm(local_file, ',', Any, '\n'; header=true)
    raw = DataFrame(x, Symbol.(vec(header)))
    DataFrame((c => Float64.(raw[!, c]) for c in cols)...)
end

url = "https://vincentarelbundock.github.io/Rdatasets/csv/causaldata/nsw_mixtape.csv"
cols = [
    :treat, :age, :educ, :black, :hisp, :marr, :nodegree,
    :re74, :re75, :re78,
]

data = read_rdatasets_csv(url, cols)

(n = nrow(data),
 treated = Int(sum(data.treat .== 1)),
 controls = Int(sum(data.treat .== 0)))
(n = 445, treated = 185, controls = 260)

Under the experimental graph, no baseline adjustment is made.

raw_difference = mean(data.re78[data.treat .== 1]) -
                 mean(data.re78[data.treat .== 0])

experimental_res = estimate_causal(
    a = [1, 0],
    data = select(data, [:treat, :re78]),
    graph = experimental_graph,
    treatment = :treat,
    outcome = :re78,
)

r = x -> round(x, sigdigits=4)

(raw_difference = r(raw_difference),
 TMLE_ACE = r(experimental_res[:TMLE].ACE),
 lower_ci = r(experimental_res[:TMLE].lower_ci),
 upper_ci = r(experimental_res[:TMLE].upper_ci))
(raw_difference = 1794.0, TMLE_ACE = 1794.0, lower_ci = 482.5, upper_ci = 3106.0)

Under the observational selection-on-observables graph, the estimator adjusts for the measured baseline variables.

observational_res = estimate_causal(
    a = [1, 0],
    data = data,
    graph = observational_graph,
    treatment = :treat,
    outcome = :re78,
)

(TMLE_ACE = r(observational_res[:TMLE].ACE),
 lower_ci = r(observational_res[:TMLE].lower_ci),
 upper_ci = r(observational_res[:TMLE].upper_ci),
 Onestep_ACE = r(observational_res[:Onestep].ACE),
 Gcomp_ACE = r(observational_res[:Gcomp].ACE),
 IPW_ACE = r(observational_res[:IPW].ACE))
(TMLE_ACE = 1638.0, lower_ci = 321.8, upper_ci = 2953.0, Onestep_ACE = 1638.0, Gcomp_ACE = 1676.0, IPW_ACE = 1613.0)

The key point is not the small numerical difference in this sample. It is that each estimate is tied to a different graph. Under the unmeasured-selection graph above, the effect is not identified, so estimate_causal should not be run.

Takeaway

The NSW example is useful because it makes the identifying assumptions visible:

  • randomized assignment implies no adjustment;
  • measured selection implies adjustment for baseline covariates;
  • unmeasured selection makes the effect non-identified in this ADMG.