Poisson models for count data often arise when units are observed for different lengths of time. A hospital admission count for a patient followed for 30 days is not directly comparable to a count for a patient followed for 365 days. The natural quantity of interest is the rate — admissions per unit time — not the raw count.
There are three related but distinct specifications, each making a different assumption:
Model 1 — Poisson with offset (rate model):
\[\log\!\bigl(E[Y]/t\bigr) = \beta_0 + \beta_1 X \tag{25.1}\]
Equivalently, \(\log E[Y] = \beta_0 + \beta_1 X + \log t\), where \(\log t\) is included as a fixed offset (coefficient constrained to 1). This is the standard exposure model: the expected count is proportional to exposure time.
Model 2 — Poisson on the ratio \(Y/t\):
\[\log\!\bigl(E[Y/t]\bigr) = \beta_0 + \beta_1 X \tag{25.2}\]
This looks similar, but it is not equivalent to the offset model, and strictly speaking \(Y/t\) is not a Poisson outcome at all: a Poisson likelihood is defined for non-negative integer counts, and the continuous ratio \(Y/t\) is not one. Because \(t\) is a fixed constant for each unit, \(E[Y/t] = E[Y]/t\) exactly (linearity of expectation), so the conditional mean is still log-linear in \(X\); that is why the point estimates are often similar. But the Poisson variance assumption fails: \(\mathrm{Var}(Y/t) = \lambda/t^2\), whereas the mean is \(E[Y/t] = \lambda/t\), so the variance is \(1/t\) times the mean rather than equal to it. (Writing it as \(\lambda/t\) would make variance equal mean and quietly restore equidispersion, which is precisely what fails here.) So fitting glm(rate ~ x, family = poisson) is a quasi-likelihood mean model rather than a correct Poisson likelihood. If you use it, treat it as a workaround: weight by exposure (weights = t) so the implied variance is right, and/or use robust standard errors. The offset model (Model 1) avoids the issue entirely by keeping the integer count as the outcome.
Model 3 — Unconstrained log-exposure:
\[\log E[Y] = \beta_0 + \beta_1 X + \gamma \log t \tag{25.3}\]
Here \(\gamma\) is estimated from the data. If \(\hat\gamma \approx 1\), the proportionality assumption is supported. If \(\hat\gamma < 1\), counts grow sub-linearly with exposure — a testable, substantive claim.
25.2 Simulation
Code
library(tidyverse)set.seed(42)n<-800# Simulate a health study: count of adverse events# Covariate: age group (0 = younger, 1 = older)age_old<-rbinom(n, 1, 0.45)# Exposure: days under observation (varies 7–365)exposure<-sample(seq(7, 365, by =7), n, replace =TRUE)# True model: rate = exp(-3 + 1.2 * age_old)# Count ~ Poisson(rate * exposure)true_rate<-exp(-3+1.2*age_old)count<-rpois(n, lambda =true_rate*exposure)df<-data.frame(count, age_old, exposure, log_exposure =log(exposure))summary(df[, c("count", "exposure")])
count exposure
Min. : 0.00 Min. : 7.0
1st Qu.: 7.00 1st Qu.:112.0
Median :13.00 Median :196.0
Mean :18.95 Mean :192.9
3rd Qu.:27.00 3rd Qu.:281.8
Max. :71.00 Max. :364.0
We simulate \(n = 800\) patients. Age group is \(\text{age\_old} \sim \text{Bernoulli}(0.45)\), and exposure is days under observation, drawn uniformly from the weekly grid 7 to 364 and independently of age. The event rate is \(\exp(-3 + 1.2\,\text{age\_old})\) per day and the count is \(\text{Poisson}(\text{rate} \times \text{exposure})\), so the true intercept is \(-3\), the true age effect is 1.2, and the true exposure elasticity is \(\gamma = 1\) by construction.
The data has a wide range of exposure times — 7 to 364 days, median 196, with counts from 0 to 71 and a mean of 19 — so ignoring them would badly confound the comparison between age groups.
25.3 Fitting the three models
Code
# Model 1: offset — coefficient on log(exposure) constrained to 1m1<-glm(count~age_old+offset(log(exposure)), family =poisson, data =df)# Model 2: Poisson on the rate Y/t (different response)df$rate<-df$count/df$exposurem2<-glm(rate~age_old, family =poisson, data =df)# Model 3: log(exposure) as a free covariatem3<-glm(count~age_old+log_exposure, family =poisson, data =df)# Collect coefficientsresults<-data.frame( Model =c("1. Offset (constrained γ=1)","2. Rate Y/t as response","3. Unconstrained log(t)"), Intercept =c(coef(m1)[1], coef(m2)[1], coef(m3)[1]), age_old =c(coef(m1)[2], coef(m2)[2], coef(m3)[2]), gamma_log_t =c(1, NA, coef(m3)[3]))knitr::kable(results, digits =3, caption ="Coefficient estimates: true intercept = −3, true age effect = 1.2, true γ = 1")
All three recover the truth. The offset model gives an intercept of \(-3.009\) and an age effect of 1.212 against true values of \(-3\) and 1.2; the rate model gives \(-2.980\) and 1.176; the unconstrained model gives \(-2.801\), 1.212 and \(\hat\gamma = 0.962\) against a true \(\gamma\) of 1.
Reading the table: Model 1 and Model 3 give nearly identical age-group coefficients because the true \(\gamma = 1\) (counts are indeed proportional to exposure time). Model 2 differs slightly in the intercept because the response variable changes. The estimated \(\hat\gamma\) from Model 3 should be close to 1.
25.4 When \(\gamma \ne 1\)
To see why the unconstrained model matters, repeat the simulation with a sub-linear count process (\(\gamma = 0.6\): an initial burst of events that doesn’t grow proportionally with long follow-up):
Code
count2<-rpois(n, lambda =true_rate*exposure^0.6)df2<-data.frame(count2, age_old, exposure, log_exposure =log(exposure))m1b<-glm(count2~age_old+offset(log(exposure)), family =poisson, data =df2)m3b<-glm(count2~age_old+log_exposure, family =poisson, data =df2)cat("Offset model age coefficient: ", round(coef(m1b)[2], 3), "\n")
Offset model age coefficient: 1.171
Code
cat("Unconstrained model age coefficient:", round(coef(m3b)[2], 3), "\n")
Two things to read off this, and the first is not what one might expect.
The offset model’s age coefficient (1.171) and the unconstrained model’s (1.168) are essentially the same, and both sit just below the true 1.2. Forcing \(\gamma = 1\) when the truth is 0.6 does not bias the treatment coefficient here — and the reason is structural: exposure is drawn by sample(...) independently of age_old, so the misspecified exposure elasticity is absorbed by the intercept and cannot contaminate a coefficient on a covariate orthogonal to it. The constraint would bite if exposure were correlated with the covariate of interest — for instance if older patients were systematically followed longer, which is common in real cohort data and is the case worth worrying about.
What the unconstrained model does buy, unambiguously, is \(\gamma\) itself: it recovers \(\hat\gamma = 0.566\) against a truth of 0.6, whereas the offset model cannot report it at all because it fixes it at 1. So Model 3’s role is diagnostic — it tells you whether the proportionality assumption holds — rather than protective of the treatment estimate.
25.5 Stata equivalent
* Model 1: offsetpoissoncount age_old, exposure(exposure)* Model 2: rate as responsegen rate = count / exposure* Stata will warn that the dependent variable is non-integer (poisson expects* integer counts bydefault) -- this is expected and can be ignored here,* since Poisson QMLE is valid for continuous non-negative outcomes too.poisson rate age_old* Model 3: unconstrained log(t)gen log_exposure = log(exposure)poissoncount age_old log_exposure
25.6 Which model to use?
Situation
Recommended model
Counts grow proportionally with time
Model 1 (offset)
Uncertain about proportionality
Model 3 (unconstrained), test \(\hat\gamma = 1\)
Rate is the natural estimand
Model 1 (offset) — it estimates a rate directly; Model 2 only as a weighted/robust quasi-model
Long follow-up with declining event rate
Model 3 or negative binomial with offset
In most health and social-science applications, Model 1 (Poisson with offset) is the default. Model 3 is the diagnostic check: fit it, look at \(\hat\gamma\), and return to Model 1 if the constraint is supported by the data.
---title: "Poisson model with exposure"date: "2024-04-20"---## Poisson with exposure or Poisson ratePoisson models for count data often arise when units are observed for differentlengths of time. A hospital admission count for a patient followed for 30 daysis not directly comparable to a count for a patient followed for 365 days. Thenatural quantity of interest is the *rate* — admissions per unit time — notthe raw count.There are three related but distinct specifications, each making a differentassumption:**Model 1 — Poisson with offset (rate model):**$$\log\!\bigl(E[Y]/t\bigr) = \beta_0 + \beta_1 X$$ {#eq-poisson-exposure-1}Equivalently, $\log E[Y] = \beta_0 + \beta_1 X + \log t$, where $\log t$ isincluded as a *fixed* offset (coefficient constrained to 1). This is thestandard exposure model: the expected count is proportional to exposure time.**Model 2 — Poisson on the ratio $Y/t$:**$$\log\!\bigl(E[Y/t]\bigr) = \beta_0 + \beta_1 X$$ {#eq-poisson-exposure-2}This looks similar, but it is not equivalent to the offset model, and strictlyspeaking $Y/t$ is **not a Poisson outcome at all**: a Poisson likelihood isdefined for non-negative integer counts, and the continuous ratio $Y/t$ is notone. Because $t$ is a fixed constant for each unit, $E[Y/t] = E[Y]/t$ exactly(linearity of expectation), so the *conditional mean* is still log-linear in$X$; that is why the point estimates are often similar. But the Poisson varianceassumption fails: $\mathrm{Var}(Y/t) = \lambda/t^2$, whereas the mean is$E[Y/t] = \lambda/t$, so the variance is $1/t$ times the mean rather than equal toit. (Writing it as $\lambda/t$ would make variance equal mean and quietly restoreequidispersion, which is precisely what fails here.) So fitting `glm(rate ~ x, family = poisson)` is a quasi-likelihood*mean model* rather than a correct Poisson likelihood. If you use it, treat itas a workaround: weight by exposure (`weights = t`) so the implied variance isright, and/or use robust standard errors. The offset model (Model 1) avoids theissue entirely by keeping the integer count as the outcome.**Model 3 — Unconstrained log-exposure:**$$\log E[Y] = \beta_0 + \beta_1 X + \gamma \log t$$ {#eq-poisson-exposure-3}Here $\gamma$ is estimated from the data. If $\hat\gamma \approx 1$, theproportionality assumption is supported. If $\hat\gamma < 1$, counts growsub-linearly with exposure — a testable, substantive claim.## Simulation```{r}#| label: sim-data#| cache: true#| warning: false#| message: falselibrary(tidyverse)set.seed(42)n <-800# Simulate a health study: count of adverse events# Covariate: age group (0 = younger, 1 = older)age_old <-rbinom(n, 1, 0.45)# Exposure: days under observation (varies 7–365)exposure <-sample(seq(7, 365, by =7), n, replace =TRUE)# True model: rate = exp(-3 + 1.2 * age_old)# Count ~ Poisson(rate * exposure)true_rate <-exp(-3+1.2* age_old)count <-rpois(n, lambda = true_rate * exposure)df <-data.frame(count, age_old, exposure, log_exposure =log(exposure))summary(df[, c("count", "exposure")])```We simulate $n = 800$ patients. Age group is$\text{age\_old} \sim \text{Bernoulli}(0.45)$, and exposure is days underobservation, drawn uniformly from the weekly grid 7 to 364 and independently ofage. The event rate is $\exp(-3 + 1.2\,\text{age\_old})$ per day and the countis $\text{Poisson}(\text{rate} \times \text{exposure})$, so the true interceptis $-3$, the true age effect is 1.2, and the true exposure elasticity is$\gamma = 1$ by construction.The data has a wide range of exposure times --- 7 to 364 days, median 196, withcounts from 0 to 71 and a mean of 19 --- so ignoring them would badlyconfound the comparison between age groups.## Fitting the three models```{r}#| label: fit-models#| cache: true#| warning: false# Model 1: offset — coefficient on log(exposure) constrained to 1m1 <-glm(count ~ age_old +offset(log(exposure)),family = poisson, data = df)# Model 2: Poisson on the rate Y/t (different response)df$rate <- df$count / df$exposurem2 <-glm(rate ~ age_old,family = poisson, data = df)# Model 3: log(exposure) as a free covariatem3 <-glm(count ~ age_old + log_exposure,family = poisson, data = df)# Collect coefficientsresults <-data.frame(Model =c("1. Offset (constrained γ=1)","2. Rate Y/t as response","3. Unconstrained log(t)"),Intercept =c(coef(m1)[1], coef(m2)[1], coef(m3)[1]),age_old =c(coef(m1)[2], coef(m2)[2], coef(m3)[2]),gamma_log_t =c(1, NA, coef(m3)[3]))knitr::kable(results, digits =3,caption ="Coefficient estimates: true intercept = −3, true age effect = 1.2, true γ = 1")```All three recover the truth. The offset model gives an intercept of $-3.009$ andan age effect of 1.212 against true values of $-3$ and 1.2; the rate model gives$-2.980$ and 1.176; the unconstrained model gives $-2.801$, 1.212 and$\hat\gamma = 0.962$ against a true $\gamma$ of 1.**Reading the table:** Model 1 and Model 3 give nearly identical age-groupcoefficients because the true $\gamma = 1$ (counts are indeed proportional toexposure time). Model 2 differs slightly in the intercept because the responsevariable changes. The estimated $\hat\gamma$ from Model 3 should be close to 1.## When $\gamma \ne 1$To see why the unconstrained model matters, repeat the simulation with asub-linear count process ($\gamma = 0.6$: an initial burst of events thatdoesn't grow proportionally with long follow-up):```{r}#| label: nonlinear-exposure#| cache: true#| warning: falsecount2 <-rpois(n, lambda = true_rate * exposure^0.6)df2 <-data.frame(count2, age_old, exposure, log_exposure =log(exposure))m1b <-glm(count2 ~ age_old +offset(log(exposure)),family = poisson, data = df2)m3b <-glm(count2 ~ age_old + log_exposure,family = poisson, data = df2)cat("Offset model age coefficient: ", round(coef(m1b)[2], 3), "\n")cat("Unconstrained model age coefficient:", round(coef(m3b)[2], 3), "\n")cat("Estimated gamma (log-exposure): ", round(coef(m3b)[3], 3)," (true = 0.6)\n")```Two things to read off this, and the first is not what one might expect.The offset model's age coefficient (1.171) and the unconstrained model's (1.168)are essentially the same, and both sit just below the true 1.2. Forcing$\gamma = 1$ when the truth is 0.6 does **not** bias the treatment coefficienthere --- and the reason is structural: `exposure` is drawn by `sample(...)`independently of `age_old`, so the misspecified exposure elasticity is absorbed bythe intercept and cannot contaminate a coefficient on a covariate orthogonal toit. The constraint would bite if exposure were *correlated* with the covariate ofinterest --- for instance if older patients were systematically followed longer,which is common in real cohort data and is the case worth worrying about.What the unconstrained model does buy, unambiguously, is $\gamma$ itself: itrecovers $\hat\gamma = 0.566$ against a truth of 0.6, whereas the offset modelcannot report it at all because it fixes it at 1. So Model 3's role is diagnostic--- it tells you whether the proportionality assumption holds --- rather thanprotective of the treatment estimate.## Stata equivalent```stata* Model 1: offsetpoisson count age_old, exposure(exposure)* Model 2: rate as responsegen rate = count / exposure* Stata will warn that the dependent variable is non-integer (poisson expects* integer counts by default) -- this is expected and can be ignored here,* since Poisson QMLE is valid for continuous non-negative outcomes too.poisson rate age_old* Model 3: unconstrained log(t)gen log_exposure = log(exposure)poisson count age_old log_exposure```## Which model to use?| Situation | Recommended model ||---|---|| Counts grow proportionally with time | Model 1 (offset) || Uncertain about proportionality | Model 3 (unconstrained), test $\hat\gamma = 1$ || Rate is the natural estimand | Model 1 (offset) — it estimates a rate directly; Model 2 only as a weighted/robust quasi-model || Long follow-up with declining event rate | Model 3 or negative binomial with offset |In most health and social-science applications, Model 1 (Poisson with offset)is the default. Model 3 is the diagnostic check: fit it, look at $\hat\gamma$,and return to Model 1 if the constraint is supported by the data.---<!-- see-also-footer -->*Systematic treatment: [R](https://xiangao.github.io/causal_econometrics_guide/poisson-iv.html) · [Julia](https://xiangao.github.io/causal_econometrics_julia/poisson-iv.html).*