Repository navigation
Simulating ACE symptom counts
An adverse-childhood-experience score is how many of a fixed list of items a person
endorses. That makes it bounded (out of k), zero-heavy, and built from binary
items that are correlated with each other. sim_symptoms() reproduces that shape by
drawing correlated Bernoulli items through a Gaussian copula, so the realized
distribution looks like real ACE data rather than an idealized one.
sim_symptoms(n, effect = 0, k = 10, p_items = NULL, rho = 0.30)| Argument | What it is |
|---|---|
n |
Sample size per group. |
effect |
Group effect as a uniform shift in every item's log-odds. 0 is the null. |
k |
Number of items, the ceiling (10 for the modern instrument, 7 for the original). |
p_items |
Per-item endorsement probabilities. Defaults to default_ace_probs(k). |
rho |
Latent (Gaussian) inter-item correlation. This is the correlation of the underlying normals; the correlation among the observed 0/1 items is lower and varies by pair. |
It returns a data frame with group, ace (the count), and k.
library(countkit)
set.seed(1)
d <- sim_symptoms(n = 1e5, effect = 0, k = 10, rho = 0.30)
c(mean = mean(d$ace), pct_zero = mean(d$ace == 0), max = max(d$ace))
#> mean pct_zero max
#> 1.49 0.35 10.00
c(p_ge1 = mean(d$ace >= 1), p_ge4 = mean(d$ace >= 4))
#> p_ge1 p_ge4
#> 0.650 0.121That is the empirical ACE shape: a mean near 1.5, about a third of people at zero, a long right tail, and roughly the population benchmark of 65% with at least one adversity. It is bounded and correlated, which is exactly why a beta-binomial, not a Poisson, is the model that respects it.
The per-item probabilities behind the default, descending from about .30 to .05. Ask for 7 items and you get the probabilities for the original instrument.
default_ace_probs(10)
#> [1] 0.30 0.25 0.20 0.18 0.15 0.12 0.10 0.08 0.06 0.05
default_ace_probs(7)
#> [1] 0.30 0.25 0.20 0.18 0.15 0.12 0.10Pass your own vector to sim_symptoms(p_items = ...) to calibrate to a different
instrument.
The true difference in expected count under a given log-odds shift, for scoring bias and coverage. The item correlation does not affect the mean, so this is just the difference of the summed marginals.
true_diff_symptoms(effect = 0.6, k = 10)
#> 0.864Simulate a task