Repository navigation
Simulating divergent thinking
Divergent-thinking tasks (list unusual uses for a brick) are the case where the analyst makes the distribution. One latent ability produces both the number of responses a person gives and how original each one is, and then the researcher chooses how to aggregate: sum the originality, average it, or take the maximum. Those choices are not three measurements of the same thing analyzed the same way. They are different variables with different shapes, and only some of them are counts.
sim_aut() builds all of them from one set of draws, so you can see the scoring
choice, and nothing else, move the statistics.
sim_aut(n, effect = 0,
fluency_mu = 8, fluency_size = 6,
orig_mean = 0.30, orig_sd = 0.15,
ability_on_fluency = 0.35, ability_on_orig = 0.10)| Argument | What it is |
|---|---|
n |
Sample size per group. |
effect |
Group difference in latent ability, in Cohen's d units. |
fluency_mu, fluency_size
|
Baseline mean and negative-binomial size for the response count. |
orig_mean, orig_sd
|
Mean and spread of per-response originality. |
ability_on_fluency, ability_on_orig
|
How strongly latent ability drives the count and the originality level. |
It returns a data frame with group, fluency (the raw count), and three scoring
choices: sum_orig, max_orig, and mean_orig.
library(countkit)
set.seed(5)
d <- sim_aut(n = 2000, effect = 0)
sapply(d[c("fluency", "sum_orig", "max_orig", "mean_orig")],
function(v) c(VMR = var(v) / mean(v), r_with_fluency = cor(v, d$fluency)))
#> fluency sum_orig max_orig mean_orig
#> VMR 3.66 2.34 0.06 0.05
#> r_with_fluency 1.00 0.93 0.61 0.44Read the two rows. Raw fluency and summed originality have variance-to-mean ratios above 1: they behave like counts. Summed originality correlates .93 with sheer response count, which is the well-known fluency confound, so summed originality is largely a fluency count in disguise. Maximum and mean originality have ratios far below 1 and are not integer-valued at all: they are bounded, roughly continuous quantities, and fitting a Poisson or negative binomial to them would be a category error.
The lesson the package is built to make visible: before you pick a count model, check whether the scoring choice even left you with a count. If it did not, model it as the continuous, bounded thing it is, and report fluency alongside any originality score so a reader can see the confound.
Simulate a task