Skip to content

Statistical intuition

BeckyW edited this page Jun 3, 2026 · 1 revision

TreeScan Statistical Intuition

The Core Statistical Question

TreeScan asks the following question:

Is the number of observed diagnoses within a specific diagnosis group and time window unusually high compared with what we would expect based on recent data?

To answer this, TreeScan compares the observed counts of diagnoses with a statistical baseline representing typical patterns in the data.


Conditioning on Time

Healthcare systems experience natural fluctuations in daily visit volume.

For example:

  • weekends may have fewer visits
  • severe weather events may increase ED volume
  • seasonal illness patterns may change overall case counts

TreeScan accounts for these fluctuations by conditioning on the total number of cases occurring each day.

This means that the total number of cases per day is treated as fixed, and the method evaluates whether diagnoses are distributed unusually across categories within those daily totals.

In other words, TreeScan evaluates whether the proportion of specific diagnoses within a given day deviates from what would normally be expected.


Conditioning on Diagnosis Groups

TreeScan also accounts for the fact that some diagnoses occur more frequently than others.

For example:

  • common respiratory diagnoses appear frequently
  • rare diagnoses appear only occasionally

TreeScan therefore conditions on the overall distribution of diagnoses across the diagnosis tree.

This ensures that common diagnoses remain expected to appear frequently, and rare diagnoses remain expected to appear rarely.

The method then evaluates whether the timing of those diagnoses is unusual relative to their typical distribution.


Searching Across Many Possible Clusters

TreeScan evaluates many possible clusters simultaneously, including:

  • different diagnosis groups in the hierarchy
  • different time windows
  • different combinations of diagnoses and time periods

Because the method searches across many possible patterns, it must carefully control for false positives.


Controlling False Positives with Monte Carlo Simulation

To determine whether a detected cluster is statistically significant, TreeScan uses Monte Carlo simulation.

Under the null hypothesis that no unusual cluster exists, the timing of diagnoses is randomized many times to generate simulated datasets.

For each simulated dataset, TreeScan repeats the scanning process and records the most unusual cluster found.

The p-value of the observed cluster is then determined by comparing its strength to the distribution of clusters observed in the simulated datasets.

This approach ensures that the probability of detecting a signal due purely to chance remains controlled despite evaluating many potential clusters.

Clone this wiki locally