-
Notifications
You must be signed in to change notification settings - Fork 2
Why TreeScan does not adjust for risk factors
TreeScan is designed to detect unexpected patterns in healthcare data, not to estimate causal relationships between risk factors and disease outcomes.
The method therefore does not adjust for variables such as:
- vaccination status
- age
- comorbidities
- demographic risk factors
- behavioral exposures
These factors are important for epidemiologic analysis, but they are not part of the statistical signal detection process.
Instead, TreeScan focuses on identifying unusual changes in diagnosis patterns over time.
TreeScan tests whether the number of observed diagnoses in a specific diagnosis group and time window is unusually high compared with recent patterns in the data.
The null hypothesis can be thought of as:
Diagnoses occur over time in the same proportions that have been observed historically.
If the number of cases in a particular diagnosis group and time window deviates strongly from this baseline expectation, the method identifies a signal.
This comparison is based entirely on patterns within the surveillance data itself, rather than external risk factors.
Including risk factors such as vaccination status would change the nature of the analysis.
TreeScan is designed to be:
- rapid
- automated
- agnostic to specific diseases
Introducing additional covariates would require detailed modeling assumptions about how those factors influence disease risk.
This would make the system:
- slower
- harder to automate
- dependent on disease-specific models
For real-time surveillance across many diagnoses, this approach would not be practical.
Risk factors such as vaccination status can become important after a signal has been detected.
Once TreeScan identifies an unusual cluster, investigators may examine additional information such as:
- vaccination status
- demographic characteristics
- geographic patterns
- travel history
- environmental exposures
These analyses help determine whether the signal represents a true public health event and what factors may explain it.
The design philosophy of TreeScan can be summarized as:
- Detect unusual patterns quickly
- Investigate signals using additional information
By separating signal detection from causal investigation, TreeScan allows surveillance systems to monitor large volumes of healthcare data efficiently.
A full pipeline for running TreeScan-based analyses using R and the TreeScan software.
This project provides a structured workflow to prepare data, run TreeScan, and process results using an R-based pipeline.
- Click the green Code button on GitHub
- Select Download ZIP
- Extract the ZIP file
- Locate the
treescan_projectsubfolder - Move
treescan_projectto your desired working directory
Download and install RStudio:
https://posit.co/download/rstudio-desktop/
Download TreeScan from:
https://www.treescan.org/download_treescan.html
- You must create an account before downloading
- Choose version based on your environment:
- Windows → if running locally
- Linux → if running on a server
- Select the NON-graphical version
- The standard (graphical) version may cause IT/access issues
After downloading:
Move the TreeScan files into the correct subfolder inside treescan_project:
| Environment | Folder |
|---|---|
| Windows | TS_windows/ |
| Linux | TS_linux/ |
- Launch RStudio
- In the bottom-right file explorer:
- Navigate to:
treescan_project/code/ - Open:
run_full_pipeline.R
- Navigate to:
Before running, update the following:
Update line 4 to match your local path:
setwd("~/TreeScan-implementation/treescan_project")Replace with wherever you saved treescan_project.
Modify these variables depending on your setup:
server <- FALSE # Set to TRUE if running on a server
first_time <- TRUE # Set to FALSE after first run- Run the script in RStudio
The pipeline will:
- Execute TreeScan
- Process outputs
- Complete the full analysis workflow
- Ensure the correct TreeScan version is placed in the matching folder (
TS_windowsorTS_linux) - Using the non-graphical version is required
- Incorrect working directory paths will cause errors