-
Notifications
You must be signed in to change notification settings - Fork 2
Why incident diagnosis are used
In this project, TreeScan is run on incident diagnoses. An incident diagnosis is a diagnosis that is treated as new for a patient at the time it appears in the study period.
This means that if a patient has multiple ED visits with the same diagnosis across time, only the first occurrence (within the relevant definition window) is counted as incident, and subsequent repeat occurrences are not counted as new cases for the purpose of signal detection.
Emergency department diagnosis data include both:
- new acute events (e.g., acute infection, injury, poisoning)
- chronic conditions (e.g., diabetes, hypertension, asthma)
- repeat visits for ongoing issues (e.g., follow-up care, persistent symptoms)
If we counted all diagnoses on all visits, then signals could be driven by:
- repeated visits by the same patients
- ongoing chronic disease burden
- differences in care-seeking behavior rather than new events
This is not ideal when the goal is to detect emerging health events.
Using incident diagnoses reduces the influence of repeat encounters and focuses detection on new occurrences of diagnoses.
Without incident diagnosis logic, a signal could arise simply because:
- the same individuals repeatedly return for the same condition
- a hospital changes follow-up scheduling behavior
- a subset of patients has frequent revisits due to social or healthcare access factors
These patterns may be important for healthcare operations, but they are not necessarily signals of new public health events.
Incident diagnoses therefore help align TreeScan signals with the concept of:
“Are we seeing more new diagnoses than expected?”
rather than:
“Are we seeing more total visits with a diagnosis than expected?”
To decide whether a diagnosis is incident, the system must determine whether the same patient had that diagnosis previously.
This requires:
- a patient identifier (e.g., unique patient ID)
- historical ED visit data prior to the study period
- logic that checks whether the diagnosis appears for that patient in the lookback window
If a patient had the diagnosis before, then the diagnosis is not treated as incident when it appears again during the study period.
Because incident diagnoses require checking prior history, the analysis must include data from before the study period.
In this project, we use a one-year lookback.
This provides a practical window for determining whether a diagnosis is plausibly “new” for surveillance purposes, while balancing feasibility and data availability.
The analysis design uses:
- 90 days as the study period (the window in which clusters are evaluated)
- 1 year of lookback prior to the start of the study period to determine incident diagnoses
Therefore, to run one analysis window, the data needed include:
- 12 months of lookback data
- 3 months (90 days) of study data
This totals approximately 15 months of ED visit data.
Using incident diagnoses is a design choice intended to make TreeScan outputs:
- more interpretable for public health surveillance
- less sensitive to repeat visits by the same individuals
- more reflective of emerging changes in new diagnoses rather than ongoing healthcare utilization patterns
This is especially important in operational settings where signals must be reviewed routinely and where false alerts caused by repeat utilization patterns would add unnecessary burden.
A full pipeline for running TreeScan-based analyses using R and the TreeScan software.
This project provides a structured workflow to prepare data, run TreeScan, and process results using an R-based pipeline.
- Click the green Code button on GitHub
- Select Download ZIP
- Extract the ZIP file
- Locate the
treescan_projectsubfolder - Move
treescan_projectto your desired working directory
Download and install RStudio:
https://posit.co/download/rstudio-desktop/
Download TreeScan from:
https://www.treescan.org/download_treescan.html
- You must create an account before downloading
- Choose version based on your environment:
- Windows → if running locally
- Linux → if running on a server
- Select the NON-graphical version
- The standard (graphical) version may cause IT/access issues
After downloading:
Move the TreeScan files into the correct subfolder inside treescan_project:
| Environment | Folder |
|---|---|
| Windows | TS_windows/ |
| Linux | TS_linux/ |
- Launch RStudio
- In the bottom-right file explorer:
- Navigate to:
treescan_project/code/ - Open:
run_full_pipeline.R
- Navigate to:
Before running, update the following:
Update line 4 to match your local path:
setwd("~/TreeScan-implementation/treescan_project")Replace with wherever you saved treescan_project.
Modify these variables depending on your setup:
server <- FALSE # Set to TRUE if running on a server
first_time <- TRUE # Set to FALSE after first run- Run the script in RStudio
The pipeline will:
- Execute TreeScan
- Process outputs
- Complete the full analysis workflow
- Ensure the correct TreeScan version is placed in the matching folder (
TS_windowsorTS_linux) - Using the non-graphical version is required
- Incorrect working directory paths will cause errors