Repository navigation
Pipeline overview: from ED data to TreeScan signals
Purpose of this page
This page provides a high-level map of the full TreeScan implementation pipeline. It is meant to help readers understand how the different components fit together, where local implementation work is required, and where common constraints arise.
This page is not a step-by-step guide. Instead, it provides a mental model of the system so you can orient yourself before diving into individual technical steps.
The mechanical backbone: the three core scripts
Script 1: Tree file construction Builds the diagnosis hierarchy (e.g., ICD-based tree and supplemental groupings) that TreeScan uses to define related diagnoses.
Script 2: Count file construction Transforms visit-level ED data into daily counts of diagnoses after applying eligibility rules, deduplication, and (optionally) incident-diagnosis logic. This is the primary data input to TreeScan.
Script 3: Parameter file creation and TreeScan execution Defines how TreeScan scans across diagnoses and time (analysis settings) and runs TreeScan to produce results and temporal graphs.
These scripts form the core mechanical pipeline, but they are not the entire system. Successful implementation depends on what happens before, between, and after these scripts.
What exists before the scripts (upstream dependencies):
Before any TreeScan scripts can run, sites must address:
-
Local ED data systems: Data may come from ESSENCE, EHRs, or local data warehouses.
-
Data access and governance: Whether visit-level data can be accessed, exported, or processed in local computing environments.
-
Data extraction and staging: Pulling ED data into a format suitable for analysis (e.g., one row per visit with diagnosis codes and dates).
-
Visit-level analytic dataset: A standardized representation of ED visits that can be used as input to the count file generation process.
These upstream steps may be where site-to-site differences occur.
What exists between the scripts (transformation layers):
Between raw ED data and TreeScan-ready inputs, several design choices are required in the pipeline:
-
Eligibility filtering: Deciding which diagnoses are in scope for surveillance.
-
Incident-diagnosis logic: Defining whether repeated diagnoses for the same person are treated as new events or excluded based on a lookback window.
-
Deduplication and severity handling: Resolving multiple diagnoses within a visit and mapping local severity fields into standardized categories (if applicable).
-
Diagnosis hierarchy mapping: Ensuring local diagnosis codes align with the tree structure used by TreeScan.
-
Formatting and schema alignment: Ensuring that count files, tree files, and parameter files use consistent formats and conventions.
These steps determine how local data are translated into the standardized representation that TreeScan can analyze.
What exists after the scripts (downstream use):
After TreeScan runs, additional work is required to make results meaningful in practice:
-
TreeScan outputs: Results files and temporal graphs that list statistically unusual clusters.
-
Initial signal review: Human review to determine whether signals are plausible, data artifacts, or expected patterns.
-
Interpretation and triage: Assessing which signals merit further investigation.
-
Operational follow-up: Deciding what actions (if any) are appropriate based on identified signals.
This downstream work is outside the scope of the scripts themselves but is essential for operational use.
Where we envisage where implementations may break:
Real-world implementations frequently diverge from the reference pipeline due to:
-
Data access constraints: Limited ability to export visit-level data or run code where data reside.
-
Identifier availability: Lack of stable person identifiers, which affects incident-diagnosis logic.
-
Data lag and backfill: Delayed reporting or retrospective updates that affect “prospective” analyses.
-
Compute environment constraints: Restrictions on installing or running TreeScan in secure environments.
-
Staffing and ownership: Limited capacity to maintain and operate the pipeline over time.
These constraints are not failures of the method. They are design constraints that must be made explicit and addressed during implementation.
Examples of where sites may get stuck:
-
The scripts run, but outputs are empty or implausible due to upstream data issues.
-
Local diagnosis codes do not map cleanly to the reference tree.
-
The pipeline works for one person but is not reproducible by others.
These are implementation challenges to think about, not user errors.
A full pipeline for running TreeScan-based analyses using R and the TreeScan software.
This project provides a structured workflow to prepare data, run TreeScan, and process results using an R-based pipeline.
- Click the green Code button on GitHub
- Select Download ZIP
- Extract the ZIP file
- Locate the
treescan_projectsubfolder - Move
treescan_projectto your desired working directory
Download and install RStudio:
https://posit.co/download/rstudio-desktop/
Download TreeScan from:
https://www.treescan.org/download_treescan.html
- You must create an account before downloading
- Choose version based on your environment:
- Windows → if running locally
- Linux → if running on a server
- Select the NON-graphical version
- The standard (graphical) version may cause IT/access issues
After downloading:
Move the TreeScan files into the correct subfolder inside treescan_project:
| Environment | Folder |
|---|---|
| Windows | TS_windows/ |
| Linux | TS_linux/ |
- Launch RStudio
- In the bottom-right file explorer:
- Navigate to:
treescan_project/code/ - Open:
run_full_pipeline.R
- Navigate to:
Before running, update the following:
Update line 4 to match your local path:
setwd("~/TreeScan-implementation/treescan_project")Replace with wherever you saved treescan_project.
Modify these variables depending on your setup:
server <- FALSE # Set to TRUE if running on a server
first_time <- TRUE # Set to FALSE after first run- Run the script in RStudio
The pipeline will:
- Execute TreeScan
- Process outputs
- Complete the full analysis workflow
- Ensure the correct TreeScan version is placed in the matching folder (
TS_windowsorTS_linux) - Using the non-graphical version is required
- Incorrect working directory paths will cause errors