Skip to content

Pipeline overview: from ED data to TreeScan signals

BeckyW edited this page Jun 3, 2026 · 1 revision

Purpose of this page

This page provides a high-level map of the full TreeScan implementation pipeline. It is meant to help readers understand how the different components fit together, where local implementation work is required, and where common constraints arise.

This page is not a step-by-step guide. Instead, it provides a mental model of the system so you can orient yourself before diving into individual technical steps.

The mechanical backbone: the three core scripts

Script 1: Tree file construction Builds the diagnosis hierarchy (e.g., ICD-based tree and supplemental groupings) that TreeScan uses to define related diagnoses.

Script 2: Count file construction Transforms visit-level ED data into daily counts of diagnoses after applying eligibility rules, deduplication, and (optionally) incident-diagnosis logic. This is the primary data input to TreeScan.

Script 3: Parameter file creation and TreeScan execution Defines how TreeScan scans across diagnoses and time (analysis settings) and runs TreeScan to produce results and temporal graphs.

These scripts form the core mechanical pipeline, but they are not the entire system. Successful implementation depends on what happens before, between, and after these scripts.

What exists before the scripts (upstream dependencies):

Before any TreeScan scripts can run, sites must address:

  • Local ED data systems: Data may come from ESSENCE, EHRs, or local data warehouses.

  • Data access and governance: Whether visit-level data can be accessed, exported, or processed in local computing environments.

  • Data extraction and staging: Pulling ED data into a format suitable for analysis (e.g., one row per visit with diagnosis codes and dates).

  • Visit-level analytic dataset: A standardized representation of ED visits that can be used as input to the count file generation process.

These upstream steps may be where site-to-site differences occur.

What exists between the scripts (transformation layers):

Between raw ED data and TreeScan-ready inputs, several design choices are required in the pipeline:

  • Eligibility filtering: Deciding which diagnoses are in scope for surveillance.

  • Incident-diagnosis logic: Defining whether repeated diagnoses for the same person are treated as new events or excluded based on a lookback window.

  • Deduplication and severity handling: Resolving multiple diagnoses within a visit and mapping local severity fields into standardized categories (if applicable).

  • Diagnosis hierarchy mapping: Ensuring local diagnosis codes align with the tree structure used by TreeScan.

  • Formatting and schema alignment: Ensuring that count files, tree files, and parameter files use consistent formats and conventions.

These steps determine how local data are translated into the standardized representation that TreeScan can analyze.

What exists after the scripts (downstream use):

After TreeScan runs, additional work is required to make results meaningful in practice:

  • TreeScan outputs: Results files and temporal graphs that list statistically unusual clusters.

  • Initial signal review: Human review to determine whether signals are plausible, data artifacts, or expected patterns.

  • Interpretation and triage: Assessing which signals merit further investigation.

  • Operational follow-up: Deciding what actions (if any) are appropriate based on identified signals.

This downstream work is outside the scope of the scripts themselves but is essential for operational use.

Where we envisage where implementations may break:

Real-world implementations frequently diverge from the reference pipeline due to:

  • Data access constraints: Limited ability to export visit-level data or run code where data reside.

  • Identifier availability: Lack of stable person identifiers, which affects incident-diagnosis logic.

  • Data lag and backfill: Delayed reporting or retrospective updates that affect “prospective” analyses.

  • Compute environment constraints: Restrictions on installing or running TreeScan in secure environments.

  • Staffing and ownership: Limited capacity to maintain and operate the pipeline over time.

These constraints are not failures of the method. They are design constraints that must be made explicit and addressed during implementation.

Examples of where sites may get stuck:

  • The scripts run, but outputs are empty or implausible due to upstream data issues.

  • Local diagnosis codes do not map cleanly to the reference tree.

  • The pipeline works for one person but is not reproducible by others.

These are implementation challenges to think about, not user errors.

Clone this wiki locally