Skip to content

What this project is

BeckyW edited this page Jun 3, 2026 · 1 revision

Welcome to the TreeScan-implementation wiki!

What this project is:

This project supports the practical implementation of TreeScan-based syndromic surveillance in public health settings. The goal is to move from a research example to a repeatable, operational workflow that health departments can use to detect unusual patterns in emergency department (ED) data in near–real time.

This wiki documents the technical and operational steps needed to go from raw ED data to TreeScan results. It is designed to support implementation across sites with different data systems, access constraints, and computing environments.

What TreeScan is:

TreeScan is a statistical surveillance tool that scans across diagnosis hierarchies and time to identify diagnosis clusters that are unusually elevated, without relying on predefined syndromes.

What baseline operational readiness means:

In this documentation, “baseline operational readiness” means that a site can:

  • Prepare TreeScan input files from its own ED data

  • Run TreeScan in its own computing environment using standard settings

  • Generate and view TreeScan outputs

  • Repeat this process reliably over time

This distinguishes between:

Demo or example runs (e.g., running TreeScan on provided test data), and

Operational use (running the pipeline on real ED data in your own environment)

This wiki focuses on helping sites reach baseline operational readiness. More advanced topics (e.g., full automation, downstream reporting products, and organizational response workflows) are (and will soon be) documented in later sections.

Who this wiki is for

This wiki is intended for people who are directly involved in implementing or operating TreeScan, such as Public Health Epidemiologists.

You do not need to be a TreeScan developer or a statistical methods expert to use this wiki. The documentation focuses on practical implementation, common failure points, and real-world constraints.

Pipeline overview

At a high level, TreeScan sits at the end of a data pipeline that transforms messy, local ED data into a standardized format that a statistical surveillance method can operate on. The major steps are:

  1. Local ED data (what you already have) Health departments start with whatever ED data they can access (e.g., ESSENCE exports, EHR tables, or a data warehouse). This data is usually not in a form that TreeScan can use directly.

  2. Standardizing visits (making data analyzable) Local ED data must be transformed into a visit-level analytic dataset with a consistent structure (one row per visit, diagnosis codes in a standard format, dates, and basic visit/admission indicators). This step is where most site-specific data cleaning and restructuring happens.

  3. Converting visits into counts (what TreeScan actually reads) TreeScan does not read raw visit records. It reads counts of diagnoses by day. The visit-level data are therefore aggregated into a count file that tells TreeScan how many times each diagnosis occurred on each day. This step encodes key epidemiologic decisions (e.g., eligibility filters, incident definitions, deduplication rules).

  4. Defining the diagnosis hierarchy (what “related diagnoses” means) TreeScan also needs a tree file that defines how diagnoses relate to one another (e.g., ICD-10 hierarchy and any supplemental groupings). This file determines how TreeScan searches across diagnosis groupings rather than just individual codes.

  5. Configuring the analysis (how TreeScan behaves) A parameter file defines how TreeScan scans across time and diagnoses (e.g., prospective vs retrospective, window sizes, conditioning choices, exclusions). This is where the statistical method is operationalized.

  6. Running TreeScan and reviewing outputs TreeScan combines the count file, tree file, and parameter file to produce results and temporal graphs. These outputs highlight statistically unusual patterns that require human review and interpretation.

This wiki follows these steps in order and highlights the points where real-world constraints commonly arise (e.g., limited data access, missing identifiers, computing restrictions, or delayed data feeds).

How to use this wiki

You can read this wiki linearly or jump to the section that matches where you are stuck:

  • If you are unsure how to extract or structure your ED data → Local Data Sources & Access

  • If you are building TreeScan input files → Visit-Level Analytic Dataset and Count File Generation

  • If TreeScan won’t run in your environment → Running TreeScan

  • If you have results but are unsure how to interpret them → Outputs & Interpretation

  • If you want to assess your site’s readiness to use TreeScan operationally → Operational Readiness Checklist

Each section includes common failure modes and practical troubleshooting guidance.

Quick start

If you are new to TreeScan implementation:

  • Review the pipeline overview above to understand the overall workflow.

  • Run the provided example pipeline end-to-end to confirm TreeScan runs in your environment.

  • Identify how ED data are accessed at your site (ESSENCE, data warehouse, etc.).

  • Map your local data to the standard visit-level analytic dataset schema.

  • Generate a count file from a small sample of real data.

  • Use the Operational Readiness Checklist to identify remaining gaps.

Clone this wiki locally