Skip to content

Architecture

Martanto edited this page Jul 3, 2026 · 17 revisions

Architecture

This page is the structural reference for eruption_forecast: every module under src/, the top-level pipeline, how the model and ensemble classes relate, what flows between stages on disk, and the utility surface that holds the rest together.


1. Package Layout

src/eruption_forecast/
β”œβ”€β”€ __init__.py            - public exports
β”œβ”€β”€ logger.py              - loguru wrapper (enable/disable/set_level/set_directory)
β”œβ”€β”€ data_container.py      - BaseDataContainer ABC for TremorData / LabelData
β”‚
β”œβ”€β”€ config/
β”‚   β”œβ”€β”€ base_config.py         - shared config primitives
β”‚   β”œβ”€β”€ constants.py           - ERUPTION_PROBABILITY_THRESHOLD, defaults
β”‚   β”œβ”€β”€ forecast_config.py     - ForecastConfig + per-stage sub-configs
β”‚   β”œβ”€β”€ training_config.py     - TrainingConfig    (standalone TrainingModel)
β”‚   β”œβ”€β”€ prediction_config.py   - PredictionConfig  (standalone PredictionModel)
β”‚   β”œβ”€β”€ evaluation_config.py   - EvaluationConfig  (standalone EvaluationModel)
β”‚   └── explanation_config.py  - ExplanationConfig (standalone ExplanationModel)
β”‚
β”œβ”€β”€ dataclass/
β”‚   β”œβ”€β”€ station_data.py                 - StationData (immutable nslc identity)
β”‚   β”œβ”€β”€ classifier_ensemble_summary.py  - ClassifierEnsembleSummary, EruptionWindow, ProbabilityPick
β”‚   └── classifier_explanation.py       - SeedExplanation, ClassifierExplanation (SHAP payloads)
β”‚
β”œβ”€β”€ decorators/
β”‚   β”œβ”€β”€ decorator_class.py     - base decorator scaffolding
β”‚   └── notify.py              - Telegram notify + send_telegram_notification
β”‚
β”œβ”€β”€ ensemble/
β”‚   β”œβ”€β”€ base_ensemble.py       - BaseEnsemble (joblib save/load mixin)
β”‚   β”œβ”€β”€ seed_ensemble.py       - SeedEnsemble (one classifier Γ— N seeds)
β”‚   β”œβ”€β”€ classifier_ensemble.py - ClassifierEnsemble (N classifiers)
β”‚   β”œβ”€β”€ metrics_ensemble.py    - MetricsEnsemble (metrics engine)
β”‚   └── explainer_ensemble.py  - ExplainerEnsemble (per-seed SHAP engine)
β”‚
β”œβ”€β”€ features/
β”‚   β”œβ”€β”€ constants.py
β”‚   β”œβ”€β”€ tremor_matrix_builder.py - TremorMatrixBuilder (windowed alignment)
β”‚   β”œβ”€β”€ features_builder.py      - FeaturesBuilder (tsfresh extraction)
β”‚   └── feature_selector.py      - FeatureSelector (tsfresh FDR or RF importance)
β”‚
β”œβ”€β”€ label/
β”‚   β”œβ”€β”€ constants.py
β”‚   β”œβ”€β”€ label_builder.py         - LabelBuilder (sliding window)
β”‚   β”œβ”€β”€ dynamic_label_builder.py - DynamicLabelBuilder (per-eruption build)
β”‚   β”œβ”€β”€ label_data.py            - LabelData (CSV wrapper)
β”‚   └── label_plots.py           - plot_label_distribution
β”‚
β”œβ”€β”€ model/
β”‚   β”œβ”€β”€ constants.py
β”‚   β”œβ”€β”€ base_model.py            - BaseModel ABC (dates, I/O, dual-mode save/load + cache identity)
β”‚   β”œβ”€β”€ forecast_model.py        - ForecastModel orchestrator
β”‚   β”œβ”€β”€ training_model.py        - TrainingModel(BaseModel)
β”‚   β”œβ”€β”€ prediction_model.py      - PredictionModel(BaseModel)
β”‚   β”œβ”€β”€ evaluation_model.py      - EvaluationModel(BaseModel)
β”‚   β”œβ”€β”€ explanation_model.py     - ExplanationModel(BaseModel)
β”‚   β”œβ”€β”€ classifier_model.py      - ClassifierModel (estimator + grid)
β”‚   └── classifier_comparator.py - ClassifierComparator (cross-classifier rank)
β”‚
β”œβ”€β”€ plots/
β”‚   β”œβ”€β”€ styles.py
β”‚   β”œβ”€β”€ tremor_plots.py          - plot_tremor
β”‚   β”œβ”€β”€ feature_plots.py         - feature-importance plots
β”‚   β”œβ”€β”€ forecast_plots.py        - plot_forecast, plot_forecast_from_file
β”‚   β”œβ”€β”€ evaluation_plots.py      - ROC, PR, confusion, threshold, importance
β”‚   └── explanation_plots.py     - SHAP waterfall / beeswarm / bar / aggregate
β”‚
β”œβ”€β”€ sources/
β”‚   β”œβ”€β”€ base.py                  - SeismicDataSource ABC
β”‚   β”œβ”€β”€ sds.py                   - Local SeisComP archive reader
β”‚   └── fdsn.py                  - FDSN client with local SDS caching
β”‚
β”œβ”€β”€ tremor/
β”‚   β”œβ”€β”€ calculate_tremor.py      - CalculateTremor (orchestrator)
β”‚   β”œβ”€β”€ rsam.py, dsar.py, shannon_entropy.py - per-metric kernels
β”‚   └── tremor_data.py           - TremorData (CSV wrapper)
β”‚
└── utils/
    β”œβ”€β”€ array.py, dataframe.py, date_utils.py
    β”œβ”€β”€ formatting.py, ml.py, pathutils.py
    β”œβ”€β”€ validation.py, window.py

70 .py files in total.


2. Pipeline Overview

       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚  Seismic     β”‚     β”‚  CalculateTremor   β”‚    β”‚   TremorData    β”‚
       β”‚  archive     β”‚ ──► β”‚  (rsam/dsar/       β”‚ ─► β”‚   (CSV wrapper) β”‚
       β”‚  (SDS|FDSN)  β”‚     β”‚   entropy/bands)   β”‚    β”‚                 β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                               β”‚
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ feature pipeline ──────────┴─────┐
       β”‚   LabelBuilder            TremorMatrixBuilder               β”‚
       β”‚   DynamicLabelBuilder ──► FeaturesBuilder (tsfresh)         β”‚
       β”‚                           FeatureSelector (FDR or RF)       β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                    β–Ό
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚     TrainingModel      β”‚
                       β”‚   build_label β†’        β”‚
                       β”‚   extract_features β†’   β”‚
                       β”‚   fit (N seeds Γ— M cv) β”‚
                       β””β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”˜
                          β”‚ writes            β”‚ assembles
                          β–Ό                   β–Ό
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ SeedEnsemble Γ— β”‚    β”‚   ClassifierEnsemble   β”‚
                β”‚ N classifiers  β”‚ ─► β”‚ (all SeedEnsembles)    β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                 β”‚
                                                 β–Ό
                                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                  β”‚      PredictionModel         β”‚
                                  β”‚   build_label β†’              β”‚
                                  β”‚   extract_features β†’         β”‚
                                  β”‚   forecast (per-seed proba)  β”‚
                                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                 β”‚
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β–Ό                                                      β–Ό
   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   β”‚   EvaluationModel    β”‚                          β”‚  forecast-results_     β”‚
   β”‚  dispatch on .kind:  β”‚                          β”‚  *.csv + forecast      β”‚
   β”‚  training | predict  β”‚  ── MetricsEnsemble ──►  β”‚  PNG/PDF               β”‚
   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚ writes (n_samples, n_seeds) y_proba / y_pred matrices
              β–Ό
   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   β”‚ ClassifierComparator β”‚   ranking_*.csv + comparison figures
   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   β”‚    ExplanationModel  (BaseModel)                               β”‚
   β”‚    dispatch on upstream model.kind: training | prediction      β”‚
   β”‚                                                                β”‚
   β”‚    ExplainerEnsemble                                           β”‚
   β”‚      ─ per-seed shap.TreeExplainer (RF / lite-rf / GB / XGB)   β”‚
   β”‚      ─ ClassifierExplanation.pkl per classifier                β”‚
   β”‚      ─ per-seed bar + beeswarm under classifiers/{Clf}/figures β”‚
   β”‚      ─ per-eruption waterfall under eruptions/{date}/          β”‚
   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

ForecastModel is the orchestrator that calls every box in sequence. The dashed arrows are also the method-chain order: fm.calculate(...).train(...).predict(...).evaluate(...).explain(...).


3. Component Details

3.1 Tremor (tremor/)

CalculateTremor reads seismic traces day-by-day from a SeismicDataSource and dispatches each day to the configured tremor kernels (rsam.py, dsar.py, shannon_entropy.py). Per-day CSVs are written to tremor/daily/, then concatenated into the merged tremor CSV at the station root. TremorData is a thin wrapper that exposes df, start_date, end_date, sampling-rate validation, and the CSV filename / basename / filetype triple.

3.2 Labels (label/)

Two builders share the same output shape (id, is_erupted) but differ in how positives are placed:

  • LabelBuilder - sliding window over the full date range; day_to_forecast controls the look-ahead window. include_eruption_date=False (default) still marks the eruption day as positive, giving day_to_forecast + 1 positive days per eruption.
  • DynamicLabelBuilder - extends LabelBuilder with a per-eruption three-phase build: (1) zero frames per eruption, (2) concat + deduplicate datetimes, (3) mark positives per eruption. Solves the issue where overlapping look-ahead windows collide in LabelBuilder.
LabelBuilder - one global window over the full date range
─────────────────────────────────────────────────────────
 include_eruption_date=False  (default)
   0 0 0 0 0 0 0 0 0 0  1  1  1  1  1  1  1
                        ↑              ↑  ↑
                    dtf start       day-before eruption
                                       eruption (also 1)
   β†’ dtf days strictly before eruption + eruption day = dtf+1 positives

 include_eruption_date=True
   0 0 0 0 0 0 0 0 0 0  0  1  1  1  1  1  1
                           ↑              ↑
                       dtf start      eruption (counted in dtf)
   β†’ exactly dtf days ending on the eruption day


DynamicLabelBuilder - per-eruption build, overlapping windows deduped
─────────────────────────────────────────────────────────────────────
 Phase 1: initiate (all zeros)
   Eruption A window           Eruption B window
   [0 0 0 0 0 0 0 0 0 0]      [0 0 0 0 0 0 0 0 0 0]

 Phase 2: concat + deduplicate datetimes
   [0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]   ← unified, sorted, unique

 Phase 3: mark positives per eruption
   Erup A (2025-03-20, dtf=2):  Mar 18–20 β†’ 1
   Erup B (2025-03-23, dtf=2):  Mar 21–23 β†’ 1
   [0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 1]
                            ↑       ↑
                          Erup A  Erup B

LabelData parses parameters (window_size, window_step, window_step_unit, day_to_forecast) directly out of the label filename so a CSV alone is enough to rehydrate the build context.

3.3 Features (features/)

        labels (id, is_erupted)         tremor_df
                β”‚                            β”‚
                β–Ό                            β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚         TremorMatrixBuilder            β”‚
        β”‚  windowed slices aligned to labels     β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚           FeaturesBuilder              β”‚
        β”‚  tsfresh extraction (per-column)       β”‚
        β”‚  training: relevance-filter on labels  β”‚
        β”‚  prediction: no filtering              β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚           FeatureSelector              β”‚
        β”‚  method="tsfresh": FDR p-value filter  β”‚
        β”‚  method="random_forest": permutation   β”‚
        β”‚                          importance    β”‚
        β”‚  β†’ top-N feature names per seed        β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

TremorMatrixBuilder.build() validates sample counts per window against minimum_completion and skips short windows so tsfresh never sees ragged input. FeaturesBuilder runs per-column independent extractions so adding a new tremor band does not invalidate the cached results for the others.

3.4 Model (model/)

The model layer follows a mixin pattern:

  • BaseModel - abstract base for every stage. Owns the date/window grid, the lazy tremor_data accessor, output_dir resolution, n_jobs clamping, the content-addressable cache identity helpers (build_identity, compute_hash, _canonicalize, tremor_fingerprint, cache_path), and the dual-mode joblib save(identity=None, path=None) / cache-only load(stage_dir, identity). When identity is supplied, save() writes to {stage_dir}/{hash}.{ClassName}.pkl (plus a .params.json sidecar). When identity is omitted, the legacy {output_dir}/{ClassName}_{basename}.pkl joblib dump is preserved for standalone manual saves. Subclasses implement set_directories, create_directories, validate, describe, to_dict, to_prompt, build_label, extract_features, and override stage_dir + build_identity when they participate in the cache.
  • TrainingModel(BaseModel) - build_label β†’ extract_features β†’ fit. fit() runs per-seed GridSearchCV in joblib.Parallel over the selected classifiers, writes a per-classifier trained-model JSON registry via save_model_json, bundles every seed into a SeedEnsemble and every classifier into a ClassifierEnsemble, then calls self.save(self.build_identity()) so the cache pickle lands at {training_dir}/{hash}.TrainingModel.pkl with a matching sidecar.
  • PredictionModel(BaseModel) - build_label β†’ extract_features β†’ forecast. Cache identity embeds the upstream training_hash (a constructor param threaded by ForecastModel.predict), so re-training automatically invalidates downstream forecasts. forecast() calls self.save(self.build_identity()); cache files live at {prediction_dir}/{hash}.PredictionModel.pkl.
  • EvaluationModel(BaseModel) - no cache; dispatches on model.kind ("training" or "prediction"). Output is namespaced under evaluation/{kind}/ so both modes can coexist.
  • ExplanationModel(BaseModel) - per-seed SHAP explanations over a fitted ClassifierEnsemble. Reuses the upstream TrainingModel or PredictionModel and dispatches on model.kind. Restricted to tree classifiers (RF / lite-rf / GB / XGB); non-tree classifiers are skipped at the ExplainerEnsemble loop with a warning. Output is namespaced under explanation/{kind}/; cache pickles land at {explanation_dir}/{hash}.ExplanationModel.pkl (already mode-namespaced so training-reuse and prediction-reuse caches never collide).
  • ForecastModel - the orchestrator. Not a BaseModel subclass - it owns CalculateTremor, builds the four stage classes lazily, and captures stage kwargs into a ForecastConfig for round-tripping.

ClassifierModel is the per-classifier descriptor (sklearn estimator + hyperparameter grid + slug). ClassifierComparator consumes the in-memory MetricsEnsemble cached on EvaluationModel to rank classifiers head-to-head.

3.5 Ensemble (ensemble/)

            BaseEnsemble (joblib save/load mixin)
                β”‚
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
       β–Ό                 β–Ό
SeedEnsemble       ClassifierEnsemble
1 classifier Γ—     N classifiers Γ—
N fitted seeds     1 SeedEnsemble each
+ feature lists    + factories (from_any, from_json,
                     from_dict, from_seed_ensembles)

           MetricsEnsemble  (standalone - not a BaseEnsemble subclass)
           wraps ClassifierEnsemble + features + y_true
           writes only (n_samples, n_seeds) y_proba / y_pred CSV matrices
           metrics / y_probas / y_preds stay in memory

           ExplainerEnsemble  (standalone - not a BaseEnsemble subclass)
           wraps ClassifierEnsemble + features
           writes per-classifier ClassifierExplanation.pkl
           + per-seed shap_values/{seed:05d}.pkl
           + per-seed bar / beeswarm + per-eruption waterfall plots

MetricsEnsemble and ExplainerEnsemble are both deliberately kept out of ensemble/__init__.py and imported via their full module paths (eruption_forecast.ensemble.metrics_ensemble, eruption_forecast.ensemble.explainer_ensemble) to keep the subpackage free of import cycles back through utils.ml and plots/.

3.6 Sources (sources/)

SeismicDataSource is the read interface: get(date) -> obspy.Stream. Two concrete implementations:

  • SDS - pure local read from {root}/{year}/{network}/{station}/{channel}.D/{file}.
  • FDSN - pulls from a remote FDSN service, then caches the downloaded MSEED into a local SDS layout (download_dir). Repeat calls with the same date hit the local cache.

3.7 Plots (plots/)

apply_nature_style() normalises every figure to a Nature/Science-friendly palette and font stack. Each plot module is a thin functional wrapper around matplotlib (and seaborn where appropriate) - see Visualization for the catalog and output paths.

3.8 Config (config/)

ForecastConfig is the round-trip record for ForecastModel. Its six sub-configs match the stage method signatures one-for-one:

ForecastConfig
β”œβ”€β”€ model:     BaseForecastConfig
β”œβ”€β”€ calculate: ForecastCalculateConfig | None
β”œβ”€β”€ train:     ForecastTrainConfig     | None
β”œβ”€β”€ predict:   ForecastPredictConfig   | None
β”œβ”€β”€ evaluate:  ForecastEvaluateConfig  | None
└── explain:   ForecastExplainConfig   | None

TrainingConfig, PredictionConfig, EvaluationConfig, and ExplanationConfig each mirror their stage model's __init__ directly and are the standalone equivalents used when the model runs outside ForecastModel. Every stage model auto-calls save_config() at the end of its main run method (fit() / forecast() / evaluate() / explain()), so a standalone run always leaves a YAML snapshot next to its artefacts. The upstream model parameter on EvaluationConfig and ExplanationConfig is intentionally omitted because it is always a live model instance.

3.9 Decorators (decorators/)

notify(label) wraps a function with start/finish/error Telegram messages. send_telegram_notification(message, files, file_caption) is the one-off helper used inside scenarios.py to ship each per-scenario plot.

3.10 Utils (utils/)

Eight focused modules that the rest of the codebase pulls from - see the table in 6.


4. Model Class Relationships

                            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                            β”‚       BaseModel         β”‚
                            β”‚   (ABC)                 β”‚
                            β”‚ β€’ dates, output_dir     β”‚
                            β”‚ β€’ tremor_data (lazy)    β”‚
                            β”‚ β€’ n_jobs clamp          β”‚
                            β”‚ β€’ save() / load()       β”‚
                            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                         β”‚ inherits
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β–Ό                β–Ό           β–Ό               β–Ό                β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚ TrainingModel β”‚ β”‚PredictionModelβ”‚ β”‚EvaluationMdl β”‚ β”‚ ExplanationModel  β”‚
  β”‚ (BaseModel)   β”‚ β”‚ (BaseModel)   β”‚ β”‚(BaseModel)   β”‚ β”‚ (BaseModel)       β”‚
  β”‚               β”‚ β”‚               β”‚ β”‚              β”‚ β”‚                   β”‚
  β”‚ build_label β†’ β”‚ β”‚ build_label β†’ β”‚ β”‚ dispatch on  β”‚ β”‚ explain β†’         β”‚
  β”‚ extract_feat β†’β”‚ β”‚ extract_feat β†’β”‚ β”‚ model.kind   β”‚ β”‚   ExplainerEns.   β”‚
  β”‚ fit (N seeds) β”‚ β”‚ forecast      β”‚ β”‚ evaluate/    β”‚ β”‚ plot β†’            β”‚
  β”‚               β”‚ β”‚               β”‚ β”‚ compare      β”‚ β”‚   per-seed +      β”‚
  β”‚               β”‚ β”‚               β”‚ β”‚              β”‚ β”‚   waterfall       β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚ produces        β”‚ consumes      β”‚ uses             β”‚ reuses
           β–Ό                 β”‚               β–Ό                  β”‚
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”‚
  β”‚  ClassifierEnsemble    β”‚β—„β”˜  β”‚   MetricsEnsemble      β”‚      β”‚
  β”‚ ───────────────────    β”‚    β”‚ β€’ (n_samples Γ— n_seeds)β”‚      β”‚
  β”‚  β€’ from_any / from_jsonβ”‚    β”‚   y_proba / y_pred CSV β”‚      β”‚
  β”‚  β€’ from_seed_ensembles β”‚    β”‚ β€’ metrics in memory    β”‚      β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β”‚
             β”‚ bundles                      β”‚ aggregates        β”‚
             β–Ό                              β–Ό                   β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚   SeedEnsemble Γ— M     β”‚    β”‚ ClassifierComparator   β”‚ β”‚ ExplainerEnsembleβ”‚
  β”‚ ─────────────────────  β”‚    β”‚ β€’ get_ranking()        β”‚ β”‚ β€’ TreeExplainer  β”‚
  β”‚ β€’ predict_proba        β”‚    β”‚ β€’ plot_all()           β”‚ β”‚   per seed       β”‚
  β”‚ β€’ predict_with_        β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β€’ ClassifierExplnβ”‚
  β”‚   uncertainty          β”‚                               β”‚   per classifier β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                               β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚ inherits
             β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚      BaseEnsemble       β”‚
 β”‚    (joblib save/load)   β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Scope cheat-sheet:

Class Scope (per …) Mixin / Inheritance Cache
BaseModel - ABC (cache identity + dual-mode save/load) self
BaseEnsemble - mixin βœ—
TrainingModel One date span BaseModel βœ“
PredictionModel One forecast window grid BaseModel βœ“
EvaluationModel One trained model BaseModel βœ—
ExplanationModel One trained ensemble BaseModel βœ“
SeedEnsemble 1 classifier Γ— N seeds BaseEnsemble βœ—
ClassifierEnsemble M classifiers Γ— N seeds BaseEnsemble βœ—
MetricsEnsemble 1 ensemble Γ— 1 dataset standalone βœ—
ExplainerEnsemble 1 ensemble Γ— 1 dataset standalone βœ—
ClassifierComparator M classifiers, post-eval standalone βœ—
ForecastModel Full pipeline standalone orchestrator via stages

5. Pipeline Data Flow

5.1 Per-stage I/O

Stage Driver class Reads Writes
Tremor CalculateTremor SeismicDataSource.get(date) tremor/daily/*.csv, merged {nslc}_{start}_{end}.csv
Label LabelBuilder Tremor index, eruption dates training/features/{cv}/features-label_*.csv
Tremor matrix TremorMatrixBuilder Tremor CSV + labels training/tremor/tremor_matrix_*.csv (+ per_method/)
Features FeaturesBuilder Tremor matrix training/features/{cv}/features-matrix_*.parquet
Feature selection FeatureSelector Features + labels training/features/{cv}/seed/{seed:05d}.csv + top_N_features.csv
Training fit TrainingModel Selected features + labels training/classifiers/{clf}/{cv}/models/*.pkl + SeedEnsemble_*.pkl + ClassifierEnsemble_*.{pkl,json}
Prediction grid PredictionModel Tremor CSV + window grid prediction/features/features-{matrix,label}_*.csv
Forecast PredictionModel.forecast Forecast features + ensemble prediction/results/{clf}/{seed:05d}.csv + forecast-results_*.csv + prediction/figures/forecast_*.{png,pdf}
Evaluation EvaluationModel.evaluate y_proba + y_true evaluation/{kind}/classifiers/{Clf}/predictions/{y_proba,y_pred}.csv + figures/aggregate/{plot}.{png,csv} + (when plot_per_seed=True) figures/{plot}/{seed:05d}.png
Compare ClassifierComparator Cached MetricsEnsemble evaluation/{kind}/comparison/metrics/ranking_*.csv + comparison/figures/*.png
Explanation ExplanationModel.explain ClassifierEnsemble + features explanation/{kind}/classifiers/{Clf}/ClassifierExplanation_*.pkl + shap_values/{seed:05d}.pkl + figures/{bar,beeswarm}/{seed:05d}.png
Waterfalls ExplainerEnsemble.plot_waterfall ClassifierExplanation + eruption dates explanation/{kind}/eruptions/{date}/{Clf}_{datetime}_seed=_index=.png

5.2 On-disk artefact graph

            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β”‚   tremor/{nslc}_{start}_{end}.csv      β”‚  ← CalculateTremor
            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β”‚ used by Training / Prediction / Evaluation
                      β–Ό
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚ training/                                                     β”‚
    β”‚  features/{cv}/                                               β”‚
    β”‚    features-matrix_*.parquet ──► features-label_*.csv         β”‚
    β”‚       β”‚                                                       β”‚
    β”‚       β–Ό                                                       β”‚
    β”‚    seed/{seed:05d}.csv  ──► resampled/{seed:05d}.csv          β”‚
    β”‚    top_{N}_features.csv  + .png                               β”‚
    β”‚                                                               β”‚
    β”‚  classifiers/                                                 β”‚
    β”‚    {clf}/{cv}/models/{seed:05d}.pkl                           β”‚
    β”‚    {clf}/{cv}/SeedEnsemble_*.pkl                              β”‚
    β”‚    ClassifierEnsemble_{cv}.{pkl,json}                         β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚ ClassifierEnsemble bundle
              β–Ό
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚ prediction/                                                   β”‚
    β”‚   features/features-matrix_*.parquet + features-label_*.csv   β”‚
    β”‚   results/{clf}/{seed:05d}.csv                                β”‚
    β”‚   figures/forecast_*.{png,pdf}                                β”‚
    β”‚ forecast-results_*.csv  (top-level dump)          β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚ ClassifierEnsemble + features + y_true (rebuilt or training-derived)
              β–Ό
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚ evaluation/{training|prediction}/                             β”‚
    β”‚   classifiers/{Clf}/                                          β”‚
    β”‚     predictions/{y_proba,y_pred}.csv   (n_samples Γ— n_seeds)  β”‚
    β”‚     figures/aggregate/{plot_name}.{png,csv}                   β”‚
    β”‚     figures/{plot_name}/{seed:05d}.png   (plot_per_seed=True) β”‚
    β”‚   labels/y_true.csv                    (prediction reuse only)β”‚
    β”‚   MetricsEnsemble.pkl                  (optional, via save()) β”‚
    β”‚   comparison/                                                 β”‚
    β”‚     metrics/ranking_*.csv                                     β”‚
    β”‚     figures/*.png                                             β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚ ClassifierEnsemble + features
              β–Ό
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚ explanation/{training|prediction}/                            β”‚
    β”‚   classifiers/{Clf}/                                          β”‚
    β”‚     ClassifierExplanation_{Clf}.pkl                           β”‚
    β”‚     shap_values/{seed:05d}.pkl                                β”‚
    β”‚     figures/{bar,beeswarm}/{seed:05d}.png                     β”‚
    β”‚   eruptions/{YYYY-MM-DD}/                                     β”‚
    β”‚     {Clf}_{datetime}_seed=_index=.png                         β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

           β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
           β”‚  Stage-internal caches (no separate cache/ subtree):       β”‚
           β”‚    training/{hash}.TrainingModel.pkl       + .params.json  β”‚  ← BaseModel.save
           β”‚    prediction/{hash}.PredictionModel.pkl   + .params.json  β”‚
           β”‚    explanation/{kind}/{hash}.ExplanationModel.pkl + sidecarβ”‚
           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

A cache hit on TrainingModel short-circuits everything in the training/ box; a cache hit on PredictionModel short-circuits the prediction/ box; a cache hit on ExplanationModel short-circuits the per-classifier SHAP pass. Evaluation is never cached - the on-disk matrices act as the cache and MetricsEnsemble.compute() is idempotent in memory once y_probas is populated.


6. Utility Modules

Module Key functions
utils/array.py detect_maximum_outlier, remove_outliers, detect_anomalies_zscore, aggregate_seed_probabilities, predict_proba_from_estimator
utils/window.py construct_windows, calculate_window_metrics
utils/date_utils.py to_datetime, normalize_dates, sort_dates, parse_label_filename, set_datetime_index, label_id_to_datetime
utils/ml.py random_under_sampler, get_significant_features, load_labels_from_csv, save_model_json, compute_seed_eruption_probability, compute_model_probabilities, get_classifier_models, compute_g_mean, compute_seed, build_y_true
utils/validation.py validate_random_state, validate_date_ranges, validate_window_step, validate_columns, check_sampling_consistency
utils/pathutils.py resolve_output_dir, ensure_dir, save_figure, save_data, load_json
utils/dataframe.py load_label_csv, DataFrame shape and column helpers
utils/formatting.py slugify, human-readable elapsed time and file sizes

utils/ml.save_model_json writes the per-classifier trained-model JSON registry (one record per seed, each with the inline top-N feature list and the path to the seed's .pkl). TrainingModel.build_seed_ensemble reads that registry via SeedEnsemble.from_any to package every seed into a SeedEnsemble, and the per-classifier SeedEnsembles are then merged into a ClassifierEnsemble (build_classifier_ensemble). All three steps run at the end of TrainingModel.fit().

utils/formatting.slugify is what turns "Scenario 1" into scenario-1 for the per-scenario output_dir used in scenarios.py.

Clone this wiki locally