Skip to content

Configuration

Martanto edited this page Jun 10, 2026 · 7 revisions

Configuration

Every ForecastModel stage auto-captures its kwargs into a ForecastConfig object, which can be saved to YAML/JSON and replayed via from_config() → run(). This page covers config persistence, the dataclass layout, Telegram notifications, and runtime logging.


ForecastConfig Lifecycle

                   ┌────────────────────────────────────────┐
                   │            ForecastModel               │
                   │   _config: ForecastConfig              │
                   └────────────┬───────────────────────────┘
                                │ stage methods auto-capture kwargs
                                ▼
   ┌─────────────────────────────────────────────────────────────┐
   │  ForecastConfig                                              │
   │   ├── version, saved_at                                      │
   │   ├── model:      BaseForecastConfig                         │
   │   ├── calculate:  ForecastCalculateConfig | None             │
   │   ├── train:      ForecastTrainConfig     | None             │
   │   ├── predict:    ForecastPredictConfig   | None             │
   │   └── evaluate:   ForecastEvaluateConfig  | None             │
   └────────────────────────┬─────────────────────────────────────┘
                            │
                ┌───────────┴───────────┐
                ▼                       ▼
         fm.save_config()         ForecastModel.from_config(path)
         → forecast.config.yaml   → new ForecastModel
                                  → fm.run() replays each non-None section
  • A stage that hasn't run yet is None in the YAML — the produced config is "partial" and can be loaded + continued.
  • fm.evaluate(...) calls save_config() automatically before returning. Call it manually at earlier points to checkpoint a partial pipeline.

Default path

{station_dir}/forecast.config.yaml      # fm.save_config()
{station_dir}/forecast.config.json      # fm.save_config(fmt="json")

{station_dir} = {output_dir}/{network}.{station}.{location}.{channel} — sibling of the per-stage cache/ directories.

Round-trip

fm.save_config("output/config.yaml")
fm2 = ForecastModel.from_config("output/config.yaml")
fm2.run()    # idempotent — replays every captured stage

ForecastConfig Schema (YAML)

# eruption-forecast ForecastModel configuration
version: "1.0"
saved_at: "2026-06-10T11:23:45"

model:
  station: OJN
  channel: EHZ
  network: VG
  location: "00"
  day_to_forecast: 2
  output_dir: null
  root_dir: null
  overwrite: false
  n_jobs: 8
  verbose: true

calculate:
  start_date: "2025-01-01"
  end_date: "2025-12-31"
  source: sds
  methods: [rsam, dsar, entropy]
  remove_outlier_method: maximum
  remove_tremor_anomalies: false
  interpolate: true
  plot_daily: true
  save_plot: true
  overwrite_plot: true
  sds_dir: "D:/Data/OJN"
  client_url: "https://service.iris.edu"
  minimum_completion_ratio: 0.3
  overwrite: false
  n_jobs: null               # null → inherit from model.n_jobs at replay
  verbose: null

train:
  start_date: "2025-01-01"
  end_date: "2025-07-26"
  eruption_dates:
    - "2025-03-20"
    - "2025-04-22"
  window_step: 6
  window_step_unit: hours
  label_builder: standard
  classifiers: [lite-rf, rf, gb, xgb]
  cv_strategy: shuffle-stratified
  cv_splits: 5
  scoring: recall
  top_n_features: 20
  include_eruption_date: true
  select_tremor_columns: [rsam_f2, rsam_f3, rsam_f4, dsar_f3-f4, entropy]
  save_tremor_matrix_per_method: true
  exclude_features: [agg_linear_trend, linear_trend_timewise, length]
  seeds: 25
  resample_method: under
  sampling_strategy: 0.75
  plot_features: true
  n_jobs: 4
  n_grids: 4
  use_cache: true

predict:
  start_date: "2025-07-27"
  end_date: "2025-08-22"
  window_step: 10
  window_step_unit: minutes
  save_seed_result: true
  plot_threshold: 0.7
  plot_pdf: true
  use_cache: false

evaluate:
  model: prediction
  plot_per_seed: true
  plot_aggregate: true

The keys mirror the kwargs accepted by each method 1:1 — see API Reference for the per-stage signatures.

None-as-inherit

For overwrite, n_jobs, and verbose, a YAML value of null means "inherit the value ForecastModel.__init__ was constructed with". This is the same semantics applied at runtime when the kwarg is omitted, so a replay behaves identically.


TrainingConfig (Standalone)

TrainingModel.__init__ captures its own kwargs into a TrainingConfig dataclass (config/training_config.py), independent of ForecastConfig. Use this when you run TrainingModel outside of ForecastModel:

tm.save_config()      # → {output_dir}/training.config.yaml

The shape mirrors the TrainingModel.__init__ signature — see Training Workflow.


config.example.yaml

A fully annotated example config ships at the repo root: config.example.yaml. Project Rule 11 keeps it in sync with forecast_config.py — when any ForecastConfig field is added, renamed, or has its default changed, the example YAML is updated in the same commit.


Telegram Notifications

eruption_forecast.decorators exposes two complementary primitives.

notify decorator

Wraps a function to send a Telegram message on success or failure:

from eruption_forecast import notify
import dotenv; dotenv.load_dotenv()

@notify("Run Forecasting")
def main():
    fm = ForecastModel(...)
    fm.calculate(...).train(...).predict(...).evaluate(...)

main()    # Telegram chat receives start, finish, and error messages

send_telegram_notification(...) helper

Used by scenarios.py to ship the per-scenario forecast plot:

from eruption_forecast import send_telegram_notification

send_telegram_notification(
    message=f"{name}: {description}",
    files=[fm.PredictionModel.forecast_plot_path],
    file_caption=f"{name}: {description}",
    send_as_document=True,        # preserves DPI — Telegram does not re-encode
)

Credentials (.env)

TELEGRAM_BOT_TOKEN=your_bot_token_here
TELEGRAM_CHAT_ID=your_chat_id_here

Both primitives degrade gracefully when the env vars are absent — they emit a warning and skip the network call instead of raising.


Logging

The package wraps loguru behind eruption_forecast.logger.

Function Purpose
enable_logging() Restore console + file handlers using the current log directory
disable_logging() Remove every active loguru handler — no console, no file
set_log_level(level) Change the console handler level ("DEBUG" / "INFO" / "WARNING" / "ERROR" / "CRITICAL")
set_log_directory(dir) Move the log file to a new directory — created if missing
from eruption_forecast import enable_logging, disable_logging
from eruption_forecast.logger import set_log_level, set_log_directory

set_log_directory("logs/2026-06-10")
set_log_level("DEBUG")        # console only — file handlers keep their level

disable_logging()
fm.calculate(...)             # silent — useful during tests
enable_logging()              # restore handlers

enable_logging, disable_logging, notify, and send_telegram_notification are exported from the package root.


Where Configuration Lives in the Filesystem

{station_dir}/
├── forecast.config.yaml                 # fm.save_config()  — full pipeline
├── training.config.yaml                 # tm.save_config()  — standalone TrainingModel
├── cache/
│   ├── TrainingModel/{hash}.params.json # CacheModel identity dumps (diff-able)
│   └── PredictionModel/{hash}.params.json
└── ...

The *.params.json files inside cache/ capture exactly what went into the cache hash. They are handy when debugging a cache miss — diff two of them to see which kwarg differed.

Clone this wiki locally