Shared ML utilities — LLM providers, evaluation metrics, visualization, and report generation.
Used as an in-house dependency across all of Sameer Maurya's ML projects and organization.
Source repo: the-forge — kept its original name;
only the installable package was renamed to meerax.
CI gates on mypy (strict = true, with 2 narrowly-scoped # type: ignore exceptions for known
third-party stub gaps) and a minimum 85% test coverage — both enforced on every PR, not just
checked locally.
Supports Python 3.9+ (requires-python = ">=3.9"). CI itself runs and is verified against
Python 3.12 only — no version matrix — but installs and the full test suite have been manually
verified clean on 3.9, 3.10, 3.11, and 3.12 before each release that touches this floor.
pip install meeraxOr pin in requirements.txt:
meerax==1.10.0
| Module | What it gives you |
|---|---|
meerax.llm |
Swap-in LLM backends — Claude, OpenAI, Ollama behind one interface, text or images |
meerax.eval.classification |
F1, AUC-ROC, precision, recall in one call |
meerax.eval.timeseries |
RMSE, MAPE, SMAPE, ADF stationarity test |
meerax.eval.text |
BLEU-4, ROUGE-L for caption / summary quality |
meerax.viz |
Dark-themed matplotlib plots (confusion matrix, ROC, forecast, decomposition) |
meerax.data |
CSV/parquet loaders with schema validation, stratified + time splits, SMOTE |
meerax.report |
Self-contained dark-themed HTML model-card report builder |
meerax.logging |
One-call structured logger factory |
meerax.vision |
Image folder dataset loader (PyTorch) + translation-grid plotting (torch or numpy/TF images) |
Every project in the ecosystem follows the same PROJECT_STANDARDS.md layout and depends on
meerax. The meerax CLI (installed alongside the package) generates or retrofits that layout:
# brand-new project
meerax new my-project --path ~/dev
# retrofit an existing, non-empty directory — additive only, never overwrites
cd ~/dev/my-existing-notebook-project
meerax initmeerax new creates the full src/{core,providers,services,utils,data} + tests/ + CI
skeleton, pins requirements.txt to the current meerax release, and runs git init. The
generated ci.yml calls this repo's reusable CI workflow
instead of embedding its own copy, so fixes to the shared CI logic reach every project that
uses it without needing to be manually reapplied. Also generates .github/dependabot.yml
(pip + github-actions, weekly) so dependency pins don't quietly go stale.
Both new and init accept --template llm-report, which adds a real, runnable example on
top of the bare skeleton — the "call an LLM, build an HTML report" shape that's shown up twice
already (trend-whisperer, pixel-drift). It's genuinely functional code with real tests, not
stub methods to fill in:
meerax new my-app --template llm-report
cd my-app && pip install -r requirements.txt
python -m src.app --prompt "Summarize this quarter's churn" --llm-provider ollamameerax init fills in whatever's missing from that same layout without touching files that
already exist, and reports any top-level files it doesn't recognize (e.g. notebooks) so you can
move them into src/ by hand.
meerax doctor checks an existing project against PROJECT_STANDARDS.md — no Python version
matrix, no committed docs/specs/docs/plans, a LICENSE file, VERSION/README/CHANGELOG
consistency, Dependabot configured, and whether the project's meerax pin is current:
cd ~/dev/my-project
meerax doctorExits non-zero if anything fails, so it's safe to run in CI.
meerax bump <version> updates VERSION, the README version badge, and inserts a dated
CHANGELOG heading in one step — the exact multi-file edit that's caused version-badge drift
more than once across this ecosystem's history:
meerax bump 1.2.0It doesn't write CHANGELOG content or a compare-link footer — those need someone who actually knows what changed.
from meerax.llm import ClaudeProvider, PromptTemplate
from meerax.eval import evaluate_classifier
from meerax.viz import apply_meerax_theme
from meerax.report import ReportBuilder, ReportSection
# LLM: swap provider without changing downstream code
llm = ClaudeProvider() # or OpenAIProvider() / OllamaProvider()
tpl = PromptTemplate("Explain {finding} to a risk manager in 3 sentences.")
response = llm.generate(tpl.render(finding="high AUC-ROC with low recall"))
# Eval
metrics = evaluate_classifier(y_true, y_pred, y_prob=probabilities)
print(metrics)
# Accuracy : 0.9823
# F1 : 0.8741
# AUC-ROC : 0.9912
# Viz + Report
apply_meerax_theme()
rb = ReportBuilder("Fraud Detection — Model Report v0.1.0")
rb.add_section(ReportSection(
title="Performance",
metrics=metrics.to_dict(),
content=response.content,
))
rb.save("reports/model_report.html")All providers implement LLMProvider.generate() and .chat(). Swap with one line:
from meerax.llm import ClaudeProvider, OpenAIProvider, OllamaProvider
llm = ClaudeProvider() # needs ANTHROPIC_API_KEY
llm = OpenAIProvider() # needs OPENAI_API_KEY
llm = OllamaProvider() # needs Ollama running locallySelf-contained benchmark scripts in benchmarks/:
| Script | Description |
|---|---|
meerax_benchmark.py |
Times meerax's own core operations — CSV loading, data splitting, classification/timeseries eval metrics, HTML report generation — across representative sizes |
kv_cache_benchmark.py |
KV caching simulation at GPT-2 Medium scale |
python benchmarks/meerax_benchmark.py # print results
python benchmarks/meerax_benchmark.py --record # also append to benchmarks/results/history.jsonl--record appends one JSON line per run, keyed by the installed meerax version, so
performance can be compared release to release. Run it as part of cutting a release,
alongside meerax bump, and commit the updated history.jsonl in the same commit. A
GitHub Actions workflow (.github/workflows/benchmarks.yml) also runs it on demand
(workflow_dispatch) and uploads the results as a build artifact — informational only,
not a required CI check, since wall-clock timings on shared runners are too noisy to
gate on.
the-forge/
├── meerax/ # Installable package
│ ├── llm/ # LLM provider abstraction
│ ├── eval/ # Evaluation metrics
│ ├── viz/ # Visualization utilities
│ ├── data/ # Data loading, splitting, resampling
│ ├── report/ # HTML report builder
│ ├── scaffold/ # Project skeleton templates + create/retrofit logic
│ ├── cli.py # `meerax new` / `init` / `doctor` / `bump` command entry point
│ ├── doctor.py # PROJECT_STANDARDS.md compliance checks
│ ├── release.py # VERSION/README badge/CHANGELOG bump helper
│ ├── vision/ # Image dataset loader + translation-grid plotting
│ └── logging.py # Structured logger
├── benchmarks/ # meerax's own perf benchmarks + standalone ML demo scripts
│ └── results/ # history.jsonl — one line per --record run, tracked across releases
├── tests/
│ ├── unit/ # 152 unit tests, zero external deps
│ └── integration/ # 4 cross-module pipeline tests
├── ARCHITECTURE.md
├── SECURITY.md
├── LICENSE
├── pyproject.toml
├── requirements.txt
└── VERSION