Human-centred demand planning intelligence.
AI agents surface evidence. You make the call. Every decision gets scored.
Built in three days for a hackathon, with Manasi Patil.
Demand planners override machine forecasts every month and write down a reason. Compass makes that loop useful:
- A panel of five specialist agents reads the real data — order books, inventory, supplier reliability, marketing campaigns, financials — and surfaces grounded evidence about the proposed override.
- A reconciler (Claude Opus 4.8) weighs all the evidence and suggests a calibrated number with a plain-English rationale.
- The planner commits a final number. The decision is stored.
- When actuals arrive, every decision is scored: did the override improve or worsen forecast accuracy? That score feeds back into the next decision.
The result: planners see exactly what the data says before they commit, and the system gets smarter every cycle.
We ran a Forecast Value Added (FVA) analysis on 23 historical Alpine planning cycles (68,126 item-months). Every figure below is recomputed by compass.fva.get_headline_stats() and cached in data/headline_stats.json:
| Metric | Machine | Planner |
|---|---|---|
| MAE | 2,013 units | 1,574 units |
| WAPE | 19.2% | 15.0% |
- Planner overrides cut MAE from 2,013 to 1,574 units, a reduction of 21.8%, and WAPE from 19.2% to 15.0%
- Planner overrides helped on 61.4% of all rows, and on 64.3% of rows where the override exceeded 5%
- Split by direction, on overrides above 5%: upward overrides helped 69.9% of the time (n=28,361), downward overrides helped 57.5% of the time (n=23,186)
- Planners add value in both directions, and more of it upward. That asymmetry is what Compass learns
pip install -r requirements.txtTwo source files are not distributed with this repository: alpine_sales_actuals.parquet and alpine_statistical_forecast.parquet. Without them scripts/setup.py cannot rebuild the FVA cache. The app still runs, because data/fva_base.parquet and data/headline_stats.json are committed outputs of a previous run.
Set your API key (needed for the live AI reconciler):
# Create .env from the template
cp .env.example .env
# Add your key: ANTHROPIC_API_KEY=sk-ant-...Without a key the app runs in demo mode with a rule-based stub.
# Initialise the database and precompute cached files (run once)
# Needs the two uncommitted source files above. Skip if you do not have them.
python scripts/setup.py
# Seed realistic historical decisions for demo (run once)
python scripts/seed_demo_data.py
# Launch
streamlit run app/main.pyApp opens at http://localhost:8501
Walk through a real decision:
- Select a product, sales org, and channel
- See the machine forecast alongside real historical context (run-rate, same month last year, stock, unit value)
- Move the slider to your proposed number and write your reason
- Hit Check my decision — five agents run, each grounded in a specific data table
- Review the evidence panel and the reconciler's suggested number with rationale
- Commit your final number
Shows what actually happened across 23 historical planning cycles. Two measured lines, machine forecast error and planner error, plus a third projected Compass line. The Compass line is a projection, not a measurement: Compass has no scored decisions of its own yet, so there is nothing to plot. It is labelled projected in the chart and in the caption above it.
These four show the most interesting stories:
| Product | ID | Org | Ch | Story |
|---|---|---|---|---|
| Premium Drills Mark II 1 | CP-0271 |
SO04 (DACH) | CH01 | Machine 29% below run-rate — clear upward case |
| Solvay CYCOM 5320-1 Tooling | SM-0545 |
SO04 (DACH) | CH04 | Conflicting signals: recent trend weak, campaign active, last year was 4,818 |
| Toray T700 Prepreg Sheet | SM-0546 |
SO16 (UK) | CH01 | Machine 85% above run-rate — downward override tension |
| Hazet 916HP Combination Wrench II | CP-0312 |
SO01 (DACH) | CH02 | Clean upward: machine 38% below three-month trend |
Use CP-0271 / SO04 / CH01 as the default — all five agents return real signals for this product.
Planner enters override + reason
│
▼
[Reason classifier → fixed taxonomy]
│
┌────┴──────────────────────────────────────┐
▼ ▼ ▼ ▼ ▼
DEMAND SUPPLY FINANCE COMMERCIAL MEMORY
orders stock biz plan campaigns scored
customers OTIF costs price chg history
returns suppliers tariffs
└────┬──────────────────────────────────────┘
│ signals with claim + direction + evidence rows
▼
RECONCILER (Claude Opus 4.8) — one call per decision
recommended_value + confidence_level + rationale
│
▼
PLANNER commits final number
│
▼
DuckDB write-back → scored when actuals arrive → feeds next decision
| Component | Choice | Why |
|---|---|---|
| Data store | DuckDB | Embedded SQL, queries parquet directly, no server |
| Agents | Python (rule-based) | Deterministic, zero model cost per signal |
| Reconciler | Claude Opus 4.8 | One call per decision — the only AI inference in the pipeline |
| Frontend | Streamlit | Python-native, no React, fast to build |
| Charts | Plotly | Interactive 3-line accuracy chart |
Not using: Neo4j, vector databases, LangChain. Every cut was made because the simpler alternative covers the actual query.
Alpine Manufacturing GmbH is the anonymised name of the industrial group whose data was supplied for the hackathon. Company, product and customer identifiers are anonymised throughout.
- 600 SKUs across Precision Components, Consumer Products, Specialty Materials
- 18 EU sales organisations (DACH, BeNeLux, Western EU, Nordics, Eastern EU, UK, Southern EU)
- 4 channels (B2B key accounts, distributors, wholesale, e-commerce)
- 2.8M sales rows (May 2023 to May 2026), in
alpine_sales_actuals.parquet, which is not committed - 23 scored planning cycles for FVA replay
The source data is not distributed with this repository. It is anonymised client data supplied for the hackathon and is not mine to republish. What is committed is the aggregate output: data/fva_base.parquet (68,126 scored rows), data/headline_stats.json and data/replay_chart_cache.parquet. Every figure quoted above recomputes from those.
This means the Replay Dashboard runs from a fresh clone and the Live Decision view does not, because the five signal agents read the per-table source files. To run Live Decision you need data/alpine-manufacturing-gmbh/ populated with tables matching the schema in compass/loader.py.
Note on parquet loading: the source files use
uint32dictionary-encoded columns that crashpandas.read_parquet. Always usecompass.loaderfunctions, which decode via PyArrow before converting to pandas.
compass/
loader.py — safe parquet loader (uint32 dict decode)
store.py — DuckDB: init, write_decision, score_decision
fva.py — FVA computation for replay dashboard
insights.py — real-data facts: run-rate, last year, value/unit, stock
classifier.py — reason text → fixed taxonomy class
embeddings.py — sentence-transformers semantic recall
memory.py — Memory agent: track record lookup + calibrated suggestion
reconciler.py — single Claude Opus call per decision
pipeline.py — run_pipeline() entry point
router.py — relevance gate (which agents run)
contracts.py — shared dataclasses (DecisionContext, Signal, MemoryContext…)
agents/
demand.py — order book, customer trends, returns
supply.py — stock cover, OTIF, production attainment
finance.py — revenue vs plan, tariff movements
commercial.py — active campaigns, recent price changes
app/
main.py — Streamlit shell, CSS, routing, theme toggle
views/
live_decision.py — 4-step decision flow
replay_dashboard.py — FVA replay chart and period drill-down
scripts/
setup.py — one-command bootstrap (DB init + cache precompute)
seed_demo_data.py — inserts 18 scored demo decisions for memory context
precompute.py — standalone cache regeneration
data/
alpine-manufacturing-gmbh/ — source parquet files (not distributed)
compass.duckdb — generated by setup.py (gitignored)
headline_stats.json — pre-computed FVA summary stats
replay_chart_cache.parquet — pre-computed cycle-by-cycle accuracy
MIT. See LICENSE.