Skip to content

yib7/STATlee

Repository files navigation

STATlee

Turn data questions into answers, in plain English.

STATlee is AI data analysis for social scientists. Upload a dataset, describe the analysis in plain English, and STATlee writes, moderates, sandboxes, runs, and explains the statistics. No Python, no R, no syntax errors.

Architecture · Pricing · Credits

Python 3.11+ Flask CI Tests LLM providers License: Elastic 2.0

STATlee end-to-end demo: upload a dataset, ask a question in plain English, watch it generate and run the code, then see the chart and a written report

Upload data, ask in plain English, get runnable code, a chart, and a written report. Try it with the included sample survey dataset.

STATlee workspace: dataset upload, plain-English analysis prompt, and the Code / Data / Results / Converse / Report / Codebook tabs


Why STATlee

Statistical software is hostile to non-coders. SPSS menus, Stata syntax, and R tibbles all stand between a researcher and a simple question like "does income predict turnout?" STATlee removes that wall: you ask in plain English and get a runnable analysis, charts, and a plain-English write-up you can defend.

Try it: the marketing landing page lives at /welcome; the app itself is at /.

What it does

  • Describe it, don't code it. Natural-language requests become real Python/R that runs and returns results, not just a code snippet.
  • Statistically valid by design. An intelligent codebook classifies each variable (nominal / ordinal / continuous) so the model won't run a linear model on a categorical outcome. Codebooks can be inferred from a PDF data dictionary or even the original survey questionnaire.
  • Secure sandbox. Generated code runs in a throwaway working directory with a secret-free environment, screened first by a non-LLM static pre-check (an AST denylist rejecting network access, process/shell spawning, os/env exfiltration, and file reads outside the run dir) and a run-guard that re-moderates any hand-edited script. The default SANDBOX_MODE=subprocess scrubs secrets from the environment but does not isolate the network or host filesystem, so those two gates are the boundary; SANDBOX_MODE=docker adds a network-less, non-root, read-only, resource-capped container per run for kernel-enforced isolation. Prefer docker mode where the host holds secrets or network access that matters.
  • Answers you can explain. Dense terminal output and p-values become plain-English Markdown (effect sizes, significance, caveats), and a debugging assistant kicks in when a run fails.
  • Bring any format. CSV, TSV, Excel (.xlsx/.xls), Stata (.dta), and SPSS (.sav); native value labels seed the codebook for free.
  • Conversational wrangling. Clean your data by chatting ("delete the notes column", "filter for age > 30") with a back-and-forth transcript, full version history (undo / redo), and a one-click revert to the original upload. Runs on the cheapest model tier (WRANGLE_ROLE=lite) to keep costs down. Plus an AI report builder grounded strictly in your real outputs and one-click project export (data + script + plots + report).
  • Pro mode toggle generates code with a larger model (gemini-3.1-pro-preview by default) for hard or unusual analyses, instead of the standard code model.

How it works

Upload data ──▶ Intelligent codebook ──▶ Describe analysis (plain English)
     │                                              │
     ▼                                              ▼
 Multi-format            moderation ▶ feature-select ▶ draft ▶ validate
 normalize to CSV                          │
                                           ▼
                          Secure sandbox (subprocess / Docker)
                                           │
                                           ▼
                     Plain-English interpretation + charts ──▶ Report / export

Every model call goes through one role-based LLM service (pro / flash / lite / draft), so swapping a model, switching LLM provider, or pointing Pro mode at a stronger model is a config change, not a code change. Per-request token usage is surfaced live in the UI.

Quickstart

cp .env.example .env          # add your GEMINI_API_KEY
docker-compose up --build     # then open http://localhost:5000

Or run locally without Docker (dev only, generated code runs on your host):

python -m venv venv && source venv/bin/activate   # Windows: venv\Scripts\activate
pip install -r requirements.txt
APP_ENV=development python wsgi.py                 # http://localhost:5000

LLM configuration

STATlee has a pluggable LLM provider: pick one with LLM_PROVIDER and supply that vendor's key. Switching providers is a one-line config change; no route, prompt, or UI code changes (every call routes through one role-based service).

LLM_PROVIDER API key Notes
gemini (default) GEMINI_API_KEY (AI Studio) Powers the multimodal paths (PDF codebook extraction, plot interpretation).
anthropic ANTHROPIC_API_KEY (Console) Claude models; also accepts a local OAuth/subscription login for keyless dev.
openai OPENAI_API_KEY (Platform) GPT models.

The selected provider's key is required in production. Each role's model id defaults sensibly per provider and can be pinned independently with MODEL_PRO / MODEL_PRO_MAX / MODEL_FLASH / MODEL_FLASH_LITE (see .env.example for the per-provider defaults).

See docs/README.md for the full setup, Docker sandbox build, and the documented .env.example.

Development

pip install -r requirements-dev.txt
ruff check .      # lint
pytest -q         # 400 tests, fully offline (deterministic fake LLM, no API key)

The test suite injects a fake LLM service, so the entire HTTP surface (uploads, codebook, wrangling, run-guard, converse, export, auth, CSRF, rate-limit keying, billing seam, Pro-mode routing) is exercised without network or API keys. CI runs ruff + byte-compile + pytest on every push (.github/workflows/ci.yml).

Architecture at a glance

The application lives in the statlee/ package (entry point: wsgi.pystatlee.app:app).

Module Responsibility
statlee/config.py Validated, env-driven configuration (one source of truth).
statlee/app.py App factory + middleware (sessions, CSRF, rate limits, ProxyFix, request-id logging).
statlee/storage.py Per-identity file isolation + dataset version control.
statlee/sandbox.py Isolated code execution (subprocess or Docker).
statlee/llm.py Role-based LLM service (pluggable Gemini / Claude / OpenAI backend): usage tracking, Pro-mode routing, response cache.
statlee/billing.py Monetization seam: check_and_debit chokepoint (off by default; atomic debit + refund + monthly top-up when enabled).
statlee/prompts.py Every prompt builder in one reviewable place.
statlee/datatools.py Multi-format ingestion + metadata profiling.
statlee/models.py SQLAlchemy models (users, runs, issue reports).
statlee/routes/ Blueprints: auth, datasets, analyze, converse, misc.

A deeper walkthrough (request lifecycle, security boundaries, data flow) is in docs/ARCHITECTURE.md.

Hosting & pricing

STATlee runs locally for free and is not currently deployed to a public website, a deliberate, zero-cost choice. Self-hosted, its only real cost is the LLM provider's API (Gemini by default), and it ships with the controls to cap that spend at a number you choose: a global monthly ceiling on premium calls, per-identity rate limits, cheapest-tier defaults, and a startup warning if billing is on without a cap. The (illustrative) monetization model and how the billing seam backs it are in Pricing.

Security

Because STATlee executes model-generated (and user-editable) code, execution safety is layered. Every script passes a non-LLM static pre-check (an AST denylist rejecting network access, process/shell spawning, os/env exfiltration, and reads outside the run directory) and an LLM run-guard that re-moderates edited scripts, before running in a throwaway, secret-free working directory. In the default SANDBOX_MODE=subprocess those two gates are the boundary: the environment is scrubbed of secrets, but the network and host filesystem remain reachable, so run that mode only where the host holds nothing sensitive. SANDBOX_MODE=docker makes the boundary kernel-enforced with a network-less, non-root, read-only, resource-capped container per run. Rate limits are keyed per-identity (client IP / account, not a resettable cookie) to protect against bill abuse. Responses carry a strict Content-Security-Policy and the usual hardening headers, CSRF is double-submit with a constant-time check, and the schema is versioned with Alembic so an existing database upgrades cleanly instead of breaking on a new column.

Found a vulnerability? See SECURITY.md for how to report it privately. Release notes are in CHANGELOG.md.

Compliance & attribution

STATlee's AI analysis is powered by your configured LLM provider, Google Gemini by default (with optional Anthropic Claude or OpenAI backends), and the in-app footer attributes the active provider. Default Gemini use is subject to the Gemini API Additional Terms and the Prohibited Use Policy, which STATlee's moderation gate enforces.

Third-party libraries and their licenses are listed in CREDITS.md.

STATlee is a research aid: always review generated code and results before relying on them.

License

STATlee is source-available under the Elastic License 2.0: you're free to use, modify, and self-host it, but you may not offer it to third parties as a hosted or managed service. This keeps the source open to read and learn from while reserving the right to run it commercially. See CREDITS.md for third-party dependency licenses.

About

AI data-analysis sandbox for social scientists: describe an analysis in plain English and STATlee generates, moderates, sandboxes, runs, and explains real Python/R. Flask app with a pluggable LLM provider (Gemini, Claude, OpenAI).

Topics

Resources

License

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Packages

 
 
 

Contributors