This project scrapes, enriches, and serves AFL player data across all clubs. It provides a FastAPI-powered JSON API and supports data persistence via SQLite for players, injuries, team line-ups, and API key data.
AFL-api originated from a need to supply AFL data to external reporting and fantasy-football tools. FastAPI makes the collected, normalised data simple to consume from clients such as Google Apps Script as well as other applications.
The service itself remains consumer-neutral: it collects AFL data from supported sources, validates and normalises it, maintains canonical identities and source authority, persists supported data safely, and exposes stable, documented APIs or exports. Consumer-specific scoring rules, fantasy-league management, presentation, and client workflows belong in separate downstream applications, not in this repository.
Current version: 0.5.0. See the v0.5.0 release notes,
upgrade runbook, and
operational source policy. Runtime code
reads the authoritative declaration in version.py.
- π Scrape player data from official AFL club sites
- π§ Enrich player info with AFL.com.au & Champion Data IDs
- π Export enriched data per club to JSON
- π Persist all data in SQLite (players, matches, injuries, lineups, clubs)
- π Scrape AFL fixture by round: team names, venue, status, and scores
- β Match team names to club codes with alias support
- π©Ό Injury list scraper with timestamped records
- π¨βπ» FastAPI admin portal with:
- Visual diff of
clubs.jsonvs DB - One-click sync (import/export)
- Flash messages for admin actions
- Visual diff of
- π API key management (hashed in SQLite with one-time key display)
- π§ͺ CLI tools for scraping, importing, and enrichment
- π§ Nickname resolution and suffix cleaning for fuzzy player matching
- β° Scheduled scraping of injuries, lineups, and live player stats using APScheduler
- π§ Lineup scrapes run at predictable times (T-1 day 5pm, Thursday 5pm, and 1h before each match)
- π Player stat scrapes run automatically 2 minutes before each match start
- π Admin page shows all upcoming scheduled jobs
For a clean database, follow the authoritative supported first-run sequence and command-selection table before starting API, Scheduler, or Admin services. The abbreviated Docker example below demonstrates the portable development stack; it does not replace database migration, season bootstrap, and verification.
- Docker & Docker Compose installed
- Ports
8000and8001available on your host, or overrideAFL_API_PORTandAFL_ADMIN_PORTin.env
git clone https://github.com/JustPlausible/afl-api.git
cd afl-api
cp .env.example .env
docker compose -f compose.example.yaml up --buildThe repository root is the Docker build context. The Python source layout remains flat at the repository root; do not cd into a nested src directory.
This repository is now intended to be the canonical application project. The root-level Dockerfile builds the app image from the repository root and installs runtime dependencies from requirements.txt. The image default command starts the public API without --reload; the example Compose file uses --reload only for local development convenience.
Dependency manifests are split by purpose:
requirements.txtcontains runtime dependencies required by the API, admin UI, scheduler, and scrapers.requirements-dev.txtextends runtime dependencies and adds test/development-only dependencies.
compose.example.yaml is a portable example/development configuration. It intentionally avoids real production paths, ports, and secrets. It uses named volumes for local data and logs state and bind-mounts .:/app only for development reload workflows.
For production, keep the real Compose file, .env, paths, ports, and secrets outside this repository. A recommended server layout is:
/opt/projects/afl-api # Git checkout and Docker build context
/opt/docker/compose/afl-api # Production Compose and .env files
/opt/docker/appdata/afl-api/data # Persistent runtime database/data
/opt/docker/appdata/afl-api/logs # Persistent runtime logs
Production should build from /opt/projects/afl-api and should not mount source over /app. Mount only persistent runtime state, for example:
services:
afl-api:
build: /opt/projects/afl-api
env_file:
- /opt/docker/compose/afl-api/.env
volumes:
- /opt/docker/appdata/afl-api/data:/app/data
- /opt/docker/appdata/afl-api/logs:/app/logsThis keeps /opt/docker/appdata/afl-api as persistent runtime state only while the application source and Docker build context live under /opt/projects/afl-api.
The intended production boundary is:
afl-apiis the only service designed for public exposure, either by publishing its API port directly or by placing a reverse proxy in front of it. The public API still requiresx-api-keyauthentication on API routes.afl-adminis an operator-facing management interface protected by HTTP Basic authentication and CSRF protection. It remains reachable through the configured${AFL_ADMIN_PORT:-8001}port on a normaladmin-accessbridge network so deployments can secure it with trusted LAN access, firewall rules, VPNs, SSH forwarding, a private reverse proxy, or equivalent controls. Do not treat it as a public unauthenticated service.afl-schedulerhas scheduler management and mutation routes such as/scheduler/refresh,/scheduler/run/{job_id}, and/scheduler/job/{job_id}. These routes must not be published to a host port or exposed by a public reverse proxy.afl-adminandafl-schedulercommunicate over the Composemanagementnetwork athttp://afl-scheduler:8000. That network is markedinternal: true, so it is the trusted service-to-service management path.afl-adminis not attached to Docker's default network in the documented Compose layout. It joinsmanagementfor private scheduler calls andadmin-accessso Docker can create a usable host-side mapping for the authenticated operator port.afl-scheduleris not attached to the default application network oradmin-access; it joins the internalmanagementnetwork for admin calls and a separatescheduler-egressbridge network only to preserve outbound internet access required by AFL scraping jobs.
Do not publish scheduler ports in production Compose files. If scheduler mutation endpoints become reachable outside the trusted management network in a supported deployment, add the smallest practical shared internal token for those mutation routes before exposing that deployment. For the documented production Compose model, a separate scheduler token is not required because the scheduler is not host-published, admin reaches it on the internal management network, admin-access is only for the authenticated admin host port, and the configured scheduler egress network is for outbound scraping rather than inbound management access.
Use the operator command selector and supported first-run sequence to choose a workflow, and the implemented parser as the complete option reference:
python cli.py --helpJSON is the preferred operational source. For example, public AFL metadata can be inspected without database writes, or bootstrapped into the canonical tables:
python cli.py --collect-afl-metadata --afl-season 2026 --print-json
python cli.py --bootstrap-afl-season 2026--bootstrap-afl-season only establishes canonical metadata, player identity,
provider crosswalks, and season membership. To initialise or refresh that
foundation and authoritative concluded match statistics in one operation,
use:
python cli.py --sync-afl-season 2026
python cli.py --sync-afl-season 2026 --round 8
python cli.py --sync-afl-season 2026 --round-from 8 --round-to 12
python cli.py --sync-afl-season 2026 --match-id 8204
python cli.py --sync-afl-season 2026 --refresh-complete --print-jsonThe command requires the same public AFL access and authenticated CFS endpoints
used by bootstrap and single-match CFS collection. The client automatically
obtains and refreshes its WMC token; there is no configured operator credential
variable for that token. The command selects
matches from the persisted canonical season: only concluded/final matches with
an explicit CD_M... mapping are collected. Existing concluded authoritative
snapshots are skipped unless --refresh-complete is given. Each match is a
separate transaction, so one failure does not undo earlier matches.
Reruns use existing metadata/player upserts and CFS snapshot-authority rules:
unchanged observations are zero-write, corrected concluded observations may be
updated, and live, stale, empty, unavailable, or unknown responses never delete
or downgrade concluded data. Statistics go only to cfs_player_stats; there is
no rendered-HTML fallback or legacy player_stats dual write.
The JSON result distinguishes skipped, unavailable/unpublished, empty, partial,
unknown, failed, already-complete, and collected matches and includes correlated
audit IDs. Default output is a concise human-readable summary; --print-json
emits the complete machine-readable result, including match-level details, to
stdout even when the command subsequently exits non-zero.
Exit status 0 means all currently actionable concluded matches were collected
or already complete; scheduled, current live/postgame, and recognised future
placeholders may be safely skipped. Exit status 1 means material actionable or
explicitly requested work was unavailable, empty, partial, rejected, unknown,
missing provider identity, absent, or failed. Exit status 2 is invalid CLI
usage. A bounded round/range returning no matches exits 1; an unbounded season
with no published fixtures reports empty_unbounded and exits 0.
Round ranges are inclusive. Repeatable --match-id values are deduplicated in
first-seen order. Combining match IDs with a round or range uses intersection
semantics, so every requested ID must also satisfy the round bound. Round 0 is
valid, negative rounds are invalid, and match IDs must be positive.
Summary units are explicit: records_received and
total_matches_discovered count selected matches; eligible_matches counts
concluded matches with provider identity, including already-complete matches;
collected_successfully counts matches actually recollected; and
already_complete_unchanged counts matches skipped as complete. rows_inserted,
rows_updated, rows_unchanged, and rows_written count CFS statistic rows,
with rows_written equal to inserted plus updated rows. Deeper relationship completeness and
reconciliation reporting remains the responsibility of Issue #107.
The existing commands remain deliberately distinct:
--bootstrap-afl-season: canonical foundation only;--sync-afl-season: whole-season persistent foundation and CFS statistics;--collect-match-player-stats: one explicitly identified CFS match;--collect-afl-data --no-database: diagnostic files only, never persistence.
Authenticated Champion Data/CFS commands require opaque provider IDs, not AFL
numeric IDs. Roster collection is read-only; match player-stat collection
persists to cfs_player_stats:
python cli.py --collect-match-rosters CD_R202601421 --print-json
python cli.py --collect-match-player-stats CD_M20260142001 --print-jsonRendered-HTML commands remain available only as explicit legacy tools; the CLI never silently falls back from JSON to HTML. Club tools, persistence behavior, identifier formats, diagnostics, prerequisites, and every legacy command are documented in the operator CLI guide.
Scheduler and Admin player-stat jobs use the same operational CFS JSON collector
and write cfs_player_stats. The explicit CLI command
--collect-match-player-stats uses that path too; the separate manual
--scrape-match command renders HTML and writes only the legacy player_stats
table. The tables are not interchangeable and these workflows do not dual-write.
The complete identifier, reader, lifecycle, fallback, and migration rules are in
the player-stat storage contract.
Injuries and operational lineups deliberately remain rendered-HTML sources:
Scheduler and Admin persist them through the shared policy, while the CLI exposes
--scrape-injuries and --scrape-lineups explicitly. Injury rows are persisted
only after canonical AFL player-ID resolution.
All API routes require an x-api-key header. Full API keys are shown only when created or renewed; SQLite stores only a non-reversible hash plus a short administrative prefix. Existing plaintext keys are migrated in place during database initialisation as described in docs/api_key_migration.md.
Manage your keys via:
PYTHONPATH=/app python3 scripts/manage_api_keys.py --add "my-label"Example request:
curl -H "x-api-key: your_key_here" http://localhost:9900/playersThe admin application is protected with HTTP Basic authentication. Configure credentials with:
ADMIN_USERNAME=admin
ADMIN_PASSWORD=change-me
SESSION_SECRET=replace-with-a-long-random-secretADMIN_PASSWORD and SESSION_SECRET must be set explicitly when ENVIRONMENT=production. The public API continues to use the x-api-key header, and disabled API keys are rejected.
Diagnostic header echoing at /api/echo-headers requires a valid API key and redacts sensitive headers such as x-api-key, authorization, cookie, and x-admin-key.
π For full interactive API docs, visit
/docswhile the app is running.
π€ Player Endpoints
| Method | Endpoint | Description |
|---|---|---|
| GET | /players |
All players |
| GET | /players?club=RIC |
Filter by club short code |
| GET | /players/club/richmond |
By club slug |
| GET | /players/{afl_id} |
Single player by AFL ID |
π©Ό Injury Endpoints
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/injuries/{afl_id} |
Injuries for a player |
| GET | /api/injuries/{afl_id}?history=1 |
All historical injuries |
π§ββοΈ Line-up Endpoints
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/lineups/{round} |
Line-up data for the full round |
| GET | /api/lineups/{round}/{afl_id} |
Line-up data for a player in a round |
| GET | /api/lineups/latest/{afl_id} |
Most recent line-up entry for a player |
π Round Endpoints
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/rounds |
All available rounds |
| GET | /api/rounds/{round_id} |
Metadata for a specific round |
π Match Endpoints
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/matches |
All matches (optionally filter by round) |
| GET | /api/matches?round_id=1155 |
Matches for a specific round |
| GET | /api/matches/{match_id} |
Details of a single match |
π Legacy Player Stats Compatibility Endpoint
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/player-stats?match_id=7043 |
Legacy HTML-scraped stats for a match |
| GET | /api/player-stats?round_id=1155 |
Legacy HTML-scraped stats for a round |
| GET | /api/player-stats?afl_id=145 |
Legacy HTML-scraped stats for an individual player |
| GET | /api/player-stats?round_id=1155&afl_id=145 |
Legacy HTML-scraped stats for a player across a round |
| GET | /api/player-stats?match_id=7043&afl_id=145 |
One legacy HTML-scraped player/match row |
This compatibility route reads player_stats; it is not an authoritative CFS
read API. New API and scoring work must follow the player-stat storage
contract.
src/
βββ api/ # FastAPI routes
βββ data/ # Scraped + enriched player data
βββ db/ # SQLite init + import logic
βββ merge/ # Data reconciliation and matching logic
βββ scraper/ # Club, player, injuries and team lineup scrapers
βββ scripts/ # API key + admin CLI tools
βββ utils/ # Logging, nicknames, config
βββ cli.py # Unified CLI (scrape, enrich, import)
βββ main.py # api entry point
The system uses APScheduler for timing recurring and per-match scraping jobs.
Jobs are registered during startup via scheduler/start.py, including:
- Daily injury scrapes (11:00 AM AWST)
- Line-up scrapes (T-1 day 5pm, Thursday 5pm, 1h before each match)
- Player stats scraping (2 minutes before match start)
Admin UI /schedule shows all currently registered jobs.
The supported production scheduler command is:
python -m uvicorn scheduler.start:app --host 0.0.0.0 --port 8000Uvicorn owns the FastAPI process lifecycle. During application startup, scheduler/start.py runs database migrations, registers scheduler jobs once, and starts APScheduler in a managed background thread. During application shutdown, it stops APScheduler and waits for executors to finish before the interpreter exits.
Standalone module execution is also supported for local operation and diagnostics:
python -m scheduler.startThis standalone mode starts the same scheduler bootstrap path, then blocks the main process in APScheduler. Stop it with normal process interruption or termination, such as Ctrl+C or SIGTERM; the signal handler shuts down APScheduler and its executors before exit.
At startup, historical one-time date jobs whose scheduled run time is already in the past are intentionally marked skipped instead of being submitted to APScheduler. This prevents old fixture data from creating a backlog of obsolete scrapes while preserving future one-time jobs and recurring cron or interval jobs.
PRs welcome. Please raise an issue first if youβd like to propose a major change.
MIT
- Documentation index
- Administrator CSRF protection
- Admin manual scheduler triggers
- API key storage migration
- SQLite database migrations
- Scheduler registry and restart recovery
- Scrape run audit records
- AFL scraper source inventory and page contracts
Plain HTTP scraper calls must use the shared synchronous client in utils.http_utils instead of calling requests, httpx, or urllib directly. Rendered-page scraping remains a separate Playwright responsibility: future modules should keep Playwright for pages whose usable content depends on JavaScript rendering, including player-stat pages unless deterministic fixtures/tests prove plain HTTP contains all required data.
The default scraper HTTP policy is intentionally conservative:
- User-Agent:
AFL-api/1.0 (+https://github.com/JustPlausible/AFL-api). - Timeout: 5 seconds to connect and 20 seconds to read by default.
- Retries: at most 3 attempts for transient failures only.
- Retryable failures: HTTP 429, HTTP 500, 502, 503, 504, connection failures, timeouts, and clearly transient transport errors such as chunked transfer or content decoding failures.
- Non-retryable failures: permanent HTTP 4xx responses are not retried unnecessarily.
- Backoff: bounded exponential backoff starting at 0.5 seconds, capped at 8 seconds, with up to 0.25 seconds of jitter.
Retry-Afteris honoured where practical and capped at 10 seconds. - Rate limiting: one thread-safe bucket per host, per Python process, with a default 1 second spacing between same-host requests. This is not a distributed or cross-container rate limiter.
The supported operational knobs are the ScraperHttpPolicy values passed to ScraperHttpClient for explicit module-specific overrides. Overrides should be rare, local to the module that needs them, and documented in code with the reason; do not introduce competing request helpers for one-off timeout, retry, or rate-limit changes. Tests can inject sessions, clocks, sleep functions, jitter functions, and rate limiters, and can reset the process default client with reset_scraper_http_client() to avoid live network access or real sleeping.