Releases: simonmak-ascent/hkex-filing-scraper
Release list
v2.4.1
v2.4.0
What's Changed
- build: live MCP gateway on Vercel (runtime config + extraction deps) by @simonplmak-cloud in #34
- fix: make the live MCP endpoint's non-POST responses informative by @simonplmak-cloud in #35
- feat: harden the live MCP gateway for third-party AI agents by @simonplmak-cloud in #36
- docs: fix inaccuracies found in the public-docs review by @simonplmak-cloud in #37
- docs: rebuild the README as a landing page and surface the MCP story by @simonplmak-cloud in #38
- chore: hosting + repo metadata hygiene (Pages off, firewall rate limit, metadata) by @simonplmak-cloud in #39
- test: guard the MCP catalogs, live endpoint, CLI flags, and llms.txt by @simonplmak-cloud in #40
- docs: rebrand the hero card and add an MCP diagram by @simonplmak-cloud in #41
- feat: list the gateway in the MCP Registry and add a What's new page by @simonplmak-cloud in #42
Full Changelog: v2.3.0...v2.4.0
v2.3.0
What's Changed
- feat: add a live HKEx MCP gateway and move docs to Vercel by @simonplmak-cloud in #33
Full Changelog: v2.2.0...v2.3.0
v2.2.0
What's Changed
- ci: publish to PyPI via Trusted Publishing by @simonplmak-cloud in #28
- revert: remove the accidentally merged opportunity-pipeline work by @simonplmak-cloud in #29
- docs: PyPI ADR, SEO, and AI/LLM discoverability by @simonplmak-cloud in #30
- feat: add a read-only MCP server over the filing corpus by @simonplmak-cloud in #32
Full Changelog: v2.1.0...v2.2.0
v2.1.0
What's Changed
- docs: flatten sink docs, drop tier language, align license metadata by @simonplmak-cloud in #9
- build: harden the supply chain and attest releases by @simonplmak-cloud in #10
- ci: drop the docker ecosystem from Dependabot by @simonplmak-cloud in #15
- fix: add real HTTP retries, pacing, and a shared session by @simonplmak-cloud in #17
- ci: bump pinned actions to their current majors by @simonplmak-cloud in #18
- feat: add a cross-sink integrity hash and harden engine safeguards by @simonplmak-cloud in #19
- docs: fix the site to a dark Bloomberg-terminal theme by @simonplmak-cloud in #20
- fix: name Material's search toggle for assistive tech by @simonplmak-cloud in #21
- fix: underline prose links and make scroll wrappers focusable by @simonplmak-cloud in #22
- fix: re-apply scroll-wrapper focus after Material inserts it by @simonplmak-cloud in #23
- test: add quality gates, offline-by-default tests, and Hypothesis coverage by @simonplmak-cloud in #24
- test: replay the HKEx API contract, add a canary, and inject sink faults by @simonplmak-cloud in #25
- feat: add --verify cross-sink reconciliation by @simonplmak-cloud in #26
- chore: publish the package to PyPI by @simonplmak-cloud in #27
Full Changelog: v2.0.0...v2.1.0
v2.0.0 — Multi-backend sinks (one scraper, any database)
v2.0.0 — One scraper, any database
HKEx Filing Scraper now writes to eight database engines through one uniform sink contract. Point DATABASE_TARGET at the store you already run and get the complete filing corpus in its native shape.
Highlights
- Multi-backend sinks —
postgres,mysql,mariadb,sqlite,duckdb,mongodb,clickhouse,neo4j,surrealdb. - Uniform architecture — one
Sinkcontract, declaredSinkCapabilities, a lazy registry, and a shared relational engine with per-dialect SQL. No ORM; every driver is an optional extra. - Graph linking on SurrealDB and Neo4j —
(company)-[has_filing]->(filing)and(filing)-[references_filing]->(company). - N-way parity reporting, per-sink schema init, per-sink failure counters, and non-zero exit when any configured sink fails.
examples/docker-compose.yml— one command to bring up PostgreSQL, MySQL, MariaDB, MongoDB, ClickHouse, Neo4j, and SurrealDB locally.
⚠️ Breaking changes
DATABASE_TARGETis now an explicit, ordered, comma-separated list of sink ids. Theboth/dualaliases and the implicitsurrealdbdefault are removed; an unset or unknown value fails fast with the valid ids.DATABASE_TARGET=both→DATABASE_TARGET=surrealdb,postgres- unset → set it explicitly (e.g.
DATABASE_TARGET=postgres)
- Reads come from the first configured sink that supports them (no per-sink read branch, no
READ_SOURCE). - Every configured sink is required: any sink's write failure now exits non-zero.
Migration table: docs/upgrading.md.
Install
pip install "git+https://github.com/simonplmak-cloud/hkex-filing-scraper@v2.0.0"
# or install the wheel belowAdd drivers as needed: .[all], .[postgres], .[mysql], .[duckdb], .[mongodb], .[clickhouse], .[neo4j]. SQLite needs no extra.
Verification
- Unit + integration:
164 passedoffline;178 passed, 0 skippedacross all nine sinks live. - CI matrix: lint/format, secret scan, CodeQL, unit tests on Python 3.10–3.13, and live integration jobs for PostgreSQL, MySQL/MariaDB, SurrealDB, dual-write, and Tier 2 (MongoDB/ClickHouse/Neo4j), plus
pip install .and.[all]install checks.
Docs
Getting started · Backend support matrix · Architecture · De-risking register
Full changelog: CHANGELOG.md
v1.1.0 — PostgreSQL Dual-Store
PostgreSQL Dual-Store — v1.1.0
The scraper can now write the same corpus to SurrealDB, PostgreSQL, or both. Point it at a PostgreSQL database with one variable and query 25+ years of HKEx filings with your existing SQL tooling — no fork, no export, no re-scrape.
Highlights
- Dual-store sink.
DATABASE_TARGET=surrealdb(default) |postgres|both. Filing metadata, extracted text, structured tables (jsonb), ticker arrays (text[]), coverage rows, and graph edges all mirror into PostgreSQL. - Automatic schema. Tables and indexes are created on startup with idempotent DDL — no manual setup.
- Idempotent writes. Parameterised
INSERT ... ON CONFLICTupserts keyed on the sharedfiling_id, so re-runs update rather than duplicate. - Failure isolation. A PostgreSQL failure never aborts or rolls back the SurrealDB write (and vice versa); per-sink success/failure counts print every run.
- New CLI:
--database-targetand--parity-report. - Optional dependency.
pip install ".[postgres]"; missing driver or DSN degrades gracefully (and fails fast only when PostgreSQL is the sole sink).
Why it matters
SurrealDB remains the graph-native store, but most downstream teams — analytics, BI, dbt, compliance — live in PostgreSQL. The same battle-tested scrape/extract pipeline now feeds either, so you can adopt the tool inside an existing data stack or run both for parity.
Quick start
# PostgreSQL only
export DATABASE_TARGET=postgres
export POSTGRES_DSN=postgresql://user:password@localhost:5432/hkex
hkex-scraper
# Dual-write + parity check
hkex-scraper --database-target both --parity-reportUpgrade notes
Non-breaking. The default sink is still SurrealDB and all existing CLI flags are unchanged. If you do nothing, behaviour is identical. Dual-write is forward-only — existing SurrealDB data is not migrated to PostgreSQL automatically (re-run a range, and --backfill-docs for documents).
Documentation
Testing
67 tests pass (unit + live PostgreSQL integration). CI runs lint, format checks, Python 3.10–3.13, and a PostgreSQL service-container integration job.
v1.0.0 — Initial Release
First public release of the HKEx Filing Scraper — an open-source Python tool to scrape and ingest 25+ years of Hong Kong Stock Exchange regulatory filings into SurrealDB.
Key Features
- Scrapes the full history of HKEx filings from April 1999 to present via an undocumented JSON API
- Downloads and extracts full text from PDF, HTML, and Excel documents
- Extracts structured tables using Camelot and substitutes them inline
- Optional graph linking:
(company)-[has_filing]->(filing)and(filing)-[references_filing]->(company) - Resilient pipeline with parallel downloads, batch upserts, and recursive retry
- Configurable via environment variables with zero hardcoded dependencies
- Modern Python package with CLI, type hints, and unit tests
Installation
pip install "hkex-filing-scraper[all]"See the README for full documentation.