Skip to content

v1.0.12

Choose a tag to compare

@github-actions github-actions released this 03 Oct 19:22
· 18 commits to main since this release
be43f86

[1.0.12] — 2026-10-04

Project-review backlog — 36 defects + 6 engineering improvements (3 October)

Source-backed review of every workspace crate and the CLI; findings and statuses in docs/project-review-2026-10-03.md.

  • Security boundaries — HTML exports escape every < in embedded graph JSON (mixed-case </SCRIPT> could terminate the data element); MCP configuration files are excluded from raw LLM enrichment (literal credentials can no longer leave the machine through a remote backend); bolt+s:///bolt+ssc:// now ride rustls with scheme-preserving transport instead of silently downgrading to plaintext TCP; HTTP MCP enforces a whole-request deadline with per-read recomputation and exact loopback matching (127.attacker.example is not loopback); yt-dlp media downloads are resolved and vetted through the SSRF policy (-J --simulate) before any byte moves; URL classification parses host/path instead of substring-matching the whole URL.
  • Graph identity & integrity — node ids are case-preserving (Foo/foo stay distinct; matching stays case-folded), with duplicate-id disambiguation rewiring edges to the surviving definition and the extraction cache version advanced to v14 so pre-fix caches invalidate cleanly; automatic global tags check existing names before allocating (no more repo duplication/replacement); merge is atomic — edges reconcile, the database swaps with rollback, artifacts stage as .new and flip only after the swap, and the generation is stamped inside the transaction; global replace runs fully transactionally with relation reconciliation after commit.
  • Freshness & publication trust — every publication mints a generation stamp (_meta.graph_generation, generation.txt, _meta in graph.json, report footer) so database, JSON, and report of one build are matchable and snapshot caches key on it; unchanged Google Workspace shortcuts re-check their remote revision instead of trusting shortcut bytes; the merge gate fails on git-detection errors instead of silently skipping its checks.
  • Semantic backend honesty — token reservations span the whole request+response with guard-based release and per-backend max_tokens sharing the same constants (the advertised budget is reserved before the call, not audited after); content needing more than the chunk cap fails loudly (ASTRIA_LLM_MAX_CHUNKS) instead of truncating and caching as success; Bedrock stopReason is checked before accepting text; derived-text lookup validates against the extraction-hash family (config-hash lookups that never matched what extraction wrote are gone); malformed LLM replies never become cached empty successes.
  • Dependency advisories (release day) — quick-xml 0.37/0.39 → 0.41 and calamine 0.34 → 0.36 close RUSTSEC-2026-0194/0195 (quadratic attribute-check and unbounded namespace-declaration DoS in XML parsing, fresh in the advisory database when the release CI first ran); office-crate call sites moved to the 0.41 decoder API.
  • Everywhere else — XLSX decompression bounds are pre-checked via a bounded <dimension> zip scan before allocation; non-ASCII document titles can't panic filename creation; watch mode covers every supported file type and directory renames; database decode errors surface instead of silently truncating graphs; Bolt 3 RUN carries its third (extra) field per the official spec; risk traversal propagates errors; the viewer bundle is drift-checked in CI; a version-to-version A/B harness ships in scripts/bench/ab/.

Follow-up review — publication, coverage, health, cache (3–4 October)

  • Cache invalidation is part of graph publication — the generation advances inside the core transaction and immediately after every derived pass that commits a content change; unchanged runs reuse the previous generation; cluster-only republishes through the same artifact workflow as full pipelines.
  • Merge-gate coverage is commit identity, not timestamps — the pipeline records _meta.git_head at publication and verifySourceCommit compares it with the current HEAD while re-hashing every manifest file (the graph's own versioned scheme); query headers disclose source drift (modified / deleted / size-changed) separately from graph age, and relative manifest paths resolve against the project root so probes work from any cwd.
  • Health-score heuristics corrected — containment and co-occurrence edges no longer count as reachability (dead-code detection finds real candidates again); hubs must clear max(10, the graph's own 95th-percentile usage degree), and test-file hubs are reported, not scored.
  • Multi-project snapshot cache — the process-wide graph cache becomes a bounded LRU keyed by database path + generation (default 3 entries, ASTRIA_SNAPSHOT_CACHE_ENTRIES 1–16); alternating MCP projects share snapshots instead of evicting each other.
  • Retrieval evidence — the paired runner's report emits exact-symbol ranking (definition recall@5 + MRR) beside file recall and delivered tokens, plus a per-case definition-miss triage list; known failure modes are ranked as evaluation priorities.
  • Documentation — architecture claims derive from the registry (language counts drift-checked by scripts/check-docs-sync.mjs), the snapshot-cache description matches the LRU, and the review backlog carries per-finding status tables.

Platform & integration — team serving, cloud backends, CI artifacts (#B1–B5)

  • MCP over HTTP + multi-project serving — astria mcp --http serves one or many project graphs over MCP Streamable HTTP (JSON responses; GET /healthz for liveness) from a single process, complementing the stdio transport. Clients select a project with the x-astria-project header or ?project= query (unknown names 404 — no silent fallback); --graph is the default project and --projects name=path adds more. Bearer auth (--token / ASTRIA_MCP_TOKEN) is mandatory whenever the server binds a non-loopback host — unauthenticated remote serving is refused at startup. No async runtime: thread-per-connection, a fresh SQLite handle per request. Transport routing, auth, and project resolution are tested without sockets.
  • First-class Azure OpenAI, AWS Bedrock, and Kimi backends — --backend azure authenticates with Azure's api-key header against {endpoint}/openai/deployments/{deployment} (+ mandatory api-version query), --backend bedrock calls the Bedrock Converse API with real SigV4 request signing (self-contained HMAC-SHA256 implementation pinned to independently computed reference signatures; static keys and AWS_SESSION_TOKEN temporary credentials both work), and --backend kimi is the OpenAI-compatible surface pointed at Moonshot with Kimi's key variables and default model. Bedrock's Converse usage format (inputTokens/outputTokens) joins the usage counter's understood wire formats. All three are documented in env-vars.md with their ASTRIA_* and vendor env names.
  • Rich PR dashboard — astria prs now pulls CI state (statusCheckRollup), review decision, mergeability, author, and diff size in one gh pr list call, maps PR branches onto git worktree list locations, ranks the review queue by urgency (failing CI, requested changes, conflicts, graph blast radius, draft penalty), and prints merge-order risk on --conflicts. --triage gives compact per-PR lines; --json emits the ranked queue with every signal.
  • Git merge driver for the graph file — astria merge-driver install wires a three-way union-merge driver into .gitattributes + merge.astria.* git config so parallel branches that both commit .astria/graph.json merge instead of conflicting: additions from both sides survive, deletions are respected, fields resolve 3-way (unchanged side takes the changed side), and communities (derived data) resolve to whichever side moved. .astria/graph_report.md gets git's built-in union driver. uninstall removes the wiring; run is the git-invoked entry point.
  • Docker distribution — a multi-stage Dockerfile in the repo root (Rust+Node builder → slim Node runtime) ships the full CLI with no toolchain inside; analyze a mounted repo or serve the HTTP MCP on exposed port 8620; --build-arg NAPI_FEATURES=--no-default-features produces a smaller image without the embedding runtime.
  • Hosted-tier OSS surface — astria merge-gate is a CI check that fails on missing/stale graphs (publish timestamp vs wall clock and last commit), health-score floors, and diff blast-radius ceilings (--json for pipelines); astria digest renders a deterministic markdown engineering brief (overview, health, hub concentration, largest communities, LLM spend) for stdout, --out, or cron. These are the same primitives the hosted tier (app.graphify.com) operates for teams.
  • Deep-clean uninstall — astria uninstall --purge removes every platform install plus the artifacts plain uninstall deliberately leaves: git hooks, merge-driver wiring, the project .astria/ data directory, and the ~/.astria global store. Explicit-flag consent, no prompt, CI-safe.

Query & graph features (#C1, #C3, #C4)

  • CJK query segmentation — the retrieval tokenizer now segments Chinese/Japanese/Korean runs with jieba (dictionary + HMM, built once per process), so "用户登录怎么处理" matches the labels that say 用户 and 登录 instead of arriving as one unmatchable character run. Non-CJK tokenization is byte-for-byte unchanged; nearest_labels suggestions now share the tokenizer (with stopword filtering), so did-you-mean works for CJK too.
  • Clustering controls — astria cluster-only --resolution <0.0–1.0> requires a minimum share of a node's neighbors to agree on the winning community before the node joins it (default 0.0 = classic propagation; higher values → more, smaller communities, same direction as Louvain's resolution), and --exclude-hubs holds high-degree hub nodes (degree ≥ max(12, 4× mean)) out of label propagation entirely so they cannot glue communities together — hubs are attached to their strongest community afterwards, so every node still lands in one. Defaults reproduce the previous behavior exactly; hub exclusion also pre-seeds the fragment-merge pass so hubs are never absorbed as fragments.
  • Cost report artifact — every run/update writes .astria/cost.json: this run's measured LLM spend (from the pipeline_runs row), lifetime totals across completed runs, backend/model identity, and — when ASTRIA_COST_INPUT_PER_MTOK/ASTRIA_COST_OUTPUT_PER_MTOK are set — a dollar estimate clearly labeled as operator-supplied rates, not vendor billing. Best-effort write: a report failure warns and never fails the build.

Ingestion & language parity expansion (#A1–A6)

  • 17 new registered languages (25 → 42): SQL (tables/views/functions + CREATE TRIGGER via tree-sitter-sequel), Julia, R, Fortran, Solidity, Groovy, Luau, Objective-C, OCaml + OCaml Interface, Common Lisp, BYOND DreamMaker (.dm), Astro (.astro); Vue + Svelte extract embedded <script> TS/JS via the JS/TS grammars (langs::embedded); VB.NET + Pascal/Delphi use regex declaration extraction (no upstream grammar crate). Every language ships its own extraction test.
  • Office documents: new astria-office crate — .docx (headings/lists/tables via word/document.xml) and .xlsx (per-sheet markdown tables, 500×30 cap) become document nodes; classified as Document.
  • Google Workspace: new astria-gws crate — .gdoc/.gsheet/.gslides shortcuts resolve a Drive file id and export via Drive API v3 (Docs→text, Sheets→CSV→markdown table, Slides→text); auth via ASTRIA_GDRIVE_ACCESS_TOKEN or gcloud ADC refresh; missing credentials skip with a notice, never fail the build (whisper-route semantics, results uncached).
  • Media URLs: astria add classifies YouTube/Vimeo/Dailymotion/Twitch links and direct media URLs as UrlKind::Media and downloads via external yt-dlp (16 kHz mono WAV with ffmpeg, raw bestaudio without) into the project for whisper transcription; yt-dlp missing degrades to a stub node with the install hint.
  • Doc formats: .qmd rides the markdown path, .html/.htm are tag-stripped (scripts/styles dropped, entities decoded), .yaml/.yml chunk as text — all classified Document.
  • Rationale comments generalize across comment styles (#, --, ;, ', !) — hash-comment languages (Ruby, Shell, Elixir, Lua, and the new Julia/R/Groovy/SQL/Luau/Common Lisp/VB.NET/Fortran) now emit rationale nodes.
  • Grammars are compile-time optional as before (new lang-* features incl. lang-ocaml-interface); EXTRACTION_HASH_VERSION unchanged — none of these types were previously ingested, so no cache invalidation is needed.

Video/audio ingestion — Whisper transcription (#82)

  • Media files join the graph — mp4/mov/webm/mkv/avi video and mp3/wav/m4a/flac/ogg/opus/aac/wma audio files are transcribed during run/update and enter the graph as transcript documents through the same markdown pipeline PDFs use. External-binary mode, like add --postgres requiring psql: transcription runs in whisper.cpp's whisper-cli, video files also need ffmpeg on PATH to demux the audio track (audio-only repos transcribe without it). Nothing is vendored, no API key is involved, and the napi binaries stay small.
  • Missing tooling degrades to a notice, not a failure — without whisper-cli, a model, or (for video) ffmpeg, media files are skipped with one actionable notice per cause per run while the rest of the graph builds normally; failed attempts are never cached, so installing the tooling is picked up on the next run even for unchanged files. Model resolution order: $ASTRIA_WHISPER_MODEL, then the first *.bin in <project>/.astria/models/, then ~/.astria/models/.
  • Plumbing — FileType gains audio (detect classifies the new extensions; the extraction cache hash bumps to v13, forcing one clean re-extraction on upgrade), and a std-only astria-audio crate joins the workspace between astria-ingest and astria-pdf.

P0 hardening — ignore-file failures surface, embed capability is disclosed, one budget default

  • A .astriaignore that cannot be loaded now fails the run. File discovery previously swallowed add_ignore errors, so an unreadable or misplaced .astriaignore silently built the graph without the user's exclusions. Malformed individual patterns still follow gitignore's lenient semantics (an unclosed [ is a literal, matching git's own behavior). Regression-tested for both the exclusion behavior and the loud failure.
  • Embed capability is disclosed, not discovered. graph_stats (CLI --json and the MCP tool) now reports embeddingsSupported / an embeddings: available | not supported in this build suffix, and stats prints the same in human output, so a darwin-x64 user (no prebuilt ONNX binaries) learns the --embed gap before hitting the runtime error. Documented in the README quick start.
  • One default query budget. The CLI (query, map, and the injected astria_query/astria_map tool presets) and the MCP server now share a single 2,000-token default (DEFAULT_QUERY_BUDGET in the MCP server, src/defaults.ts in the CLI, guarded by a structure test); the injected astria_query preset previously fell back to 3,000 while everywhere else used 2,000.

P1 hardening — module splits, feature-gated grammars, CI gates, integration tests

  • astria-query and astria-semantic split into domain modules — the 4,300-line query god file becomes store/scoring/render modules behind an unchanged public API; the semantic crate separates shared types/prompt/chunking/http from one module per backend (Claude, OpenAI-compatible, Gemini, Jev). Pure extraction, zero behavior change; every test passes untouched.
  • Language walker branches move into langs/ — Rust doc-comment/test-attribute handling, Python overload semantics, the JavaScript name/binding/doc walker, and PHP route-label synthesis now live beside their language configs; walkers.rs keeps only language-neutral machinery.
  • Tree-sitter grammars are feature-gated — each language gets a lang-* cargo feature (lang-all remains the default), and the engine skips files of languages compiled out with a one-time warning instead of failing. --no-default-features --features "lang-python,lang-javascript" now produces a slim extraction build. The three ad-hoc grammar pins move into [workspace.dependencies].
  • astria-napi stops being a monolith — the seven export formats move to a new astria-export crate, the health/risk diagnostics join astria-analyze, and the Neo4j push joins astria-bolt. All moved code was napi-free; astria_napi::export_wiki::… paths re-export unchanged.
  • CI gates — a new MSRV workflow compiles the workspace on the declared Rust 1.88 floor; a Coverage workflow publishes per-crate llvm-cov numbers (informational until a baseline exists); CONTRIBUTING documents the versioning policy (features minor, fixes patch — no cargo-semver-checks because no crate is published).
  • Integration tests — the MCP stdio loop is now a transport-independent serve_loop with framing tests (response-per-line, malformed-line skip, stop-at-EOF), and astria-bolt gains #[ignore]d live-Neo4j tests (handshake, RUN/PULL round trip, auth-failure-is-an-error) runnable against Docker.

What's Changed

  • chore(homebrew): sync packaging formula to 1.0.11 by @erictong0602 in #97
  • Release 1.0.12 — the trust release (review backlog + platform expansion) by @erictong0602 in #101
  • fix(test): budget/reservation test race (release-CI flake in v1.0.12 run) by @erictong0602 in #113
  • fix(release): skip platform publishes without binaries (musl ENEEDAUTH abort) by @erictong0602 in #114
  • fix(release): tolerate missing trusted-publisher bindings for never-published packages by @erictong0602 in #115

Full Changelog: v1.0.11...v1.0.12