Repository navigation
v1.0.12
[1.0.12] — 2026-10-04
Project-review backlog — 36 defects + 6 engineering improvements (3 October)
Source-backed review of every workspace crate and the CLI; findings and statuses in docs/project-review-2026-10-03.md.
- Security boundaries — HTML exports escape every
<in embedded graph JSON (mixed-case</SCRIPT>could terminate the data element); MCP configuration files are excluded from raw LLM enrichment (literal credentials can no longer leave the machine through a remote backend);bolt+s:///bolt+ssc://now ride rustls with scheme-preserving transport instead of silently downgrading to plaintext TCP; HTTP MCP enforces a whole-request deadline with per-read recomputation and exact loopback matching (127.attacker.exampleis not loopback);yt-dlpmedia downloads are resolved and vetted through the SSRF policy (-J --simulate) before any byte moves; URL classification parses host/path instead of substring-matching the whole URL. - Graph identity & integrity — node ids are case-preserving (
Foo/foostay distinct; matching stays case-folded), with duplicate-id disambiguation rewiring edges to the surviving definition and the extraction cache version advanced to v14 so pre-fix caches invalidate cleanly; automatic global tags check existing names before allocating (no more repo duplication/replacement); merge is atomic — edges reconcile, the database swaps with rollback, artifacts stage as.newand flip only after the swap, and the generation is stamped inside the transaction; global replace runs fully transactionally with relation reconciliation after commit. - Freshness & publication trust — every publication mints a generation stamp (
_meta.graph_generation,generation.txt,_metaingraph.json, report footer) so database, JSON, and report of one build are matchable and snapshot caches key on it; unchanged Google Workspace shortcuts re-check their remote revision instead of trusting shortcut bytes; the merge gate fails on git-detection errors instead of silently skipping its checks. - Semantic backend honesty — token reservations span the whole request+response with guard-based release and per-backend
max_tokenssharing the same constants (the advertised budget is reserved before the call, not audited after); content needing more than the chunk cap fails loudly (ASTRIA_LLM_MAX_CHUNKS) instead of truncating and caching as success; BedrockstopReasonis checked before accepting text; derived-text lookup validates against the extraction-hash family (config-hash lookups that never matched what extraction wrote are gone); malformed LLM replies never become cached empty successes. - Dependency advisories (release day) —
quick-xml0.37/0.39 → 0.41 andcalamine0.34 → 0.36 close RUSTSEC-2026-0194/0195 (quadratic attribute-check and unbounded namespace-declaration DoS in XML parsing, fresh in the advisory database when the release CI first ran); office-crate call sites moved to the 0.41 decoder API. - Everywhere else — XLSX decompression bounds are pre-checked via a bounded
<dimension>zip scan before allocation; non-ASCII document titles can't panic filename creation; watch mode covers every supported file type and directory renames; database decode errors surface instead of silently truncating graphs; Bolt 3 RUN carries its third (extra) field per the official spec; risk traversal propagates errors; the viewer bundle is drift-checked in CI; a version-to-version A/B harness ships inscripts/bench/ab/.
Follow-up review — publication, coverage, health, cache (3–4 October)
- Cache invalidation is part of graph publication — the generation advances inside the core transaction and immediately after every derived pass that commits a content change; unchanged runs reuse the previous generation;
cluster-onlyrepublishes through the same artifact workflow as full pipelines. - Merge-gate coverage is commit identity, not timestamps — the pipeline records
_meta.git_headat publication andverifySourceCommitcompares it with the current HEAD while re-hashing every manifest file (the graph's own versioned scheme); query headers disclose source drift (modified / deleted / size-changed) separately from graph age, and relative manifest paths resolve against the project root so probes work from any cwd. - Health-score heuristics corrected — containment and co-occurrence edges no longer count as reachability (dead-code detection finds real candidates again); hubs must clear max(10, the graph's own 95th-percentile usage degree), and test-file hubs are reported, not scored.
- Multi-project snapshot cache — the process-wide graph cache becomes a bounded LRU keyed by database path + generation (default 3 entries,
ASTRIA_SNAPSHOT_CACHE_ENTRIES1–16); alternating MCP projects share snapshots instead of evicting each other. - Retrieval evidence — the paired runner's report emits exact-symbol ranking (definition recall@5 + MRR) beside file recall and delivered tokens, plus a per-case definition-miss triage list; known failure modes are ranked as evaluation priorities.
- Documentation — architecture claims derive from the registry (language counts drift-checked by
scripts/check-docs-sync.mjs), the snapshot-cache description matches the LRU, and the review backlog carries per-finding status tables.
Platform & integration — team serving, cloud backends, CI artifacts (#B1–B5)
- MCP over HTTP + multi-project serving —
astria mcp --httpserves one or many project graphs over MCP Streamable HTTP (JSON responses;GET /healthzfor liveness) from a single process, complementing the stdio transport. Clients select a project with thex-astria-projectheader or?project=query (unknown names 404 — no silent fallback);--graphis the default project and--projects name=pathadds more. Bearer auth (--token/ASTRIA_MCP_TOKEN) is mandatory whenever the server binds a non-loopback host — unauthenticated remote serving is refused at startup. No async runtime: thread-per-connection, a fresh SQLite handle per request. Transport routing, auth, and project resolution are tested without sockets. - First-class Azure OpenAI, AWS Bedrock, and Kimi backends —
--backend azureauthenticates with Azure'sapi-keyheader against{endpoint}/openai/deployments/{deployment}(+ mandatoryapi-versionquery),--backend bedrockcalls the Bedrock Converse API with real SigV4 request signing (self-contained HMAC-SHA256 implementation pinned to independently computed reference signatures; static keys andAWS_SESSION_TOKENtemporary credentials both work), and--backend kimiis the OpenAI-compatible surface pointed at Moonshot with Kimi's key variables and default model. Bedrock's Converse usage format (inputTokens/outputTokens) joins the usage counter's understood wire formats. All three are documented in env-vars.md with their ASTRIA_* and vendor env names. - Rich PR dashboard —
astria prsnow pulls CI state (statusCheckRollup), review decision, mergeability, author, and diff size in onegh pr listcall, maps PR branches ontogit worktree listlocations, ranks the review queue by urgency (failing CI, requested changes, conflicts, graph blast radius, draft penalty), and prints merge-order risk on--conflicts.--triagegives compact per-PR lines;--jsonemits the ranked queue with every signal. - Git merge driver for the graph file —
astria merge-driver installwires a three-way union-merge driver into.gitattributes+merge.astria.*git config so parallel branches that both commit.astria/graph.jsonmerge instead of conflicting: additions from both sides survive, deletions are respected, fields resolve 3-way (unchanged side takes the changed side), and communities (derived data) resolve to whichever side moved..astria/graph_report.mdgets git's built-inuniondriver.uninstallremoves the wiring;runis the git-invoked entry point. - Docker distribution — a multi-stage Dockerfile in the repo root (Rust+Node builder → slim Node runtime) ships the full CLI with no toolchain inside; analyze a mounted repo or serve the HTTP MCP on exposed port 8620;
--build-arg NAPI_FEATURES=--no-default-featuresproduces a smaller image without the embedding runtime. - Hosted-tier OSS surface —
astria merge-gateis a CI check that fails on missing/stale graphs (publish timestamp vs wall clock and last commit), health-score floors, and diff blast-radius ceilings (--jsonfor pipelines);astria digestrenders a deterministic markdown engineering brief (overview, health, hub concentration, largest communities, LLM spend) for stdout,--out, or cron. These are the same primitives the hosted tier (app.graphify.com) operates for teams. - Deep-clean uninstall —
astria uninstall --purgeremoves every platform install plus the artifacts plain uninstall deliberately leaves: git hooks, merge-driver wiring, the project.astria/data directory, and the~/.astriaglobal store. Explicit-flag consent, no prompt, CI-safe.
Query & graph features (#C1, #C3, #C4)
- CJK query segmentation — the retrieval tokenizer now segments Chinese/Japanese/Korean runs with jieba (dictionary + HMM, built once per process), so "用户登录怎么处理" matches the labels that say 用户 and 登录 instead of arriving as one unmatchable character run. Non-CJK tokenization is byte-for-byte unchanged;
nearest_labelssuggestions now share the tokenizer (with stopword filtering), so did-you-mean works for CJK too. - Clustering controls —
astria cluster-only --resolution <0.0–1.0>requires a minimum share of a node's neighbors to agree on the winning community before the node joins it (default 0.0 = classic propagation; higher values → more, smaller communities, same direction as Louvain's resolution), and--exclude-hubsholds high-degree hub nodes (degree ≥ max(12, 4× mean)) out of label propagation entirely so they cannot glue communities together — hubs are attached to their strongest community afterwards, so every node still lands in one. Defaults reproduce the previous behavior exactly; hub exclusion also pre-seeds the fragment-merge pass so hubs are never absorbed as fragments. - Cost report artifact — every
run/updatewrites.astria/cost.json: this run's measured LLM spend (from the pipeline_runs row), lifetime totals across completed runs, backend/model identity, and — whenASTRIA_COST_INPUT_PER_MTOK/ASTRIA_COST_OUTPUT_PER_MTOKare set — a dollar estimate clearly labeled as operator-supplied rates, not vendor billing. Best-effort write: a report failure warns and never fails the build.
Ingestion & language parity expansion (#A1–A6)
- 17 new registered languages (25 → 42): SQL (tables/views/functions +
CREATE TRIGGERviatree-sitter-sequel), Julia, R, Fortran, Solidity, Groovy, Luau, Objective-C, OCaml + OCaml Interface, Common Lisp, BYOND DreamMaker (.dm), Astro (.astro); Vue + Svelte extract embedded<script>TS/JS via the JS/TS grammars (langs::embedded); VB.NET + Pascal/Delphi use regex declaration extraction (no upstream grammar crate). Every language ships its own extraction test. - Office documents: new
astria-officecrate —.docx(headings/lists/tables viaword/document.xml) and.xlsx(per-sheet markdown tables, 500×30 cap) become document nodes; classified asDocument. - Google Workspace: new
astria-gwscrate —.gdoc/.gsheet/.gslidesshortcuts resolve a Drive file id and export via Drive API v3 (Docs→text, Sheets→CSV→markdown table, Slides→text); auth viaASTRIA_GDRIVE_ACCESS_TOKENor gcloud ADC refresh; missing credentials skip with a notice, never fail the build (whisper-route semantics, results uncached). - Media URLs:
astria addclassifies YouTube/Vimeo/Dailymotion/Twitch links and direct media URLs asUrlKind::Mediaand downloads via externalyt-dlp(16 kHz mono WAV with ffmpeg, raw bestaudio without) into the project for whisper transcription; yt-dlp missing degrades to a stub node with the install hint. - Doc formats:
.qmdrides the markdown path,.html/.htmare tag-stripped (scripts/styles dropped, entities decoded),.yaml/.ymlchunk as text — all classifiedDocument. - Rationale comments generalize across comment styles (
#,--,;,',!) — hash-comment languages (Ruby, Shell, Elixir, Lua, and the new Julia/R/Groovy/SQL/Luau/Common Lisp/VB.NET/Fortran) now emitrationalenodes. - Grammars are compile-time optional as before (new
lang-*features incl.lang-ocaml-interface);EXTRACTION_HASH_VERSIONunchanged — none of these types were previously ingested, so no cache invalidation is needed.
Video/audio ingestion — Whisper transcription (#82)
- Media files join the graph —
mp4/mov/webm/mkv/avivideo andmp3/wav/m4a/flac/ogg/opus/aac/wmaaudio files are transcribed duringrun/updateand enter the graph as transcript documents through the same markdown pipeline PDFs use. External-binary mode, likeadd --postgresrequiringpsql: transcription runs in whisper.cpp'swhisper-cli, video files also needffmpegon PATH to demux the audio track (audio-only repos transcribe without it). Nothing is vendored, no API key is involved, and the napi binaries stay small. - Missing tooling degrades to a notice, not a failure — without
whisper-cli, a model, or (for video)ffmpeg, media files are skipped with one actionable notice per cause per run while the rest of the graph builds normally; failed attempts are never cached, so installing the tooling is picked up on the next run even for unchanged files. Model resolution order:$ASTRIA_WHISPER_MODEL, then the first*.binin<project>/.astria/models/, then~/.astria/models/. - Plumbing —
FileTypegainsaudio(detect classifies the new extensions; the extraction cache hash bumps to v13, forcing one clean re-extraction on upgrade), and a std-onlyastria-audiocrate joins the workspace betweenastria-ingestandastria-pdf.
P0 hardening — ignore-file failures surface, embed capability is disclosed, one budget default
- A
.astriaignorethat cannot be loaded now fails the run. File discovery previously swallowedadd_ignoreerrors, so an unreadable or misplaced.astriaignoresilently built the graph without the user's exclusions. Malformed individual patterns still follow gitignore's lenient semantics (an unclosed[is a literal, matching git's own behavior). Regression-tested for both the exclusion behavior and the loud failure. - Embed capability is disclosed, not discovered.
graph_stats(CLI--jsonand the MCP tool) now reportsembeddingsSupported/ anembeddings: available | not supported in this buildsuffix, andstatsprints the same in human output, so a darwin-x64 user (no prebuilt ONNX binaries) learns the--embedgap before hitting the runtime error. Documented in the README quick start. - One default query budget. The CLI (
query,map, and the injectedastria_query/astria_maptool presets) and the MCP server now share a single 2,000-token default (DEFAULT_QUERY_BUDGETin the MCP server,src/defaults.tsin the CLI, guarded by a structure test); the injectedastria_querypreset previously fell back to 3,000 while everywhere else used 2,000.
P1 hardening — module splits, feature-gated grammars, CI gates, integration tests
- astria-query and astria-semantic split into domain modules — the 4,300-line query god file becomes store/scoring/render modules behind an unchanged public API; the semantic crate separates shared types/prompt/chunking/http from one module per backend (Claude, OpenAI-compatible, Gemini, Jev). Pure extraction, zero behavior change; every test passes untouched.
- Language walker branches move into
langs/— Rust doc-comment/test-attribute handling, Python overload semantics, the JavaScript name/binding/doc walker, and PHP route-label synthesis now live beside their language configs;walkers.rskeeps only language-neutral machinery. - Tree-sitter grammars are feature-gated — each language gets a
lang-*cargo feature (lang-allremains the default), and the engine skips files of languages compiled out with a one-time warning instead of failing.--no-default-features --features "lang-python,lang-javascript"now produces a slim extraction build. The three ad-hoc grammar pins move into[workspace.dependencies]. - astria-napi stops being a monolith — the seven export formats move to a new
astria-exportcrate, the health/risk diagnostics joinastria-analyze, and the Neo4j push joinsastria-bolt. All moved code was napi-free;astria_napi::export_wiki::…paths re-export unchanged. - CI gates — a new MSRV workflow compiles the workspace on the declared Rust 1.88 floor; a Coverage workflow publishes per-crate llvm-cov numbers (informational until a baseline exists); CONTRIBUTING documents the versioning policy (features minor, fixes patch — no
cargo-semver-checksbecause no crate is published). - Integration tests — the MCP stdio loop is now a transport-independent
serve_loopwith framing tests (response-per-line, malformed-line skip, stop-at-EOF), andastria-boltgains#[ignore]d live-Neo4j tests (handshake, RUN/PULL round trip, auth-failure-is-an-error) runnable against Docker.
What's Changed
- chore(homebrew): sync packaging formula to 1.0.11 by @erictong0602 in #97
- Release 1.0.12 — the trust release (review backlog + platform expansion) by @erictong0602 in #101
- fix(test): budget/reservation test race (release-CI flake in v1.0.12 run) by @erictong0602 in #113
- fix(release): skip platform publishes without binaries (musl ENEEDAUTH abort) by @erictong0602 in #114
- fix(release): tolerate missing trusted-publisher bindings for never-published packages by @erictong0602 in #115
Full Changelog: v1.0.11...v1.0.12