lci-codegraph walks a source checkout and, from one tree-sitter parse per file, produces
semantic chunks and a structural call graph with cross-file resolution. It is a pure extractor — no
database, no cluster dependency — the caller decides where the graph and (unembedded) chunks go. The
one deliberate exception: when semantic embeddings are configured, the crate
itself makes a blocking HTTP call to an OpenAI-compatible /embeddings endpoint, because embedding is
inherently a round-trip to a model — there is no in-process model this crate could run instead. That
call is off by default and stays off until an operator sets OPENAI_BASE_URL.
[dependencies]
lci-codegraph = "0.1"MSRV: Rust 1.88. The crate is edition 2024 (floor 1.85), but lopdf — pulled in through
pdf-extract, and floored at a version that patches
RUSTSEC-2026-0187 — uses let-chains, stabilised
in 1.88. CI checks this floor on every build.
use std::path::Path;
use lci_codegraph::{WalkOptions, walk_checkout};
fn main() -> anyhow::Result<()> {
let root = Path::new(".");
let options = WalkOptions::builder().build_graph(true).build();
let output = walk_checkout(root, &options)?;
// Chunks: ready to hand to an embedding model.
for chunk in &output.chunks {
println!(
"{} [{}] {:?} L{}-{}",
chunk.file_path, chunk.chunk_type, chunk.symbol_name, chunk.start_line, chunk.end_line
);
}
// Graph: definitions and the calls between them, resolved across files.
for edge in &output.graph.edges {
println!("{} --{}--> {}", edge.source, edge.relation, edge.target);
}
Ok(())
}Both come out of the same walk — the tree is parsed once per file and fed to the chunker and the
graph builder together (build_graph: false, the default, skips graph extraction and returns an
empty Graph, so a caller that only wants chunks pays nothing extra).
The crate indexes bytes, not files. A RawInput
is just a logical path plus content; nothing below it cares where those bytes came from. walk_checkout
above is a convenience built on a filesystem reader (FsSource), but it is one reader among several
possible ones — a host that already has content in memory (a git object store, a tarball stream, an
HTTP fetch, open editor buffers, rows pulled from a database) never has to materialise a checkout on
disk just to use this crate. It can build RawInputs directly and hand them to the indexing core in
input.
For a caller that has every input up front, index_inputs
takes an iterator of RawInput and an IndexOptions:
use lci_codegraph::{IndexOptions, RawInput, index_inputs};
fn main() {
let inputs = vec![
RawInput::text("src/main.rs", "fn main() { helper(); }\n"),
RawInput::text("src/helper.rs", "pub fn helper() {}\n"),
// No filename extension to detect from — a DB row, a gist, a paste — so say the language
// explicitly instead.
RawInput::text("gist:a1b2c3", "def greet():\n pass\n").with_language("python"),
];
let options = IndexOptions::builder().build_graph(true).build();
let output = index_inputs(inputs, &options);
for chunk in &output.chunks {
println!("{} [{}]", chunk.file_path, chunk.chunk_type);
}
// main() --calls--> helper(), resolved across the two files above.
for edge in &output.graph.edges {
println!("{} --{}--> {}", edge.source, edge.relation, edge.target);
}
}For a caller that does not have every input up front — reading tarball entries one at a time,
say — drive the streaming Indexer
directly instead: push each input as it arrives, then finish once:
use lci_codegraph::{IndexOptions, Indexer, RawInput};
fn index_streaming(sources: impl Iterator<Item = (String, Vec<u8>)>) -> lci_codegraph::IndexOutput {
let options = IndexOptions::builder().build_graph(true).build();
let mut indexer = Indexer::new(options);
for (path, bytes) in sources {
indexer.push(RawInput::new(path, bytes));
}
indexer.finish()
}FsSource itself is public and directly usable as an Iterator<Item = RawInput>, not just through
walk_checkout — so a caller can mix filesystem inputs with in-memory ones in a single Indexer:
let mut indexer = Indexer::new(IndexOptions::from(&walk_options));
for input in FsSource::new(root, &walk_options)? {
indexer.push(input);
}
indexer.push(RawInput::text("scratch/notes.md", "# TODO\n"));
let output = indexer.finish();What the indexer does and does not do: Indexer::push applies the byte cap (MAX_INPUT_BYTES at the
reader level, then the tighter per-kind cap once it knows source vs PDF), the UTF-8/binary content
sniff, language detection (RawInput::language, falling back to lang::from_path), and the PDF size
bound. It does not filter paths — deciding which inputs are even worth handing over is the
reader's job, because only the reader can avoid paying to produce an input that will just be
discarded. FsSource is where that filtering happens for a filesystem checkout, composing the repo's
own .gitignore with the operator glob layer (see "The ignore model" below).
A Chunk is one embeddable
unit of source: file_path, language, chunk_type ("function", "class", "impl", "window",
…), an optional symbol_name, a 0-based start_line/end_line line range, and the content text.
Structured languages get tree-sitter-extracted items (functions, structs, classes, impls, methods);
everything else — or a file too large to parse, or a language with no grammar — falls back to
fixed-size overlapping line windows.
Two more fields exist for the semantic embeddings step and stay None unless
it runs: embedding: Option<Vec<f32>>, the vector returned by the embeddings endpoint, and
embed_input: Option<String>, the exact text that was (or would be) sent for it — the chunk's
content plus a graph-aware header, truncated to the configured cap. Both are
#[serde(skip_serializing_if = "Option::is_none")], so JSON output is unchanged when embedding is
off.
A Graph is a flat list of
[GraphNode]s (node_id, label, source_file, start_line) and [GraphEdge]s (source,
target, relation). Three relations are emitted:
contains— a file → its top-level definitions, and a container definition (mod/struct/trait/enum/class/…) → the definitions nested inside it.method— a type container (impl/trait/struct/enum/class/interface) → a callable it defines directly. This is a specialisation ofcontainskept as its own relation.calls— a caller definition → a callee definition, resolved across files: a call recorded in file A can resolve to a definition in file B.
Two conventions to know when reading node ids and labels:
- Line numbers in the graph are 1-based (
start_line: 1is the file's first line) — unlike aChunk's 0-basedstart_line/end_line. - Callable labels carry a
()suffix: a function namedaddgets the labeladd(); a non-callable definition (a struct, a class, animplblock) keeps its bare name.
Here is a real excerpt of the committed golden (tests/golden/sample-repo.graph.json) — a Rust
fixture with a main.rs that calls into math.rs:
{
"nodes": [
{ "node_id": "src/main.rs", "label": "main.rs", "source_file": "src/main.rs", "start_line": 1 },
{ "node_id": "src/main.rs#3:main", "label": "main()", "source_file": "src/main.rs", "start_line": 3 },
{ "node_id": "src/math.rs#2:add", "label": "add()", "source_file": "src/math.rs", "start_line": 2 },
{ "node_id": "src/math.rs#7:print_result", "label": "print_result()", "source_file": "src/math.rs", "start_line": 7 }
],
"edges": [
{ "source": "src/main.rs", "target": "src/main.rs#3:main", "relation": "contains" },
{ "source": "src/main.rs#3:main", "target": "src/math.rs#2:add", "relation": "calls" },
{ "source": "src/main.rs#3:main", "target": "src/math.rs#7:print_result", "relation": "calls" }
]
}main() in src/main.rs calling add() and print_result() in src/math.rs are the two
cross-file calls edges — the part a per-file extractor cannot produce on its own.
The same fixture's shapes.rs, drawn as a graph:
graph LR
file["src/shapes.rs"] -- contains --> Circle["Circle"]
file -- contains --> implCircle["impl (Circle)"]
implCircle -- method --> new["new()"]
implCircle -- method --> area["area()"]
file -- contains --> Square["Square"]
file -- contains --> implSquare["impl (Square)"]
implSquare -- method --> describe["describe()"]
describe -- calls --> area
file -- contains --> muc["make_unit_circle()"]
muc -- calls --> new
flowchart LR
A["reader (e.g. FsSource)<br/>produces RawInput"] --> B["Indexer::push:<br/>parse (one tree-sitter pass per input)"]
B --> C["chunk: tree-sitter items,<br/>windowed fallback"]
B --> D["extract_file:<br/>per-file defs + call sites"]
D --> E["resolve:<br/>cross-file name resolution"]
E --> F["canonical Graph<br/>(sorted + deduped)"]
C --> G[Vec of Chunk]
| Language | Extensions | Chunking | Graph extractor |
|---|---|---|---|
| Rust | .rs |
tree-sitter | Native node-kind extractor (interesting_node + call_expression navigation) — kept separate so the committed golden stays byte-stable |
| Python | .py |
tree-sitter | The grammar's bundled tags.scm (tree-sitter-python::TAGS_QUERY) |
| JavaScript | .js, .jsx, .mjs, .cjs |
tree-sitter | The grammar's bundled tags.scm (tree-sitter-javascript::TAGS_QUERY) |
| TypeScript | .ts |
tree-sitter | JavaScript's tags.scm composed with TypeScript's tags.scm — the TS query alone only covers TS-specific constructs (signatures, interfaces, modules), not concrete class/function/method/call |
| TSX | .tsx |
tree-sitter (JSX-aware grammar) | Same composed JS+TS tags.scm, run against the dedicated TSX grammar (the plain TypeScript grammar cannot parse JSX) |
| Java | .java |
tree-sitter | The grammar's bundled tags.scm (tree-sitter-java::TAGS_QUERY) |
For every one of these, chunking and graph extraction share the same parse of the file
(WalkOptions::build_graph).
A few more extensions are recognised as a language label with no tree-sitter grammar in this crate
— .go, .c/.h, .cpp/.cc/.cxx/.hpp, and a generic text bucket for .md/.txt/.toml/
.yaml/.yml/.json. These are chunked via the windowed-line fallback only (no structured chunks,
no graph). Any file whose extension is not recognised at all is skipped entirely — not chunked, not
graphed.
Adding a language means implementing one LanguageSupport in src/lang/<language>.rs and adding it
to the registry; see docs/architecture.md.
WalkOptions (bon builder) configures the filesystem reader (FsSource /
walk_checkout). Three of its six fields are content-level — they simply pass through to
IndexOptions via
impl From<&WalkOptions> for IndexOptions, so a non-filesystem reader configures the identical
behaviour by building IndexOptions directly. Two are FS-reader-only: they decide which paths
FsSource hands over in the first place and have no IndexOptions equivalent at all — a reader with
no filesystem to walk has nothing to plug them into. The last, embed, is walk_checkout-only in
a different sense: it is not a path-filtering decision either, it is a post-processing step
walk_checkout runs after Indexer::finish returns (see Semantic embeddings),
so it has no IndexOptions equivalent for the same reason a non-filesystem caller wanting embeddings
calls embed::embed_output
itself rather than finding an IndexOptions field for it.
| Field | Default | Level | Meaning |
|---|---|---|---|
tuning |
IndexTuning::default() |
content (IndexOptions::tuning) |
Chunking/window sizing (below) |
build_graph |
false |
content (IndexOptions::build_graph) |
Build the structural graph. Off by default: a caller that only wants chunks pays no graph-extraction cost |
extract_pdfs |
true |
content (IndexOptions::extract_pdfs) |
Extract text from PDFs and chunk it |
respect_gitignore |
true |
FS-reader-only | Honour the repo's own .gitignore (and nested/parent ignore files) |
extra_ignore_globs |
[] |
FS-reader-only | Operator-supplied gitignore-syntax globs, layered on top of the built-in defaults |
embed |
None |
walk_checkout-only |
Some(EmbedConfig) embeds every chunk against an OpenAI-compatible endpoint after Indexer::finish returns (see Semantic embeddings); None dials out to nothing. walk_checkout_from_env sets this from EmbedConfig::from_env() |
A caller driving the raw-inputs core directly (index_inputs/Indexer) builds
IndexOptions instead —
the same tuning/build_graph/extract_pdfs three fields, with the same defaults, and no ignore or
embed fields at all (path filtering isn't its job, and embedding is a driver-level step layered on top
via embed::embed_output; see "Raw inputs: the filesystem is one reader" above).
IndexTuning fields, each readable from an environment variable via
IndexTuning::from_env
(unset or unparseable falls back to the default; every value is clamped to >= 1):
| Field | Env var | Default | Meaning |
|---|---|---|---|
embed_batch_size |
INDEX_EMBED_BATCH_SIZE |
32 |
Chunks per embedding round-trip |
max_chunk_lines |
INDEX_MAX_CHUNK_LINES |
150 |
Max lines a structured chunk may span before falling back to windowing |
window_size |
INDEX_WINDOW_SIZE |
100 |
Windowed-fallback window size, in lines |
window_step |
INDEX_WINDOW_STEP |
50 |
Windowed-fallback step, in lines (overlap = window_size - window_step) |
walk_checkout_from_env(root, build_graph)
is a convenience that builds WalkOptions from the environment: IndexTuning::from_env() for
tuning, and LCI_CODEGRAPH_IGNORE_GLOBS (newline- or comma-separated) for extra_ignore_globs.
build_graph itself is a plain function argument, not read from the environment.
The operator ignore layer composes with the repo's own .gitignore — it does not replace it.
walk_checkout drives the file walk with ignore::WalkBuilder, which
honours the repo's .gitignore (and nested/parent ignore files) natively when respect_gitignore is
true; IgnoreList is then applied as an additional filter on top, so a junk directory that slipped
past the repo's own rules (or a repo with no .gitignore at all) still gets skipped. Every skip is
logged at debug/info so an over-broad glob is diagnosable rather than silently hiding real files.
DEFAULT_IGNORE_GLOBS — the built-in defaults, always included unless a caller builds IgnoreConfig
directly with include_defaults(false):
target/ node_modules/ .git/ dist/ build/ vendor/ .venv/ venv/ .next/ __pycache__/
Every run — walk_checkout, index_inputs, or a manual Indexer — returns
IndexStats on
output.stats, and logs the same counters at info on finish. They exist because, with raw inputs,
"nothing got indexed" can mean several different things from the outside, and the counters tell them
apart:
| Field | Meaning |
|---|---|
files_chunked |
Inputs that produced at least one chunk |
paths_ignored |
Inputs a reader pruned before they ever reached Indexer::push (e.g. FsSource skipping a .gitignored or operator-ignored path) |
pdfs_extracted |
PDFs successfully text-extracted |
pdfs_skipped |
PDFs over the byte cap or that failed extraction |
files_skipped_binary |
Content rejected as binary: either it fails UTF-8 decoding outright, or it decodes fine but trips the content sniff (a NUL byte in the first 512 bytes) |
files_skipped_too_large (new) |
Inputs over chunk::MAX_FILE_BYTES, rejected before any parse/decode work |
files_skipped_unsupported (new) |
Inputs with no determinable language: RawInput::language was None and lang::from_path couldn't classify the path either (and it wasn't a PDF) |
files_skipped_too_large and files_skipped_unsupported are new with the raw-inputs core: a raw
input often has no meaningful filename or a size nobody validated up front — a DB row, a gist, an
editor buffer — so both failure modes needed their own counter instead of silently landing in
files_skipped_binary or vanishing. files_skipped_binary itself now also covers UTF-8 decode
failures; previously that path incremented nothing at all.
Repos carry documentation as PDFs; those get bounded text extraction and are fed to the same windowed-chunk path plain text files take. PDF parsing over untrusted repo input is a crash/OOM/ hang surface, so extraction is bounded in layers:
- Input bytes are capped at the I/O level (
MAX_PDF_BYTES, 5 MiB) before the parser ever sees the file — a multi-gigabyte "PDF" never lands in memory whole. - Before the real parser runs, every
FlateDecodecontent stream is pre-flighted through a bounded inflate (MAX_PDF_DECOMPRESSED_BYTES, 256 MiB cumulative budget) that never materialises more than a small buffer, rejecting a decompression bomb before it can trigger an (uncatchable) allocation failure. - The real parse runs on a worker thread under a 15s (
PDF_PARSE_TIMEOUT) watchdog. - The parser call is wrapped in
catch_unwind(pdf-extractcan panic on malformed input), and extracted text is truncated toMAX_PDF_TEXT_BYTES(2 MiB).
Honest residual limits, documented in src/pdf.rs: the decompression guard only covers FlateDecode
streams found by a syntactic scan — a bomb behind a non-Flate or cascaded filter, or a blow-up in
font/glyph tables rather than stream inflation, is not pre-flighted. The wall-clock watchdog is the
only backstop for those, and on timeout the worker thread is abandoned, not killed — its memory is
not reclaimed. A hard per-parse memory ceiling (subprocess + RLIMIT_AS) is not implemented here; a
caller running this over fully untrusted input at scale should isolate the process accordingly.
Set OPENAI_BASE_URL and walk_checkout embeds every chunk it produces against an OpenAI-compatible
/embeddings endpoint — no local/in-process model, no Cargo feature gate. Leave it unset and nothing
changes: no request is ever made, Chunk::embedding stays None on every chunk, and the walk costs
exactly what it cost before this existed.
The embed step itself, embed::embed_output,
takes an IndexOutput
rather than anything filesystem-specific, so it isn't only for walk_checkout — a caller driving
index_inputs/Indexer directly (see "Raw inputs: the filesystem is one reader" above) gets the same
graph-aware embedding by calling it themselves once their own indexing run finishes. walk_checkout
is simply the one caller that does this for you, wired through WalkOptions::embed.
OPENAI_BASE_URL is the switch: EmbedConfig::from_env
returns None when it is unset or blank, and walk_checkout_from_env
wires that straight into WalkOptions::embed. The API key is a separate, optional knob —
OPENAI_API_KEY absent (or blank) means an unauthenticated endpoint, which is a legitimate
configuration, not a missing one: a local gateway, vLLM, or Ollama's OpenAI-compatible shim typically
needs no key at all.
| Env var | Maps to | Default | Meaning |
|---|---|---|---|
OPENAI_BASE_URL |
EmbedConfig::base_url |
unset | The switch. Unset or blank → nothing is embedded, zero extra requests. Set → the API base, e.g. https://api.openai.com/v1; the request URL is {base_url}/embeddings |
OPENAI_API_KEY |
EmbedConfig::api_key |
None |
Sent as Authorization: Bearer <key> only when set. Blank/whitespace-only counts as unset |
OPENAI_EMBEDDING_MODEL |
EmbedConfig::model |
text-embedding-3-small |
Falls back to the default when unset |
OPENAI_EMBEDDING_DIMENSIONS |
EmbedConfig::dimensions |
None |
Passed through as the request's dimensions field only when set (not every server accepts it). Unparseable → None, not an error |
OPENAI_EMBEDDING_TIMEOUT_SECS |
EmbedConfig::timeout |
30 |
Per-request timeout, in seconds. Unparseable or 0 → default |
OPENAI_EMBEDDING_MAX_RETRIES |
EmbedConfig::max_retries |
3 |
Retries on 429 / 5xx / transport error, exponential backoff starting at 500ms. Unparseable → default; 0 is legal (no retry) |
OPENAI_EMBEDDING_MAX_INPUT_CHARS |
EmbedConfig::max_input_chars |
8000 |
Each input is truncated to this many chars (not bytes) before being sent. Unparseable or 0 → default |
There is deliberately no OPENAI_EMBEDDING_BATCH_SIZE: batch size reuses the existing
INDEX_EMBED_BATCH_SIZE / IndexTuning::embed_batch_size knob (default 32, see
Configuration) rather than adding a second name for the same setting.
EmbedConfig::max_context_refs (default 8, the max callees and max callers listed in the context
header, each side capped independently) has no environment variable — set it through the EmbedConfig
builder directly.
The text sent to the model is graph-aware, not just chunk.content on its own:
embed::embed_chunks builds a small index over the resolved Graph once per walk and prepends a
deterministic header — enclosing container, callees, callers — ahead of each chunk's own source.
Below is the header produced for add in a three-file Rust checkout where src/math.rs defines
impl Calculator { pub fn add(..) }, add calls helper/log in src/util.rs, and main in
src/main.rs calls add:
// file: src/math.rs
// language: rust
// function: add
// within: impl
// calls: helper() [src/util.rs], log() [src/util.rs]
// called by: main() [src/main.rs]
pub fn add(&self, x: i32) -> i32 {
helper(x) + log(x)
}
Two things that example is showing honestly rather than prettily. The within: line reads impl,
not impl Calculator — a Rust impl block's graph label is the bare node kind, because the block
has no name of its own to take (see the label rule in src/graph/emit.rs), so for Rust the
within: line tells the model that a method sits in an impl but not which type's. A class in
Python/TypeScript/Java, which does have a name, renders as // within: Calculator. And the content
keeps the source's original interior indentation while starting flush at pub — the chunk is the
node's exact byte range, not a re-indented copy.
file: and language: are always present. The third line is {chunk_type}: {symbol_name} when the
chunk has a symbol name, else the bare {chunk_type} (a windowed chunk with no symbol renders as
// window). within:, calls:, and called by: are each omitted entirely when there's nothing to
say — a window chunk, a PDF chunk, or any chunk from a walk with build_graph: false maps to no
graph node at all, so its header is just the first three lines. calls:/called by: entries are
label [source_file], sorted and deduplicated, capped at max_context_refs per side; a truncated list
gets a trailing marker, e.g. // calls: aaa() [b.rs], target() [b.rs] (+1 more). Every header uses //
regardless of the chunk's actual language — the header is never compiled, only read by the embedding
model, so a per-language comment token would buy nothing. The whole thing (header + content) is then
truncated to max_input_chars chars, preserving the header and cutting content first; only a
header that alone exceeds the cap gets cut itself.
The result lands on the chunk: Chunk::embed_input holds the exact text that was sent (or would be,
once truncation applies), and Chunk::embedding holds the vector that came back.
- A failed embedding call fails the whole walk.
walk_checkoutreturnsErrrather than an output with some chunks embedded and others not — configured means required here, because a half-embedded index that looks complete is worse than a walk that visibly failed. - No graph, degraded header. With
embedconfigured butbuild_graph: false, every chunk maps to no graph node, so the header isfile:/language:/symbol only — nowithin:,calls:, orcalled by:lines. The walk does not forcebuild_graphon to compensate — that would silently change the cost profile of a walk a caller configured without it — it logs atracing::warn!instead. - The endpoint sees your source code. Every chunk's content (up to
max_input_chars) goes out in the request body to whateverOPENAI_BASE_URLpoints at. Pointing it at a third-party host is a data-egress decision the operator is making, not one this crate can make safe on its behalf.
cargo testruns the unit tests (each module) and the integration suite tests/parity.rs, which walks a
committed fixture repo (tests/fixtures/sample-repo) and asserts the canonicalised graph is
byte-identical to the committed golden (tests/golden/sample-repo.graph.json) — the regression guard
for the graph engine. Regenerate the golden intentionally with:
UPDATE_GOLDEN=1 cargo test --test paritycargo test --features container-testsadditionally runs the Docker-backed suites — each needs a running Docker daemon:
container_neo4j— loads the emitted nodes/edges into a real Neo4j with the same generic:Symbol+[:REL {relation}]write a downstream host performs, then runs the retrieval queries against it, proving the downstream retrieval contract end to end.container_build— builds and runs the crate inside Linux glibc and musl containers.container_repos— clones pinned real-world repositories inside a container and asserts the walk holds its invariants on input nobody wrote for the tests.
The graph returned by walk_checkout/walk_checkout_from_env is canonicalised: nodes and edges are
sorted and deduplicated before being returned (Graph's resolve step). Running the same walk twice
over the same checkout produces byte-identical output — stable to snapshot in a golden test, and
stable to submit downstream without spurious diffs.
Cross-file calls resolution is precision-favouring, not best-effort: when a bare callee name matches
more than one definition and no qualifier narrows it to exactly one, the call is dropped, not
guessed — it is never fanned out to every same-named candidate and never resolved to an arbitrary one.
Concretely: a name defined in two files with no importing/qualifying context to tell them apart
produces no calls edge for that call site. A qualifier is recovered only from the immediate receiver
in the source (Foo.bar() → qualifier Foo) — there is no type inference, so self/this/cls/
super receivers, and calls through a variable of unknown type, carry no qualifier and resolve on the
bare name alone (a single match still resolves; multiple still drop). This trades recall for not
mis-attributing a call to the wrong definition.
Exported from vymalo/lightbridge-code-intelligence.
Design rationale: ADR-0086. Licensed under
MIT.