Repository navigation
Releases: knight22-21/CodePrism
Release list
v0.1.11 — import-aware call links, gitignore support
Upgrading: existing indexes are rebuilt once automatically (index format 2).
Fixed
- Call links follow imports (Python). Cross-file calls were linked by name alone, so
asyncio.run()could link to your ownrun,str()to a projectstr, and a method call to
a field of the same name. Calls now follow the caller's imports (including package
re-exports and function-local imports); builtins and third-party calls link to nothing; calls
on objects link only when there is exactly one candidate in a file the caller imports. On a
new cross-file benchmark, precision went from 0.68-0.96 to 1.00 on requests, flask, httpx
and CodePrism, with recall equal or better. Other languages still resolve by name. (#33) - Gitignored files are no longer indexed (vendored checkouts, build output). Uses git's own
file list, with a plain-walk fallback outside git. Opt out withrespect_gitignore = false.
(#31) - Ambiguous paths are reported, not "not indexed". A short path matching several files
(utils.py) now returns the candidate paths and a hint. (#32)
Added
benchmarks/run_call_precision.py: cross-file call accuracy against a Pythonastoracle.
Install
pip install -U codeprism-ai==0.1.11v0.1.10 — tools accept project-relative paths
Fixed
- Tools now accept project-relative file paths. Files are stored with absolute paths, so
get_callers("src/app.py", ...),get_module_summary,get_context,search_symbol's
project_pathand the other file tools silently returned empty results for the relative
paths agents normally use. Relative paths,./, either separator and (on Windows) any case
now resolve against the served project root; an ambiguous name (twoutils.py) is never
guessed. (#30) get_callers/get_calleesreport errors for an unindexed file or unknown symbol instead
of an empty list, and include each caller's file path, not just its internal id. (#30)- No duplicate file records:
codeprism index .stored relative paths while the MCP server
stored absolute ones. Paths are now always absolute; an old relative index heals itself on the
next index run. (#30)
Install
pip install -U codeprism-ai==0.1.10v0.1.9 — graph sync & scale, working Claude Code setup, Codex + AGENTS.md
Graph sync and scale fixes, a working Claude Code setup, Codex support, and
AGENTS.md as the shared agent guide.
Upgrading: existing indexes are rebuilt automatically on the next
codeprism index
or server start (new index format version), which also clears stale rows left by the
re-index bugs fixed below. Re-runcodeprism setup <agent>to move to the new config
locations and getAGENTS.md.
Added
codeprism setup codex: registers[mcp_servers.codeprism]in.codex/config.toml
(or~/.codex/config.tomlwith--global). Comments and other settings are preserved.AGENTS.mdis the shared guide. Every setup writes the CodePrism usage guide to
AGENTS.md(read by Codex, Cursor, Windsurf, Zed, Copilot, ...).CLAUDE.mdbecomes a thin
@AGENTS.mdimport. Continue.dev gets.continue/rules/codeprism.md. (#27)- One user-level config for every project.
codeprism servewithout a path serves the
project containing the current directory (nearest.git/.codeprism.toml) and builds or
refreshes its index in the background at startup.setup claude|codex --globaluse it.
get_graph_statsreportsindex_status. (#29) - Parallel parsing.
codeprism indexand the MCP server parse in a process pool
(--workers/-j,parse_workersconfig), and SQLite is tuned for bulk writes.
Full index of a 2M LOC corpus: 89.6s → 42.2s. (#24) - Index format version. Indexes from an older format are fully re-parsed once instead of
being trusted by checksum. (#25)
Fixed
- Claude Code setup never actually registered the server. MCP entries were written to
.claude/settings.json, which Claude Code doesn't read. Now.mcp.json(project, pre-approved
insettings.local.json) or~/.claude.json(--global); stale entries are migrated, and
setup files go to--projectinstead of the current directory. (#26) - Java, C, C++, Ruby and PHP files were never indexed, and the watcher ignored them (and
.rs). All 10 languages are now indexed and watched by default. (#20) - Re-indexing left stale data: renamed symbols lingered and edges were duplicated when line
numbers shifted;--forcedidn't remove deleted files. (#21) update_filedropped callers from other files from the in-memory graph. (#22)- Single-file updates did whole-repo work: 2.6s → ~55ms at 2M LOC. (#28)
search_symbolkind filter was case-sensitive;kind="Function"returned nothing.
Thanks @MilindLate. (#19, fixes #18)
Install
pip install -U codeprism-ai==0.1.9v0.1.8 — Java, C/C++, Ruby, PHP parsers (10-language support)
What's new in v0.1.8
10-language parser support
CodePrism now indexes 10 languages out of the box:
| Parser | Languages | New in this release |
|---|---|---|
| Java (L1) | .java |
✅ Classes, interfaces, methods, annotations, generics |
| C / C++ (L2) | .c, .h, .cpp, .cc, .cxx, .hpp, .hh |
✅ Functions, structs, classes, inheritance, namespaces |
| Ruby (L3) | .rb, .rake, .gemspec |
✅ Modules, classes, instance/singleton methods, visibility |
| PHP (L4) | .php, .php5, .phtml |
✅ Namespaces, classes, traits, interfaces, methods |
Previously supported: Python, JavaScript, TypeScript, Go, Rust.
Documentation
- README Supported Languages table updated (all 10 languages)
- Mermaid benchmark charts added to README and
docs/benchmark-results.md INTEGRATIONS.mdnow has a Supported Languages section with extensions
Install
pip install codeprism-ai==0.1.8Full changelog
See docs/changelog.md.
v0.1.7 — Corpus expansion + get_dependencies fix
What's new
Added
- Benchmark corpus expansion: pallets/flask 3.0.3 and encode/httpx 0.27.2 (10 tasks each). Token reduction: flask 91.3%, httpx 93.0%. Overall average across 3 production codebases: 91%.
Fixed
get_dependenciesoutput format: was returning imported symbol names instead of source modules. Root cause:source_modulestored inextra{}(excluded from SQLite). Fix: store in thesignaturefield (persisted). Re-index required for updated output.- Benchmark ground truths for
requests_004,requests_006,requests_008: updated to match actual tool output.
Benchmark summary
| Corpus | Baseline | CodePrism | Reduction |
|---|---|---|---|
| psf/requests v2.32.3 | 6,407 tokens | 738 tokens | 88.5% |
| pallets/flask 3.0.3 | 9,558 tokens | 828 tokens | 91.3% |
| encode/httpx 0.27.2 | 12,685 tokens | 894 tokens | 93.0% |
Full results: docs/benchmark-results.md
v0.1.6 - Level 2 latency benchmark + get_dependencies fix
What's new in v0.1.6
Level 2 latency benchmark
benchmarks/run_latency_benchmark.py: measures p50/p95/p99 query latency with warmup runs- Engine held open across all tasks — times pure query latency, not indexing overhead
- Results on psf/requests v2.32.3 (20 reps, 3 warmup discarded):
get_contextp50=1.2ms, p95=1.8msget_callersp50=1.2ms, p95=1.8msget_impactp50=2.0ms, p95=3.4msget_dependenciesp50=3.8ms, p95=5.1ms (after fix)- Overall p50=1.9ms, p95=2.5ms
All queries under 5ms p95 — noise relative to LLM API latency (500-3000ms).
Performance fix: get_dependencies
- Was doing
SELECT * FROM symbols(full table scan) on every call - Fixed with
SELECT DISTINCT name WHERE kind != 'import'(name-only projection) - Latency dropped from 10.7ms to 3.8ms p50 on the requests corpus (64% faster)
- All 341 tests pass
v0.1.5 - Benchmark harness + cross-file call edge fix + docs
What's new in v0.1.5
Indexer fix
- Cross-file call edges now resolve correctly.
BaseParser.resolve_intrafile_refswas consuming CALLS refs that matched import stubs, preventing them from reaching cross-file resolution. Fixed by skipping CALLS-to-import-stub wiring intra-file. Impact:run_payment -> compute_checksumand all similar cross-file calls now appear correctly inget_callers().
Level 1 benchmark harness
- Token reduction measurement via tiktoken (no API key required)
- LLM-as-judge accuracy eval via Ollama cloud (
gpt-oss:120b) - Fixture corpus (10 tasks) + psf/requests v2.32.3 corpus (10 tasks)
- CI workflow: runs on every push/PR, token-only, results as build artifact
Measured results
- 88.5% token reduction on psf/requests v2.32.3 (real-world, 6k-12k token files)
- Accuracy: baseline 0.81 / CodePrism 0.66 on requests corpus
Documentation
docs/benchmark-results.md- full methodology and all run resultsdocs/architecture.md- 5-layer system design and data flowdocs/changelog.md- version history- README: Benchmarks section + Documentation table
Run the benchmark yourself
pip install -e ".[bench]" tree-sitter-python
python -m benchmarks.setup_repos
python -m benchmarks.run_token_benchmark --tasks benchmarks/tasks/requests_tasks.jsonv0.1.4 - Property-based tests (Phase B)
Phase B: Hypothesis Property-Based Tests
341 tests total (up from 324). Added 17 property-based tests using Hypothesis that assert invariants hold for any valid input.
Coverage
- SecurityScanner: status always in {PASS,WARN,BLOCK}; BLOCK iff BLOCK issue present; identical diff always PASS
- SecretsDetector: any api_key/password pattern fires; all findings carry fix_suggestion
- InjectionDetector: f-string SQL always flagged; severity always BLOCK or WARN
- WeakCryptoDetector: findings are always WARN severity
- GraphEngine: N unique symbols produce N nodes; removal makes node unreachable; edge roundtrip; caller/callee symmetry
- SearchMatch: score always in [0.0, 1.0]
v0.1.3 - Semantic search via embeddings (Phase A)
Phase A: Semantic Search
search_symbol now supports embedding-based semantic search.
New features
- codeprism index --embeddings: builds ChromaDB vector index alongside the graph
- index_project MCP tool: new embeddings param
- search_symbol finds symbols by meaning, not just name substring
- Falls back to substring search silently if embeddings not built
- enable_embeddings = true in .codeprism.toml wires semantic search into MCP server
Usage
pip install 'codeprism-ai[embeddings]'
codeprism index /path/to/project --embeddings
v0.1.2 — Spec compliance fixes
What's fixed
- search_symbol: added project_path param to scope results by directory prefix
- get_graph_stats: path filter now actually filters stats instead of being silently ignored
- get_dependencies: circular_deps[] is now populated via NetworkX cycle detection
- pyproject.toml: bumped version, fixed author email
- CODE_OF_CONDUCT / CONTRIBUTING: updated contact email