-
Notifications
You must be signed in to change notification settings - Fork 0
Home
Welcome to the bicdb wiki!
A database built for AI agents.
Repository: https://github.com/xuji755/bicdb · License: Apache-2.0
Current release: v0.1.0 (2026-10-06) — first runnable version
Status: Storage, transactions, crash recovery and an SQL front end are built and tested. Memory, search, graph and asset features are designed but not built yet.
- What is bicdb?
- Why does it exist?
- Who is it for?
- What does it do?
- Try it right now
- What works today — and what doesn't
- What makes it different
- What it is not
- Where the project is going
- License
bicdb is a database designed specifically for AI agents — not for websites, not for business reporting, but for the software that runs on your behalf, remembers things, looks things up, and picks up where it left off.
Think of it as the private memory and storage layer for an AI assistant. Every user gets their own sealed-off space. Inside that space live the things an agent needs to do its job: what it has learned, what it remembers about you, how it is configured, how things relate to each other, and the big files it works with.
Today, teams building agents glue together four or five different products to do this — a vector database for search, a key-value store for session state, a relational database for records, a file store for documents, and something hand-rolled for memory. bicdb's bet is that all of it belongs in one place, under one set of guarantees.
And it is no longer just a plan: as of October 2026 there is a real, running database you can build and poke at.
AI agents generate a peculiar kind of data, and existing databases were not designed for it.
| What agents need | What ordinary databases give you |
|---|---|
| Remember a conversation last week and find it again | Session state that is lost on restart, or plain text files |
| Store knowledge that can be forgotten and revoked | Rows you can only delete, with no notion of "stop citing this" |
| Distinguish a guess from a confirmed fact | No concept of confidence, provenance or contradiction |
| Keep long-term personal facts and short-term task progress | One kind of storage, one retention policy |
| Search by meaning and by exact keyword, together | Either/or — you pick one engine |
| Keep different users' data genuinely separate | Permission systems that are easy to get wrong |
| Store a 50 MB PDF next to a note about it | Either blow up the database, or bolt on a file store |
The result today is usually a stack where each piece has different rules about durability, isolation and deletion. An agent that "forgets" something often forgets it in one place and keeps it in three others.
bicdb's answer is to put all of it behind one transaction and one isolation boundary, so "delete this memory" or "this fact has been corrected" means the same thing everywhere.
- People building AI assistants, agents and copilots that need long-term memory
- Teams that want to self-host their agent's data rather than rent five cloud services
- Anyone who needs strong per-user privacy — one user's data must never leak into another's
- Projects that want memory, search, knowledge graphs and file storage from a single system
bicdb is designed to hold five kinds of data in one place, under one set of rules.
Every user gets a fully separate workspace. Nobody can reach anyone else's data — not by guessing IDs, not through a shared cache, not by crafting a request. There is no "share this with another user" feature and there never will be.
Alongside these sits one shared public area, readable by everyone and writable by nobody except an administrator. It is where common knowledge lives — the equivalent of a published library next to everyone's private notebooks.
This is the heart of the project. bicdb distinguishes two kinds of memory:
Short-term (working) memory — the state of a task in progress. What the goal is, what has been decided so far, what step it reached, what is still to do. If the agent crashes or the process restarts, the task can be resumed instead of starting over.
Long-term memory — facts, preferences, experience and knowledge about a person or a domain. Here bicdb carries ideas you will not find in ordinary databases:
- Versions — when a fact changes, the old version is kept, not overwritten. You can see how understanding evolved.
- Provenance — every memory can point back to where it came from: which conversation, which document, which model produced it.
- Confidence as data — a generated guess and a user-confirmed fact are different things, and bicdb records which is which. Being confident does not make something true, and the database does not pretend otherwise.
- Time that makes sense — when something happened, when it was recorded, and when it was actually true are three different questions, and bicdb stores all three.
- Contradictions allowed — two conflicting memories can coexist, linked as conflicting, instead of the newer one silently burying the older one.
- Forgettability — memories can expire on a schedule, or be revoked. Revoking something also stops everything derived from it from being returned.
- Nothing is invented by the database — bicdb never decides on its own that something is important, and never calls an AI model itself. Deciding what matters is the application's job.
Agents constantly need to answer "have we dealt with this before?" bicdb is designed to support:
- Keyword search — matched text, names, error codes, technical identifiers
- Semantic search — "find things that mean this", using vector similarity
- Both at once — results from keyword and semantic search combined into one ranked list, rather than you having to choose
- Relations — data connectable as a knowledge graph, so you can ask "what is connected to this, and how far does it go"
- Freshness you can control — a query can demand "guaranteed up to date, even if slower", or accept "as up to date as the index currently is", and know which one it got
Crucially, searches never mix one user's data with another's, and never return something that has been deleted or revoked — even when the underlying search is approximate.
Prompt templates, tool configurations, model settings and workflow definitions are stored with immutable versions. Changing a setting creates a new version. Switching back is instant, because rollback only changes which version is active.
Secrets are stored as references, never as plain text sitting inside a config blob. And bicdb only ever stores these things — it never loads code or executes instructions found in configuration.
Documents, images, model weights and generated artefacts live as managed files next to the database, tracked by the database. They are written once and then read-only — to change one, you save a new one and point at it. This keeps big files out of the way of the transaction system while still making them first-class citizens: they belong to a user, they count against a quota, and they are cleaned up when nothing references them any more.
Everything above shares the same transactional guarantees — data is never half-written, committed data survives a crash, and uncommitted data disappears cleanly. This part is no longer a promise: crash recovery is implemented and tested, including killing the process at arbitrary points.
Importantly, bicdb is explicit about which parts of your data come with which guarantees, instead of quietly promising more than it delivers:
- Data that needs full protection gets it.
- Data you have deliberately marked as expendable (like raw chat logs) is clearly marked as such — it is faster, but it can be lost on a crash, and the project says so out loud rather than hiding it.
- Backups and point-in-time restore are supported for protected data, and the documentation states plainly which data is not covered.
Because a workspace is a self-contained package, the design allows it to be cloned almost instantly into an independent copy — a snapshot of a user's entire memory and knowledge that evolves separately. Useful for testing, experimentation, or creating a sandbox from real data. Merging those copies back is deliberately not supported, because doing it correctly is a different and much harder problem.
You need a Rust toolchain. Then:
cargo build --release -p bicdb-cli
./target/release/bicdb init ./demo
./target/release/bicdb sql ./demo "CREATE TABLE t (id NUMBER NOT NULL, name VARCHAR2(32))"
./target/release/bicdb sql ./demo "INSERT INTO t VALUES (1, 'alpha'); INSERT INTO t VALUES (2, 'beta')"
./target/release/bicdb sql ./demo "CREATE UNIQUE INDEX t_pk ON t (id)"
./target/release/bicdb sql ./demo "SELECT id, name FROM t WHERE id >= 1 ORDER BY id DESC LIMIT 10"
./target/release/bicdb shell ./demo # interactive; `;` ends a statementinit creates a real on-disk database, not a toy: dictionary file, undo segment, log directory and two copies of the control file. Every command recovers from a crash on open and shuts down cleanly with a full checkpoint on close — so if you kill the process mid-write, nothing that was committed is lost. You can try that yourself.
bicdb is being built in stages, and it is refreshingly clear about where it has got to. As of v0.1.0:
| Private workspaces | Isolation, identity routing, per-workspace directories and quotas |
| The storage engine | Pages, rows, heap tables, overflow handling for large values, space allocation |
| Transactions | Begin / commit / rollback, statement snapshots, row locking, deadlock detection |
| Crash safety | Write-ahead logging, checkpoints, and three-phase recovery that actually runs on every open |
| Indexes | B+Tree indexes, including unique indexes, plus fast bulk index building |
| SQL (basic) |
CREATE TABLE, CREATE INDEX, DROP, INSERT, and SELECT with WHERE, ORDER BY and LIMIT
|
| Command line | Create an instance, run SQL, or use an interactive shell |
The following are designed in detail but not implemented — and when you try them, the database says so by name rather than silently ignoring you:
-
UPDATEandDELETE - Aggregates (
COUNT,SUM, …), joins,DISTINCT, set operations, table aliases - Memory — versions, provenance, expiry, revocation
- Search — keyword, semantic and hybrid retrieval
- Knowledge graph — relations and traversal
- File assets — storing documents and images alongside data
- Server / remote access — everything is local and single-user for now; there is no network protocol or SDK yet
So: the engine is real and tested. The agent-specific features that make bicdb interesting — memory, search, graph, assets — are still ahead. One workspace per instance, for now.
| A typical stack today | bicdb | |
|---|---|---|
| Storage | Vector DB + KV store + relational DB + file store | One system, five kinds of data |
| Memory | Application code bolted on top | Built in, with versions, provenance and expiry |
| Deletion | Delete in 4 places, hope you got them all | One transaction, one meaning |
| Isolation | Permissions you configure and audit | Physically separate, cannot be reached |
| Search | Pick keyword or semantic | Both, fused into one ranked result |
| Honesty | "Everything is durable" | Documented per data type, including what can be lost |
| Best fit | General business applications | AI agents specifically |
The project is unusually clear about its boundaries, which is a good sign:
- Not a general-purpose database. It is tuned for agents: read-heavy, short transactions, lots of search.
- Not a cloud service. It runs on a single machine that you control. Self-hosted by design.
- Not distributed. No clustering, no multi-node failover, no global transactions across machines.
- Not an AI model runtime. It stores configurations and prompts; it never runs models or executes user code.
- Not a sharing platform. There is no cross-user sharing, no transfer of ownership, no share links. This is permanent, not a missing feature.
- Not Oracle. It matches the behaviour of six common data types so that migrating code feels familiar, but it is an independent design with its own storage format.
- Not finished. It is an early v0.1.0 — a solid foundation, not yet a complete product.
Development proceeds in stages. The recent release completed the first four.
| Stage | What it delivers | State |
|---|---|---|
| Foundation | Private workspaces, isolation, testing setup | Done |
| Storage | Durable data, crash recovery, indexes | Done |
| Concurrency | Transactions, row locks, B+Tree indexes | Done |
| SQL | Language front end, catalogue, command line | In progress |
| Memory & assets | Private memory with lifecycle, stored files | Next |
| Search | Keyword and semantic search, then fast approximate vector search | Planned |
| Multi-model | Knowledge graphs alongside everything else | Planned |
| Release | Reliability hardening, backup, security testing, controlled trial | Planned |
The next milestone is usable private memory and file assets — the point at which the database starts being useful for actual agents rather than just for storing rows.
Calendar time never substitutes for a stage actually passing its tests.
Apache License 2.0 — permissive, business-friendly, and requires attribution.
This is the short, non-technical introduction, updated for v0.1.0. For the full design — storage layout, transaction protocol, isolation mechanisms, platform baselines and the complete requirement list — see the companion document bicdb-WIKI.md, and the repository's own README.md, CHANGELOG.md and docs/.
Note: the companion bicdb-WIKI.md was written against the pre-implementation design freeze and still describes the project as unbuilt. Its technical explanations remain accurate; its status chapter does not.