The backend agents run on.
Durable memory that records where every fact came from, a graph nobody had to model, sandboxes to run agent code in, and one line that touches the outside world.
- Memory that keeps being wrong. Nothing is rewritten. A correction is a later fact and the earlier one still answers where it was written, so "what did it believe on Tuesday" is a question rather than a log search.
- Every fact knows what made it. Provenance is a slot in the row, not a convention: an answer either came from outside or names the code that produced it, and there is no third option to forget.
- A graph you did not have to model. An edge is a fact whose value is another id. No node type, no edge type, no second store to keep in step.
- Code runs where it cannot reach. Agent code runs in Lua (Luerl, inside the BEAM) with no clock, no network and no filesystem. Isolation is the absence of anything to reach, so there is no rule to misconfigure.
- One line touches the outside world. A job is the only thing handed the network and the only thing a schedule attaches to, so what an agent did is a list you can read, with its failures on it.
- It tells you when something changed.
watchis the same question asked again as things land, not a second mechanism.
ada.height = 180
ada.friend = grace -- an edge is a field holding another entity
print(ada.friend.height)
for p in each { height = 180 } do print(p.id) end
ada.height = nil -- unsaid; at(42).ada.height still answers 180
That is the whole surface. An entity is a table, a field is a field, and Lua is the query language — so counting, grouping and joining need no vocabulary of their own. Fact, attribute, snapshot and pattern are how this is built; none of them is something you have to learn to use it.
DESIGN.md is why it is shaped like this — the decisions the code is made of, what each one costs, what was deliberately not built, and the four things we got wrong badly enough to leave a test behind.
The vocabulary and the reasoning live in .monty/ontology.db: fourteen words
and twenty-four numbered doctrines, linted against the code by just check.
Those words name what this is made of, which is a different question from what
you type — the surface above is Lua and nothing else. There are no other design
documents: claims are doctrine in that database, choices are in the commit that
made them, and mechanism is in the moduledoc beside the code. A document
describing the system is a second source that drifts.
just check # what must pass before a commit
just test
mix test --include crash # a real process, SIGKILL'd mid-write
mix test --include load # what it costs at size
mix test --include gcp # talks to Cloud KMS
run world, lua -> what it returned, and the name it read at
watch world, what -> answered again as things land (websocket)
There is one door. run opens a snapshot, evaluates the chunk against it, and
appends whatever the chunk wrote — three steps that used to be three operations
a caller had to know the vocabulary for.
A caller holds the snapshot's name, never its bytes. Because an answer at a
name never changes, a client caches on {name, source} and never invalidates —
there is no cache-coherence protocol because there is nothing to cohere.
claim takes a world name and grants it to whoever claimed it; that is the
only other thing the wire does.
A world was called a ledger until recently. The storage kept the old spelling
on purpose — LEDGER_DIR, /data/ledgers, the ledgers/ backup prefix — and so
did the open_ledgers reading, because live facts and live objects already sit
under those names. Renaming them would have orphaned every one while the tests
stayed green, since a test writes with the same code it reads.
Authorization is which worlds a caller may name. Not row rules, not predicates. Every operation that names a world is checked, because a snapshot name is a plain map and a caller can write one by hand.
Everything that can differ between deployments is read at boot, so one artefact runs anywhere and carries no secret.
| Variable | Meaning |
|---|---|
SECRET_KEY_BASE |
Required in production. The release refuses to boot without it. |
PORT |
Defaults to 4000. |
LEDGER_DIR |
Where facts live. Must be persistent storage. |
LEDGER_SYNC |
true fsyncs every transaction: durable and slow. |
KEY_DIR |
Where key-encryption keys live. Must be persistent storage. |
KMS_KEY |
A Cloud KMS key. Setting it selects the KMS-backed keyring. |
GOOGLE_APPLICATION_CREDENTIALS |
A service account key, for the KMS. |
BACKUP_BUCKET |
With BACKUP_ENDPOINT, BACKUP_ACCESS_KEY_ID, BACKUP_SECRET_ACCESS_KEY, and optionally BACKUP_REGION and BACKUP_PREFIX. |
BACKUP_DIR |
A directory to copy into instead — a second disk, or a test. |
BACKUP_EVERY |
Seconds between runs. Defaults to 900. |
DRILL_EVERY |
Seconds between restore drills. Defaults to 21600 — six hours. 0 turns the drill off. |
DRILL_DIR |
Where a drill stages what it restores. Defaults to the container's temp space, and must never be LEDGER_DIR. |
DRILL_MAX_BYTES |
The largest world a drill will pull down. Defaults to 512 MB. |
Without LEDGER_DIR a world is in memory, which is right for a test and
wrong for everything else. Without KMS_KEY the keyring keeps its keys in a
file under a master from the environment — right for development, wrong in
front of real users, because a file can come back from a restore and erasure
has to be irreversible.
A fact's value is sealed under a key belonging to whoever its entity belongs
to. Erasing destroys that key: the bytes stay and become noise, no segment is
rewritten, and an old name still answers — it answers :erased.
Three tiers, because a KMS key version costs about six cents a month and one per subject prices itself out at exactly the scale erasure starts to matter. A single KMS key wraps a master, the master protects per-subject keys, and a per-fact data key wrapped by the subject key travels in the fact.
Erasure also writes a tombstone — an ordinary fact — and the keyring reconciles against those whenever it opens, so a key store restored from before an erasure is corrected rather than trusted.
The one thing it cannot do is stated plainly: a fact written before its subject was declared is not covered. Subject is decided at write time or not at all.
Copying the facts somewhere else is a job, because it reaches outside. That
is not a technicality: it means a backup has a cadence, writes what it copied
as ordinary facts, writes a failure as an ordinary fact, and answers "when did
this last succeed" with Job.last_run/2 like anything else. None of that had
to be built.
A world is append-only, so a run copies the byte range it has not copied yet and names the segment for the range it holds. The cost of a backup is what changed rather than what exists, which is what lets the cadence be minutes. A copy stops at the last complete record, because the tail of a live log may be half-written and half a record is not a fact.
Checkpoints are not copied — they are derived, and opening without one is correct and merely slower. Keys are, whole, every run: facts without them are noise. What lands in the bucket is already encrypted under a master the KMS holds, so whoever owns the bucket cannot open it, and an old key store cannot resurrect an erased subject because the keyring reconciles against erasure tombstones every time it opens.
Backup.run(...) copy what has not been copied
Backup.verify(...) what the target holds against what is here
Backup.restore(...) pull it back — refuses to land on top of live facts
verify compares readable local bytes, and what a target holds is the
contiguous run from zero rather than the highest segment boundary. Those are
the same number until a put fails after a later one succeeds, and then the
difference is a backup claiming bytes it cannot give back. Stopping at the hole
costs a re-copy and heals it.
Restoring refuses rather than overwrites, and names what is in the way. Pass
only: :keys or only: :worlds to restore one half — losing a key store
while the facts are fine is a real thing, and not the same operation as
restoring a machine.
Two limits, stated rather than hidden. verify always names $backup itself,
because a run copies and then records what it copied; the next run catches it
up. And no target can delete, but the credentials a deployment holds usually
can — so versioning and a retention window belong on the bucket, out of reach
of a node that has been taken over.
A backup is only ever proven by the last restore, and a restore a human did once by hand is a rumour with a timestamp. So the restore is a job too: every six hours it pulls one world back out of the backup, opens it, asks it for every fact it holds, and asks the live world the same question. Equal answers, or the run fails and the failure is a fact.
The comparison is made at the restored world's transaction, never the live
one's. A backup is always a prefix — a run copies and the live world goes on
being written to — so the honest question is whether the copy answers at
transaction n exactly what the original answers there. That is a snapshot, and
an answer at a snapshot is the same answer forever, which is why this can be an
equality rather than a tolerance. Byte counts are what verify compares, and a
broken restore has plenty of those.
One world per run, and the one picked is the one drilled longest ago, read out
of the drill's own past facts — so a redeploy resumes the rotation and there is
no state to reconcile. With n worlds each is proven inside n cadences.
Restoring everything every cadence would cost more than the backup it checks,
every time, and a check nobody can afford is a check somebody turns off. Only the
sampled world's segments cross the wire: the drill hands Backup.restore/1 a
narrowed view of the target rather than a second restore path of its own.
It stages into a scratch directory, restores only: :worlds, opens the copy
under a name of its own, and removes both the directory and the name on the way
out whether it proved anything or raised. Opening the copy under the live name
would hand back the live world — World.open/2 does that by design — and the
drill would compare it against itself and pass forever.
proven_at when we last proved we could restore
drilled which world this run gave back
proven_tx the transaction it was proven to
compared_facts how many facts had to agree
too_big a world over the ceiling, skipped and named
Four limits, stated rather than hidden. Only worlds that are open are
drilled, because comparing means asking the live one and a drill must never
start a live world. A world over DRILL_MAX_BYTES is skipped and named rather
than passed over. Keys are not restored — pointing the keyring at a restored
key store would mean swapping a live global for the length of a drill, and both
sides of the comparison reveal through the same keyring anyway. And a run with
nothing to drill records zero rather than failing: proven_at simply stops
advancing, which is the thing to watch.
A two-stage image: build with the toolchain, ship without it. Facts and keys live on a mounted volume, never in the image — a container is replaced on every deploy and a world is not.
CI runs the gate a machine that has never seen the repo can run: formatting,
warnings-as-errors, the suite, the SIGKILL crash tests, and the vocabulary
lint. The deploy workflow builds a release, boots it, and refuses to ship one
that answers anything but 401 to an unauthenticated request.
Apache 2.0. See LICENSE.
