AI safety audit · local-first tooling · evidence-backed engineering. The connective tissue across the repos Nelson ships under
taipei49314.
This is not a portfolio. Each repo here earns its place by either (a) implementing a piece of the AI-safety-audit stack, (b) being the tool the audit stack itself was built with, or (c) being a public subject the stack is pointed at. Subjects are not flagships. Nothing ships here that a reproducible harness has not produced on a clean checkout. Nothing claims a status it has not held under that harness.
A living index. Not a monorepo. Not a curated list. It exists so that the relationship between the linked project repos is legible in one read, and so that anyone auditing Nelson's work can navigate from the principles down to the evidence in one pass.
Last reconciled against GitHub repository, archive, release, and CI metadata: 2026-08-17.
The headline principle across everything here:
Declarations are not evidence. The workload does not judge itself. Models may propose; only the verifier decides.
UNKNOWN/INCOMPLETEis better than a false pass.
That sentence is from
RepoPassport. It is also the
operating principle of every other repo on this page.
Nelson's primary research target is how to audit an AI-driven system in a way that the system itself cannot fake. Claims are not trusted until measured. The active front of the stack is therefore admission, then measurement: refuse a session that never entered a frozen contract, refuse a journey the examinee wrote, score the checkout, refuse unmeasured phase advances, then run the six audit questions.
| # | Audit question | Primary repo | Supporting repos |
|---|---|---|---|
| 00 | Did this session enter a frozen task contract? | walkaround |
— |
| 00b | Was the journey allowed to count as an exam? | charterlock |
RepoPassport (declared journey); unasked (who may say verified) |
| 0 | What does the checkout score from local evidence? | trust-meter |
phaseledger (fresh PASS required to advance) |
| 0b | Can this phase advance without a measurer verdict? | phaseledger |
trust-meter |
| 0c | Did a pre-registered decision beat chance? | nullbench |
branchback (belief-at-the-time replay) |
| 1 | Did the declared journey work? | RepoPassport |
unasked (non-certifying investigation) |
| 2 | Did the workload stay within its declared capabilities? | RepoPassport |
— |
| 3 | Was the result reproducible? | stateweaver |
RepoPassport (deterministic plans) |
| 4 | Was cleanup complete? | RepoPassport |
stateweaver (reality replay) |
| 5 | What evidence exists, and who signed it? | stateweaver |
RepoPassport (attestation bundles) |
| 6 | Is that evidence still current? | tomorrowci · tomorrowci-lab |
RepoPassport (verdict staleness) |
walkaround and charterlock sit in front of measurement. walkaround asks
whether the session entered a frozen contract — done without entry is
BYPASSED; a receipt is not a verification of the work. charterlock asks
whether that journey was allowed to be the exam — same key writing and sitting
it is CHARTER_COLLAPSED; two MAC keys do not prove two people.
trust-meter and phaseledger then score first and advance only on a fresh
deterministic PASS. nullbench asks whether a pre-registered decision beat
chance without backfill. RepoPassport answers questions 1, 2, and 4 — the
workload-side invariants. stateweaver answers questions 3 and 5 — the
verifier-side invariants. tomorrowci answers question 6 — will this
evidence still be valid tomorrow?
evidence-workbench is the
authority-preserving control plane plus artifact/execution shell that catalogs
the twelve active cells below. It pins source and runtime candidates, imports
exact artifacts, records native envelopes, and executes only separately
admitted operations whose required boundary is implemented. EWB itself is not
globally read-only. The NelsonCode integration is a deliberately read-only
catalog adapter, but it is optional and outside the audit mainline. It carries
no authority handoff. Exact bytes and digests move between cells only through
separately admitted handoffs; registry presence alone does not perform a
handoff. EWB does not certify a workload, reinterpret a native status, or
synthesize an aggregate PASS.
An arrow means “the next cell may independently admit this exact artifact.” It does not claim that bytes have already moved, that the upstream cell granted authority to the downstream one, or that every investigation must pass through one linear pipeline.
flowchart LR
EWB["Evidence Workbench<br/>control plane / transport"]
WA["01 walkaround<br/>session admission"]
CL["02 charterlock<br/>exam admission"]
TM["03 trust-meter<br/>advisory measure"]
PL["04 phaseledger<br/>freshness gate"]
NB["05 nullbench<br/>chance baseline"]
RP["06 RepoPassport<br/>workload invariants"]
GW["07 Greenwash<br/>concrete detector"]
SL["08 smallestlie<br/>authorized falsification"]
UA["09 Unasked<br/>non-certifying research"]
NS["10 NormShift<br/>domain-transfer branch"]
SW["11 StateWeaver<br/>replay / reality"]
TC["12 TomorrowCI Lab<br/>freshness horizon"]
EWB -. "exact pins + native envelopes" .-> WA
WA -->|"frozen session receipt"| CL
CL -->|"admitted exam artifact"| TM
TM -->|"native measure"| PL
PL -->|"selected branch"| NB
PL -->|"selected branch"| RP
PL -->|"selected branch"| GW
PL -->|"selected branch"| SL
PL -->|"selected branch"| UA
PL -->|"selected branch"| NS
NB -->|"preregistered score"| SW
RP -->|"attestation bundle"| SW
GW -->|"finding artifact"| SW
SL -->|"authorized finding"| SW
UA -->|"research evidence"| SW
NS -->|"domain evidence"| SW
SW -->|"replay evidence"| TC
TC -. "native freshness observation" .-> EWB
The workbench registry is allowed to be ahead of execution. A pinned cell whose
interpreter, dependency closure, authorization contract, or containment is not
bound remains fail_closed; catalog visibility is not execution readiness.
Suppose an agent says done, writes its own exam, then makes CI green by weakening a test. The audit chain is:
walkaroundadmits the session only if it entered a frozen contract. Done without entry isBYPASSED. A receipt is unsigned at M4 and is not a verification of the work.charterlockrefuses the exam if the same key wrote and sat it (CHARTER_COLLAPSED) or if the subject'smustset is a proper subset of the charter (CHARTER_NARROWED).trust-meterscores the checkout from local evidence.phaseledgerrefuses a phase advance unless that measure is a freshPASS.greenwashdetects the tampering at the diff level — assertion strength weakened, golden file rewritten, CI runner script quieted. Zero LLM, zero network, byte-identical verdict.RepoPassportre-runs the declared scenario in a sandbox where the workload cannot self-judge, and produces a structured verdict (functional / capability / cleanup).stateweaverreplays the finding against a clean root and demands the patched build blocks the same path; only then doesSYNTHETIC_REPRODUCEDadvance.tomorrowciforecasts the earliest concrete breakage horizon — the moment the patched build's dependencies or runtime stop supporting the verification path.
A finding that survives the chain is publishable. Any stage that fails
must be re-run from the previous stage's clean root. smallestlie is the
authorized adversarial complement: find the smallest lie the repo still
accepts. nullbench is the chance-baseline complement: pre-register the
claim, then score it against chance — never backfill.
The flagship is this spine. It is not a demonstration chamber.
cell-shift (display name
CELL//SHIFT) is the first CI-backed public subject the spine is pointed
at. Spatial agent toy. Not a cancer model. Not a claim that the audit loop
is finished.
| What | Where |
|---|---|
| Frozen journey | JOURNEY.md — npm test → npx tsc -b → npx vite build → npm run audit |
| Walkaround contract | journey/contract.json |
| Declared RepoPassport scenario | repo-passport.yml (public-ci) |
| Host CI | Actions |
| External receipts | artifacts/spine/ |
The subject's own test suite is a lock on chamber geometry. It is not
admission, not an exam charter, and not a trust-meter score. First
host-side spine pass (2026-08-17): walkaround ADMITTED with an empty
organ set; trust-meter 100.0 is vacuous (no Python files); RepoPassport
validate holds and inspect fails closed on 開始艙室.bat. Live
verify is not claimed: the alpha path is a dependency-free Linux
sandbox, and this subject installs npm packages.
Every repo is pre-alpha or pre-release unless otherwise noted. Status badges on each repo are authoritative; this map is secondary.
“Active indexed project repos” means public repositories owned by
taipei49314 that are explicitly listed in
this map and whose GitHub isArchived flag is false. The count excludes this
nelson-stack index and the taipei49314 profile repository. The map currently
contains 31 active projects plus one archived historical prototype; together
with those two meta repositories, that reconciles to all 34 public owner repos.
| Repo | Language | Role in the stack |
|---|---|---|
walkaround |
Python | Session admission kernel. Done without entry is BYPASSED. No release; receipts unsigned; no VERIFIED. |
charterlock |
Python | Exam-admission measurer. Same key writing and sitting the exam is CHARTER_COLLAPSED. No release; independence_claim is always not_claimed. |
trust-meter |
Python | Measure-first scorer. Claims are not trusted until measured. GitHub Latest: v0.2.1. |
phaseledger |
Python | Phase ledger. Advance only on a fresh deterministic measurer PASS. |
nullbench |
Python | Pre-register decisions; score against chance; never backfill. GitHub Latest: v0.7.0. |
RepoPassport |
Go | Workload-side audit: capabilities, cleanup, attestation bundles. Working v1alpha1 slice; 37-row acceptance registry is machine-checked; observer coverage remains incomplete. No release. |
stateweaver |
Python | Verifier-side audit: deterministic replays, oracle verdicts, signed evidence. Source-only pre-alpha; M6–M8 implementation gates exist; trusted Reality proof is not claimed. No release. |
tomorrowci · tomorrowci-lab |
Rust · Python | Time-side audit: dependency / runtime breakage forecasting. Newest lab prerelease: v0.2.0-alpha.1 (CANDIDATE_ONLY_NOT_RELEASE_AUTHORIZED). Neither repo advertises a GitHub “Latest” release: the rejected v0.1.0-grok-session tags are prerelease-marked historical evidence (the tomorrowci one carries orphan history with no common ancestor to master), and tags on both repos are ruleset-protected against update or deletion. |
greenwash |
Python | Diff-level detector for AI agent tampering with verification layers. GitHub Latest: v0.1.42. At the 2026-08-16 reconciliation, main was the v0.1.42 release merge commit, content-identical to that tag. |
unasked |
Python | Evidence-gated repository investigation; non-certifying alpha. GitHub Latest: v0.4.0. Public result remains M0_NOT_DEMONSTRATED. |
smallestlie |
Python | Authorized adversarial harness: smallest lie a repo still accepts. |
branchback |
TypeScript | Local-first decision replay lab — belief-at-the-time vs knowledge-now. v2.0.0. |
constraint-deck |
Python | Session-first authorial constraint deck; measure first; contract over vibes. |
persona-consistency-checker |
Python | Archived historical prototype. PersonaChain experiments for persona drift under adversarial prompts. |
null-city |
TypeScript | Deterministic, partially observable crisis-response sandbox for agent eval. |
NormShift |
Python | Evidence-backed semantic diff for technical standards (M0 local HTML slice). |
| Repo | Language | What it does |
|---|---|---|
evidence-workbench |
Rust | Authority-preserving control plane: exact pins, native envelopes, and fail-closed artifact transport. Not a verifier and not an aggregate judge. |
nelsoncode-ide |
TypeScript / Electron | Off-mainline personal AI coding preview; timeline as backbone; reversible sessions. External security audit remains NO-GO for untrusted use. |
md-brain |
Python | Model-independent continuity runtime for AI memory. |
github-radar |
Python | GitHub research with measured uncertainty; can submit hashed findings to Frontier Atlas. |
receiptradar |
Rust | Local receipt → ledger CLI. No cloud, no account. |
nelson-release-studio |
Python | Music creation, asset management, and release workbench — Windows-first. |
music-lab |
Python | Deterministic local music toolkit; analysis first; no cloud account. |
tw-stock-lab |
Python | Active local TW stock research lab (v0.2.0). Research simulation, not investment advice. |
aurora |
Python | Finds unnamed industries from evidence — deterministic, no LLM at runtime, no stock tips. |
FutureShow-pet |
Python | Personal fork of HKUDS/FutureShow: local desktop pet (Taiwan news + GitHub AI-repo tracker) on Ollama / Qwen. |
universe-explorer |
Python | Epistemically honest science knowledge system — separates known from unknown. |
cell-shift |
TypeScript | Public subject, not flagship. CELL//SHIFT — deterministic 3D tissue chamber. Maintenance-only. Host CI is green; spine verify is not claimed. Not biology. |
why-ledger |
Docs | Why Ledger / 依據本 — justified sovereign decisions (WJSD). Documentation-first. |
editorial-doll-engineering-preview |
TypeScript | Public M0–M3 engineering preview of a deterministic editorial styling engine. |
vibe-oracle |
TypeScript | Explicit anti-evidence foil. Admits the theater; pure vibe, not evidence. |
These are not aspirations. Every repo in the stack has been measured against
them on a clean checkout, and the failures are kept in the repo as evidence
(FAILURES.md, benchmarks/decoy/, CYCLE*-VERDICT.md, etc.).
- Deterministic over LLM-judged. Whenever a verdict can be reached by deterministic analysis of the diff, the artifact, or the replay, it is. LLM judges are advisory.
- Measured, not asserted. Every headline number — false-positive rate, detection count, replay latency — comes out of a reproducible harness. Nothing is hand-typed into a README.
- Honest about out-of-sample. The detectors and verifiers in this stack are reported to perform worse on corpora they have never seen. That degradation is published, not hidden.
- Fail-closed by default. A missing observation is
incomplete. A capability violation outranks a functional pass. A signature without a trust key isunknown, notvalid. - Local-first, zero network. Tools in this stack run on the operator's hardware. The audit trail does not phone home.
- Trust boundaries written into the repo.
SECURITY.md,AGENTS.md,ABUSE_POLICY.md,THREATMODEL.md, and traceability matrices are not optional. They are part of the deliverable. - Verdicts are immutable history. A
VERDICT.mdfile is rewritten only by a later cycle that explicitly supersedes it. Earlier verdicts are kept so the next reader can see the dead ends.
This index repo is itself pre-release. Specifically:
- The grouping above is the author's current model of how the repos relate. It is open to revision if a repo's own README claims a different role.
- "Primary repo" / "supporting repo" in the audit table reflects what is shipped today, not what is intended. Future repos may swap the responsibilities.
- No repo here has been independently audited by a third party, except where
its own README states otherwise
(
nelsoncode-idev0.2.0 has an external audit with a NO-GO verdict, which is kept on the README). - External adoption is still minimal and is not used as a quality claim. Stars, downloads, and README assertions do not replace reproducible evidence.
- Some public releases are marked prerelease and therefore have no GitHub
“Latest” badge even though a release tag exists (
md-brain,github-radar,FutureShow-pet,null-city,tomorrowci,tomorrowci-labv0.2.0-alpha.1). That is intentional honesty, not absence. walkaroundandcharterlockjoined the map after the 2026-08-12 reconciliation. They are pre-alpha admission cells, not later-stage verifiers.cell-shiftwas already on the 2026-08-16 map as a local-first demo. On 2026-08-17 it is classified as a subject: the spine measures it. Pointing the spine at it does not make the spine finished.
What you can rely on: every link above resolves to a real repo, every repo's README states its own status honestly, and the principles above are demonstrably applied — including in the cases where the measurement showed the principle was not yet met.
If you are auditing Nelson's work, the recommended reading order is:
walkaroundandcharterlock— admission before measurement. Neither has a GitHub Release.ADMITTEDis not verified work;CHARTER_SPLITis not two humans.trust-meterandphaseledger— measure first; no phase advance without a freshPASS.trust-meterreachedv0.2.1;phaseledgerhas no GitHub Release yet.nullbench— for chance baselines and pre-registered decision scoring.RepoPassport— for the workload-side invariants and the attestation model.stateweaver— for the verifier-side model (state before chat, reality as final oracle).greenwash— for a concrete worked example of how a single detected failure is reported. GitHub Latest isv0.1.42; at the 2026-08-16 reconciliation,mainwas the release merge commit, content-identical to that tag.tomorrowci·tomorrowci-lab— for how time-horizon forecasts are produced. Neither repo advertises a GitHub “Latest” release; the rejected historical tags are prerelease-marked evidence, and the newest lab prerelease isv0.2.0-alpha.1, still candidate-only.
NelsonCode is deliberately not part of this audit reading order. It remains an active indexed project under local-first tools as an off-mainline personal preview; its optional EWB adapter is read-only catalog access, not an authority handoff.
If you are using Nelson's work, start with the quickstart in the repo that
matches your target question. The audit table above tells you which one.
If you want to see the first public subject the spine measures, open
cell-shift — then come back
here for the measurers.
This index repo is a directory. Pull requests that correct the mapping (a repo's role has changed, a repo has been retired, a new repo has joined the stack) are welcome. Pull requests that soften the honest-status section are not.
This index repo is documentation only. Its own files are Apache-2.0.
Linked project repos are not uniformly licensed. A public repo is not a
grant of rights. Read each repo's LICENSE and README.
| Repo | Current public license signal |
|---|---|
walkaround |
Apache-2.0 |
charterlock |
Apache-2.0 |
trust-meter |
MIT |
phaseledger |
Apache-2.0 |
nullbench |
MIT |
greenwash |
Apache-2.0 |
RepoPassport |
Apache-2.0 |
stateweaver |
Apache-2.0 |
tomorrowci |
Apache-2.0 |
tomorrowci-lab |
LICENSE text is Apache-2.0; GitHub still classifies the repo as NOASSERTION |
unasked |
Publicly readable; copyright reserved; not an open-source license |
smallestlie |
MIT |
null-city |
MIT |
NormShift |
Apache-2.0 |
constraint-deck · branchback |
MIT |
cell-shift |
MIT |
| Local-first tools | Mix of Apache-2.0 and MIT; music-lab has no LICENSE file. See each repo |
Archived or private repos may differ. This table is a reconciliation, not a license grant.

