Skip to content

Repository files navigation

Agent evaluation write-ups

Public-safe measurement notes from a private sealed evaluation lab for local coding agents, constrained generation, RAG routes, skill/instruction docs, and live completion-honesty probes.

Author: Cameron Sanderson

This package is reports, essay, and method language only. It does not include private sealed fixtures, raw agent transcripts, or full lab code.

Live surfaces

Surface URL
Guided tour (GitHub Pages) https://camerontjs-dot.github.io/agent-eval-notes/
This repo https://github.com/camerontjs-dot/agent-eval-notes
Runnable honesty demo https://github.com/camerontjs-dot/verified-done

Tour order: Start → Methodology → Labels → H1 rejection → Essay → Verify ablation → Multi-path → Transfer → RAG → Skills → Try the method → Limitations → Resume bullets.

Local preview:

python3 -m http.server 8080 --directory docs

Public surface map

Surface Role
This repo + Pages tour Methods, locked numbers, harness disclosure, essay
verified-done Runnable honesty demo (tasks + selftest + live evidence summary)
Related apparatus claim-audit-lab, evidence-bundler, apparatus-contracts
Workspace OS MainFrame (public Stage 1b cut; MindGraph nested)

Operator publish logs and internal program notes stay in the private workspace. They are not part of this repo.

What you will find

Path Contents
docs/ GitHub Pages guided tour
reports/ Six self-contained Markdown write-ups (01–06)
essays/ Longform H1 rejection essay
briefs/ One-pagers, harness isolation protocol, public harness descriptions
distribution/ LinkedIn paste body + post log
pdf/ Print-ready PDFs (H1, multi-path, transfer)
numbers-lock.md Locked headline numbers + cross-repo links
METHODOLOGY.md How measurement is defined
LIMITATIONS.md What these notes do not claim
proof-points.md Resume / interview talking points (numbers-locked)
distribution/linkedin-post-draft.md Ready-to-post LinkedIn draft
notes/rag-adversarial-next.md RAG harder-slice plan (no new numbers yet)

Reports (pick one story)

Report One-line story Best for
01 - H1 rejection Better aggregate still not promoted AI safety / eval
02 - Multi-path coding Multi-file is the discriminating case Local coding agents
03 - Task-family transfer Coding winner fails constrained prose Eval design
04 - RAG routes Use-case routes, not one Elo RAG / retrieval
05 - Agent-as-evaluator Skill edge matrix + thin smoke receipts Prompt / skills
06 - Verify tool ablation run_verify flips honesty (Haiku + local coder-14b); abstainer control Agent reliability

Essay: The harness that scored better and was not promoted

Honesty rules (binding)

  1. Say measured / exploratory, not validated or proven, unless a confirmatory protocol is named.
  2. State n with every headline metric.
  3. Say not promoted, not thrown out, for the H1 decision.
  4. Name the stack (local open-weight vs live frontier) when it matters.
  5. DEV routing only unless a report documents a graduated path.
  6. Do not blend coding, prose, RAG, and skill scores into one model quality number.

Related public work

License

MIT

About

Public-safe agent evaluation write-ups: harness gates, multi-path coding screens, task-family transfer, RAG routes, agent-as-evaluator. Guided Pages tour.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages