Skip to content

Code-role v0.3.1: Real project case studies

Choose a tag to compare

@Magicray1217 Magicray1217 released this 31 Jul 01:50

Two projects, two failure modes, one milestone-control loop

This release adds sanitized bilingual case studies from two private AI engineering projects:

  • DeepBrain: strong local evidence included 1,750 passing unit tests, S50 at 50/50, LongMemEval-S at 499/500, and 100/100 grounded source joins. Independent Evaluation still held milestone completion and Reviewer routing at 0 because fair comparison, raw benchmark reruns, repair proof, clean reproducibility, and production cost/SLO evidence were incomplete.
  • Leaper Agent: the Project Manager rejected a professional-looking evaluation baseline before Engineering started because task artifacts, holdout isolation, grader execution, runtime conditions, and integrity evidence were not yet executable. The correction replaced declarations with concrete task, holdout, calibration, command, and integrity artifacts.

Neither case is presented as a completed product milestone. The case studies show the core Code-role value: preserve proven progress while preventing unsupported claims and invalid handoffs.

Included

  • README case-study proof points.
  • Bilingual DeepBrain and Leaper Agent case studies.
  • A ready-to-publish two-case technical launch story.
  • Updated launch copy and public disclosure boundaries.

Verification

  • 69 repository tests passed.
  • All relative Markdown links resolve.
  • Private source code, repository links, local paths, customer data, and implementation details are excluded.