Skip to content

Agent harness experiment: n=300, 3 arms, pre-registered kill-criterion #137

Description

@nikolay-e

Context

The experiment that decides the primitive thesis. External bar: FastContext (trained 4B–30B explorer) reports up to −60% main-agent tokens and up to +5.5% resolve on Mini-SWE-Agent. Our claim only needs −25–40% tokens at non-worse resolve with ZERO model cost — deterministic, sub-second, no weights.

Design (pre-registered before first run)

  • Harness: mini-swe-agent (FastContext comparability) + one pinned model; n=300 SWE-bench Verified + multilingual slice
  • Arms: (a) baseline; (b) diffctx MCP available, neutral prompt; (c) forced self-diff impact call each patch iteration
  • Metrics: resolve rate, total tokens (main agent), tool calls, wall-clock, opened-file localization recall
  • Stats: paired bootstrap + permutation, seeds 42/43/44; publish traces + per-instance CSVs like paper-v2 artifacts

Tasks

  • Pre-registration doc committed (arms, metrics, gates) before any run
  • Harness integration + checkpoint/resume; per-instance timeout policy
  • Run, analyze, commit artifacts

Kill-criterion (explicit)

  • If token savings <10% AND resolve delta ≤0 across all arms at n=300: the agent-primitive thesis is falsified → pivot focus to CI-review niche; record the result in the paper either way.

Depends on: 01-rrf-fusion-mode or 05-learned-fusion-ltr, 03-mcp-single-tool, 10-serve-daemon-incremental, 11-self-diff-impact-oracle

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions