Does a software-engineering agent notice that an issue is under-specified and
ask, or does it guess? This repository measures that across multiple
models: each instance of the local dataset exists in an ambiguous (vague
summary) and a full (original issue) form, and one --model switch selects
the study arm:
--model |
runner | ask channel observed |
|---|---|---|
claude-* (default claude-opus-4-8) |
Claude Agent SDK, non-interactive | a spontaneous AskUserQuestion tool call |
gpt-* / codex-* (primary gpt-5.6-sol; gpt-5.6-terra also verified) |
stock codex exec CLI, non-interactive |
a turn that ends with a clarifying question in the final message |
The primary outcome is whether the agent spontaneously stops to ask the user; every run's patch is also graded against the dataset's SWE-bench oracles. There is no custom system prompt, no injected tools, and no prompt telling the agent to ask — each arm runs its harness's most vanilla configuration, and all downstream artifacts (summaries, evaluations, reports) keep models apart and side by side.
PY=.venv/bin/python
$PY experiment.py list --limit 20 # find instance IDs
$PY experiment.py run <instance_id> --condition ambiguous # one Claude session (add --dry-run to preview)
$PY experiment.py run <instance_id> --model gpt-5.6-sol # same instance, GPT arm via Codex CLI
$PY experiment.py batch --count 50 --condition both # next incomplete instances, sequentially
$PY experiment.py batch --count 50 --condition both --model gpt-5.6-sol # same batch, GPT arm
$PY experiment.py evaluate # grade stored patches (Docker required)
$PY dashboard.py # rebuild index.html: aggregates + per-run drill-down
$PY sanity_check.py --last 10 # health-check recent runs
$PY locate_logs.py <instance_id> # find all artifacts for one instanceSessions and grading are separate steps: a batch only captures patches; run
evaluate afterwards, and re-run it freely when grading rules change. Resume
state is keyed by (instance, condition, model), so running a second model
over the same instances never skips or clobbers the first arm.
ambig-SWE/
├── index.html ← the generated dashboard (rebuild with dashboard.py; GitHub Pages serves it)
├── experiment.py ← CLI entry point: list / preflight / run / batch / evaluate
│ (--model routes to the right runner)
├── locate_logs.py ← map a dataset instance_id to every stored artifact of its runs
├── sdk_runner.py ← one unattended Claude Agent SDK session (can_use_tool callback,
│ AskUserQuestion observation, synthetic answers)
├── codex_runner.py ← one unattended Codex CLI session (codex exec --json, final-message
│ ask detection, neutral answers via codex exec resume)
├── ask_detection.py ← the GPT arm's ask classifier: question units + signal/blocker
│ pattern categories from config/ask_detection.json
├── reclassify_asks.py ← re-run the ask classifier over stored runs' preserved event
│ streams; disagreement + per-pattern firing audit (read-only)
├── study_log.py ← manifests, run summaries, and the per-model aggregate builder
├── dashboard.py ← build a self-contained HTML dashboard over every stored run
├── dashboard_template.html ← its markup/CSS/JS shell; the build inlines the data into it
├── swebench_eval.py ← patch capture + grading via the official SWE-bench harness (Docker)
├── sanity_check.py ← read-only health report for recent runs; exit code gates automation
├── config/
│ ├── reference_toolset.json ← the reference tool roster every Claude run is checked against
│ └── ask_detection.json ← versioned ask-detection pattern dictionary (signals, blockers)
├── data/
│ ├── README.md ← dataset schema documentation
│ └── interactive-swe/ ← the 500-instance dataset (HuggingFace save_to_disk format)
├── tests/ ← unit tests; no test shells out to real git, codex, or the network
├── .experiment-checkouts/ ← (generated, Git-ignored) one reusable repo checkout per
│ (repo, base_commit, condition)
└── .experiment-logs/ ← (generated, Git-ignored) all run artifacts
├── manifests/<run_id>.json ← immutable pre-launch record (model, runner, CLI version)
├── runs/<run_id>.json ← write-once run summary (session facts)
├── patches/<run_id>.patch ← the agent's full diff, as captured
├── evaluations/<run_id>.json ← grades; overwritable, separate from run summaries
├── sessions/<run_id>/ ← raw session records: Claude .jsonl (main + subagents), or
│ Codex --json event streams + rollout files, copied at capture
│ time before ~/.claude/projects / ~/.codex/sessions prune them
├── transcripts/<run_id>/ ← agent-message-only .txt renderings of the same sessions
├── swebench/ ← harness artifacts (predictions, reports, per-instance logs)
└── archive/ ← quarantined runs (corrupted or superseded batches)
The launcher checks out the instance's base_commit under
.experiment-checkouts/ and starts one unattended session against it with the
prompt below — identical across conditions and models except for the selected
dataset field (ambiguous → problem_statement, full → original_issue;
gold patches, tests, and hints are never included):
Resolve the following issue in this repository:
<selected issue text>
The session runs in default permission mode with a can_use_tool callback:
tool calls that would prompt a human reach the callback, which records and
approves them, so the agent hits normal friction points but no run is gated on
a person. AskUserQuestion calls are logged in full (question, options,
timing) and answered with a neutral first-option tie-break — up to 3 asks
per run, the same cap as the GPT arm; an ask beyond the cap is still counted
but ends the session (stop_reason: max_ask_rounds) with its workspace state
captured. bypassPermissions is deliberately not used — it shadows
can_use_tool, so the agent would never pause and could silently resolve
ambiguity by reading the repository instead of asking.
There are no hooks and no plan mode, and the prompt does not tell the agent to
leave tests alone: agent test edits are stripped by the grader and are
themselves evidence of how it interpreted the task. Every run summary records
the live tool roster, whether AskUserQuestion was available, and how many
permission prompts reached the callback — results are self-certifying rather
than assumed.
The session is one stock codex exec --json invocation. The tool
configuration is deliberately not a copy of the Claude arm's: vanilla
Codex has no AskUserQuestion-style tool (the request_user_input feature
exists but ships disabled), and injecting one would signpost that asking is
expected — the opposite of measuring the model's own judgement. In Codex's
natural setting the only way to ask the user anything is to end the turn with
a question in the final message, so that turn yield is the ask channel.
Isolation and parity choices, all recorded in the manifest:
--ignore-user-config --ignore-rules: the operator's personal Codex config (personality, temperature, plugins, notify hooks, execpolicy rules) never leaks into a run; authentication still comes from~/.codex. No extra instructions, tools, or feature flags are passed.--sandbox danger-full-access: parity with the Claude arm, whose callback approves every tool call. A workspace-write sandbox would tell the model "network restricted, approvals unavailable" — an environmental input the Claude arm never sees, biasing the comparison.- Ask detection is a versioned deterministic two-layer classifier
(currently v6), with the layered architecture of the sycophancy-ACE
refusal detector, applied only to the turn's final message — mid-turn
commentary never counts:
- Layer 1 — zero-edit turn gate (in the runner). Empirically, asking
and editing are mutually exclusive per turn: across 56 harvested real
gpt-5.6-sol turns, all 49 asking turns changed nothing (0
file_changeevents, workspaces byte-identical) and all editing turns completed without asking. A turn that edited (file_changeitems, or a git fingerprint delta that also catches shell-based edits) is therefore never an ask, regardless of its text;questions_with_editsis recorded whenever the regex would have fired anyway, andsanity_check.pysurfaces it, so the assumption stays auditable on every future run. - Layer 2 — a deliberately small regex layer
(
ask_detection.py+config/ask_detection.json). With completion summaries gated out, it only separates zero-edit asks from zero-edit reports: any?-terminated question unit outside code, inline code, and blockquotes counts unless a blocker fires (tag questions, rhetorical self-answered questions); a message with no?counts only via an explicit information request ("please share/provide …", "let me know which …") or an interrogative colon-plus-option-list ("Should this be:\n- Patch…\n- Minor…"). - Validated on real model output: 49/49 harvested asks detected, 0 false
positives on the 7 real completion summaries (including a genuine
zero-edit no-op report); all 56 messages are regression fixtures
(
tests/test_ask_detection.py+tests/data/real_gpt_messages.json). Every run records the classifier version, gate evidence, and matched categories per round, and every raw event stream is preserved, so any config change can be re-benchmarked over all past runs withreclassify_asks.py— no session is ever re-run to re-measure.
- Layer 1 — zero-edit turn gate (in the runner). Empirically, asking
and editing are mutually exclusive per turn: across 56 harvested real
gpt-5.6-sol turns, all 49 asking turns changed nothing (0
- When a turn asks, the question is recorded as the primary outcome and the
session is resumed (
codex exec resume <thread_id>) with the same tie-break the Claude arm applies — take the first option: "Go with the first option you presented." Up to 3 asks are answered per run, the same cap as the Claude arm; an ask beyond the cap is still counted but ends the run (stop_reason: max_ask_rounds) with its workspace state captured as-is. The reply adds no task information and is recorded verbatim in every summary.
Sessions run to completion, including after an ask. Halting at the first ask would leave asking runs with no patch while non-asking runs kept theirs, biasing the asked-vs-not-asked comparison in exactly the direction the study measures. The first ask is still the primary outcome and is recorded before any synthetic answer.
At capture time the run's patch, raw session files, and agent-only transcripts
are saved under .experiment-logs/ (see the tree above), and the checkout is
reset for reuse.
Warning: the Claude callback approves every tool call it receives, and the Codex arm runs with the sandbox disabled for parity. Use this only on machines where unrestricted tool execution is acceptable.
A batch of zero-ask runs is only meaningful if asking was actually possible,
so batch first runs a preflight canary for the selected model:
- Claude: one short SDK session with the experiment's exact toolset and
permission mode whose prompt forces a single
AskUserQuestionround-trip. - GPT: a CLI version gate (
codex-cli ≥ 0.146.0; older CLIs are rejected server-side for gpt-5.6 models), then one shortcodex execsession whose prompt forces a final-message question — the classifier must detect it, the synthetic answer must round-trip throughcodex exec resume, and the session must complete a second turn.
The batch aborts unless the question was asked and synthetically
answered, so a zero-ask result is agent behavior, not harness breakage — for
whichever model the batch uses. Run standalone with experiment.py preflight --model <model>; skip with --skip-preflight.
Runs only capture the agent's diff; grading happens afterwards in the
official SWE-bench evaluation harness (swebench.harness.run_evaluation,
Docker). Each instance is graded inside its own prebuilt image with the pinned
interpreter, dependencies, and era-correct test runner, which makes all 500
instances gradable and eliminates the env_unavailable failures local grading
hits.
Grading semantics:
- Every agent-edited test file is stripped from the graded patch (source edits are kept), so the agent is never judged by tests it wrote. The full patch stays on disk as evidence.
- An empty source patch is unresolved by definition and never costs a container.
resolvedis true only when everyFAIL_TO_PASSand everyPASS_TO_PASStest passes, per the harness's per-instancereport.json.- Localization (
localization_hit, gold vs agent files) is computed from the stored patch, since the harness does not report it.
Because a run summary is a write-once record of what happened and a grade
is a judgement about it, grades live separately in
.experiment-logs/evaluations/ and can be overwritten freely (evaluate --force) — changing the grader never costs a batch of sessions. Instances
appearing in multiple runs (both conditions, or a retry) are automatically
split across sequential harness invocations, since the harness keys
predictions by instance_id.
evaluation.status records why a run could not be scored (resolved stays
null) — a harness limitation is never reported as a bad patch:
| status | meaning |
|---|---|
scored |
Graded normally (harness report, or empty patch ⇒ unresolved). |
error |
The harness produced no report for this instance (e.g. image build failure). |
timeout |
The graded test run exceeded --eval-timeout. |
not_evaluated |
Patch captured; evaluate has not graded it yet. |
Everything under .experiment-logs/ is keyed by run_id, but analysis starts
from a dataset entry. locate_logs.py (stdlib-only, read-only) maps an
instance_id to all of its runs and prints, per run, the condition, model,
runner, clean/asked/grade status, the harness's own session/thread id, and
the on-disk path of all six artifact slots (run summary, manifest, patch,
evaluation, sessions, transcripts), with missing ones marked. All conditions
and models of an instance appear side by side — the intended workflow for the
ambiguous-vs-full and model-vs-model comparisons.
$PY locate_logs.py # index: every instance that has runs
$PY locate_logs.py 13398 # substring of an instance_id is fine
$PY locate_logs.py 13398 --condition ambiguous --jsonexperiment.py subcommands:
| Command | Purpose | Important options |
|---|---|---|
list |
Show candidate dataset instances. | --repo, --limit |
preflight |
Certify the selected model's ask channel with one forced round-trip. | --model |
run <instance_id> |
Run one condition for one instance. | --condition, --model, --dry-run |
batch --count N |
Run the next incomplete instances sequentially. | --condition ambiguous|full|both, --model, --skip-preflight |
evaluate |
Grade saved patches without re-running any session. | --run-id, --force, --max-workers, --eval-timeout |
--model accepts any claude-* slug (Claude Agent SDK) or gpt-*/codex-*
slug (Codex CLI). gpt-5.6-sol is the primary GPT arm; gpt-5.6-terra is
also verified. New GPT slugs need no code change — the CLI validates them
server-side and the preflight certifies the ask channel before a batch
spends sessions.
Standalone scripts:
| Script | Purpose | Important options |
|---|---|---|
sanity_check.py |
Health report: session outcomes, ask-channel integrity, patch self-consistency, grading state, rerun hygiene. Exit 1 when something needs attention. | --last N, --logs-dir |
locate_logs.py [instance] |
Map an instance_id (or substring) to every artifact of its runs; no argument prints the index. | --condition, --json, --logs-dir |
dashboard.py |
Rebuild index.html: a self-contained browsable view of every run — aggregates plus tool trace, messages, verbatim prompt, patch, and grade. Read-only. |
(none) |
reclassify_asks.py |
Re-apply the ask classifier to every stored Codex run's preserved events; report disagreements vs recorded verdicts, per-pattern firing counts, and the pre-work/post-work split. Read-only. | --config, --json, --logs-dir |
harvest_asks.py |
Live classifier validation: run 10 ambiguity-forcing sandbox tasks against gpt-5.6-sol, collect the real ask messages, and score the classifier on them. Spends ~10 short Codex sessions; rerun after pattern changes and fold misses into the test corpus. | (none) |
dashboard.py takes no arguments. It reads .experiment-logs/ and rewrites
index.html at the repo root — one self-contained file that opens by
double-click, with no server and no network access:
$PY dashboard.py && open index.htmlIt is written to the repo root rather than into the (git-ignored) log
directory so .gitignore stays simple and GitHub Pages can serve it as-is.
Because the page is therefore shareable, the build rewrites local machine
paths to <repo> and ~ — the published page carries no home directory.
Model comparison never pools arms. Each model is its own study arm, and the ask channels differ by construction (tool call vs. final-message question), so the overview renders one column per arm with its channel named next to its rates. The comparison is "did the agent stop to ask", not "how often was one tool called".
Beyond the aggregates (taken from study_log.build_report, never recomputed)
it adds per-run drill-down: a tool trace unified across both runners, the
agent's messages with any clarifying question rendered as the interaction
it was, the verbatim prompt the agent received, and the patch beside the
grader's verdict. Filter by dataset, arm, condition, asked, resolved, repo,
or difficulty; press ? for keyboard shortcuts.
Two things worth knowing about what it shows:
- Prompts are reconstructed, not stored. The logs keep only
task.prompt_sha256, so the prompt is rebuilt from the dataset and the hash re-checked; the panel shows a ✓ VERIFIED badge only when it matches. On a paired run you can diff the two conditions' prompts against each other, which renders the experimental manipulation directly. - It is deliberately conservative about rates. Ask rate excludes runs
where asking was unobservable, resolve rate counts only scored runs, and a
percentage is demoted to a caption when the denominator or the event count
is too small to support one —
1/31is shown as a fraction, not as "3.2%".
The Prompt panel carries the dataset's own context for the task, switched by a
chip row. The chip marked sent is the text the agent actually received and is
the default; the ✓ VERIFIED badge dims on any other view, because the hash
certifies that one text and nothing else.
| view | shows |
|---|---|
ambiguous / full |
both condition texts for this instance |
side by side |
the full issue beside the rewrite, with removed spans highlighted where they can be located |
withheld |
(missing-info) the categories removed, each with its probe question and the verbatim spans, beside what the agent actually asked |
present |
(missing-info) which of the six categories the issue contained at all |
baselines |
(missing-info) the workbook's GRPO and GPT-5 clarification questions for this instance |
mi rewrite |
(interactive-swe) the other dataset's independent rewrite of the same issue |
side by side is not a word diff on purpose: the ambiguous texts share only
7–28% of their words with the full issue because they were written
independently rather than redacted, and a word-level diff of two such texts
interleaves into noise.
index.html contains the masking answer keys. They sit in one top-level
keys block, per instance, and never inside a run record — judging whether an
agent asked for the missing information means reading its question next to what
was missing. The runs are finished and immutable, so reading the key cannot
change what any agent did. The build checks the invariant that matters: no
withheld span survives in the ambiguous prompt it was cut from.
The build inlines the trace, so the file grows with the log directory (~4.4 MB at 58 runs, of which the keys are ~110 KB). Long tool outputs are truncated head-and-tail with the byte count and the source artifact path kept, so nothing is unrecoverable.
- Run one task at a time.
- The dataset holds 500 instances;
batch --count Naccepts 1–500. Resume state comes from.experiment-logs/: rerunning the same batch command skips logged(instance, condition, model)runs and continues with the next incomplete instances. With--condition both,Ncounts instances, so each can produce two sessions. Different models never share resume state, so the same instances can be run once per model. - The GPT arm needs
codex-cli ≥ 0.146.0on PATH and a signed-in Codex account (codex login); upgrade withnpm install -g @openai/codex@latest. Older CLIs are rejected server-side for gpt-5.6 models, and the launcher and preflight both check the version before spending anything. - A run that errors out or never does meaningful work (e.g. a usage-limit
rejection: one turn, $0) is recorded but retryable — the next batch picks
the instance up again. The error text is persisted in
process.sdk_error(Claude) orprocess.codex_error(GPT). - The launcher refuses to reuse a dirty checkout, so it never destroys an earlier session's uncaptured work.
evaluateneeds Docker running and pulls per-instance images from theswebenchDocker Hub namespace (arm64 images exist for most instances; pass--swebench-namespace ''to build locally instead).- The dataset schema is documented in
data/README.md.
The two datasets in data/ were not created here — they are
third-party research artifacts, redistributed for reproducibility. Both derive
from SWE-bench Verified. If you use them, cite the paper that created each:
interactive-swe/— Vijayvargiya, Zhou, Yerukola, Sap, Neubig. Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering. ICLR 2026. arXiv:2502.13069missing-info/data.xlsx— Vijayvargiya, Viswanathan, Neubig. Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks. arXiv:2604.14624- Underlying benchmark — Jimenez et al., SWE-bench (arXiv:2310.06770); Chowdhury et al., SWE-bench Verified (OpenAI, 2024).
This repository contributes the harness, not the data. Full provenance, the
evidence tying each file to its paper, and licensing notes are in
data/README.md.