Bug Hunt Bench: 105 real bugs in two production repos, frontier coding models (GPT-6, Claude, Grok, Gemini, DeepSeek...) find and fix them in their own CLI, graded blind. Live leaderboard + every receipt.
benchmark leaderboard bug-fixing llm-evaluation coding-agents ai-coding claude-code codex-cli llm-benchmark gpt-6 software-engineering-agents
-
Updated
Sep 5, 2026 - JavaScript