Skip to content

docs(benchmarks): Add production and Kimi results - #517

Merged
gricha merged 1 commit into
mainfrom
feat/benchmark-comparisons
Aug 22, 2026
Merged

docs(benchmarks): Add production and Kimi results#517
gricha merged 1 commit into
mainfrom
feat/benchmark-comparisons

Conversation

@gricha

@gricha gricha commented Aug 22, 2026

Copy link
Copy Markdown
Member

Publish four standardized security-review corpus results for the current
production model stack, Kimi K2.6, and Kimi K3 at low and high effort. The
production row uses Warden 0.46.2 on Pi with Grok 4.5 high for scanning and
GPT 5.6 Luna high for auxiliary calls; it records 37 of 86 known findings.

Add a source-linked production profile and require stable benchmark data to
contain exact matches for both model lanes. This keeps the matrix from
labeling a stale run or a primary-only match as current production. The
committed result files are sanitized summaries; raw JSONL remains withheld
pending sensitive-data review.

All six production shards completed without repair. The final auxiliary merge
dropped two exact scan-time corpus matches, so the published 37/86 is
intentionally the end-to-end system score rather than primary-model recall.
The Kimi rows preserve the completed and repaired shard provenance needed to
interpret their costs.

Validation passed with pnpm lint, pnpm build, pnpm test,
pnpm --filter warden-docs check, and git diff --check.

Record standardized production, Kimi K2.6, and Kimi K3 effort runs. Track exact production model lanes so published comparisons do not inherit stale labels.

Co-Authored-By: GPT-5.6 Sol <noreply@anthropic.com>

@sentry-junior sentry-junior Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me. The production profile matches the pinned getsentry/.github warden.toml, the new validator correctly requires a stable row for both model lanes, and the published numbers line up with the committed result files (including the July vs current overlap and the Kimi turn/token figures). Nice call documenting the aux-merge drop so 37/86 reads as an end-to-end system score.

@gricha
gricha marked this pull request as ready for review August 22, 2026 15:23
@gricha
gricha merged commit 543c603 into main Aug 22, 2026
22 checks passed
@gricha
gricha deleted the feat/benchmark-comparisons branch August 22, 2026 16:46
@gricha gricha changed the title docs(benchmarks): Add production and Kimi results docs(benchmarks): Add production, Kimi, and Grok 4.6 results Aug 23, 2026
@gricha gricha changed the title docs(benchmarks): Add production, Kimi, and Grok 4.6 results docs(benchmarks): Add production and Kimi results Aug 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant