Repository navigation
Releases: Kendr-AI/LLM-Benchmark
Release list
LLM Benchmark Protocol v1.0.3 - KGBP 1.0 Research Release
LLM Benchmark Protocol v1.0.3
Release date: 2026-08-08
Researcher: Dr. Prashant Kumar Dey
Project steward: Kendr
Purpose of this release
Version 1.0.3 publishes the dated Kendr current-frontier evaluation and its
separate preview companion. It adds release-grade human-readable handouts,
machine-readable aggregates, provenance and privacy checks, and GitHub release
assets without changing the KGBP 1.0 protocol profile.
This is a research publication. It is not a universal model ranking,
certification, declaration of global acceptance, or standards-body decision.
Current-frontier GA publication
The frozen callable-subset matrix used 15 objective questions: three from each
of five LiveBench task strata, one generation per endpoint-question cell, and
score-weighted operational goodput as its primary metric. Five generally
available candidates were ranked; GPT-5.5 was an explicitly declared,
unranked baseline.
| Rank | GA candidate | Kendr endpoint | Goodput | 95% interval | Availability |
|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | kc-gpt-5.6-sol |
79.36% | 60.00%–95.60% | 86.67% |
| 2 | Grok 4.5 | kc-grok-4.5 |
72.00% | 46.67%–92.00% | 73.33% |
| 3 | Claude Opus 5 | kc-claude-opus-5 |
68.91% | 46.71%–88.89% | 80.00% |
| 4 | Gemini 3.6 Flash | kc-google-gemini-3-6-flash |
33.11% | 12.00%–57.33% | 100.00% |
| 5 | DeepSeek V4 Flash 0731 | kc-ollama-deepseek-v4-flash-0731 |
13.33% | 0.00%–33.33% | 13.33% |
The unranked GPT-5.5 baseline recorded 68.89% goodput and 73.33%
availability. Six other GA targets remain explicit N/A entries because their
frozen identity, access, maturity, or preflight gate was not satisfied. N/A is
not a zero-capability score.
Six scored endpoints produced 15 paired comparisons. Two separated after Holm
family-wise correction: GPT-5.6 Sol versus DeepSeek V4 Flash 0731, and Claude
Opus 5 versus DeepSeek V4 Flash 0731. GPT-5.6 Sol's point estimate exceeded the
GPT-5.5 baseline by 10.4667 percentage points, but the paired interval was
−2.6667 to +28.7333 points (p = 0.5; Holm-adjusted p = 1.0). This run
therefore established neither a GPT-5.6/GPT-5.5 difference nor practical
equivalence.
Authoritative materials:
- execution handout;
- market and Kendr coverage register;
- GPT-5.6 versus GPT-5.5 analysis;
- GA aggregate JSON,
CSV,
generated Markdown,
and bundle checksum manifest.
GA matrix ID:
20260808T070202Z-frontier-market-kendr-20260808-cfec3672.
Preview companion publication
Preview and limited-access configurations were not pooled into the GA rank
sequence. A separately frozen companion matrix used the same 15 question IDs
and sample hash.
Gemini 3.1 Pro Preview (kc-gemini-3.1-pro-preview) completed 15/15 cells with
39.78% operational goodput, a 16.67% to 65.11% 95% interval,
39.78% conditional quality, and 100.00% availability.
It receives no ordinal rank because it is the companion's only scored
endpoint. No pairwise test exists for a one-endpoint comparison family, and
the result must not be read as a rank against the GA candidates. Qwen 3.8 Max
Preview (kc-qwen3.8-max-preview) remains N/A because the dedicated paid Model
Studio Token Plan endpoint and credential were not configured.
Authoritative companion materials:
Companion matrix ID:
20260808T083825Z-frontier-preview-kendr-20260808-d910f1e1.
Version and provenance boundaries
Publication version and execution-software version are intentionally separate:
v1.0.3is the software and publication release described here;- the GA current-frontier matrix records execution software
1.0.2; - the preview companion records execution software
1.0.2; - the earlier 35-endpoint catalog pilot records execution software
1.0.0.
The version bump does not retroactively relabel any execution. The GA and
preview-companion JSON, CSV, generated Markdown, and nested SHA256SUMS files
remain byte-identical to their frozen publication inputs.
Privacy and integrity boundary
The public frontier bundles contain aggregate metrics, public configuration
identity, explicit N/A states, source hashes, and bounded scientific claims.
They exclude raw prompts, raw responses, provider request identifiers,
provider error messages, credentials, and machine-local paths.
The offline release verifier pins the execution versions, matrix IDs, scope,
row identity and order, scoring content, no-rank companion treatment,
provenance hashes, privacy declarations, CSV/Markdown consistency, and bundle
checksums. Checksums detect byte drift; they are not a substitute for release
attestation or independent replication.
The LiveBench adapter also has a bounded grading-recovery path: when every
planned answer is current but a successful answer lacks a judgment, it retries
only that missing local judgment once with serial grading. It never replays a
provider inference and still fails closed if grading remains incomplete.
Tagged release assets
The GitHub tag workflow retains the existing package, white paper, catalog
pilot, protocol-audit, SBOM, brand, and release-wide checksum assets. Version
1.0.3 additionally includes:
- the frontier execution, coverage, and GPT-5.6/GPT-5.5 handouts;
- the GA JSON, CSV, and generated Markdown;
- the preview-companion JSON, CSV, and generated Markdown;
- both nested bundle manifests under unique GA and preview-companion filenames.
The workflow then generates an outer release-wide SHA256SUMS and provenance
attestations for the assembled assets.
Limitations and permitted claim
Both frontier runs are small, English-oriented, one-generation,
endpoint-as-served snapshots. Preview behavior can change, provider defaults
were not fully normalized, rank intervals were not estimated, and the study
does not cover the multilingual, multimodal, safety, repeated-generation,
multi-region, load, independent-review, or external-replication requirements
needed for a broad global claim.
Permitted summary:
In the dated 2026-08-08 Kendr callable-subset matrix, GPT-5.6 Sol had the
highest operational-goodput point estimate among five scored GA candidates;
most pairwise comparisons remained unresolved after correction. A separate
one-endpoint preview companion measured Gemini 3.1 Pro Preview at 39.78%
goodput with 100% availability and assigned no rank. Qwen 3.8 Max Preview
remained N/A.
Citation
Use CITATION.cff, cite release v1.0.3, identify
Dr. Prashant Kumar Dey as the researcher and Kendr as project steward,
and include the exact matrix ID for every reused result set.
LLM Benchmark Protocol v1.0.2 - KGBP 1.0 Research Release
LLM Benchmark Protocol v1.0.2
Release date: 2026-08-08
Researcher: Dr. Prashant Kumar Dey
Project steward: Kendr
Purpose of this patch
Version 1.0.2 is a publication-design correction. The technical white paper
now uses the same core color system as Kendr's public design language:
- Ink
#151412 - Saffron
#E2712A - Paper
#FAF8F4 - Warm grey
#8A8378
The PDF adapts those colors for a long technical report while preserving the
Kendr mark's geometry, colors, and clear space.
Publication design
- The cover and running header use Ink, with Saffron rails and rules and Paper
typography. - Body pages use a warm Paper canvas, Ink text, Saffron structural accents,
and dark table headers. - Charts use Saffron bars with a darker Saffron outline.
- Quotes and rank highlights use a pale Saffron tint; code blocks and alternate
table rows use a neutral warm tint. - Small orange text uses derived deep Saffron
#9A5022, and secondary small
text uses derived muted Ink#615C54. These choices avoid the insufficient
contrast of raw Saffron or Warm grey on Paper. - Links remain distinguishable without relying on color alone because they are
both underlined and rendered in deep Saffron. - The previous blue, cyan, and navy drawing colors are absent from the PDF.
No gradients, shadows, logo recoloring, logo rotation, or substitute wordmark
were introduced.
Verification evidence
- All 50 A4 pages were rasterized at 150 dpi and visually inspected, including
the cover, contents, chart, dense result tables, appendices, and references. - The PDF contains the canonical Ink, Saffron, Paper, and Warm-grey drawing
colors and no legacy blue, cyan, or navy drawing operations. - The cover embeds the canonical Kendr logo and credits Dr. Prashant Kumar
Dey. - PDF SHA-256:
1e73cf2b629168e49fe837134ab1103b1218eaa5f49fb4a0f1ba3b00deef41f1 - Resolved Markdown SHA-256:
5eb99199a3cef536f63d0ca03a43195d01ed751b7bed5d742e10ac4b6a102113
The resolved Markdown digest is unchanged from v1.0.1, confirming that this
patch changes presentation rather than scientific content.
Frozen benchmark provenance
No provider calls were rerun. No prompts, responses, judgments, model labels,
endpoint identities, scores, intervals, costs, availability values, ranking
rows, pairwise tests, or conclusions changed.
The pilot remains a narrow, English-oriented, 15-item, one-generation study
with zero Holm-significant pairwise differences among 595 comparisons. Its
execution-software version remains 1.0.0; v1.0.2 is the publication release.
Historical release policy
The v1.0.0 and v1.0.1 tags and their attested assets remain unchanged.
This correction uses a new tag and new artifacts instead of moving an existing
tag or overwriting a published PDF, package, checksum, SBOM, or attestation.
Citation
Use CITATION.cff, cite release v1.0.2, identify
Dr. Prashant Kumar Dey as the researcher, and include the exact benchmark
matrix identifier when citing the frozen catalog pilot.
LLM Benchmark Protocol v1.0.1 - KGBP 1.0 Research Release
LLM Benchmark Protocol v1.0.1
Release date: 2026-08-08
Researcher: Dr. Prashant Kumar Dey
Project steward: Kendr
Purpose of this patch
Version 1.0.1 is a branding and attribution correction. It standardizes the
organization name as Kendr, adds the supplied Kendr mark to the public
project materials, and identifies Dr. Prashant Kumar Dey as the researcher
working on LLM Benchmark Protocol.
The GitHub owner slug remains Kendr-AI inside repository URLs because it is a
technical address, not the public organization name.
What changed
- Added the canonical 512 x 512 transparent Kendr mark under
assets/brand/. - Added the mark and researcher attribution to the README.
- Added the mark to the white-paper cover on Kendr Paper with the prescribed
clear space, without recoloring, rotation, gradients, or shadows. - Updated the white-paper source, cover byline, PDF author metadata, PDF creator
metadata, citation file, package metadata, governance, protocol card, data
attribution, copyright notice, and project authorship page. - Added automated checks for the canonical logo, image dimensions, researcher
attribution, legacy display-brand text, PDF cover image, and PDF metadata. - Included the brand assets in source distributions and GitHub release assets.
- Changed release automation to reject attempts to replace an existing release.
Frozen benchmark provenance
No provider calls were rerun. No prompts, responses, judgments, model labels,
endpoint identities, scores, intervals, costs, availability values, ranking
rows, pairwise tests, or conclusions changed.
The published pilot bundle continues to record execution software version
1.0.0. That is intentional: v1.0.1 is the publication/correction release,
not a retroactive claim that the 2026-08-07 experiment ran newer software.
The scientific claim boundary is unchanged. The pilot remains a narrow,
English-oriented, 15-item, one-generation study with zero Holm-significant
pairwise differences among 595 comparisons.
Historical release policy
The v1.0.0 tag and its attested assets remain unchanged as historical
evidence. This patch uses a new tag and new artifacts instead of moving the old
tag or overwriting its PDF, packages, data, checksums, SBOM, or attestations.
Citation
Use CITATION.cff, cite release v1.0.1, identify
Dr. Prashant Kumar Dey as the researcher, and include the exact benchmark
matrix identifier when citing the frozen catalog pilot.
LLM Benchmark Protocol v1.0.0 - KGBP 1.0 Research Release
LLM Benchmark Protocol v1.0.0
Release date: 2026-08-08
LLM Benchmark Protocol v1.0.0 is the first versioned public release of the
KGBP 1.0 specification, its Python reference harness, machine-readable evidence
contracts, and a privacy-reviewed catalog-pilot result bundle.
This is a release of evaluation methodology and software. It is not a claim
that KGBP is an accredited standard, that the included pilot is globally
representative, or that its descriptive row order is a resolved universal
ranking of language models.
Highlights
- Claim-first protocol design with ten non-compensatory dimensions and separate
design, execution, evidence, and publication gates. - Explicit system types for fixed endpoints, routed systems, research systems,
agents, ensembles, multimodal systems, and applications. - Frozen sampling, deterministic interleaving, repeated-generation planning,
failure-aware scoring, hierarchical bootstrap estimates, paired comparisons,
multiplicity control, practical-equivalence decisions, operational goodput,
and router counterfactuals. - Versioned JSON Schema 2020-12 contracts covering systems, items, schedules,
attempts, answers, judgments, observations, scorecards, and evidence
manifests. - Resumable LiveBench execution that preserves the first captured provider
trial and supports offline finalization without silently replaying inference. - Public
llm-benchmark-*commands with the historicalkendr-*aliases and
kendr_benchimport namespace retained for compatibility. - Offline examples, release verification, CI, governance, contribution,
security, data-license, citation, and adoption materials.
Start here
- Quick start
- Protocol card
- Adoption guide
- Schema reference
- Provider-adapter guide
- Full protocol
- Technical white paper
Install a source checkout with Python 3.11 or newer:
python -m venv .venv
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"
python scripts/verify_release.py --expected-version 1.0.0
python -m pytestProvider-backed benchmark commands can incur API charges. The quick start
begins with offline audit and toy-scoring commands and requires explicit paid-run
confirmation before provider inference.
Public catalog pilot
The bundled ranking handout reports a frozen campaign
against the Kendr API catalog snapshot captured on 2026-08-07:
- 37 catalog entries assessed;
- 35 text-compatible endpoints ranked on endpoint-as-served operational
goodput; - two non-text entries reported as not applicable, rather than forced into the
text ranking; - 15 selected LiveBench questions across five task categories;
- one generation per endpoint-question cell;
- 595 paired endpoint comparisons and zero Holm-adjusted rejections.
The machine-readable exports are:
Raw prompts, model responses, provider request identifiers, credentials, and
local paths are not included in the public bundle. See the
evidence-bundle guide and data license
before publishing additional run artifacts.
Scientific interpretation
The catalog campaign is a descriptive pilot. It does not satisfy KGBP global
publication gates because:
- all selected questions are English-language items dated in 2024;
- the sample is too small for narrow inference;
- generation-to-generation variation was not estimated;
- no multi-region load phase was run;
- no private rolling holdout was used;
- independent review and external replication are absent.
No pairwise comparison survived the declared Holm family-wise correction. A
failure to reject is not proof of equality, while a zero endpoint-as-served
score can reflect availability or capability negotiation rather than intrinsic
model incapability. Use the results interpretation guide
and statistical analysis plan before citing or
extending the pilot.
The strict protocol audit scores the reference design above 9 on all ten
dimensions. That design score evaluates whether controls are specified; it does
not replace execution evidence, conformance review, external replication,
certification, or standards-body adoption.
Compatibility
- Python 3.11, 3.12, and 3.13 are the declared test matrix.
- The distribution name is
llm-benchmark-protocol. - The import namespace remains
kendr_bench. - New command names are
llm-benchmark,llm-benchmark-livebench,
llm-benchmark-matrix, andllm-benchmark-protocol. - Existing
kendr-bench,kendr-livebench,kendr-benchmark-matrix, and
kendr-protocolcommands remain aliases in this release. - Protocol identity, software version, and benchmark-round identity are
independently versioned. This software release implements the KGBP 1.0
profile and publishes the immutable matrix ID named in the ranking bundle.
Known limitations
- The bundled public pilot is not a freshness, multilingual, multimodal,
longitudinal, safety, or production-load study. - Provider aliases and routed candidate sets can change after capture; results
apply only to the recorded endpoint identities and observation window. - Some costs are lower bounds when failed calls lacked complete usage
telemetry. - The reference harness cannot establish legal compliance, social benefit, or
field effectiveness without the complementary evidence described by the
protocol. - The optional Kendr dependency is pinned to a Git revision and therefore
requires Git when that extra is installed.
Governance and security
Protocol changes follow governance, including public change
control and the appeals and corrections process.
Security issues should follow the private reporting instructions in
SECURITY.md, not a public issue containing exploit or
credential details. Threat assumptions are documented in the
threat model.
Release integrity
Maintainers must follow RELEASING.md and run:
python scripts/verify_release.py --expected-version 1.0.0
python -m pytest
python -m buildFor the tag-triggered workflow, python scripts/verify_release.py --tag v1.0.0 additionally checks exact tag-to-package version alignment. The
verifier is offline and does not create, push, tag, or publish anything.
After the annotated tag is pushed, the release workflow repeats the gates,
audits the isolated runtime dependency set, generates a CycloneDX SBOM and
SHA-256 manifest, attests build provenance, and publishes the GitHub release.
Citation
Use CITATION.cff, cite software version 1.0.0 and protocol
profile KGBP 1.0, and include the exact benchmark-round identifier when citing
the catalog results.
The complete change list is in CHANGELOG.md.