Skip to content

Releases: Kendr-AI/LLM-Benchmark

LLM Benchmark Protocol v1.0.3 - KGBP 1.0 Research Release

Choose a tag to compare

@github-actions github-actions released this 08 Aug 09:21

LLM Benchmark Protocol v1.0.3

Release date: 2026-08-08

Researcher: Dr. Prashant Kumar Dey

Project steward: Kendr

Purpose of this release

Version 1.0.3 publishes the dated Kendr current-frontier evaluation and its
separate preview companion. It adds release-grade human-readable handouts,
machine-readable aggregates, provenance and privacy checks, and GitHub release
assets without changing the KGBP 1.0 protocol profile.

This is a research publication. It is not a universal model ranking,
certification, declaration of global acceptance, or standards-body decision.

Current-frontier GA publication

The frozen callable-subset matrix used 15 objective questions: three from each
of five LiveBench task strata, one generation per endpoint-question cell, and
score-weighted operational goodput as its primary metric. Five generally
available candidates were ranked; GPT-5.5 was an explicitly declared,
unranked baseline.

Rank GA candidate Kendr endpoint Goodput 95% interval Availability
1 GPT-5.6 Sol kc-gpt-5.6-sol 79.36% 60.00%–95.60% 86.67%
2 Grok 4.5 kc-grok-4.5 72.00% 46.67%–92.00% 73.33%
3 Claude Opus 5 kc-claude-opus-5 68.91% 46.71%–88.89% 80.00%
4 Gemini 3.6 Flash kc-google-gemini-3-6-flash 33.11% 12.00%–57.33% 100.00%
5 DeepSeek V4 Flash 0731 kc-ollama-deepseek-v4-flash-0731 13.33% 0.00%–33.33% 13.33%

The unranked GPT-5.5 baseline recorded 68.89% goodput and 73.33%
availability. Six other GA targets remain explicit N/A entries because their
frozen identity, access, maturity, or preflight gate was not satisfied. N/A is
not a zero-capability score.

Six scored endpoints produced 15 paired comparisons. Two separated after Holm
family-wise correction: GPT-5.6 Sol versus DeepSeek V4 Flash 0731, and Claude
Opus 5 versus DeepSeek V4 Flash 0731. GPT-5.6 Sol's point estimate exceeded the
GPT-5.5 baseline by 10.4667 percentage points, but the paired interval was
−2.6667 to +28.7333 points (p = 0.5; Holm-adjusted p = 1.0). This run
therefore established neither a GPT-5.6/GPT-5.5 difference nor practical
equivalence.

Authoritative materials:

GA matrix ID:
20260808T070202Z-frontier-market-kendr-20260808-cfec3672.

Preview companion publication

Preview and limited-access configurations were not pooled into the GA rank
sequence. A separately frozen companion matrix used the same 15 question IDs
and sample hash.

Gemini 3.1 Pro Preview (kc-gemini-3.1-pro-preview) completed 15/15 cells with
39.78% operational goodput, a 16.67% to 65.11% 95% interval,
39.78% conditional quality, and 100.00% availability.

It receives no ordinal rank because it is the companion's only scored
endpoint. No pairwise test exists for a one-endpoint comparison family, and
the result must not be read as a rank against the GA candidates. Qwen 3.8 Max
Preview (kc-qwen3.8-max-preview) remains N/A because the dedicated paid Model
Studio Token Plan endpoint and credential were not configured.

Authoritative companion materials:

Companion matrix ID:
20260808T083825Z-frontier-preview-kendr-20260808-d910f1e1.

Version and provenance boundaries

Publication version and execution-software version are intentionally separate:

  • v1.0.3 is the software and publication release described here;
  • the GA current-frontier matrix records execution software 1.0.2;
  • the preview companion records execution software 1.0.2;
  • the earlier 35-endpoint catalog pilot records execution software 1.0.0.

The version bump does not retroactively relabel any execution. The GA and
preview-companion JSON, CSV, generated Markdown, and nested SHA256SUMS files
remain byte-identical to their frozen publication inputs.

Privacy and integrity boundary

The public frontier bundles contain aggregate metrics, public configuration
identity, explicit N/A states, source hashes, and bounded scientific claims.
They exclude raw prompts, raw responses, provider request identifiers,
provider error messages, credentials, and machine-local paths.

The offline release verifier pins the execution versions, matrix IDs, scope,
row identity and order, scoring content, no-rank companion treatment,
provenance hashes, privacy declarations, CSV/Markdown consistency, and bundle
checksums. Checksums detect byte drift; they are not a substitute for release
attestation or independent replication.

The LiveBench adapter also has a bounded grading-recovery path: when every
planned answer is current but a successful answer lacks a judgment, it retries
only that missing local judgment once with serial grading. It never replays a
provider inference and still fails closed if grading remains incomplete.

Tagged release assets

The GitHub tag workflow retains the existing package, white paper, catalog
pilot, protocol-audit, SBOM, brand, and release-wide checksum assets. Version
1.0.3 additionally includes:

  • the frontier execution, coverage, and GPT-5.6/GPT-5.5 handouts;
  • the GA JSON, CSV, and generated Markdown;
  • the preview-companion JSON, CSV, and generated Markdown;
  • both nested bundle manifests under unique GA and preview-companion filenames.

The workflow then generates an outer release-wide SHA256SUMS and provenance
attestations for the assembled assets.

Limitations and permitted claim

Both frontier runs are small, English-oriented, one-generation,
endpoint-as-served snapshots. Preview behavior can change, provider defaults
were not fully normalized, rank intervals were not estimated, and the study
does not cover the multilingual, multimodal, safety, repeated-generation,
multi-region, load, independent-review, or external-replication requirements
needed for a broad global claim.

Permitted summary:

In the dated 2026-08-08 Kendr callable-subset matrix, GPT-5.6 Sol had the
highest operational-goodput point estimate among five scored GA candidates;
most pairwise comparisons remained unresolved after correction. A separate
one-endpoint preview companion measured Gemini 3.1 Pro Preview at 39.78%
goodput with 100% availability and assigned no rank. Qwen 3.8 Max Preview
remained N/A.

Citation

Use CITATION.cff, cite release v1.0.3, identify
Dr. Prashant Kumar Dey as the researcher and Kendr as project steward,
and include the exact matrix ID for every reused result set.

LLM Benchmark Protocol v1.0.2 - KGBP 1.0 Research Release

Choose a tag to compare

@github-actions github-actions released this 08 Aug 04:16

LLM Benchmark Protocol v1.0.2

Release date: 2026-08-08

Researcher: Dr. Prashant Kumar Dey

Project steward: Kendr

Purpose of this patch

Version 1.0.2 is a publication-design correction. The technical white paper
now uses the same core color system as Kendr's public design language:

  • Ink #151412
  • Saffron #E2712A
  • Paper #FAF8F4
  • Warm grey #8A8378

The PDF adapts those colors for a long technical report while preserving the
Kendr mark's geometry, colors, and clear space.

Publication design

  • The cover and running header use Ink, with Saffron rails and rules and Paper
    typography.
  • Body pages use a warm Paper canvas, Ink text, Saffron structural accents,
    and dark table headers.
  • Charts use Saffron bars with a darker Saffron outline.
  • Quotes and rank highlights use a pale Saffron tint; code blocks and alternate
    table rows use a neutral warm tint.
  • Small orange text uses derived deep Saffron #9A5022, and secondary small
    text uses derived muted Ink #615C54. These choices avoid the insufficient
    contrast of raw Saffron or Warm grey on Paper.
  • Links remain distinguishable without relying on color alone because they are
    both underlined and rendered in deep Saffron.
  • The previous blue, cyan, and navy drawing colors are absent from the PDF.

No gradients, shadows, logo recoloring, logo rotation, or substitute wordmark
were introduced.

Verification evidence

  • All 50 A4 pages were rasterized at 150 dpi and visually inspected, including
    the cover, contents, chart, dense result tables, appendices, and references.
  • The PDF contains the canonical Ink, Saffron, Paper, and Warm-grey drawing
    colors and no legacy blue, cyan, or navy drawing operations.
  • The cover embeds the canonical Kendr logo and credits Dr. Prashant Kumar
    Dey
    .
  • PDF SHA-256: 1e73cf2b629168e49fe837134ab1103b1218eaa5f49fb4a0f1ba3b00deef41f1
  • Resolved Markdown SHA-256:
    5eb99199a3cef536f63d0ca03a43195d01ed751b7bed5d742e10ac4b6a102113

The resolved Markdown digest is unchanged from v1.0.1, confirming that this
patch changes presentation rather than scientific content.

Frozen benchmark provenance

No provider calls were rerun. No prompts, responses, judgments, model labels,
endpoint identities, scores, intervals, costs, availability values, ranking
rows, pairwise tests, or conclusions changed.

The pilot remains a narrow, English-oriented, 15-item, one-generation study
with zero Holm-significant pairwise differences among 595 comparisons. Its
execution-software version remains 1.0.0; v1.0.2 is the publication release.

Historical release policy

The v1.0.0 and v1.0.1 tags and their attested assets remain unchanged.
This correction uses a new tag and new artifacts instead of moving an existing
tag or overwriting a published PDF, package, checksum, SBOM, or attestation.

Citation

Use CITATION.cff, cite release v1.0.2, identify
Dr. Prashant Kumar Dey as the researcher, and include the exact benchmark
matrix identifier when citing the frozen catalog pilot.

LLM Benchmark Protocol v1.0.1 - KGBP 1.0 Research Release

Choose a tag to compare

@github-actions github-actions released this 08 Aug 03:48

LLM Benchmark Protocol v1.0.1

Release date: 2026-08-08

Researcher: Dr. Prashant Kumar Dey

Project steward: Kendr

Purpose of this patch

Version 1.0.1 is a branding and attribution correction. It standardizes the
organization name as Kendr, adds the supplied Kendr mark to the public
project materials, and identifies Dr. Prashant Kumar Dey as the researcher
working on LLM Benchmark Protocol.

The GitHub owner slug remains Kendr-AI inside repository URLs because it is a
technical address, not the public organization name.

What changed

  • Added the canonical 512 x 512 transparent Kendr mark under assets/brand/.
  • Added the mark and researcher attribution to the README.
  • Added the mark to the white-paper cover on Kendr Paper with the prescribed
    clear space, without recoloring, rotation, gradients, or shadows.
  • Updated the white-paper source, cover byline, PDF author metadata, PDF creator
    metadata, citation file, package metadata, governance, protocol card, data
    attribution, copyright notice, and project authorship page.
  • Added automated checks for the canonical logo, image dimensions, researcher
    attribution, legacy display-brand text, PDF cover image, and PDF metadata.
  • Included the brand assets in source distributions and GitHub release assets.
  • Changed release automation to reject attempts to replace an existing release.

Frozen benchmark provenance

No provider calls were rerun. No prompts, responses, judgments, model labels,
endpoint identities, scores, intervals, costs, availability values, ranking
rows, pairwise tests, or conclusions changed.

The published pilot bundle continues to record execution software version
1.0.0. That is intentional: v1.0.1 is the publication/correction release,
not a retroactive claim that the 2026-08-07 experiment ran newer software.

The scientific claim boundary is unchanged. The pilot remains a narrow,
English-oriented, 15-item, one-generation study with zero Holm-significant
pairwise differences among 595 comparisons.

Historical release policy

The v1.0.0 tag and its attested assets remain unchanged as historical
evidence. This patch uses a new tag and new artifacts instead of moving the old
tag or overwriting its PDF, packages, data, checksums, SBOM, or attestations.

Citation

Use CITATION.cff, cite release v1.0.1, identify
Dr. Prashant Kumar Dey as the researcher, and include the exact benchmark
matrix identifier when citing the frozen catalog pilot.

LLM Benchmark Protocol v1.0.0 - KGBP 1.0 Research Release

Choose a tag to compare

@github-actions github-actions released this 07 Aug 19:52

LLM Benchmark Protocol v1.0.0

Release date: 2026-08-08

LLM Benchmark Protocol v1.0.0 is the first versioned public release of the
KGBP 1.0 specification, its Python reference harness, machine-readable evidence
contracts, and a privacy-reviewed catalog-pilot result bundle.

This is a release of evaluation methodology and software. It is not a claim
that KGBP is an accredited standard, that the included pilot is globally
representative, or that its descriptive row order is a resolved universal
ranking of language models.

Highlights

  • Claim-first protocol design with ten non-compensatory dimensions and separate
    design, execution, evidence, and publication gates.
  • Explicit system types for fixed endpoints, routed systems, research systems,
    agents, ensembles, multimodal systems, and applications.
  • Frozen sampling, deterministic interleaving, repeated-generation planning,
    failure-aware scoring, hierarchical bootstrap estimates, paired comparisons,
    multiplicity control, practical-equivalence decisions, operational goodput,
    and router counterfactuals.
  • Versioned JSON Schema 2020-12 contracts covering systems, items, schedules,
    attempts, answers, judgments, observations, scorecards, and evidence
    manifests.
  • Resumable LiveBench execution that preserves the first captured provider
    trial and supports offline finalization without silently replaying inference.
  • Public llm-benchmark-* commands with the historical kendr-* aliases and
    kendr_bench import namespace retained for compatibility.
  • Offline examples, release verification, CI, governance, contribution,
    security, data-license, citation, and adoption materials.

Start here

Install a source checkout with Python 3.11 or newer:

python -m venv .venv
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"
python scripts/verify_release.py --expected-version 1.0.0
python -m pytest

Provider-backed benchmark commands can incur API charges. The quick start
begins with offline audit and toy-scoring commands and requires explicit paid-run
confirmation before provider inference.

Public catalog pilot

The bundled ranking handout reports a frozen campaign
against the Kendr API catalog snapshot captured on 2026-08-07:

  • 37 catalog entries assessed;
  • 35 text-compatible endpoints ranked on endpoint-as-served operational
    goodput;
  • two non-text entries reported as not applicable, rather than forced into the
    text ranking;
  • 15 selected LiveBench questions across five task categories;
  • one generation per endpoint-question cell;
  • 595 paired endpoint comparisons and zero Holm-adjusted rejections.

The machine-readable exports are:

Raw prompts, model responses, provider request identifiers, credentials, and
local paths are not included in the public bundle. See the
evidence-bundle guide and data license
before publishing additional run artifacts.

Scientific interpretation

The catalog campaign is a descriptive pilot. It does not satisfy KGBP global
publication gates because:

  • all selected questions are English-language items dated in 2024;
  • the sample is too small for narrow inference;
  • generation-to-generation variation was not estimated;
  • no multi-region load phase was run;
  • no private rolling holdout was used;
  • independent review and external replication are absent.

No pairwise comparison survived the declared Holm family-wise correction. A
failure to reject is not proof of equality, while a zero endpoint-as-served
score can reflect availability or capability negotiation rather than intrinsic
model incapability. Use the results interpretation guide
and statistical analysis plan before citing or
extending the pilot.

The strict protocol audit scores the reference design above 9 on all ten
dimensions. That design score evaluates whether controls are specified; it does
not replace execution evidence, conformance review, external replication,
certification, or standards-body adoption.

Compatibility

  • Python 3.11, 3.12, and 3.13 are the declared test matrix.
  • The distribution name is llm-benchmark-protocol.
  • The import namespace remains kendr_bench.
  • New command names are llm-benchmark, llm-benchmark-livebench,
    llm-benchmark-matrix, and llm-benchmark-protocol.
  • Existing kendr-bench, kendr-livebench, kendr-benchmark-matrix, and
    kendr-protocol commands remain aliases in this release.
  • Protocol identity, software version, and benchmark-round identity are
    independently versioned. This software release implements the KGBP 1.0
    profile and publishes the immutable matrix ID named in the ranking bundle.

Known limitations

  • The bundled public pilot is not a freshness, multilingual, multimodal,
    longitudinal, safety, or production-load study.
  • Provider aliases and routed candidate sets can change after capture; results
    apply only to the recorded endpoint identities and observation window.
  • Some costs are lower bounds when failed calls lacked complete usage
    telemetry.
  • The reference harness cannot establish legal compliance, social benefit, or
    field effectiveness without the complementary evidence described by the
    protocol.
  • The optional Kendr dependency is pinned to a Git revision and therefore
    requires Git when that extra is installed.

Governance and security

Protocol changes follow governance, including public change
control and the appeals and corrections process.
Security issues should follow the private reporting instructions in
SECURITY.md, not a public issue containing exploit or
credential details. Threat assumptions are documented in the
threat model.

Release integrity

Maintainers must follow RELEASING.md and run:

python scripts/verify_release.py --expected-version 1.0.0
python -m pytest
python -m build

For the tag-triggered workflow, python scripts/verify_release.py --tag v1.0.0 additionally checks exact tag-to-package version alignment. The
verifier is offline and does not create, push, tag, or publish anything.

After the annotated tag is pushed, the release workflow repeats the gates,
audits the isolated runtime dependency set, generates a CycloneDX SBOM and
SHA-256 manifest, attests build provenance, and publishes the GitHub release.

Citation

Use CITATION.cff, cite software version 1.0.0 and protocol
profile KGBP 1.0, and include the exact benchmark-round identifier when citing
the catalog results.

The complete change list is in CHANGELOG.md.