Skip to content

docs(QUANT-EXL3): align support and usage with code - #3176

Draft
localai-org-maint-bot wants to merge 3 commits into
mudler:mainfrom
localai-org-maint-bot:row/QUANT-EXL3-DOCS
Draft

docs(QUANT-EXL3): align support and usage with code#3176
localai-org-maint-bot wants to merge 3 commits into
mudler:mainfrom
localai-org-maint-bot:row/QUANT-EXL3-DOCS

Conversation

@localai-org-maint-bot

@localai-org-maint-bot localai-org-maint-bot commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator

Row

QUANT-EXL3, public documentation only.
Tracks ISSUE-LOCAL-01M2CBGG169HDX0DWKHQWCBBGN in the canonical local issue file.

Before starting

  • Searched open upstream PRs, current issues, and claims at e030f1b90. The existing EXL3 implementation claim owns source and its implementation spec. This contribution owns public documentation only.
  • One PR carries the spec-first commit and implementation, using the unattended policy default.
  • The quantization matrix keeps QUANT-EXL3 at ACTIVE. No implementation lifecycle changes.
  • Inspected the loader, shared dispatch, CUDA instantiations, CPU dispatch regression, and committed generation evidence. Exact anchors are in .agents/specs/exl3-public-docs.md.

What changed

Replace the old EXL3 news with the current CUDA Qwen3.8-27B and DFlash2 capability. Shorten the EXL3 feature summary and five artifact entries. Correct the obsolete DeepSeek-V4 loader and tokenizer limitations while keeping its real-artifact gate open. Preserve filenames, sizes, revisions, and hashes, and archive removed prose verbatim. Link benchmark conditions without adding speed or correctness claims.

Evidence

Independent final review: PASS, no findings, on bdfc4d9b038991fea23a634bf4b8172ce1d8ee72. The submitted head has the identical tree. The operator reran the focused gate.

  • scripts/agent-preflight.sh passes. The broad local run exited 1 with failures and skips. Missing CMake/readelf/numpy and extracted Python affect this container; shell/tool suites also fail. The run captured a stale origin/main base before refresh. The record, release-workflow, trailer, and compile-scope checks were rerun successfully against the correct base. No full-suite green is claimed.
  • python3 scripts/check-readme-structure.py
  • python3 scripts/check-supported-models.py (44 registered architectures)
  • python3 scripts/check-quickstart-recipes.py
  • python3 scripts/check-benchmark-index.py
  • python3 scripts/check-agent-record.py
  • python3 -m unittest discover -s tests/scripts -p 'test_check_readme_structure.py' (19 tests)
  • python3 scripts/check-tree-compiles.py --base upstream/main: no source, header, or build file is in scope.
  • python3 scripts/check-commit-trailers.py --range upstream/main..HEAD --filled and check-commit-style.py --range upstream/main..HEAD.
  • Direct comparison against e030f1b90: all five edited artifact entries retain their first five cells byte-for-byte.
  • Public docs change only for their owned user-facing facts. No source, tests, checkers, benchmark values, or unrelated model rows change.

Speed claims

  • This PR makes no new speed claim. Existing benchmark numbers and dispositions are unchanged.

Honest gaps

This is documentation-only work on a CPU-only host. No GPU execution, model inference, or new runtime measurement was performed. Source checks and retained evidence support the descriptions. The sampled HumanEval comparison is explicitly not a correctness gate. Full local preflight is not green, and hosted CI must supply the unavailable tooling. The local issue remains open until the PR lands. No upstream merge is requested.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-6 [Codex]

The public EXL3 description retains obsolete limits and attempt history.
Define a source-checked editorial correction without changing runtime or
benchmark claims.

Tracks ISSUE-LOCAL-01M2CBGG169HDX0DWKHQWCBBGN.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-6 [Codex]
The EXL3 feature row mixes current capabilities with obsolete attempts.
Describe the CUDA dispatch and artifact limits against the current source.
Preserve removed prose in the dated archive and keep artifact pins intact.

Issue: ISSUE-LOCAL-01M2CBGG169HDX0DWKHQWCBBGN
Branch: row/QUANT-EXL3-DOCS-IMPLEMENT
Review: mudler#3176

Focused documentation gates pass, including 19 README checker tests.
The operator owns broad preflight and its environment-failure comparison.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:gpt-6 [Codex]
The MTP width applies to trellis weights, not companion tensors.
The ROCm record includes a later BF16 control, so name the remaining
clock-attribution and discrete-GPU gaps instead.

Issue: ISSUE-LOCAL-01M2CBGG169HDX0DWKHQWCBBGN

All seven focused documentation gates pass, including 19 README tests.
Artifact filenames, sizes, revisions, and hashes remain unchanged.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:gpt-6 [Codex]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant