Objective
Run a reproducible end-to-end serving qualification sweep for the fork's gfx1151 path after runtime blockers and image provenance are resolved.
Related fork issues
Test matrix
- Context-length scalability through the selected model's supported range.
- Deterministic quality probes plus a named evaluation suite where applicable.
- Cold/warm TTFT, prefill throughput, decode throughput, and concurrency scaling.
- Peak unified memory, graph-reserved memory, KV-cache capacity, and host-memory headroom.
- Backend-selection evidence proving intended gfx1151 AITER/Triton/HIP paths run without CPU fallback.
Record exact model revision, quantization, ROCm/Torch/vLLM revisions, launch settings, prompt/output lengths, repetitions, and summary statistic.
Acceptance criteria
- All blocking correctness and crash issues are linked to minimized reproducers.
- Results are captured in a machine-readable artifact and concise comparison table.
- No hidden CPU fallback or unsupported FP8/FP4/MLA path is presented as acceleration.
- Any serving migration proposal names its reference backend and requires API/quality compatibility plus equal-or-better latency/throughput on the agreed workload.
Objective
Run a reproducible end-to-end serving qualification sweep for the fork's gfx1151 path after runtime blockers and image provenance are resolved.
Related fork issues
Test matrix
Record exact model revision, quantization, ROCm/Torch/vLLM revisions, launch settings, prompt/output lengths, repetitions, and summary statistic.
Acceptance criteria