Releases: tsaipifong/whirl-llm
Release list
WHIRL v0.1.2
WHIRL v0.1.2
English · 繁體中文
A small usability update. Kernels and numerics are unchanged, so outputs are bit-identical to v0.1.1.
What's new
- The RAM tier of the prefix cache now sizes itself to your machine. By default it uses about 1/4 of physical RAM, at least 8 GB and at most 32 GB, and never more than half the memory that is free at startup. On a 64 GB machine that comes to 16 GB, up from a fixed ~9 GB before, so long agent sessions keep their history in RAM longer before entries fall back to SSD.
--kv-ram-mb N/WHIRL_KV_RAM_MBstill override the default, and0turns the tier off. The startup log shows the size it chose and why. Startup takes about 1.5 s longer for the larger pinned pool. The tier gate passes 24/24. - Usage recipes in docs/recipes.md (繁中). They cover:
- connecting agents and chat front ends
- long agent sessions with subagents (with real numbers)
- long context, several users and images
- how to read the server log
- troubleshooting
See CHANGELOG.md.
Download
whirl-0.1.2-windows-x64.zip. Requirements are the same as before: an AMD Radeon AI PRO R9700, Windows 11 and AMD Software Adrenalin 26.8.1 or newer.
SHA-256: 3c56e4857d82aa324eb32ec2da91db2cf40497129a5e7dbcdb53f1bd959123fb
The executables are unsigned. See docs/windows_security.md.
繁體中文
小幅改善易用性。kernel 和數值計算都沒有改動,輸出和 v0.1.1 逐位元相同。
RAM 快取大小改為依電腦自動決定
- 預設約為實體記憶體的 1/4,最少 8 GB、最多 32 GB,也不會超過啟動時可用記憶體的一半。
- 64 GB 的電腦會用 16 GB(原本固定約 9 GB),長時間跑 agent 時,對話歷史能在 RAM 裡留得更久。
- 仍可用
--kv-ram-mb手動指定,設 0 就關閉。 - 啟動 log 會列出選了多少、為什麼選這個大小。
- 啟動時間多約 1.5 秒;tier 快取檢查 24/24 通過。
使用情境指南
新增 docs/recipes_zh-TW.md,內容包括:
- 接上 agent 或聊天介面
- 長時間 agent 工作與子代理,附實測數據
- 長上下文、多人使用、圖片
- 怎麼看懂伺服器 log
- 疑難排解
下載
whirl-0.1.2-windows-x64.zip,SHA-256 見上方。
WHIRL v0.1.1
WHIRL v0.1.1
English · 繁體中文
A small update driven by real agent use (Hermes agent with subagents on Swift-1.5 27B MXFP4). Models, kernels and numerics are unchanged from v0.1.0. Outputs are bit-identical.
What's new
-
Decode floor:
--decode-min-tps N(envWHIRL_DECODE_MIN_TPS, default 20,0= off). When subagents send long prompts while another conversation is streaming, the streaming request used to almost stop until their prefill finished. Now every streaming request keeps at least N tok/s, and the server adapts each cycle. When nothing is streaming, prefill runs at full speed as before.Measured on the R9700: one stream at 25.7k context plus three ~17k-token prompts.
Off (N=0) Default (N=20) Swift-1.5 27B MXFP4, stream tok/s during prefill 3.3 23.0 Longest pause 1.22 s 0.47 s Mean TTFT of the three prompts 17.8 s 18.6 s Ornith-1.5 35B-A3B MXFP4, stream tok/s 6.7 31.5 -
Compatibility endpoints for tools that auto-detect the server type and context length:
GET /propsandGET /v1/props: a llama.cpp-style subset (n_ctx,total_slots, model alias and more).GET /version.
LM Studio and Ollama probe paths still return 404, but each is logged only once. They no longer fill the log with warnings.
-
Reasoning effort from more clients. The server now accepts the
"reasoning": {"effort": ..., "enabled": ...}object. It also maps aliases:max/ultra/xhigh/high→ xhigh,medium→ medium,low/minimal→ low,none→ thinking off. An unknown value logs a warning and uses the default instead of failing the request.
See CHANGELOG.md for details.
Download
whirl-0.1.1-windows-x64.zip. Requirements are the same as v0.1.0: an AMD Radeon AI PRO R9700, Windows 11 and AMD Software Adrenalin 26.8.1 or newer. Nothing else is needed.
SHA-256: 37fe79d0fe2f15170290abebf026d52e4daa7290fb6b666de5a5bd6e623f5cf6
The executables are unsigned. See docs/windows_security.md for SmartScreen and Smart App Control.
繁體中文
這是一次小更新,來自實際用 Hermes agent 加子代理跑 Swift-1.5 27B MXFP4 時發現的問題。模型、kernel 和數值計算都和 v0.1.0 相同,輸出逐位元一致。
新功能
-
Decode 保底速度
--decode-min-tps N(環境變數WHIRL_DECODE_MIN_TPS;預設 20,設 0 關閉)- 以前子代理送進長 prompt 時,正在串流輸出的對話會幾乎停住,要等它們 prefill 完才繼續。
- 現在每個正在串流的請求都至少保有 N tok/s,伺服器每一輪自動調整。
- 沒有請求在串流時,prefill 和以前一樣全速跑。
- R9700 實測數據見上方英文表格。
-
相容端點:新增
/props、/v1/props、/version。Hermes 這類工具可以自動偵測上下文長度。LM Studio 和 Ollama 的探測網址還是回 404,但同一個網址只記一次 log。 -
思考層級:支援
"reasoning": {"effort": ...}寫法,以及 ultra、max、minimal、none 等別名。遇到不認得的值,只記一行警告並改用預設值,不會讓請求失敗。
下載
whirl-0.1.1-windows-x64.zip,需求和 v0.1.0 相同。SHA-256 見上方。
WHIRL v0.1.0
WHIRL v0.1.0
English · 繁體中文
WHIRL (Windows HIP Inference for RDNA LLMs) is a native Windows LLM inference engine for the AMD
Radeon AI PRO R9700, written in C++ and HIP. It runs directly on the AMD graphics driver: no WSL,
no Linux VM, no llama.cpp runtime. It ships two programs: whirl.exe (chat, bench, device tools)
and whirl-server.exe (OpenAI-compatible HTTP server).
Requirements
- AMD Radeon AI PRO R9700 (RDNA 4, gfx1201), single GPU
- Windows 11, 64-bit
- AMD Software: Adrenalin Edition 26.8.1 or newer. Nothing else: no HIP SDK, no ROCm, no Visual C++ runtime
- A
qwen35(dense) orqwen35moe(MoE) GGUF, for example
Swift-1.5-Qwen3.8-27b-MXFP4-GGUF (variant A recommended)
Highlights
- Hand-tuned RDNA 4 kernels: int8 GEMV at the measured memory-bandwidth limit, WMMA prefill, MXFP4 × fp8 matrix paths.
- MTP + n-gram speculative decoding whose output is bit-identical to plain greedy decoding.
- OpenAI-compatible server with continuous batching and prefix caching over a multi-tier KV cache
(VRAM → pinned RAM → SSD); sessions survive a server restart. - Image input with a Qwen3-VL-style mmproj (
--mmproj).
WHIRL 0.1.0 vs llama.cpp b11214 on the R9700, greedy decoding, same prompts (full method in docs/benchmarks.md):
| R9700, greedy | Ornith MXFP4 (MoE) | Swift MXFP4-A (dense 27B) | Qwen3.8-27B Q4_K_M (dense) |
|---|---|---|---|
| Prefill 8k tok/s | 10,858 vs 4,637 (2.34×) | 3,278 vs 1,338 (2.45×) | 1,689 vs 1,223 (1.38×) |
| Prefill 32k tok/s | 7,978 vs 3,778 (2.11×) | 2,595 vs 1,174 (2.21×) | 1,479 vs 1,086 (1.36×) |
| Decode, zh coding, WHIRL MTP+n-gram vs llama.cpp fastest | 244.5 vs 118.8 (2.06×) | 107.5 vs 60.8 (1.77×) | 98.5 vs 56.0 (1.76×) |
| Decode, zh coding, no MTP vs llama.cpp plain | 166.7 vs 118.8 (1.40×) | 37.5 vs 33.4 (1.13×) | 34.4 vs 30.9 (1.11×) |
| Server, 4 concurrent users, aggregate tok/s | 396.7 vs 191.9 (2.07×) | 198.6 vs 65.8 (3.02×) | 134.1 vs 58.5 (2.29×) |
Decode: 7 Chinese coding prompts, 800 tokens, median of 3 rounds; "fastest" is llama.cpp's best of
plain / MTP / MTP + n-gram for that model. Test GPU connected as a USB4 eGPU.
Known limitations
- Unsigned executables. SmartScreen may show "Windows protected your PC" (More info → Run anyway,
orUnblock-Filethe zip after checking its SHA-256). With Smart App Control On, Windows may block
the programs outright. Seedocs/windows_security.md. - One GPU only: the R9700 (gfx1201). No multi-GPU; the Radeon 8060S (Ryzen AI Max+ 395) version is planned, not included.
- Two architectures only:
qwen35andqwen35moe. Other architectures and tensor types without
WHIRL kernels (for example NVFP4) are refused at load time; there is no generic fallback. - Plain decoding without speculation is memory-bandwidth bound, so WHIRL's lead there is small on the
dense models (1.1–1.2×); 88-token prefill on Q4_K_M is at parity; WHIRL's CLI uses 3–5 GiB more
VRAM than llama-bench on the dense models at the same context. - The first run with a model tunes the GPU kernels once (about 1.5 minutes for a 27B Q4_K_M file),
cached in%LOCALAPPDATA%\whirl. - The Ornith-1.5-35B-A3B MXFP4 GGUF is available at https://huggingface.co/tsaipifong/Ornith-1.5-35B-A3B-MXFP4-GGUF
Download
whirl-0.1.0-windows-x64.zip from https://github.com/tsaipifong/whirl-llm/releases
SHA-256: b5ee46100ba6bd22ab2f4dcd824189dba972202aa3208a51cb979d6c4c8170f9
(Get-FileHash .\whirl-0.1.0-windows-x64.zip -Algorithm SHA256).Hash.ToLower()License: Apache-2.0 (see LICENSE, NOTICE, THIRD_PARTY_NOTICES.md).
繁體中文
WHIRL(Windows HIP Inference for RDNA LLMs)是為 AMD Radeon AI PRO R9700 打造的原生 Windows LLM 推論引擎,
以 C++ 與 HIP 撰寫,直接跑在 AMD 顯示卡驅動程式上:不用 WSL、不用 Linux 虛擬機,底下也沒有 llama.cpp。
發行版包含兩支程式:whirl.exe(對話、效能測試、裝置工具)與 whirl-server.exe(相容 OpenAI API 的 HTTP 伺服器)。
系統需求
- AMD Radeon AI PRO R9700(RDNA 4,gfx1201),僅支援單張 GPU
- Windows 11 64 位元
- AMD Software: Adrenalin Edition 26.8.1 或更新版本;其他都不需要(不用 HIP SDK、ROCm、Visual C++ runtime)
qwen35(dense)或qwen35moe(MoE)架構的 GGUF,例如
Swift-1.5-Qwen3.8-27b-MXFP4-GGUF(建議用 A 版)
重點
- 針對 RDNA 4 手工調校的 kernel:達到實測記憶體頻寬上限的 int8 GEMV、WMMA prefill、MXFP4 × fp8 矩陣路徑。
- MTP + n-gram 推測解碼,輸出與一般 greedy 解碼逐位元相同。
- 相容 OpenAI API 的伺服器:continuous batching,加上多層 KV 快取(VRAM → pinned RAM → SSD)的前綴快取;伺服器重啟後對話仍可還原。
- 搭配 Qwen3-VL 形式的 mmproj 支援圖片輸入(
--mmproj)。
WHIRL 0.1.0 與 llama.cpp b11214 在 R9700 上的比較(greedy 解碼、相同提示詞;完整方法見 docs/benchmarks.md):
| R9700,greedy | Ornith MXFP4(MoE) | Swift MXFP4-A(dense 27B) | Qwen3.8-27B Q4_K_M(dense) |
|---|---|---|---|
| Prefill 8k tok/s | 10,858 vs 4,637(2.34×) | 3,278 vs 1,338(2.45×) | 1,689 vs 1,223(1.38×) |
| Prefill 32k tok/s | 7,978 vs 3,778(2.11×) | 2,595 vs 1,174(2.21×) | 1,479 vs 1,086(1.36×) |
| Decode,中文程式題,WHIRL MTP+n-gram vs llama.cpp 最快設定 | 244.5 vs 118.8(2.06×) | 107.5 vs 60.8(1.77×) | 98.5 vs 56.0(1.76×) |
| Decode,中文程式題,不用 MTP vs llama.cpp 一般解碼 | 166.7 vs 118.8(1.40×) | 37.5 vs 33.4(1.13×) | 34.4 vs 30.9(1.11×) |
| 伺服器,4 位同時使用者,總 tok/s | 396.7 vs 191.9(2.07×) | 198.6 vs 65.8(3.02×) | 134.1 vs 58.5(2.29×) |
Decode:7 題中文程式提示詞、800 token、3 輪取中位數;「最快設定」指 llama.cpp 在該模型上 plain/MTP/MTP + n-gram 三者中最快者。測試用 GPU 以 USB4 外接。
已知限制
- 執行檔未經程式碼簽章。 SmartScreen 可能顯示「Windows 已保護您的電腦」(按「其他資訊 → 仍要執行」,
或在核對 SHA-256 後對 zip 執行Unblock-File)。若「智慧型應用程式控制」設為「開啟」,Windows 可能直接封鎖程式。詳見docs/windows_security_zh-TW.md。 - 只支援一張 GPU: R9700(gfx1201)。不支援多 GPU;Radeon 8060S(Ryzen AI Max+ 395)版本已在規劃中,本版未包含。
- 只支援兩種架構:
qwen35與qwen35moe。其他架構,以及 WHIRL 沒有 kernel 的張量型別(例如 NVFP4),載入時就會被拒絕,沒有通用的備援路徑。 - 不用推測解碼時受記憶體頻寬限制,dense 模型上 WHIRL 的領先幅度很小(1.1–1.2×);Q4_K_M 的 88 token prefill 與 llama.cpp 持平;
在相同 context 下,dense 模型的 CLI 比 llama-bench 多用 3–5 GiB VRAM。 - 每個模型第一次執行時會調校一次 GPU kernel(27B Q4_K_M 約 1.5 分鐘),結果快取在
%LOCALAPPDATA%\whirl。 - Ornith-1.5-35B-A3B MXFP4 GGUF 已發布於 https://huggingface.co/tsaipifong/Ornith-1.5-35B-A3B-MXFP4-GGUF
下載
從 https://github.com/tsaipifong/whirl-llm/releases 下載 whirl-0.1.0-windows-x64.zip
SHA-256:b5ee46100ba6bd22ab2f4dcd824189dba972202aa3208a51cb979d6c4c8170f9
授權:Apache-2.0(見 LICENSE、NOTICE、THIRD_PARTY_NOTICES.md)。