Skip to content

Releases: tsaipifong/whirl-llm

WHIRL v0.1.2

Choose a tag to compare

@tsaipifong tsaipifong released this 03 Oct 10:33

WHIRL v0.1.2

English · 繁體中文

A small usability update. Kernels and numerics are unchanged, so outputs are bit-identical to v0.1.1.

What's new

  • The RAM tier of the prefix cache now sizes itself to your machine. By default it uses about 1/4 of physical RAM, at least 8 GB and at most 32 GB, and never more than half the memory that is free at startup. On a 64 GB machine that comes to 16 GB, up from a fixed ~9 GB before, so long agent sessions keep their history in RAM longer before entries fall back to SSD. --kv-ram-mb N / WHIRL_KV_RAM_MB still override the default, and 0 turns the tier off. The startup log shows the size it chose and why. Startup takes about 1.5 s longer for the larger pinned pool. The tier gate passes 24/24.
  • Usage recipes in docs/recipes.md (繁中). They cover:
    • connecting agents and chat front ends
    • long agent sessions with subagents (with real numbers)
    • long context, several users and images
    • how to read the server log
    • troubleshooting

See CHANGELOG.md.

Download

whirl-0.1.2-windows-x64.zip. Requirements are the same as before: an AMD Radeon AI PRO R9700, Windows 11 and AMD Software Adrenalin 26.8.1 or newer.

SHA-256: 3c56e4857d82aa324eb32ec2da91db2cf40497129a5e7dbcdb53f1bd959123fb

The executables are unsigned. See docs/windows_security.md.


繁體中文

小幅改善易用性。kernel 和數值計算都沒有改動,輸出和 v0.1.1 逐位元相同。

RAM 快取大小改為依電腦自動決定

  • 預設約為實體記憶體的 1/4,最少 8 GB、最多 32 GB,也不會超過啟動時可用記憶體的一半。
  • 64 GB 的電腦會用 16 GB(原本固定約 9 GB),長時間跑 agent 時,對話歷史能在 RAM 裡留得更久。
  • 仍可用 --kv-ram-mb 手動指定,設 0 就關閉。
  • 啟動 log 會列出選了多少、為什麼選這個大小。
  • 啟動時間多約 1.5 秒;tier 快取檢查 24/24 通過。

使用情境指南

新增 docs/recipes_zh-TW.md,內容包括:

  • 接上 agent 或聊天介面
  • 長時間 agent 工作與子代理,附實測數據
  • 長上下文、多人使用、圖片
  • 怎麼看懂伺服器 log
  • 疑難排解

下載

whirl-0.1.2-windows-x64.zip,SHA-256 見上方。

WHIRL v0.1.1

Choose a tag to compare

@tsaipifong tsaipifong released this 03 Oct 09:03

WHIRL v0.1.1

English · 繁體中文

A small update driven by real agent use (Hermes agent with subagents on Swift-1.5 27B MXFP4). Models, kernels and numerics are unchanged from v0.1.0. Outputs are bit-identical.

What's new

  • Decode floor: --decode-min-tps N (env WHIRL_DECODE_MIN_TPS, default 20, 0 = off). When subagents send long prompts while another conversation is streaming, the streaming request used to almost stop until their prefill finished. Now every streaming request keeps at least N tok/s, and the server adapts each cycle. When nothing is streaming, prefill runs at full speed as before.

    Measured on the R9700: one stream at 25.7k context plus three ~17k-token prompts.

    Off (N=0) Default (N=20)
    Swift-1.5 27B MXFP4, stream tok/s during prefill 3.3 23.0
    Longest pause 1.22 s 0.47 s
    Mean TTFT of the three prompts 17.8 s 18.6 s
    Ornith-1.5 35B-A3B MXFP4, stream tok/s 6.7 31.5
  • Compatibility endpoints for tools that auto-detect the server type and context length:

    • GET /props and GET /v1/props: a llama.cpp-style subset (n_ctx, total_slots, model alias and more).
    • GET /version.

    LM Studio and Ollama probe paths still return 404, but each is logged only once. They no longer fill the log with warnings.

  • Reasoning effort from more clients. The server now accepts the "reasoning": {"effort": ..., "enabled": ...} object. It also maps aliases: max / ultra / xhigh / high → xhigh, medium → medium, low / minimal → low, none → thinking off. An unknown value logs a warning and uses the default instead of failing the request.

See CHANGELOG.md for details.

Download

whirl-0.1.1-windows-x64.zip. Requirements are the same as v0.1.0: an AMD Radeon AI PRO R9700, Windows 11 and AMD Software Adrenalin 26.8.1 or newer. Nothing else is needed.

SHA-256: 37fe79d0fe2f15170290abebf026d52e4daa7290fb6b666de5a5bd6e623f5cf6

The executables are unsigned. See docs/windows_security.md for SmartScreen and Smart App Control.


繁體中文

這是一次小更新,來自實際用 Hermes agent 加子代理跑 Swift-1.5 27B MXFP4 時發現的問題。模型、kernel 和數值計算都和 v0.1.0 相同,輸出逐位元一致。

新功能

  • Decode 保底速度 --decode-min-tps N(環境變數 WHIRL_DECODE_MIN_TPS;預設 20,設 0 關閉)

    • 以前子代理送進長 prompt 時,正在串流輸出的對話會幾乎停住,要等它們 prefill 完才繼續。
    • 現在每個正在串流的請求都至少保有 N tok/s,伺服器每一輪自動調整。
    • 沒有請求在串流時,prefill 和以前一樣全速跑。
    • R9700 實測數據見上方英文表格。
  • 相容端點:新增 /props、/v1/props、/version。Hermes 這類工具可以自動偵測上下文長度。LM Studio 和 Ollama 的探測網址還是回 404,但同一個網址只記一次 log。

  • 思考層級:支援 "reasoning": {"effort": ...} 寫法,以及 ultra、max、minimal、none 等別名。遇到不認得的值,只記一行警告並改用預設值,不會讓請求失敗。

下載

whirl-0.1.1-windows-x64.zip,需求和 v0.1.0 相同。SHA-256 見上方。

WHIRL v0.1.0

Choose a tag to compare

@tsaipifong tsaipifong released this 03 Oct 04:02

WHIRL v0.1.0

English · 繁體中文

WHIRL (Windows HIP Inference for RDNA LLMs) is a native Windows LLM inference engine for the AMD
Radeon AI PRO R9700, written in C++ and HIP. It runs directly on the AMD graphics driver: no WSL,
no Linux VM, no llama.cpp runtime. It ships two programs: whirl.exe (chat, bench, device tools)
and whirl-server.exe (OpenAI-compatible HTTP server).

Requirements

  • AMD Radeon AI PRO R9700 (RDNA 4, gfx1201), single GPU
  • Windows 11, 64-bit
  • AMD Software: Adrenalin Edition 26.8.1 or newer. Nothing else: no HIP SDK, no ROCm, no Visual C++ runtime
  • A qwen35 (dense) or qwen35moe (MoE) GGUF, for example
    Swift-1.5-Qwen3.8-27b-MXFP4-GGUF (variant A recommended)

Highlights

  • Hand-tuned RDNA 4 kernels: int8 GEMV at the measured memory-bandwidth limit, WMMA prefill, MXFP4 × fp8 matrix paths.
  • MTP + n-gram speculative decoding whose output is bit-identical to plain greedy decoding.
  • OpenAI-compatible server with continuous batching and prefix caching over a multi-tier KV cache
    (VRAM → pinned RAM → SSD); sessions survive a server restart.
  • Image input with a Qwen3-VL-style mmproj (--mmproj).

WHIRL 0.1.0 vs llama.cpp b11214 on the R9700, greedy decoding, same prompts (full method in docs/benchmarks.md):

R9700, greedy Ornith MXFP4 (MoE) Swift MXFP4-A (dense 27B) Qwen3.8-27B Q4_K_M (dense)
Prefill 8k tok/s 10,858 vs 4,637 (2.34×) 3,278 vs 1,338 (2.45×) 1,689 vs 1,223 (1.38×)
Prefill 32k tok/s 7,978 vs 3,778 (2.11×) 2,595 vs 1,174 (2.21×) 1,479 vs 1,086 (1.36×)
Decode, zh coding, WHIRL MTP+n-gram vs llama.cpp fastest 244.5 vs 118.8 (2.06×) 107.5 vs 60.8 (1.77×) 98.5 vs 56.0 (1.76×)
Decode, zh coding, no MTP vs llama.cpp plain 166.7 vs 118.8 (1.40×) 37.5 vs 33.4 (1.13×) 34.4 vs 30.9 (1.11×)
Server, 4 concurrent users, aggregate tok/s 396.7 vs 191.9 (2.07×) 198.6 vs 65.8 (3.02×) 134.1 vs 58.5 (2.29×)

Decode: 7 Chinese coding prompts, 800 tokens, median of 3 rounds; "fastest" is llama.cpp's best of
plain / MTP / MTP + n-gram for that model. Test GPU connected as a USB4 eGPU.

Known limitations

  • Unsigned executables. SmartScreen may show "Windows protected your PC" (More info → Run anyway,
    or Unblock-File the zip after checking its SHA-256). With Smart App Control On, Windows may block
    the programs outright. See docs/windows_security.md.
  • One GPU only: the R9700 (gfx1201). No multi-GPU; the Radeon 8060S (Ryzen AI Max+ 395) version is planned, not included.
  • Two architectures only: qwen35 and qwen35moe. Other architectures and tensor types without
    WHIRL kernels (for example NVFP4) are refused at load time; there is no generic fallback.
  • Plain decoding without speculation is memory-bandwidth bound, so WHIRL's lead there is small on the
    dense models (1.1–1.2×); 88-token prefill on Q4_K_M is at parity; WHIRL's CLI uses 3–5 GiB more
    VRAM than llama-bench on the dense models at the same context.
  • The first run with a model tunes the GPU kernels once (about 1.5 minutes for a 27B Q4_K_M file),
    cached in %LOCALAPPDATA%\whirl.
  • The Ornith-1.5-35B-A3B MXFP4 GGUF is available at https://huggingface.co/tsaipifong/Ornith-1.5-35B-A3B-MXFP4-GGUF

Download

whirl-0.1.0-windows-x64.zip from https://github.com/tsaipifong/whirl-llm/releases

SHA-256: b5ee46100ba6bd22ab2f4dcd824189dba972202aa3208a51cb979d6c4c8170f9

(Get-FileHash .\whirl-0.1.0-windows-x64.zip -Algorithm SHA256).Hash.ToLower()

License: Apache-2.0 (see LICENSE, NOTICE, THIRD_PARTY_NOTICES.md).


繁體中文

WHIRL(Windows HIP Inference for RDNA LLMs)是為 AMD Radeon AI PRO R9700 打造的原生 Windows LLM 推論引擎,
以 C++ 與 HIP 撰寫,直接跑在 AMD 顯示卡驅動程式上:不用 WSL、不用 Linux 虛擬機,底下也沒有 llama.cpp。
發行版包含兩支程式:whirl.exe(對話、效能測試、裝置工具)與 whirl-server.exe(相容 OpenAI API 的 HTTP 伺服器)。

系統需求

  • AMD Radeon AI PRO R9700(RDNA 4,gfx1201),僅支援單張 GPU
  • Windows 11 64 位元
  • AMD Software: Adrenalin Edition 26.8.1 或更新版本;其他都不需要(不用 HIP SDK、ROCm、Visual C++ runtime)
  • qwen35(dense)或 qwen35moe(MoE)架構的 GGUF,例如
    Swift-1.5-Qwen3.8-27b-MXFP4-GGUF(建議用 A 版)

重點

  • 針對 RDNA 4 手工調校的 kernel:達到實測記憶體頻寬上限的 int8 GEMV、WMMA prefill、MXFP4 × fp8 矩陣路徑。
  • MTP + n-gram 推測解碼,輸出與一般 greedy 解碼逐位元相同。
  • 相容 OpenAI API 的伺服器:continuous batching,加上多層 KV 快取(VRAM → pinned RAM → SSD)的前綴快取;伺服器重啟後對話仍可還原。
  • 搭配 Qwen3-VL 形式的 mmproj 支援圖片輸入(--mmproj)。

WHIRL 0.1.0 與 llama.cpp b11214 在 R9700 上的比較(greedy 解碼、相同提示詞;完整方法見 docs/benchmarks.md):

R9700,greedy Ornith MXFP4(MoE) Swift MXFP4-A(dense 27B) Qwen3.8-27B Q4_K_M(dense)
Prefill 8k tok/s 10,858 vs 4,637(2.34×) 3,278 vs 1,338(2.45×) 1,689 vs 1,223(1.38×)
Prefill 32k tok/s 7,978 vs 3,778(2.11×) 2,595 vs 1,174(2.21×) 1,479 vs 1,086(1.36×)
Decode,中文程式題,WHIRL MTP+n-gram vs llama.cpp 最快設定 244.5 vs 118.8(2.06×) 107.5 vs 60.8(1.77×) 98.5 vs 56.0(1.76×)
Decode,中文程式題,不用 MTP vs llama.cpp 一般解碼 166.7 vs 118.8(1.40×) 37.5 vs 33.4(1.13×) 34.4 vs 30.9(1.11×)
伺服器,4 位同時使用者,總 tok/s 396.7 vs 191.9(2.07×) 198.6 vs 65.8(3.02×) 134.1 vs 58.5(2.29×)

Decode:7 題中文程式提示詞、800 token、3 輪取中位數;「最快設定」指 llama.cpp 在該模型上 plain/MTP/MTP + n-gram 三者中最快者。測試用 GPU 以 USB4 外接。

已知限制

  • 執行檔未經程式碼簽章。 SmartScreen 可能顯示「Windows 已保護您的電腦」(按「其他資訊 → 仍要執行」,
    或在核對 SHA-256 後對 zip 執行 Unblock-File)。若「智慧型應用程式控制」設為「開啟」,Windows 可能直接封鎖程式。詳見 docs/windows_security_zh-TW.md。
  • 只支援一張 GPU: R9700(gfx1201)。不支援多 GPU;Radeon 8060S(Ryzen AI Max+ 395)版本已在規劃中,本版未包含。
  • 只支援兩種架構: qwen35 與 qwen35moe。其他架構,以及 WHIRL 沒有 kernel 的張量型別(例如 NVFP4),載入時就會被拒絕,沒有通用的備援路徑。
  • 不用推測解碼時受記憶體頻寬限制,dense 模型上 WHIRL 的領先幅度很小(1.1–1.2×);Q4_K_M 的 88 token prefill 與 llama.cpp 持平;
    在相同 context 下,dense 模型的 CLI 比 llama-bench 多用 3–5 GiB VRAM。
  • 每個模型第一次執行時會調校一次 GPU kernel(27B Q4_K_M 約 1.5 分鐘),結果快取在 %LOCALAPPDATA%\whirl。
  • Ornith-1.5-35B-A3B MXFP4 GGUF 已發布於 https://huggingface.co/tsaipifong/Ornith-1.5-35B-A3B-MXFP4-GGUF

下載

從 https://github.com/tsaipifong/whirl-llm/releases 下載 whirl-0.1.0-windows-x64.zip

SHA-256:b5ee46100ba6bd22ab2f4dcd824189dba972202aa3208a51cb979d6c4c8170f9

授權:Apache-2.0(見 LICENSE、NOTICE、THIRD_PARTY_NOTICES.md)。