Skip to content

v0.1.1

Latest

Choose a tag to compare

@joa joa released this 06 Oct 19:23
· 27 commits to main since this release

Highlights

  • The expert cache follows the whole card's free memory (NVML): it shrinks when another program moves in and grows back when it leaves.
  • scripts/agent_bench.py --squeeze replays an agent session under memory pressure. Results are in BENCHMARKS.md.

What's Changed

  • chat: --prefill-thinking opens the think block on a fixed prefill by @joa in #12

Full Changelog: v0.1.0...v0.1.1