Skip to content

v1.1.5

Choose a tag to compare

@mann1x mann1x released this 26 Dec 17:41
· 35 commits to master since this release

v1.1.5

  • New ps Command - Show running models and their status
    • Display all models currently loaded in memory
    • Works with both local and remote Ollama servers
    • Shows NAME, ID, SIZE, VRAM USAGE, CONTEXT, UNTIL in formatted table
    • Same output format as run/chat command preload display
  • Chat Model Preloading - Models are now automatically loaded into memory before first chat input
    • Sends empty chat request to preload model
    • Displays loaded model status table after preload
    • Shows model name, ID (shortened digest), size, VRAM usage, context length, and expiration time
  • Process Status Display - New formatted table showing all loaded models via /api/ps
    • NAME: Model name with truncation for long names
    • ID: First 12 characters of model digest (Docker-style)
    • SIZE: Disk size combined with parameter count (e.g., "4.54 GB (8.0B)")
    • VRAM USAGE: Memory allocated in VRAM
    • CONTEXT: Context window size (e.g., 4096)
    • UNTIL: Human-readable expiration time (e.g., "2 minutes from now", "About a minute from now")
  • Improved Chat Performance - Fixed streaming response buffering for remote servers
    • Uses HttpCompletionOption.ResponseHeadersRead for immediate streaming
    • Remote chat now responds as fast as local chat
    • No more delays waiting for full response buffering