v1.1.5
v1.1.5
- New
psCommand - Show running models and their status- Display all models currently loaded in memory
- Works with both local and remote Ollama servers
- Shows NAME, ID, SIZE, VRAM USAGE, CONTEXT, UNTIL in formatted table
- Same output format as run/chat command preload display
- Chat Model Preloading - Models are now automatically loaded into memory before first chat input
- Sends empty chat request to preload model
- Displays loaded model status table after preload
- Shows model name, ID (shortened digest), size, VRAM usage, context length, and expiration time
- Process Status Display - New formatted table showing all loaded models via
/api/ps- NAME: Model name with truncation for long names
- ID: First 12 characters of model digest (Docker-style)
- SIZE: Disk size combined with parameter count (e.g., "4.54 GB (8.0B)")
- VRAM USAGE: Memory allocated in VRAM
- CONTEXT: Context window size (e.g., 4096)
- UNTIL: Human-readable expiration time (e.g., "2 minutes from now", "About a minute from now")
- Improved Chat Performance - Fixed streaming response buffering for remote servers
- Uses
HttpCompletionOption.ResponseHeadersReadfor immediate streaming - Remote chat now responds as fast as local chat
- No more delays waiting for full response buffering
- Uses