v1.2.6
v1.2.6
- QcView Multiple Output Formats - Export results to various formats
- Markdown (.md) - Tables formatted for GitHub and documentation
- HTML (.html) - Interactive report with dark/light theme toggle, collapsible Q&A sections, color-coded scores
- PDF (.pdf) - Professional report using QuestPDF library with:
- Summary tables with color-coded scores for all metrics
- Scores by category table
- Rankings tables (by score, eval speed, perplexity, best answers)
- Full Q&A pages for each quantization with judgment details
- Use
-Fo md,-Fo html, or-Fo pdfto select format
- QcView Repository URL - New
--repoargument to specify model source repository- URL is displayed in output headers and included in JSON export
- Can be saved during
qctesting and overridden inqcview
- Headless/Background Mode Fix - Fixed console errors when running qc command in background
- Console.WindowWidth and Console.ReadKey now properly handled in headless environments
- Prevents "The handle is invalid" errors when running without a terminal
- New Version Command - Added
osync version(alias-v) to display version info- Shows osync version number and build timestamp
--verboseflag displays detailed info: binary path, installation status, shell type/version, tab completion status- Detects bash, zsh, PowerShell (Core/Desktop), and cmd shells
- Smart installation detection: when running a different binary, compares version AND build timestamp with installed version
- Reports if installed version is older/newer (e.g.,
installed v1.2.6 (b20260110-1156) is older) - Fixed tab completion detection to match actual script markers in profiles
- Model Digest Tracking - QC results now include SHA256 digest for each tested model
- Full digest (
Digest) and short digest (ShortDigest, first 12 chars) stored in results JSON - Automatically populated from local Ollama or HuggingFace registry
- Backfill: missing digests are automatically retrieved when loading existing results files
- Full digest (
- Fixed Model ID Display - IDs now show first 12 chars of manifest SHA256 (matches
ollama ls)osync lsand manage TUI compute SHA256 of manifest file content- ID column width increased from 8 to 12 characters
- Consistent with
ollama lsoutput for easy cross-reference
- Improved osync ps Output - Dynamic console width and better model name display
- Detects console width and adjusts column sizes dynamically
- Model names now truncated from beginning to preserve full tag (e.g.,
...0B-A3B-Instruct-GGUF:Q4_K_S) - Better visibility of quantization tags for HuggingFace models with long paths
- Load Command Timing - Shows elapsed time and API-reported load duration
- Displays total elapsed time and Ollama's
load_durationfrom response - Example:
✓ Model 'model:tag' loaded successfully (2m 15s) (API: 2m 5s)
- Displays total elapsed time and Ollama's
- QCView Table Alignment Fix - Tag and Quant columns now left-aligned instead of centered
- Timeout Handling Improvements - Better handling of HTTP timeouts during testing
- Timeouts are now properly distinguished from user cancellation (no longer shows "Operation cancelled by user")
- Timeouts trigger retry with exponential backoff instead of immediate failure
- After retry attempts exhausted, prompts user: y=cancel, n=double timeout and retry
- Allows recovery from slow model responses without losing progress
- Improved On-Demand Model Cleanup - Fixed critical bug where models were deleted during testing
- Models with incomplete test results are NEVER cleaned up (preserves for resume)
- Fixed cleanup to protect incomplete models regardless of error type (timeout, cancellation, etc.)
- On-demand status tracking is now consistent when resuming interrupted tests
- Fixed HuggingFace Wildcard Tag Detection - Now detects all quantization formats
- Added support for XL variants (Q2_K_XL, Q3_K_XL, Q4_K_XL, Q5_K_XL, Q6_K_XL, Q8_K_XL)
- Added support for TQ ternary quantization (TQ1_0)
- Fixed HuggingFace Model Quant Column - QC results now correctly show quantization type
- Ollama returns
"quantization_level": "unknown"for HuggingFace models - Now extracts quantization type from model name/tag when API returns "unknown"
- Ollama returns
- Enhanced Quantization Display with Tensor Analysis - Quant column now shows dominant tensor quantization
- Analyzes transformer block weight tensors only (excludes embeddings, output, and norms)
- Calculates weighted percentage by tensor size (elements × bits per weight)
- Displays format like
Q4_0 (87%)orQ6_K (81% Q8_0)showing actual tensor distribution - Uses Ollama API
verbose=trueto fetch tensor metadata - Fixed: Extract quant type from model name before tensor analysis for correct formatting
- Fixed: Filter to transformer weights only (Q8_0 embeddings/output were skewing results)
- Fixed: Unknown tensor types shown with "?" suffix (e.g.,
Q3_K?) to indicate uncertainty - Supports all quantization types: Q*_K variants, IQ (importance matrix), and TQ (ternary)
- Fixed QC Model Validation - Relaxed overly strict parameter size comparison
- Parameter size formatting varies between models (e.g., "999.89M" vs "1,000M" for same model)
- Now only warns on family mismatch instead of blocking testing
- Testing continues even with warnings
- Improved Judge API Retry Strategy - More resilient handling of judge server errors
- Increased retry attempts from 5 to 25 for judge API calls
- Delay ramps from 5 seconds to 30 seconds progressively
- Shows warning and skips judgment only after all retries exhausted (instead of failing)
- Better handles overloaded or slow judge servers (HTTP 500 errors)
- Fixed Base Model Re-Pull When Adding Quants - Skip base model if results already exist
- When adding new quants to existing test results without
-b, no longer tries to pull the base model - If results file contains any base model results (even partial), the base is skipped entirely
- Improved base model detection: automatically identifies base by common patterns (fp16, f16, bf16, etc.)
- Use
--forceto re-run the base model if needed
- When adding new quants to existing test results without
- Improved osync ls Wildcard Handling - Better shell expansion handling on Linux/macOS
- Default behavior:
osync ls codematches models starting with "code" (prefix match, same ascode*) - Suffix match:
osync ls *q4_k_mfinds all models ending with "q4_k_m" (useful for finding by quantization) - Contains match:
osync ls *code*finds models containing "code" anywhere in the name - Shell expansion handling: detects when shell expanded unquoted wildcards and shows helpful warning
- Suggests using quotes to prevent expansion:
osync ls 'gemma*'
- Default behavior:
- Wildcard Tag Expansion for osync pull - Pull multiple models with tag patterns
- Supports wildcards in tags:
osync pull gemma3:1b-it-q*pulls all matching tags - Works with HuggingFace:
osync pull hf.co/unsloth/gemma-3-1b-it-GGUF:IQ2* - Works with remote servers:
osync pull -d http://server:11434 gemma3:1b-it-q* - Automatically resolves available tags from Ollama registry or HuggingFace API
- Shows list of matching tags before pulling
- Supports wildcards in tags:
- Judge Best Answer Tracking - QC judge now evaluates which response is qualitatively better
- Judge model returns
bestanswer: A (base better), B (quant better), or AB (tie) - Verbose output shows best answer for each judgment:
Score: 75% (27/50 54%) Best: AB - Handles edge cases: normalizes various formats (ResponseA, Response_A, Tie, identical, etc.)
- Results automatically re-judged if
--judgeis active andbestansweris missing
- Judge model returns
- QcView Judge Best Column - New column showing quant win statistics
- Format:
67% (B:10 A:5 =:3)showing quant won 67% of non-tie comparisons - B = quant better, A = base better, = = tie
- Best percentage excludes ties (only counts decisive wins/losses)
- Color-coded: green (>=60%), yellow (40-60%), red (<40%)
- Format:
- Enhanced JSON Output - Additional statistics in JSON export
- Per-question
BestAnswerfield (A/B/AB) - Per-quantization:
BestCount,WorstCount,TieCount,BestPercentage,WorstPercentage,TiePercentage - Per-category:
CategoryBestStatswith counts and percentages
- Per-question
- QcView Metrics-Only Mode - New
--metricsonlyargument to ignore judgment data- Shows only metrics-based scores (token similarity, logprobs divergence, perplexity, length consistency)
- Useful for comparing pure model output quality without judge influence
- Works with all output formats (table, json, md, html, pdf)
- Automatic Judge Context Length - Judge model context is now auto-calculated by default
- When
--judge-ctxsizeis 0 (new default), calculates: test_ctx × 2 + 2048 - Ensures judge has enough context for both base and quantized responses plus prompt
- Can still be manually overridden with explicit value
- When
- PDF Generation Progress Bar - Visual progress indicator when generating PDF reports
- Shows progress through Q&A pages for each quantization
- Useful for large test results files with many questions
- PDF Layout Improvements - Better page break handling in PDF reports
- Ranking tables use ShowEntire() to prevent splitting across pages
- Speed columns simplified to show only percentage (removed tok/s to prevent wrapping)
- Category scores section moves entirely to next page if it won't fit
- Rankings organized into paired rows (Final Score + Eval Speed, Perplexity + Prompt Speed, Best Answers)
- Added Prompt Speed ranking table with vs Base percentage column
- Manage TUI Batch Delete Fix - Fixed multi-selection delete not working
- Delete now properly handles multiple selected models (Ctrl+D with checkmarks)
- Confirmation dialog shows count and lists all models to be deleted
- Dialog title shows model count (e.g., "Confirm Delete (3 models)")
- Success message shows count of deleted models
- Partial success handling when some deletions fail
v1.2.5
- QC Resume Bug Fixes - Fixed critical issues with resuming from saved results files
- Fixed model name parsing when resuming: full model paths (e.g.,
hf.co/namespace/repo:tag) are now preserved correctly instead of being incorrectly derived from the-Margument - Fixed base model handling when resuming: the stored full model name is now used instead of just the tag portion
- Fixed verification loop to use stored model names from results file
- Fixed model name parsing when resuming: full model paths (e.g.,
- Cancellation Confirmation Prompt - Added y/n confirmation before cancelling QC tests
- First Ctrl+C now prompts "Cancel testing? (y/n)" instead of immediately cancelling
- Prevents accidental cancellation of long-running tests
- Second Ctrl+C still force exits immediately