What's Changed
- feat(models): add GLM-5 Next quantization support by @ZX-ModelCloud in #3045
- bench(marlin): add tile-padding performance by @ZX-ModelCloud in #3046
- [MODEL] support
apertus1p5by @ZX-ModelCloud in #3047 - fix: normalize multimodal calibration inputs by @ZX-ModelCloud in #3049
- fix: improve quantization and runtime correctness by @ZX-ModelCloud in #3055
- Shared-input Hessian dedup: plan metadata, E2E validation, and telemetry by @Qubitium in #3052
- support lm-head/embed requant by @ZX-ModelCloud in #3050
- fix: isolate JIT extension cache fingerprint from unrelated kernel changes by @Qubitium in #3056
- Sync native GGUF support with current upstream types by @Qubitium in #3058
- feat: checkpointing and resuming quantization process from last checkpoint by @okdshin in #3057
- Update version.py by @Qubitium in #3060
- Fix Triton patch for pre-existing autotuner instances by @Qubitium in #3062
- [MODEL] support
ouroandspark2_5by @ZX-ModelCloud in #3061 - Sync README release news and model support for 7.4.0 by @Qubitium in #3063
- Document quantization checkpoint and resume workflows by @Qubitium in #3064
New Contributors
Full Changelog: v7.3.6...v7.4.0