Release 0.11
AMD Quark for PyTorch
AMD Quark 0.11 is tested against PyTorch 2.9, and compatible with upstream transformers==4.57.
Fused "rotation" and "quarot" algorithms in a single interface
The pre-quantization algorithms "rotation" and "quarot" are fused together into a single rotation algorithm. It can be configured using RotationConfig. By default, only R1 rotation is applied, corresponding to the previous quant_algo="rotation" behavior.
Quark Torch Quantization Config Refactor
-
The quantization configuration classes have been renamed for better clarity and consistency:
QuantizationSpecis deprecated in favor ofQTensorConfig.QuantizationConfigis deprecated in favor ofQLayerConfig.Configis deprecated in favor ofQConfig.
-
The deprecated class names (
QuantizationSpec,QuantizationConfig,Config) are still available as aliases for backward compatibility, but will be removed in a future release. -
Before Refactor:
from quark.torch.quantization.config.config import Config, QuantizationConfig, QuantizationSpec quant_spec = QuantizationSpec(dtype=Dtype.int8, ...) quant_config = QuantizationConfig(weight=quant_spec, ...) config = Config(global_quant_config=quant_config, ...)
-
After Refactor:
from quark.torch.quantization.config.config import QConfig, QLayerConfig, QTensorConfig quant_spec = QTensorConfig(dtype=Dtype.int8, ...) quant_config = QLayerConfig(weight=quant_spec, ...) config = QConfig(global_quant_config=quant_config, ...)
quark torch-llm-ptq CLI Refactor and Simplification
The CLI has been significantly refactored to use the new LLMTemplate interface and remove redundant features:
- Removed model-specific algorithm configuration files (e.g.,
awq_config.json,gptq_config.json,smooth_config.json). Algorithm configurations are now automatically handled byLLMTemplate. - Removed unnecessary CLI arguments, retaining only a dozen or so essential arguments.
- Simplified export: The CLI now only exports to Hugging Face safetensors format.
- Simplified evaluation: Evaluation now uses perplexity (PPL) on wikitext-2 dataset instead of the previous multi-task evaluation framework.
Code Organization and Examples Refactor
Moved common utilities to quark.torch.utils:
model_preparation.pyanddata_preparation.pyare now available inquark.torch.utilsfor easier reuse across examples and applications.module_replacementutilities are now located inquark.torch.utils.module_replacement.
Moved LLM evaluation code to quark.contrib:
- The
llm_evalmodule has been moved toquark.contrib.llm_evalandexamples/contrib/llm_eval. - Perplexity evaluation (
ppl_eval) is now shared between CLI and examples viaquark.contrib.llm_eval.
Reorganized example scripts:
- Removed model-specific algorithm configuration files (e.g.,
awq_config.json,gptq_config.json,smooth_config.json). Algorithm configurations are now automatically handled byLLMTemplate.
Extended quantize_quark.py example script and quark torch-llm-ptq CLI with new features:
- Support for custom model templates and quantization schemes registration (example script only).
- Support for per-layer quantization scheme configuration via
--layer_quant_schemeargument. - Support for custom algorithm configurations via
--quant_algo_config_fileargument (example script only). - Simplified quantization scheme naming, directly use the built-in scheme names (see breaking changes below).
Setting log level with QUARK_LOG_LEVEL
Logging level can now be set with the environment variable QUARK_LOG_LEVEL, e.g. QUARK_LOG_LEVEL=debug or QUARK_LOG_LEVEL=warning or QUARK_LOG_LEVEL=error or QUARK_LOG_LEVEL=critical.
Support for online rotations (online hadamard transform)
The rotation algorithm supports online rotations, such that:
where
Online rotations can be enabled using online_r1_rotation=True in RotationConfig. Please refer to its documentation and to the user guide for more details.
Support for rotation / SmoothQuant scales fine-tuning (SpinQuant/OSTQuant)
We support fine-tuning joint rotations and smoothing scales as a non-destructive transformation
The support is well tested for llama, qwen3, qwen3_moe and gpt_oss architectures.
Rotation fine-tuning and online rotations are compatible with other algorithms as GPTQ or Qronos.
Please refer to the documentation of RotationConfig, the example and the user guide for more details.
Minor changes and bug fixes
- Fix memory duplication and OOM issues when loading
gpt_ossmodels for quantization. ModelQuantizer.freezebehavior is changed to permanently quantize weights. Weights are still in high precision, but QDQ (quantize + dequantize) is run on them. This allows to avoid to rerun QDQ on static weights at each subsequent call.scaled_fake_quantizeoperator, which is used for QDQ, is now by default compiled withtorch.compile, allowing significant speedups depending on the quantization scheme (1x - 8x).- An efficient MXFP4 dynamic quantization kernel is used for activations when quantizing models, fusing scale computation and QDQ operations.
- Batching support is fixed in
lm-evaluation-harnessintegration in the examples, correctly passing the user-provided--eval_batch_size. - CPU/GPU communication is removed in quantization observers, allowing for faster quantization and runtime during e.g. the evaluation of models.
Deprecations and breaking changes
-
Quantization scheme names in
examples/torch/language_modeling/llm_ptq/quantize_quark.pyandquark torch-llm-ptqCLI have been simplified and renamed:w_int4_per_group_symis deprecated in favor ofint4_wo_32,int4_wo_64,int4_wo_128(depending on group size).w_uint4_per_group_asymis deprecated in favor ofuint4_wo_32,uint4_wo_64,uint4_wo_128(depending on group size).w_int8_a_int8_per_tensor_symis deprecated in favor ofint8.w_fp8_a_fp8is deprecated in favor offp8.w_mxfp4_a_mxfp4is deprecated in favor ofmxfp4.w_mxfp4_a_fp8is deprecated in favor ofmxfp4_fp8.w_mxfp6_e3m2_a_mxfp6_e3m2is deprecated in favor ofmxfp6_e3m2.w_mxfp6_e2m3_a_mxfp6_e2m3is deprecated in favor ofmxfp6_e2m3.w_bfp16_a_bfp16is deprecated in favor ofbfp16.w_mx6_a_mx6is deprecated in favor ofmx6.
-
The
--group_sizeand--group_size_per_layerarguments inexamples/torch/language_modeling/llm_ptq/quantize_quark.pyandquark torch-llm-ptqCLI have been removed. Group size is now embedded in the scheme name (e.g.,int4_wo_32,int4_wo_64,int4_wo_128). -
The
--layer_quant_schemeargument format inexamples/torch/language_modeling/llm_ptq/quantize_quark.pyandquark torch-llm-ptqCLI has changed to repeated arguments with pattern and scheme pairs (e.g.,--layer_quant_scheme lm_head int8 --layer_quant_scheme '*down_proj' fp8). -
The token counter used count the number of tokens seen by each expert during calibration is now disabled by default, and requires the environment variable
QUARK_COUNT_OBSERVED_SAMPLES=1. -
The export format
"quark_format"is removed, following deprecation in AMD Quark 0.10. Additionally,quark.torch.export.api.ModelExporterandquark.torch.export.api.ModelImporterare removed, please refer to the 0.10 release notes and to the documentation for the current API.
AMD Quark for ONNX
New Features
-
Auto Search Pro
- Hierarchical Search: Support for conditional and nested hyperparameter trees for advanced search strategies.
- Custom Objectives: Support custom evaluation logic that perfectly aligns with specific needs.
- Sampler Flexibility: Various samplers ('TPE', 'Grid Search', etc) are available .
- Parallel search: Take advantage of parallelization to run multiple searches simultaneously, reducing time to solution.
- Checkpoint: Resume interrupted hyperparameter optimization from the last checkpoint.
- Visualization: View real-time visualizations that show your optimization performance and feature importance, making it easier to interpret results.
- Output Saving: Automatically save the best configuration, study database, and generated plots for your analysis.
-
Latency and memory usage profiling
-
Latency Profiling: Each quantization stage performs specific operations that contribute to the overall quantization pipeline, and their individual latency are reported in the profiling results.
-
Memory profiling
- CPU Memory Profiling: By wrapping the Python script with mprof, we can record detailed memory traces during execution.
- ROCM GPU Memory Profiling: For workflows involving ROCMExecutionProvider or any GPU-based quantization step, Quark ONNX offers a lightweight tool to monitor ROCm GPU memory usage in real time.
-
-
ONNX Adapter: It is a graph transformation tool that can perform graph transformation of preprocessing like constant folding, operator fusion, removal of redundant nodes, streamlining input and output nodes, and optimizing the graph structure.
-
Support 20 preprocessing features
- Convert BatchNormalization operations to Conv operations.
- Convert Clip operations to Relu operations.
- Convert models from FP16 to FP32.
- Convert models from NCHW to NHWC.
- Convert opset version of models.
- Convert ReduceMean operations to GlobalAveragePool operations.
- Convert Split operations to Slice operations.
- Duplicate initializers for shared Bias.
- Duplicate initializers for shared ones.
- Fix shapes for models with dynamic shapes.
- Fold BatchNormalization operations.
- Fold BatchNormalization operations after Concat operations.
- Fuse Gelu operations.
- Fuse InstanceNormalization operations.
- Fuse LpNormalization operations.
- Fuse LayerNormalization operations.
- Optimize models with ONNXRuntime.
- Remove initializers from model inputs.
- Simplify models with OnnxSlim.
- Split GlobalAveragePool operations.
-
Enhancements
-
Support Python 3.12 for Quark ONNX and remove dependency on CMake < 4.0.
-
Enhance tensor-wise mixed precision for integer quantization data types
- Enable the option
TensorQuantOverridesto replace originalMixedPrecisionTensor. - Add support for setting per-tensor or per-channel quantization.
- Add support for setting symmetric or asymmetric quantization.
- Add support for setting more parameters, such as scale, zero_point and etc.
- Prioritize the mixed precision setting when there are multiple settings on the same tensor.
- Enable the option
-
Refactor the codebase to make the quantizer easier to maintain and more reliable in operation
-
Replace ONNX Simplifier with OnnxSlim in preprocessing process before quantization.
-
Allow specific inputs or outputs to be converted from NCHW to NHWC.
-
Refactor the import paths
-
Before refactor:
from quark.onnx import ModelQuantizer from quark.onnx.quantization import QConfig from quark.onnx.quantization.config.spec import QLayerConfig, Int8Spec from quark.onnx.quantization.config.algorithm import CLEConfig, AdaRoundConfig quantization_config = QConfig( # This is a global quantization configuration example using Int8 for activation, weight and bias. If the quantization for the bias is not specified, it will automatically follow the same quantization as the weights. global_config=QLayerConfig(activation=Int8Spec(), weight=Int8Spec()), # For example, quantize the activation, weight, and bias of the two specified nodes using Int16. specific_layer_config={Int16: ["/layer.0/Conv_0", "/layer.11/Conv_2"]}, # For example, quantize the activation, weight, and bias of the all MatMul nodes using Int16 and exclude all Gemm nodes to quantize. layer_type_config={Int16: ["MatMul"], None: ["Gemm"]}, )
-
After refactor:
# All configurations are now imported uniformly from quark.onnx from quark.onnx import ModelQuantizer, QConfig, QLayerConfig, Int8Spec, CLEConfig, AdaRoundConfig quantization_config = QConfig( # Rename activation to input_tensors in QLayerConfig global_config=QLayerConfig(input_tensors=Int8Spec(), weight=Int8Spec()), # Compared to before, it is now to specify the quantization for each tensor of a node. # For example, keep the input_tensors as Int8, quantize the weight and bias using Int16 for two specified nodes. specific_layer_config={QLayerConfig(weight=Int16Spec(), bias=Int16Spec()): ["/layer.0/Conv_0", "/layer.11/Conv_2"]}, # Compared to before, it is now to specify the quantization for each tensor of all nodes of specific operation types. # For example, keep the input_tensors and bias as Int8, only quantize the weight using Int16 for all MatMul nodes and exclude all Gemm nodes to quantize. layer_type_config={QLayerConfig(weight=Int16Spec()): ["MatMul"], None: ["Gemm"]}, )
-
-
Reduce the memory consumption of the default mode of MinMSE to prevent OOM
-
Significantly speedup the calibration process using parallel computation
-
Fixed seed for Fast Finetune
Documentation
- Removed ONNXRuntime dependency from Quark for simplified environment setup.
Bug fixes and minor improvements
-
Fixed percentile value selection for LayerwisePercentile
-
Fixed the out-of-bounds axis issue when weight or bias is a scalar in BFP and MX quantization
-
Fixed bug for replacing clip with ReLU operator.