Skip to content

Latest commit

 

History

History
executable file
·
470 lines (352 loc) · 32 KB

File metadata and controls

executable file
·
470 lines (352 loc) · 32 KB

NVIDIA Model Optimizer Changelog (Linux)

0.41 (2026-01-19)

Bug Fixes

  • Fix Megatron KV Cache quantization checkpoint restore for QAT/QAD (device placement, amax sync across DP/TP, flash_decode compatibility).

New Features

  • Add support for Transformer Engine quantization for Megatron Core models.
  • Add support for Qwen3-Next model quantization.
  • Add support for dynamically linked TensorRT plugins in the ONNX quantization workflow.
  • Add support for KV Cache Quantization for vLLM FakeQuant PTQ script. See examples/vllm_serve/README.md for more details.
  • Add support for subgraphs in ONNX autocast.
  • Add support for parallel draft heads in Eagle speculative decoding.
  • Add support to enable custom emulated quantization backend. See :meth:`register_quant_backend <modelopt.torch.quantization.nn.modules.tensor_quantizer.register_quant_backend>`` for more details. See an example in tests/unit/torch/quantization/test_custom_backend.py.
  • Add examples/llm_qad for QAD training with Megatron-LM.

Deprecations

  • Deprecate num_query_groups parameter in Minitron pruning (mcore_minitron). You can use ModelOpt 0.40.0 or earlier instead if you need to prune it.

Backward Breaking Changes

  • Remove torchprofile as a default dependency from ModelOpt as its used only for flops-based FastNAS pruning (computer vision models). It can be installed separately if needed.

0.40 (2025-12-12)

Bug Fixes

  • Fix a bug in FastNAS pruning (computer vision models) where the model parameters were sorted twice messing up the ordering.
  • Fix Q/DQ/Cast node placements in 'FP32 required' tensors in custom ops in the ONNX quantization workflow.

New Features

  • Add MoE (e.g. Qwen3-30B-A3B, gpt-oss-20b) pruning support for num_moe_experts, moe_ffn_hidden_size and moe_shared_expert_intermediate_size parameters in Minitron pruning (mcore_minitron).
  • Add specdec_bench example to benchmark speculative decoding performance. See examples/specdec_bench/README.md for more details.
  • Add FP8/NVFP4 KV cache quantization support for Megatron Core models.
  • Add KL Divergence loss based auto_quantize method. See auto_quantize API docs for more details.
  • Add support for saving and resuming auto_quantize search state. This speeds up the auto_quantize process by skipping the score estimation step if the search state is provided.
  • Add flag trt_plugins_precision in ONNX autocast to indicate custom ops precision. This is similar to the flag already existing in the quantization workflow.
  • Add support for PyTorch Geometric quantization.
  • Add per tensor and per channel MSE calibrator support.
  • Added support for PTQ/QAT checkpoint export and loading for running fakequant evaluation in vLLM. See examples/vllm_serve/README.md for more details.

Documentation

Misc

  • NVIDIA TensorRT Model Optimizer is now officially rebranded as NVIDIA Model Optimizer. GitHub will automatically redirect the old repository path (NVIDIA/TensorRT-Model-Optimizer) to the new one (NVIDIA/Model-Optimizer). Documentation URL is also changed to nvidia.github.io/Model-Optimizer.
  • Bump TensorRT-LLM docker to 1.2.0rc4.
  • Bump minimum recommended transformers version to 4.53.
  • Replace ONNX simplification package from onnxsim to onnxslim.

0.39 (2025-11-11)

Deprecations

  • Deprecated modelopt.torch._deploy.utils.get_onnx_bytes API. Please use modelopt.torch._deploy.utils.get_onnx_bytes_and_metadata instead to access the ONNX model bytes with external data. see examples/onnx_ptq/download_example_onnx.py for example usage.

New Features

  • Add flag op_types_to_exclude_fp16 in ONNX quantization to exclude ops from being converted to FP16/BF16. Alternatively, for custom TensorRT ops, this can also be done by indicating 'fp32' precision in trt_plugins_precision.
  • Add LoRA mode support for MCore in a new peft submodule: modelopt.torch.peft.update_model(model, LORA_CFG).
  • Support PTQ and fakequant in vLLM for fast evaluation of arbitrary quantization formats. See examples/vllm_serve for more details.
  • Add support for nemotron-post-training-dataset-v2 and nemotron-post-training-dataset-v1 in examples/llm_ptq. Default to a mix of cnn_dailymail and nemotron-post-training-dataset-v2 (gated dataset accessed using HF_TOKEN environment variable) if no dataset is specified.
  • Allow specifying calib_seq in examples/llm_ptq to set the maximum sequence length for calibration.
  • Add support for MCore MoE PTQ/QAT/QAD.
  • Add support for multi-node PTQ and export with FSDP2 in examples/llm_ptq/multinode_ptq.py. See examples/llm_ptq/README.md for more details.
  • Add support for Nemotron Nano VL v1 & v2 models in FP8/NVFP4 PTQ workflow.
  • Add flags nodes_to_include and op_types_to_include in AutoCast to force-include nodes in low precision, even if they would otherwise be excluded by other rules.
  • Add support for torch.compile and benchmarking in examples/diffusers/quantization/diffusion_trt.py.
  • Enabled native Modelopt quantization support for FP8 and NVFP4 formats in SGLang. See SGLang quantization documentation for more details.
  • Added modelopt quantized checkpoints in vLLM/SGLang CI/CD pipelines (PRs are under review).
  • Add support for exporting QLoRA checkpoint fintuned using ModelOpt.
  • Update NVFP4 AWQ checkpoint export. It now fuses scaling factors of o_proj and down_proj layers into the model when possible to facilitate deployment.

Documentation

0.37 (2025-10-08)

Deprecations

  • Deprecated ModelOpt's custom docker images. Please use the PyTorch, TensorRT-LLM or TensorRT docker image directly or refer to the installation guide for more details.
  • Deprecated quantize_mode argument in examples/onnx_ptq/evaluate.py to support strongly typing. Use engine_precision instead.
  • Deprecated TRT-LLM's TRT backend in examples/llm_ptq and examples/vlm_ptq. Tasks build and benchmark support are removed and replaced with quant. engine_dir is replaced with checkpoint_dir in examples/llm_ptq and examples/vlm_ptq. For performance evaluation, please use trtllm-bench directly.
  • --export_fmt flag in examples/llm_ptq is removed. By default we export to the unified Hugging Face checkpoint format.
  • Deprecated examples/vlm_eval as it depends on the deprecated TRT-LLM's TRT backend.

New Features

  • high_precision_dtype default to fp16 in ONNX quantization, i.e. quantized output model weights are now FP16 by default.
  • Upgrade TensorRT-LLM dependency to 1.1.0rc2.
  • Support Phi-4-multimodal and Qwen2.5-VL quantized HF checkpoint export in examples/vlm_ptq.
  • Support storing and restoring Minitron pruning activations and scores for re-pruning without running the forward loop again.
  • Add Minitron pruning example for Megatron-LM framework. See Megatron-LM/examples/post_training/modelopt for more details.

0.35 (2025-09-04)

Deprecations

  • Deprecate torch<2.6 support.
  • Deprecate NeMo 1.0 model support.

Bug Fixes

  • Fix attention head ranking logic for pruning Megatron Core GPT models.

New Features

  • ModelOpt now supports PTQ and QAT for GPT-OSS models. See examples/gpt_oss for end-to-end PTQ/QAT example.
  • Add support for QAT with HuggingFace + DeepSpeed. See examples/gpt_oss for an example.
  • Add support for QAT with LoRA. The LoRA adapters can be folded into the base model after QAT and deployed just like a regular PTQ model. See examples/gpt_oss for an example.
  • ModelOpt provides convenient trainers such as :class:`QATTrainer`, :class:`QADTrainer`, :class:`KDTrainer`, :class:`QATSFTTrainer` which inherits from Huggingface trainers. ModelOpt trainers can be used as drop in replacement of the corresponding Huggingface trainer. See usage examples in examples/gpt_oss, examples/llm_qat or examples/llm_distill.
  • (Experimental) Add quantization support for custom TensorRT op in ONNX models.
  • Add support for Minifinetuning (MFT; https://arxiv.org/abs/2506.15702) self-corrective distillation, which enables training on small datasets with severely mitigated catastrophic forgetting.
  • Add tree decoding support for Megatron Eagle models.
  • For most VLMs, we now explicitly disable quant on the vision part so we add them to the excluded_modules during HF export.
  • Add support for mamba_num_heads, mamba_head_dim, hidden_size and num_layers pruning for Megatron Core Mamba or Hybrid Transformer Mamba models in mcore_minitron (previously mcore_gpt_minitron) mode.
  • Add example for QAT/QAD training with LLaMA Factory. See examples/llm_qat/llama_factory for more details.
  • Upgrade TensorRT-LLM dependency to 1.0.0rc6.
  • Add unified HuggingFace model export support for quantized NVFP4 GPT-OSS models.

0.33 (2025-07-14)

Backward Breaking Changes

  • PyTorch dependencies for modelopt.torch features are no longer optional and pip install nvidia-modelopt is now same as pip install nvidia-modelopt[torch].

New Features

  • Upgrade TensorRT-LLM dependency to 0.20.
  • Add new CNN QAT example to demonstrate how to use ModelOpt for QAT.
  • Add support for ONNX models with custom TensorRT ops in Autocast.
  • Add quantization aware distillation (QAD) support in llm_qat example.
  • Add support for BF16 in ONNX quantization.
  • Add per node calibration support in ONNX quantization.
  • ModelOpt now supports quantization of tensor-parallel sharded Huggingface transformer models. This requires transformers>=4.52.0.
  • Support quantization of FSDP2 wrapped models and add FSDP2 support in the llm_qat example.
  • Add NeMo 2 Simplified Flow examples for quantization aware training/distillation (QAT/QAD), speculative decoding, pruning & distillation.
  • Fix a Qwen3 MOE model export issue.

0.31 (2025-06-04)

Backward Breaking Changes

  • NeMo and Megatron-LM distributed checkpoint (torch-dist) stored with legacy version can no longer be loaded. The remedy is to load the legacy distributed checkpoint with 0.29 and store a torch checkpoint and resume with 0.31 to convert to a new format. The following changes only apply to storing and resuming distributed checkpoint.
  • auto_quantize API now accepts a list of quantization config dicts as the list of quantization choices.
    • This API previously accepts a list of strings of quantization format names. It was therefore limited to only pre-defined quantization formats unless through some hacks.
    • With this change, now user can easily use their own custom quantization formats for auto_quantize.
    • In addition, the quantization_formats now exclude None (indicating "do not quantize") as a valid format because the auto_quantize internally always add "do not quantize" as an option anyway.
  • Model export config is refactored. The quant config in hf_quant_config.json is converted and saved to config.json. hf_quant_config.json will be deprecated soon.

Deprecations

  • Deprecate Python 3.9 support.

New Features

  • Upgrade LLM examples to use TensorRT-LLM 0.19.
  • Add new model support in the llm_ptq example: Qwen3 MoE.
  • ModelOpt now supports advanced quantization algorithms such as AWQ, SVDQuant and SmoothQuant for cpu-offloaded Huggingface models.
  • Add AutoCast tool to convert ONNX models to FP16 or BF16.
  • Add --low_memory_mode flag in the llm_ptq example support to initialize HF models with compressed weights and reduce peak memory of PTQ and quantized checkpoint export.
  • Support NemotronHForCausalLM, Qwen3ForCausalLM, Qwen3MoeForCausalLM Megatron Core model import/export (from/to HuggingFace).

0.29 (2025-05-08)

Backward Breaking Changes

  • Refactor SequentialQuantizer to improve its implementation and maintainability while preserving its functionality.

Deprecations

  • Deprecate torch<2.4 support.

New Features

0.27 (2025-04-03)

Deprecations

New Features

  • Add new model support in the llm_ptq example: OpenAI Whisper. Experimental support: Llama4, QwQ, Qwen MOE.
  • Add blockwise FP8 quantization support in unified model export.
  • Add quantization support to the Transformer Engine Linear module.
  • Add support for SVDQuant. Currently, only simulation is available; real deployment (for example, TensorRT deployment) support is coming soon.
  • Store modelopt_state in Megatron Core distributed checkpoint (used in NeMo and Megatron-LM) differently to support distributed checkpoint resume expert-parallel (EP). The legacy modelopt_state in the distributed checkpoint generated by previous modelopt version can still be loaded in 0.27 and 0.29 but will need to be stored in the new format.
  • Add triton-based NVFP4 quantization kernel that delivers approximately 40% performance improvement over the previous implementation.
  • Add a new API :meth:`mtq.compress <modelopt.torch.quantization.compress>` for model compression for weights after quantization.
  • Add option to simplify ONNX model before quantization is performed.
  • Add FP4 KV cache support for unified HF and TensorRT-LLM export.
  • Add speculative decoding support to Multi-Token Prediction (MTP) in Megatron Core models.
  • (Experimental) Improve support for ONNX models with custom TensorRT op:
    • Add support for --calibration_shapes flag.
    • Add automatic type and shape tensor propagation for full ORT support with TensorRT EP.

Known Issues

  • Quantization of T5 models is broken. Please use nvidia-modelopt==0.25.0 with transformers<4.50 meanwhile.

0.25 (2025-03-03)

Deprecations

  • Deprecate Torch 2.1 support.
  • Deprecate humaneval benchmark in llm_eval examples. Please use the newly added simple_eval instead.
  • Deprecate fp8_naive quantization format in llm_ptq examples. Please use fp8 instead.

New Features

0.23 (2025-01-29)

Backward Breaking Changes

  • Support TensorRT-LLM to 0.17. Examples (e.g. benchmark task in llm_ptq) may not be fully compatible with TensorRT-LLM 0.15.
  • Nvidia Model Optimizer has changed its LICENSE from NVIDIA Proprietary (library wheel) and MIT (examples) to Apache 2.0 in this first full OSS release.
  • Deprecate Python 3.8, Torch 2.0, and Cuda 11.x support.
  • ONNX Runtime dependency upgraded to 1.20 which no longer supports Python 3.9.
  • In the Huggingface examples, the trust_remote_code is by default set to false and require users to explicitly turning it on with --trust_remote_code flag.

New Features

  • Added OCP Microscaling Formats (MX) for fake quantization support, including FP8 (E5M2, E4M3), FP6 (E3M2, E2M3), FP4, INT8.
  • Added NVFP4 quantization support for NVIDIA Blackwell GPUs along with updated examples.
  • Allows export lm_head quantized TensorRT-LLM checkpoint. Quantize lm_head could benefit smaller sized models at a potential cost of additional accuracy loss.
  • TensorRT-LLM now supports Moe FP8 and w4a8_awq inference on SM89 (Ada) GPUs.
  • New models support in the llm_ptq example: Llama 3.3, Phi 4.
  • Added Minitron pruning support for NeMo 2.0 GPT models.
  • Exclude modules in TensorRT-LLM export configs are now wildcards
  • The unified llama3.1 FP8 huggingface checkpoints can be deployed on SGLang.

0.21 (2024-12-03)

Backward Breaking Changes

New Features

Bug Fixes

  • Improve Minitron pruning quality by avoiding possible bf16 overflow in importance calculation and minor change in hidden_size importance ranking.

Misc

  • Added deprecation warnings for Python 3.8, torch 2.0, and CUDA 11.x. Support will be dropped in the next release.

0.19 (2024-10-23)

Backward Breaking Changes

New Features

  • ModelOpt is compatible for SBSA aarch64 (e.g. GH200) now! Except ONNX PTQ with plugins is not supported.
  • Add effective_bits as a constraint for :meth:`mtq.auto_qauntize <modelopt.torch.quantization.model_quant.auto_quantize>`.
  • lm_evaluation_harness is fully integrated to modelopt backed by TensorRT-LLM. lm_evaluation_harness benchmarks are now available in the examples for LLM accuracy evaluation.
  • A new --perf flag is introduced in the modelopt_to_tensorrt_llm.py example to build engines with max perf.
  • Users can choose the execution provider to run the calibration in ONNX quantization.
  • Added automatic detection of custom ops in ONNX models using TensorRT plugins. This requires the tensorrt python package to be installed.
  • Replaced jax with cupy for faster INT4 ONNX quantization.
  • :meth:`mtq.auto_quantize <modelopt.torch.quantization.model_quant.auto_quantize>` now supports search based automatic quantization for NeMo & MCore models (in addition to HuggingFace models).
  • Add num_layers and hidden_size pruning support for NeMo / Megatron-core models.

0.17 (2024-09-11)

Backward Breaking Changes

New Features

Misc

0.15 (2024-07-25)

Backward Breaking Changes

New Features

  • Added quantization support for torch RNN, LSTM, GRU modules. Only available for torch>=2.0.
  • modelopt.torch.quantization now supports module class based quantizer attribute setting for :meth:`mtq.quantize <modelopt.torch.quantization.model_quant.quantize>` API.
  • Added new LLM PTQ example for DBRX model.
  • Added new LLM (Gemma 2) PTQ and TensorRT-LLM checkpoint export support.
  • Added new LLM QAT example for NVIDIA NeMo framework.
  • TensorRT-LLM dependency upgraded to 0.11.0.
  • (Experimental): :meth:`mtq.auto_quantize <modelopt.torch.quantization.model_quant.auto_quantize>` API which quantizes a model by searching for the best per-layer quantization formats.
  • (Experimental): Added new LLM QLoRA example with NF4 and INT4_AWQ quantization.
  • (Experimental): modelopt.torch.export now supports exporting quantized checkpoints with packed weights for Hugging Face models with namings aligned with its original checkpoints.
  • (Experimental) Added support for quantization of ONNX models with TensorRT plugin.

Misc

  • Added deprecation warning for torch<2.0. Support will be dropped in next release.

0.13 (2024-06-14)

Backward Breaking Changes

New Features

  • Adding TensorRT-LLM checkpoint export support for Medusa decoding (official MedusaModel and Megatron Core GPTModel).
  • Enable support for mixtral, recurrentgemma, starcoder, qwen in PTQ examples.
  • Adding TensorRT-LLM checkpoint export and engine building support for sparse models.
  • Import scales from TensorRT calibration cache and use them for quantization.
  • (Experimental) Enable low GPU memory FP8 calibration for the Hugging Face models when the original model size does not fit into the GPU memory.
  • (Experimental) Support exporting FP8 calibrated model to VLLM deployment.
  • (Experimental) Python 3.12 support added.

0.11 (2024-05-07)

Backward Breaking Changes

  • [!!!] The package was renamed from ammo to modelopt. The new full product name is Nvidia Model Optimizer. PLEASE CHANGE ALL YOUR REFERENCES FROM ammo to modelopt including any paths and links!
  • Default installation pip install nvidia-modelopt will now only install minimal core dependencies. Following optional dependencies are available depending on the features that are being used: [deploy], [onnx], [torch], [hf]. To install all dependencies, use pip install "nvidia-modelopt[all]".
  • Deprecated inference_gpus arg in modelopt.torch.export.model_config_export.torch_to_tensorrt_llm_checkpoint. User should use inference_tensor_parallel instead.
  • Experimental modelopt.torch.deploy module is now available as modelopt.torch._deploy.

New Features

  • modelopt.torch.sparsity now supports sparsity-aware training (SAT). Both SAT and post-training sparsification supports chaining with other modes, e.g. SAT + QAT.
  • modelopt.torch.quantization natively support distributed data and tensor parallelism while estimating quantization parameters. The data and tensor parallel groups needs to be registered with modelopt.torch.utils.distributed.set_data_parallel_group and modelopt.torch.utils.distributed.set_tensor_parallel_group APIs. By default, the data parallel group is set as the default distributed group and the tensor parallel group is disabled.
  • modelopt.torch.opt now supports chaining multiple optimization techniques that each require modifications to the same model, e.g., you can now sparsify and quantize a model at the same time.
  • modelopt.onnx.quantization supports FLOAT8 quantization format with Distribution calibration algorithm.
  • Native support of modelopt.torch.opt with FSDP (Fully Sharded Data Parallel) for torch>=2.1. This includes sparsity, quantization, and any other model modification & optimization.
  • Added FP8 ONNX quantization support in modelopt.onnx.quantization.
  • Added Windows (win_amd64) support for ModelOpt released wheels. Currently supported for modelopt.onnx submodule only.

Bug Fixes

  • Fixed the compatibility issue of modelopt.torch.sparsity with FSDP.
  • Fixed an issue in dynamic dim handling in modelopt.onnx.quantization with random calibration data.
  • Fixed graph node naming issue after opset conversion operation.
  • Fixed an issue in negative dim handling like dynamic dim in modelopt.onnx.quantization with random calibration data.
  • Fixed allowing to accept .pb file for input file.
  • Fixed copy extra data to tmp folder issue for ONNX PTQ.