-
Notifications
You must be signed in to change notification settings - Fork 2.8k
[None][docs] retarget ModelOpt llm_ptq docs paths to hf_ptq #18960
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -74,7 +74,7 @@ Follow this step-by-step guide to quantize a model: | |
|
|
||
| ```bash | ||
| git clone https://github.com/NVIDIA/Model-Optimizer.git | ||
| cd Model-Optimizer/examples/llm_ptq | ||
| cd Model-Optimizer/examples/hf_ptq | ||
| scripts/huggingface_example.sh --model <huggingface_model_card> --quant fp8 | ||
| ``` | ||
|
|
||
|
|
@@ -84,7 +84,7 @@ To generate the checkpoint for NVFP4 KV cache: | |
|
|
||
| ```bash | ||
| git clone https://github.com/NVIDIA/Model-Optimizer.git | ||
| cd Model-Optimizer/examples/llm_ptq | ||
| cd Model-Optimizer/examples/hf_ptq | ||
| scripts/huggingface_example.sh --model <huggingface_model_card> --quant fp8 --kv_cache_quant nvfp4 | ||
| ``` | ||
|
|
||
|
|
@@ -141,4 +141,4 @@ FP8 block wise scaling GEMM kernels for sm100/103 are using MXFP8 recipe (E4M3 a | |
|
|
||
| - [KV Cache Compression](kv-cache-compression.md) | ||
| - [Pre-quantized Models by ModelOpt](https://huggingface.co/collections/nvidia/model-optimizer-66aa84f7966b3150262481a4) | ||
| - [ModelOpt Support Matrix](https://nvidia.github.io/Model-Optimizer/guides/0_support_matrix.html) | ||
| - [ModelOpt Support Matrix](https://nvidia.github.io/Model-Optimizer/guides/0_support_matrix.html) | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This line is byte-identical to before — the only change is removing the trailing newline. |
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
🔎 Supported by static analysis
🏁 Script executed:
Repository: NVIDIA/TensorRT-LLM
Length of output: 4359
🏁 Script executed:
Repository: NVIDIA/TensorRT-LLM
Length of output: 9852
🤖 get_repo_knowledge executed:
get_repo_knowledge NVIDIA/Model-OptimizerLength of output: 129
🌐 Web query:
site:github.com/NVIDIA/Model-Optimizer hf_ptq examples huggingface_example.sh export_fmt💡 Result:
The term hf_ptq refers to the Post-Training Quantization (PTQ) example directory within the NVIDIA Model Optimizer repository (formerly often referred to as llm_ptq) [1]. This directory contains tools and scripts for quantizing Hugging Face models for deployment, particularly with TensorRT-LLM [2][3]. The huggingface_example.sh script is a primary entry point in this directory used to automate the quantization and export process [2][3]. It delegates the core quantization logic to the hf_ptq.py script [3]. Regarding the export_fmt parameter: In the context of the NVIDIA Model Optimizer, export_fmt is an argument used in the quantization process to define the output format of the quantized model [4][5]. While the script often handles export paths automatically, users may specify formats to ensure compatibility with downstream deployment engines, such as TensorRT-LLM or specific Hugging Face-compatible structures [4][5]. Key features of these tools include: - Recipe-driven quantization: Users are encouraged to use --recipe to load predefined quantization configurations, which can be more robust than manually specifying individual parameters [2][6]. - Quantization Formats: Supported formats include fp8, int8, int4_awq, and newer formats like nvfp4, which are designed for specific hardware architectures (e.g., Blackwell GPUs) [2][3]. - Workflow: The general workflow involves providing a model path (local or Hugging Face ID), choosing a quantization format or recipe, and specifying an export path [2][3]. If an enc-dec model or specific sparse quantization is detected, the script may advise the user to use specific deployment tools like trtllm-build [3]. For more information, consult the official repository documentation at https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq [2].
Citations:
🏁 Script executed:
Repository: NVIDIA/TensorRT-LLM
Length of output: 14577
🏁 Script executed:
Repository: NVIDIA/TensorRT-LLM
Length of output: 1867
Update remaining Model Optimizer commands to the current contract.
The current
examples/hf_ptqparser accepts--model,--quant, and related options, but not--export_fmt. Update the Exaone, Qwen, disaggregated, and other remainingexamples/llm_ptqreferences to useexamples/hf_ptqand remove the obsolete option.📍 Affects 2 files
docs/source/features/quantization.md#L77-L77(this comment)docs/source/features/quantization.md#L87-L87docs/source/torch/features/quantization.md#L16-L18🤖 Prompt for AI Agents