From 201a6d48dcc2e561c84b89546f6b6de6ca4c10d9 Mon Sep 17 00:00:00 2001 From: ADou Date: Wed, 9 Sep 2026 09:20:25 +0000 Subject: [PATCH] [None][docs] retarget ModelOpt llm_ptq docs paths to hf_ptq Signed-off-by: ADou --- docs/source/features/quantization.md | 6 +++--- docs/source/torch/features/quantization.md | 6 +++--- 2 files changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/source/features/quantization.md b/docs/source/features/quantization.md index 295314662f24..87d7bfcdf643 100644 --- a/docs/source/features/quantization.md +++ b/docs/source/features/quantization.md @@ -74,7 +74,7 @@ Follow this step-by-step guide to quantize a model: ```bash git clone https://github.com/NVIDIA/Model-Optimizer.git -cd Model-Optimizer/examples/llm_ptq +cd Model-Optimizer/examples/hf_ptq scripts/huggingface_example.sh --model --quant fp8 ``` @@ -84,7 +84,7 @@ To generate the checkpoint for NVFP4 KV cache: ```bash git clone https://github.com/NVIDIA/Model-Optimizer.git -cd Model-Optimizer/examples/llm_ptq +cd Model-Optimizer/examples/hf_ptq scripts/huggingface_example.sh --model --quant fp8 --kv_cache_quant nvfp4 ``` @@ -141,4 +141,4 @@ FP8 block wise scaling GEMM kernels for sm100/103 are using MXFP8 recipe (E4M3 a - [KV Cache Compression](kv-cache-compression.md) - [Pre-quantized Models by ModelOpt](https://huggingface.co/collections/nvidia/model-optimizer-66aa84f7966b3150262481a4) -- [ModelOpt Support Matrix](https://nvidia.github.io/Model-Optimizer/guides/0_support_matrix.html) +- [ModelOpt Support Matrix](https://nvidia.github.io/Model-Optimizer/guides/0_support_matrix.html) \ No newline at end of file diff --git a/docs/source/torch/features/quantization.md b/docs/source/torch/features/quantization.md index 47cc745165b4..a74e53cac430 100644 --- a/docs/source/torch/features/quantization.md +++ b/docs/source/torch/features/quantization.md @@ -13,6 +13,6 @@ Or you can try the following commands to get a quantized model by yourself: ```bash git clone https://github.com/NVIDIA/Model-Optimizer.git -cd Model-Optimizer/examples/llm_ptq -scripts/huggingface_example.sh --model --quant fp8 --export_fmt hf -``` +cd Model-Optimizer/examples/hf_ptq +scripts/huggingface_example.sh --model --quant fp8 +``` \ No newline at end of file