diff --git a/docs/source/features/quantization.md b/docs/source/features/quantization.md index 295314662f24..87d7bfcdf643 100644 --- a/docs/source/features/quantization.md +++ b/docs/source/features/quantization.md @@ -74,7 +74,7 @@ Follow this step-by-step guide to quantize a model: ```bash git clone https://github.com/NVIDIA/Model-Optimizer.git -cd Model-Optimizer/examples/llm_ptq +cd Model-Optimizer/examples/hf_ptq scripts/huggingface_example.sh --model --quant fp8 ``` @@ -84,7 +84,7 @@ To generate the checkpoint for NVFP4 KV cache: ```bash git clone https://github.com/NVIDIA/Model-Optimizer.git -cd Model-Optimizer/examples/llm_ptq +cd Model-Optimizer/examples/hf_ptq scripts/huggingface_example.sh --model --quant fp8 --kv_cache_quant nvfp4 ``` @@ -141,4 +141,4 @@ FP8 block wise scaling GEMM kernels for sm100/103 are using MXFP8 recipe (E4M3 a - [KV Cache Compression](kv-cache-compression.md) - [Pre-quantized Models by ModelOpt](https://huggingface.co/collections/nvidia/model-optimizer-66aa84f7966b3150262481a4) -- [ModelOpt Support Matrix](https://nvidia.github.io/Model-Optimizer/guides/0_support_matrix.html) +- [ModelOpt Support Matrix](https://nvidia.github.io/Model-Optimizer/guides/0_support_matrix.html) \ No newline at end of file diff --git a/docs/source/torch/features/quantization.md b/docs/source/torch/features/quantization.md index 47cc745165b4..a74e53cac430 100644 --- a/docs/source/torch/features/quantization.md +++ b/docs/source/torch/features/quantization.md @@ -13,6 +13,6 @@ Or you can try the following commands to get a quantized model by yourself: ```bash git clone https://github.com/NVIDIA/Model-Optimizer.git -cd Model-Optimizer/examples/llm_ptq -scripts/huggingface_example.sh --model --quant fp8 --export_fmt hf -``` +cd Model-Optimizer/examples/hf_ptq +scripts/huggingface_example.sh --model --quant fp8 +``` \ No newline at end of file