Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -357,7 +357,7 @@ recordings with the same evaluation scope.

- **MOSS-Transcribe-Diarize** brings long-form ASR, timestamps, and anonymous speaker labels to FunASR services, Docker, Kubernetes, vLLM/SGLang workflows, and FunClip. [Deploy MOSS ->](./docs/moss_transcribe_diarize.md)
- **FunASR 1.4.15** adds tested NumPy 2 compatibility and fixes streaming KWS/VAD boundaries and checkpoint ranking. Install with `python -m pip install -U "funasr==1.4.15"`. [Release and verification scope ->](https://github.com/modelscope/FunASR/releases/tag/v1.4.15)
- **Production deployment** now includes faster, more resilient realtime serving and verified llama.cpp archives for ten Linux, macOS, and Windows targets. [GPU services ->](./docs/vllm_guide.md) · [CPU/edge packages ->](https://www.funasr.com/en/deploy/llama-cpp.html)
- **Native Transformers:** Fun-ASR-Nano [merged upstream](https://github.com/huggingface/transformers/pull/46180). Use the official `-hf` checkpoint and pinned source; stable 5.16.1 does not include it. [Installation and inference ->](./docs/transformers_native.md)

> See [GitHub Releases](https://github.com/modelscope/FunASR/releases) for the complete changelog and downloadable assets.

Expand Down
2 changes: 1 addition & 1 deletion README_ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,7 @@ CER/WER をそろえて比較してください。オフラインのスループ

- **MOSS-Transcribe-Diarize** を FunASR service、Docker、Kubernetes、vLLM/SGLang workflow、FunClip に統合し、長時間 ASR、timestamp、匿名 speaker label を一度に処理できます。[MOSS をデプロイ ->](./docs/moss_transcribe_diarize.md)
- **FunASR 1.4.15** はテスト済みの NumPy 2 互換性を追加し、ストリーミング KWS/VAD の境界処理と checkpoint の順位付けを修正します。`python -m pip install -U "funasr==1.4.15"`。[リリースと検証範囲 ->](https://github.com/modelscope/FunASR/releases/tag/v1.4.15)
- **Production deployment** に、より高速で安定した realtime serving と、Linux、macOS、Windows の 10 target 向け llama.cpp package を追加しました。[GPU service ->](./docs/vllm_guide.md) · [CPU / edge package ->](https://www.funasr.com/en/deploy/llama-cpp.html) · [v0.2.6 binaries ->](https://github.com/modelscope/FunASR/releases/tag/runtime-llamacpp-v0.2.6)
- **ネイティブ Transformers:** Fun-ASR-Nano が[上流にマージ](https://github.com/huggingface/transformers/pull/46180)されました。公式 `-hf` checkpoint と固定ソースを使用します。安定版 5.16.1 には未収録です。[導入ガイド(英語) ->](./docs/transformers_native.md)

> 完全な変更履歴と download asset は [GitHub Releases](https://github.com/modelscope/FunASR/releases) を参照してください。

Expand Down
2 changes: 1 addition & 1 deletion README_ko.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,7 +104,7 @@ checkpoint/revision, 오디오 집합, 하드웨어, 배치, 워밍업, 측정

- **MOSS-Transcribe-Diarize**를 FunASR service, Docker, Kubernetes, vLLM/SGLang workflow, FunClip에 통합해 긴 오디오 ASR, timestamp, 익명 speaker label을 한 번에 처리합니다. [MOSS 배포 ->](./docs/moss_transcribe_diarize.md)
- **FunASR 1.4.15**는 테스트를 거친 NumPy 2 호환성을 추가하고 스트리밍 KWS/VAD 경계 처리와 checkpoint 순위 산정을 수정합니다. `python -m pip install -U "funasr==1.4.15"`. [릴리스 및 검증 범위 ->](https://github.com/modelscope/FunASR/releases/tag/v1.4.15)
- **Production deployment**에 더 빠르고 안정적인 realtime serving과 Linux, macOS, Windows 10개 target용 llama.cpp package를 추가했습니다. [GPU service ->](./docs/vllm_guide.md) · [CPU / edge package ->](https://www.funasr.com/en/deploy/llama-cpp.html) · [v0.2.6 binaries ->](https://github.com/modelscope/FunASR/releases/tag/runtime-llamacpp-v0.2.6)
- **네이티브 Transformers:** Fun-ASR-Nano가 [업스트림에 병합](https://github.com/huggingface/transformers/pull/46180)되었습니다. 공식 `-hf` checkpoint와 고정 소스를 사용하세요. 안정 버전 5.16.1에는 아직 포함되지 않았습니다. [설치 가이드(영문) ->](./docs/transformers_native.md)

> 전체 변경 기록과 download asset은 [GitHub Releases](https://github.com/modelscope/FunASR/releases)에서 확인할 수 있습니다.

Expand Down
2 changes: 1 addition & 1 deletion README_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -150,7 +150,7 @@ checkpoint/revision、音频集、硬件、批量大小、预热、计时范围

- **MOSS-Transcribe-Diarize** 已接入 FunASR 服务、Docker、Kubernetes、vLLM/SGLang 工作流和 FunClip,一次完成长音频转写、时间戳与匿名说话人标注。[部署 MOSS ->](./docs/moss_transcribe_diarize_zh.md)
- **FunASR 1.4.15** 新增经过测试的 NumPy 2 兼容支持,修复流式 KWS/VAD 边界处理和 checkpoint 排序。升级命令:`python -m pip install -U "funasr==1.4.15"`。[发布说明与验证范围 ->](https://github.com/modelscope/FunASR/releases/tag/v1.4.15)
- **工业部署** 新增更快、更稳定的实时服务,以及覆盖 Linux、macOS、Windows 十种目标的 llama.cpp 预编译包。[GPU 服务 ->](./docs/vllm_guide_zh.md) · [CPU/端侧包 ->](https://www.funasr.com/deploy/llama-cpp.html)
- **原生 Transformers:** Fun-ASR-Nano [已合入上游](https://github.com/huggingface/transformers/pull/46180)。使用官方 `-hf` checkpoint 与固定源码;稳定版 5.16.1 尚未包含。[安装与推理 ->](./docs/transformers_native_zh.md)

> 完整改动记录和可下载资产请查看 [GitHub Releases](https://github.com/modelscope/FunASR/releases)。

Expand Down
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@
| Install and check the Python environment | [Installation](installation/installation.md) / [中文](installation/installation_zh.md) |
| Choose a model, language, or timestamp capability | [Model selection](model_selection.md) / [中文](model_selection_zh.md) |
| Run the first transcription | [Quickstart](tutorial/README.md) / [中文](tutorial/README_zh.md) / [CLI](cli.md) |
| Use native Hugging Face Transformers APIs | [Fun-ASR-Nano native guide](transformers_native.md) / [中文](transformers_native_zh.md) |
| Configure AutoModel, generate, and streaming cache | [Python SDK](python_api.md) / [中文](python_api_zh.md) |
| Move from Whisper or a cloud API | [Migration](migration_from_whisper.md) / [中文](migration_from_whisper_zh.md) |

Expand Down
1 change: 1 addition & 0 deletions docs/deployment_matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ Use this page to choose the shortest deployment path for a product, demo, benchm

| Path | Best for | Start here | Operational notes |
|---|---|---|---|
| Native Transformers | Python applications using Hugging Face processors and generation | [Native Fun-ASR-Nano guide](./transformers_native.md) | Official `-hf` checkpoint and pinned source; stable 5.16.1 lacks support at the 2026-09-09 check. Python inference, not an HTTP or realtime server. |
| Colab notebook | Browser smoke tests, first evaluation, shareable demos | [Colab quickstart](../examples/colab/) | No local setup; first run downloads model files, GPU runtime is faster. |
| Python API | Notebooks, offline jobs, first model evaluation | [README quick start](../README.md#quick-start) | Lowest ceremony; caller owns batching, retries, and files. |
| OpenAI-compatible API | Private speech API, agents, Dify/LangChain/AutoGen-style clients | [OpenAI API example](../examples/openai_api/) | Easiest integration for apps that already support OpenAI audio APIs. |
Expand Down
1 change: 1 addition & 0 deletions docs/deployment_matrix_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@

| 路径 | 适合场景 | 从这里开始 | 运维提示 |
|---|---|---|---|
| 原生 Transformers | 已使用 Hugging Face 处理器与生成接口的 Python 应用 | [原生 Fun-ASR-Nano 指南](./transformers_native_zh.md) | 官方 `-hf` checkpoint 与固定源码;2026-09-09 核验的稳定版 5.16.1 尚未包含。是 Python 推理入口,不是 HTTP 或实时服务器。 |
| Colab Notebook | 浏览器 smoke test、首次评估、可分享 demo | [Colab 快速体验](../examples/colab/README_zh.md) | 不需要本地环境;首次运行会下载模型,GPU runtime 更快。 |
| Python API | Notebook、离线任务、首次模型评测 | [README 快速开始](../README_zh.md#快速开始) | 最简单;调用方自己负责批处理、重试和文件管理。 |
| OpenAI 兼容 API | 私有语音 API、Agent、Dify/LangChain/AutoGen 风格客户端 | [OpenAI API 示例](../examples/openai_api/README_zh.md) | 已支持 OpenAI audio API 的应用最容易接入。 |
Expand Down
208 changes: 208 additions & 0 deletions docs/transformers_native.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,208 @@
# Fun-ASR-Nano with native Transformers

[简体中文](./transformers_native_zh.md) | English

Use this path when your Python application already works with Hugging Face
processors and generation APIs. It runs Fun-ASR-Nano through native Transformers
classes, without the FunASR toolkit or checkpoint-provided remote Python code.
For a managed HTTP service, start with the [deployment matrix](./deployment_matrix.md);
this guide does not add an OpenAI endpoint, request queue or realtime server.

## What was merged, and what was released?

On **2026-09-09**, [Transformers PR #46180](https://github.com/huggingface/transformers/pull/46180)
was merged as `fc501343edfccdc840eb8594a6cafa8185c2de53`.
At verification time that day, the latest stable **5.16.1 did not contain this
model**: both its release tag and actual wheel were inspected. A merge is not
a stable release. The source build identifies itself as `5.17.0.dev0`; that is
not a promised release version or date. Pin the tested source below, or check
that a later released package really contains `fun_asr_nano` before migrating.

## Choose the checkpoint and the interface together

| Path | Checkpoint | Entry point |
| --- | --- | --- |
| FunASR toolkit | `FunAudioLLM/Fun-ASR-Nano-2512` | `funasr.AutoModel`; separate documented split-engine path |
| Native Transformers | `FunAudioLLM/Fun-ASR-Nano-2512-hf` | `AutoProcessor` + `AutoModelForSpeechSeq2Seq` |
| Native vLLM | `FunAudioLLM/Fun-ASR-Nano-2512-vllm` | [Official native vLLM guide](./vllm_official_native_validation.md) |
| Native C++ | Converted GGUF files for the chosen runtime | [llama.cpp guide](../runtime/llama.cpp/README.md) |

These are different artifacts and API contracts. Renaming a model ID or passing
`backend="vllm"` does not convert weights. Native Transformers support does not
make the `-hf` checkpoint a vLLM, GGUF or FunASR HTTP-server checkpoint.

The official [native HF checkpoint](https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-hf/tree/d93b302ee7fd505e1b3576120fc142fc6f7820e1)
is public and ungated at revision `d93b302ee7fd505e1b3576120fc142fc6f7820e1`.
Its configuration maps to `FunAsrNanoForConditionalGeneration`. Its feature
extractor is embedded in `processor_config.json`; a separate
`preprocessor_config.json` is not required for this snapshot.

## Independent CPU environment

Use a new directory for the virtual environment; do not upgrade an existing
FunASR or vLLM service environment in place. The following is **Linux x86-64,
Python 3.12, CPU-only**, not a CUDA or macOS installation recipe.

```bash
python3.12 -m venv .venv-funasr-native
. .venv-funasr-native/bin/activate
python -m pip install --index-url https://download.pytorch.org/whl/cpu \
'torch==2.10.0+cpu' 'torchaudio==2.10.0+cpu'
python -m pip install 'numpy==1.26.4' 'librosa==0.11.0' 'soundfile==0.13.1' \
'huggingface-hub==1.30.0' 'tokenizers==0.23.2' \
'https://github.com/huggingface/transformers/archive/fc501343edfccdc840eb8594a6cafa8185c2de53.zip'
python -m pip check
```

The verification environment installed a wheel built from the exact source
commit above. Matching **torch and torchaudio** are required for this native
feature extractor: it calls `torchaudio.compliance.kaldi.fbank`. The toolkit's
optional audio dependencies do not imply that torchaudio is optional here.
Older prepared environments failed dependency gates; do not disable those gates
or mix incompatible Hub/tokenizers versions to get an import to succeed.
This is a tested environment, not a complete lock of every transitive dependency;
record `python -m pip freeze`, platform details and model revision in your own run.

## Check preprocessing without loading weights

This downloads the public configuration, tokenizer and chat template, **not
model weights**. One second of synthetic silence checks the entry point; it is
not a speech-recognition quality test.

<!-- native-example: processor -->
```python
import numpy as np
from transformers import AutoProcessor

model_id = "FunAudioLLM/Fun-ASR-Nano-2512-hf"
revision = "d93b302ee7fd505e1b3576120fc142fc6f7820e1"
processor = AutoProcessor.from_pretrained(
model_id, revision=revision, trust_remote_code=False, token=False
)
inputs = processor.apply_transcription_request(
audio=np.zeros(16000, dtype=np.float32),
language="en",
processor_kwargs={
"return_tensors": "pt",
"audio_kwargs": {"sampling_rate": 16000},
},
)
print({name: tuple(value.shape) for name, value in inputs.items()})
```

In the verified environment the text tensors have shape `(1, 42)`, audio features
`(1, 17, 560)`, and feature mask `(1, 17)`. There are 17 audio placeholder tokens
and 17 valid feature frames. These shapes depend on input length and template;
do not hard-code them for real recordings.

## Transcribe one recording

Prepare a non-empty **mono 16 kHz WAV** named `audio.wav` in your working directory.
The code deliberately rejects a different sample rate or channel layout rather
than silently assigning 16 kHz to arbitrary samples. Resample/downmix with an
explicit policy before this step and preserve the original recording for review.

The first model load downloads the weights unless the pinned snapshot is already
cached. It needs substantially more memory and time than preprocessing. This
example uses CPU float32 explicitly; GPU placement, mixed precision, attention
backends and service concurrency need separate validation.

<!-- native-example: transcribe -->
```python
import soundfile as sf
import torch
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

model_id = "FunAudioLLM/Fun-ASR-Nano-2512-hf"
revision = "d93b302ee7fd505e1b3576120fc142fc6f7820e1"
audio, sample_rate = sf.read("audio.wav", dtype="float32")
if sample_rate != 16000 or audio.ndim != 1 or audio.size == 0:
raise ValueError("Use a non-empty mono 16 kHz WAV file")

processor = AutoProcessor.from_pretrained(
model_id, revision=revision, trust_remote_code=False, token=False
)
model = AutoModelForSpeechSeq2Seq.from_pretrained(
model_id, revision=revision, trust_remote_code=False, token=False,
dtype=torch.float32,
).to("cpu").eval()
inputs = processor.apply_transcription_request(
audio=audio, language="zh",
processor_kwargs={
"return_tensors": "pt",
"audio_kwargs": {"sampling_rate": sample_rate},
},
)
with torch.inference_mode():
generated = model.generate(**inputs, max_new_tokens=128, do_sample=False)
new_tokens = generated[:, inputs.input_ids.shape[1]:]
print(processor.batch_decode(new_tokens, skip_special_tokens=True)[0])
```

`generated` includes the input prompt. Slice off `inputs.input_ids.shape[1]`
before decoding so prompt framing is not mistaken for recognized speech.
`max_new_tokens=128` bounds this short-file example; hitting the limit can
truncate output. A longer limit is not proof of complete coverage.
Use `language="en"` or `language="ja"` for the corresponding target language;
language selection is a transcription instruction, not a promise of translation.

## Context, hotwords and batches

`apply_transcription_request` accepts `prompt` for context and `keywords` for
hotwords. These are native processor arguments, not the toolkit's `hotword`
parameter or a vLLM HTTP field. The source API accepts one language for the batch
or a language list matching its length; per-recording prompt and nested keyword
lists must also match the batch. Keep the decoded results in input order.

Do not infer quality gains from a template accepting a keyword. Evaluate names,
numbers and unusual terms against authorized, representative recordings.
Batch padding, memory use and generated-text correctness must be tested together
before turning a single-file example into a production batch worker.

## Evidence and limits

The documented single-recording recipe was also executed on the official Chinese
sample, followed by an English single recording, a Chinese keyword request and a
Chinese/English batch. All returned non-empty text and EOS before the 128-token
limit. These are functional checks on two short public samples, not a CER/WER or
capacity evaluation. CPU: Intel Xeon Platinum 8480+, float32, four Torch threads;
matching torch/torchaudio 2.10.0+cpu and librosa 0.11.0. Model weights were loaded
from a previously downloaded official cache whose SHA/size were rechecked.

Samples came from the original model's
[fixed example directory](https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512/tree/272c57b82523ada6fd87095e955f8e29100979ab/example),
not from an invented -hf example directory. The mono 48 kHz Chinese MP3 was
explicitly resampled with librosa's default soxr_hq and written as float32 16 kHz
WAV before running the unmodified recipe. Its raw text was
`开饭时间早上九点至下午五点。`; adding the candidate keyword `开放时间` did not force that
spelling into the output. A keyword is a hint, not an enforced vocabulary.

Synthetic processor tests exercised language aliases, context/keyword lists and
two-length batches. The empty-list case still raises an upstream `IndexError`
before its intended validation; reject empty batches in your application.
The guide's explicit non-empty waveform check avoids that entry path. Successful
cases do not erase this observed limitation. Initial resampling also exposed the
old librosa 0.10.1 dependency on removed `pkg_resources`; the final environment
uses 0.11.0 and was retested.

The 2026-09-09 isolated CPU verification passed package dependency checks,
native configuration/class resolution and synthetic preprocessing with the
pinned official revision. Installed configuration, modeling, processing and
feature-extraction files were byte-identical to the merged upstream source.
A preprocessing pass is **not an accuracy or capacity result**. It does not
establish word timestamps, diarization, realtime behavior, GPU compatibility,
or performance equivalence with vLLM.

For anonymous speakers and segment timestamps, see the third-party OpenMOSS
[MOSS guide](./moss_transcribe_diarize.md); for other choices see the
[Model Zoo](../model_zoo/readme.md). Do not infer timestamp or identity support
from a model returning text. Preserve input hashes, versions, raw responses and
an authorized human review when assessing an application. Never attach private
recordings or credentials to a public issue.

## Sources and next steps

- [Merged implementation and usage documentation](https://github.com/huggingface/transformers/blob/fc501343edfccdc840eb8594a6cafa8185c2de53/docs/source/en/model_doc/fun_asr_nano.md).
- [Native processor](https://github.com/huggingface/transformers/blob/fc501343edfccdc840eb8594a6cafa8185c2de53/src/transformers/models/fun_asr_nano/processing_fun_asr_nano.py) and [audio feature extractor](https://github.com/huggingface/transformers/blob/fc501343edfccdc840eb8594a6cafa8185c2de53/src/transformers/models/fun_asr_nano/feature_extraction_fun_asr_nano.py).
- [Why the checkpoint suffix matters](https://www.funasr.com/en/blog/fun-asr-nano-transformers.html), an application-oriented walkthrough.
- [Fun-ASR model project](https://github.com/QwenAudio/Fun-ASR) and [FunASR toolkit](https://github.com/modelscope/FunASR).
Loading
Loading