Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
2857 commits
Select commit Hold shift + click to select a range
e047900
more silent_token
aluminumbox Dec 30, 2025
f030a3f
more silent_token
aluminumbox Dec 30, 2025
8fc255f
update tokens
aluminumbox Dec 30, 2025
a2abd89
update tokens
aluminumbox Dec 30, 2025
df30970
update tokens
aluminumbox Dec 30, 2025
d292002
Merge branch 'main' into main
aluminumbox Dec 31, 2025
db0fadb
Merge branch 'main' into main
aluminumbox Dec 31, 2025
089d681
Merge branch 'main' into main
aluminumbox Dec 31, 2025
358718a
Merge pull request #1640 from Jzz1943/main
aluminumbox Dec 31, 2025
9267cbb
Merge pull request #1640 from Jzz1943/main
aluminumbox Dec 31, 2025
d4059a4
Merge pull request #1640 from Jzz1943/main
aluminumbox Dec 31, 2025
d731cd3
update readme
aluminumbox Dec 31, 2025
0c18489
update readme
aluminumbox Dec 31, 2025
9b3a2a1
update readme
aluminumbox Dec 31, 2025
8981fe4
Merge pull request #1758 from orbisai0security/fix/V-005-pickle-deser…
aluminumbox Dec 31, 2025
f981f60
Merge pull request #1758 from orbisai0security/fix/V-005-pickle-deser…
aluminumbox Dec 31, 2025
3cda348
Merge pull request #1758 from orbisai0security/fix/V-005-pickle-deser…
aluminumbox Dec 31, 2025
01f7aa6
fix bug
aluminumbox Jan 5, 2026
6410121
fix bug
aluminumbox Jan 5, 2026
adb0e0f
fix bug
aluminumbox Jan 5, 2026
a2d19cd
Merge branch 'main' of github.com:FunAudioLLM/CosyVoice into main
aluminumbox Jan 5, 2026
764cf45
Merge branch 'main' of github.com:FunAudioLLM/CosyVoice into main
aluminumbox Jan 5, 2026
a001745
Merge branch 'main' of github.com:FunAudioLLM/CosyVoice into main
aluminumbox Jan 5, 2026
34354bd
fix sequence logic
aluminumbox Jan 7, 2026
1be8229
fix sequence logic
aluminumbox Jan 7, 2026
8fb7b13
fix sequence logic
aluminumbox Jan 7, 2026
94b719f
fix padding
aluminumbox Jan 7, 2026
68b9fcc
fix padding
aluminumbox Jan 7, 2026
6dd2afe
fix padding
aluminumbox Jan 7, 2026
c313f54
Update README with citation and license information
ZhenghuaBao Jan 9, 2026
e27d2b8
Format LICENSE file for better readability
ZhenghuaBao Jan 9, 2026
bffbee5
add japanese
aluminumbox Jan 12, 2026
59fcae7
add japanese
aluminumbox Jan 12, 2026
0fc3198
add japanese
aluminumbox Jan 12, 2026
c3d16f5
Update model name from Chroma-4B to Chroma 1.0
ZhenghuaBao Jan 15, 2026
d677fc3
Merge pull request #1250 from raivisdejus/add-latvian-community-model
raivisdejus Jan 17, 2026
2b7309e
update higgs audio v2.5 in readme (#169)
iamtonymwt Jan 18, 2026
63115f1
Update README.md
ZhenghuaBao Jan 19, 2026
11a14ff
fix fm train bug
aluminumbox Jan 19, 2026
341186c
fix fm train bug
aluminumbox Jan 19, 2026
ed370e8
fix fm train bug
aluminumbox Jan 19, 2026
6b6a65e
Fix Technical Report link in README
ZhenghuaBao Jan 19, 2026
a56e001
Fix link for Technical Report in README.md
ZhenghuaBao Jan 19, 2026
9f4db25
Enhance README with getting started section
ZhenghuaBao Jan 19, 2026
74cfe0e
Update Chroma section to link to Voice Agents
ZhenghuaBao Jan 19, 2026
ab30dab
Update contact information in README
ZhenghuaBao Jan 19, 2026
6337b41
Adding support for hf:// links on CLI (#1252)
raivisdejus Jan 21, 2026
0f0a8c1
Fix epoch updates count logic as drop_last
SWivid Jan 21, 2026
8ba0857
Formatting
SWivid Jan 21, 2026
a3b62e4
Update README.md
yishi888 Jan 21, 2026
544889d
change prepare_csv_wavs from relative path to absolute path and get d…
ZhikangNiu Jan 22, 2026
b4f4645
fix many tensorboard writer and only log in main_process
ZhikangNiu Jan 22, 2026
eb92f88
doc: Remove outdated playground
ZhenghuaBao Jan 22, 2026
6eecadd
fig: Update model architecture
ZhenghuaBao Jan 22, 2026
ac678b1
feat: compact latest transformers
Jan 22, 2026
8a93ac1
add tqdm in convert text to pinyin
ZhikangNiu Jan 22, 2026
afcbef3
update: example
kaishen-Dotc Jan 22, 2026
2cf4bc2
update: torch requirements
kaishen-Dotc Jan 22, 2026
b048ed9
Initial commit
wangxiongts Jan 22, 2026
947c6d5
Merge pull request #1256 from ZhikangNiu/main
SWivid Jan 22, 2026
db7c5d2
fix: depandency updates
terrencechentr Jan 22, 2026
4ed3a72
fix: hf speech_tokenizer download bug.
wangxiongts Jan 22, 2026
9cc9591
Update: readme quick start
sailorjs0804 Jan 23, 2026
a432575
fix and update: fix finetuning/prepare_data.py and update citation
wangxiongts Jan 23, 2026
838a849
update: modify default non_streaming_mode
wangxiongts Jan 23, 2026
9ad399d
fixup wrong flags
vasqu Jan 23, 2026
4136752
fix loading logic
vasqu Jan 23, 2026
7e23766
Merge pull request #15 from vasqu/fix-fa-flags
wangxiongts Jan 24, 2026
1582846
update: update readme
wangxiongts Jan 24, 2026
8d7257d
Optimizing 14GB Model on 4GB VRAM
agentifyanchor Jan 24, 2026
b4abb3b
Increase default max_duration from 4096 to 65536 in cfm.py (#1260)
SWivid Jan 25, 2026
f46b482
update: update readme
wangxiongts Jan 25, 2026
d62d8eb
Ignore padding at the end of the ground truth mel spectrogram when tr…
ZhikangNiu Jan 26, 2026
db8aaba
doc: Add social media links to README
ZhenghuaBao Jan 26, 2026
697cbed
Merge pull request #1261 from ZhikangNiu/main
SWivid Jan 26, 2026
a5e0269
doc: Uploaded Demo
ZhenghuaBao Jan 28, 2026
4794309
Update eval README with ctranslate2 installation instructions
SWivid Jan 28, 2026
ee6a5e7
online feature
aluminumbox Jan 28, 2026
da98124
online feature
aluminumbox Jan 28, 2026
adef3a1
online feature
aluminumbox Jan 28, 2026
1ad4bd1
Update package versions in requirements.txt
agentifyanchor Jan 28, 2026
beeb59d
update
aluminumbox Jan 29, 2026
bf64c7e
update
aluminumbox Jan 29, 2026
507bdc5
update
aluminumbox Jan 29, 2026
8e346f4
update
aluminumbox Jan 29, 2026
64a1d3f
update
aluminumbox Jan 29, 2026
e7aa3a7
update
aluminumbox Jan 29, 2026
35160c0
update
aluminumbox Jan 29, 2026
21a7e04
update
aluminumbox Jan 29, 2026
d11b83e
update
aluminumbox Jan 29, 2026
a0bf61b
[BUG FIX] 使用 float64 避免精度误差问题,弃用 CPU 计算,避免拖累性能
hexisyztem Jan 30, 2026
800f6ac
[BUG FIX] 使用 float64 避免精度误差问题,弃用 CPU 计算,避免拖累性能
hexisyztem Jan 30, 2026
3ee3150
[BUG FIX] 使用 float64 避免精度误差问题,弃用 CPU 计算,避免拖累性能
hexisyztem Jan 30, 2026
0c085cb
Update run_asr_wer method in utils_eval.py for compat with jiwer>=4.0.0
SWivid Feb 2, 2026
a8bdff9
fix: adjust finetune script for stable training.
wangxiongts Feb 4, 2026
02a279a
fix ras
aluminumbox Feb 4, 2026
fa8fcf1
fix ras
aluminumbox Feb 4, 2026
c1a5b77
fix ras
aluminumbox Feb 4, 2026
9307563
Merge pull request #1814 from hexisyztem/main
aluminumbox Feb 4, 2026
ac752ea
Merge pull request #1814 from hexisyztem/main
aluminumbox Feb 4, 2026
ce24b7c
Merge pull request #1814 from hexisyztem/main
aluminumbox Feb 4, 2026
28e2d8b
fix: padding bugs in 12Hz tokenizer decode.
wangxiongts Feb 5, 2026
b369e89
fix: padding value process bug in tokenizer decode
wangxiongts Feb 6, 2026
9b3b127
keep offline embedding/token extraction for compatibale
aluminumbox Feb 11, 2026
1a543cf
keep offline embedding/token extraction for compatibale
aluminumbox Feb 11, 2026
935e52b
keep offline embedding/token extraction for compatibale
aluminumbox Feb 11, 2026
0714299
simply code
aluminumbox Feb 11, 2026
a965211
simply code
aluminumbox Feb 11, 2026
1a82bdd
simply code
aluminumbox Feb 11, 2026
ec05f94
fix vllm yaml version
aluminumbox Feb 11, 2026
4eb32be
fix vllm yaml version
aluminumbox Feb 11, 2026
cba07ad
fix vllm yaml version
aluminumbox Feb 11, 2026
21c520d
Use torch.utils.checkpoint in mmdit forward loop when enabled to redu…
ZhikangNiu Feb 14, 2026
4426300
Merge pull request #1265 from ZhikangNiu/main
SWivid Feb 14, 2026
7be59ae
Make wandb project/run_name/resume_id configurable via Hydra yaml, ba…
ZhikangNiu Feb 16, 2026
c190f69
Optimize DiT text embedding with batched per-sample seq handling
QingyuLiu0521 Feb 16, 2026
bb09d83
Apply ruff formatting
QingyuLiu0521 Feb 16, 2026
df34adf
Merge pull request #1266 from ZhikangNiu/main
SWivid Feb 16, 2026
2ac4c5c
Unify seq_len naming in DiT get_input_embed
QingyuLiu0521 Feb 16, 2026
d909862
Merge pull request #1267 from QingyuLiu0521/qyl/pr-dit-only
SWivid Feb 16, 2026
4b4e0c5
Bump version from 1.1.15 to 1.1.16
SWivid Feb 16, 2026
68ee421
feat:add mmdit flash attn support
ZhikangNiu Feb 23, 2026
a9c46bb
when use flash_attn, log_sample should under autocast context
ZhikangNiu Feb 25, 2026
2a67653
Merge pull request #1269 from ZhikangNiu/main
SWivid Feb 26, 2026
f09101d
Add show_info parameter to preprocess_ref_audio_text
mlxu995 Mar 4, 2026
de08e8f
Merge pull request #1271 from mlxu995/patch-1
SWivid Mar 4, 2026
ade043c
Merge pull request #1270 from ZhikangNiu/main
ZhikangNiu Mar 4, 2026
11244f7
Bump version from 1.1.16 to 1.1.17
SWivid Mar 4, 2026
bd6fee8
Remove pydantic<=2.10.6 restriction to fit latest gradio version
SWivid Mar 7, 2026
df9a222
add cosyvoice3
Mar 16, 2026
ed9be05
add cosyvoice3
Mar 16, 2026
a26f2ec
add cosyvoice3
Mar 16, 2026
3fdb5f5
rename files
Mar 16, 2026
ffabefc
rename files
Mar 16, 2026
c365957
rename files
Mar 16, 2026
ddcb033
update results
Mar 16, 2026
58fcff5
update results
Mar 16, 2026
3b9c7bc
update results
Mar 16, 2026
b5358f0
update results
Mar 16, 2026
7027b11
update results
Mar 16, 2026
6bfbf52
update results
Mar 16, 2026
c700a16
fix lint
yuekaizhang Mar 16, 2026
092ddeb
fix lint
yuekaizhang Mar 16, 2026
d151834
fix lint
yuekaizhang Mar 16, 2026
afbc130
Merge pull request #1850 from yuekaizhang/cosy3_pr
aluminumbox Mar 16, 2026
abcd47b
Merge pull request #1850 from yuekaizhang/cosy3_pr
aluminumbox Mar 16, 2026
893044a
Merge pull request #1850 from yuekaizhang/cosy3_pr
aluminumbox Mar 16, 2026
c6b603a
Add Arabic model details to SHARED.md (#1279)
karimouda Mar 16, 2026
e58a068
fix finetuning bug
wangxiongts Mar 17, 2026
06f4fe4
F5TTS v1 Small + LibriTTS training config
ZhikangNiu Mar 23, 2026
7248a60
Merge pull request #1280 from ZhikangNiu/main
SWivid Mar 23, 2026
0177e80
remove ineffective ThreadPoolExecutor in infer_batch_process
zhuxiaoxuhit Mar 24, 2026
b6c6224
Merge pull request #1281 from zhuxiaoxuhit/fix/remove-ineffective-thr…
SWivid Mar 24, 2026
1e8dde1
Several fixes for utils_infer.py; separate streaming and non-streamin…
SWivid Mar 24, 2026
fa79f21
Bump version from 1.1.17 to 1.1.18
SWivid Mar 24, 2026
cb5d9fc
Init commit
zhu-han Apr 1, 2026
8464824
update paper
zhu-han Apr 2, 2026
a414406
update link
zhu-han Apr 2, 2026
490700f
Fix/mps cloning interactive (#13)
Talpik Apr 3, 2026
a07899b
Fix single gpu fine-tuning (#23)
zhu-han Apr 3, 2026
c8e6e6c
reuse resamplers and cache vocos MelSpectrogram instances, it will re…
ZhikangNiu Apr 4, 2026
615754e
Merge pull request #1285 from ZhikangNiu/main
SWivid Apr 4, 2026
449eaf7
Fix non-verbal generation (#35)
zhu-han Apr 4, 2026
fbf598d
Bump version from 0.1.1 to 0.1.2
zhu-han Apr 4, 2026
4b62ec1
fix torchaudio backend
zhu-han Apr 7, 2026
06aed41
fix text processing for non-verbal and special symbols
zhu-han Apr 7, 2026
4af4797
add no_asr and instruct options to demo
zhu-han Apr 7, 2026
7d2ce24
relax pytorch version requirements
zhu-han Apr 7, 2026
8db07d6
Add tips in readme
zhu-han Apr 7, 2026
c751d2b
Merge pull request #61 from zhu-han/fix_20260407
zhu-han Apr 7, 2026
3c48642
fix: instruct in infer_batch
Pastells Apr 9, 2026
b86ff09
Merge pull request #72 from Pastells/fix/instruct-batch
zhu-han Apr 9, 2026
b4e72bb
fix: batch_inference without ref_text or ref_audio_path (#70)
Pastells Apr 11, 2026
96efd9c
docs: add omnivoice-server to community projects (#80)
maemreyo Apr 13, 2026
cc9af56
fix infer_batch.py for mixed modes
zhu-han Apr 13, 2026
d71a3c0
use soundfile+librosa instead of torchaudio to avoid issues on some d…
zhu-han Apr 13, 2026
bf2f87b
replace Chinese parentheses with English ones in text
zhu-han Apr 13, 2026
faff117
add more tips in README
zhu-han Apr 13, 2026
c696a8f
add google colab examples
zhu-han Apr 13, 2026
45463a9
Bump version to 0.1.4
zhu-han Apr 13, 2026
eae5c3a
restore community projects
zhu-han Apr 13, 2026
0417096
docs: add omnivoice-rs to community projects (#94)
FerrisMind Apr 13, 2026
8d69b02
set load_asr=True in colab notebook
zhu-han Apr 13, 2026
8e9a721
remove language_name docs as it is not used by the model
zhu-han Apr 13, 2026
6329250
add issue templates
zhu-han Apr 14, 2026
b7fdfec
switch to torchaudio resampling for consistency with training
zhu-han Apr 14, 2026
02d5a13
Fixes #1287 #1288
SWivid Apr 16, 2026
299cb54
Bump version from 1.1.18 to 1.1.19
SWivid Apr 16, 2026
02cc149
fix: cap Gradio version to <6.11 to prevent UI freeze
will422-l Apr 17, 2026
de23852
Merge pull request #1290 from will422-l/fix/gradio-version-cap
SWivid Apr 17, 2026
76654f5
Merge pull request #16 from agentifyanchor/main
yi-here Apr 17, 2026
68b519f
v1.1.20: refactor cache handling in DiT, MMDiT, and UNetT classes (la…
SWivid Apr 20, 2026
fc63c4e
fix model loading without internet
zhu-han Apr 20, 2026
bae1d21
docs: add omnivoice-trtllm to community projects (#110)
tlitech Apr 20, 2026
04ff39e
fix unmask schedule
zhu-han Apr 20, 2026
66efc3d
add more community projects
zhu-han Apr 20, 2026
a2fd9d2
Update README
zhu-han Apr 22, 2026
bb438ba
Support training with SDPA
zhu-han Apr 22, 2026
8256496
update docs
zhu-han Apr 22, 2026
5f9033a
add more tips
zhu-han Apr 28, 2026
9dc27ce
Bump version to 0.1.5
zhu-han Apr 28, 2026
c5c015c
Revise disclaimer wording in README
pkufool Apr 30, 2026
90da463
Merge pull request #136 from k2-fsa/pkufool-patch-1
pkufool Apr 30, 2026
b3bbbdd
docs: Add Auris to community projects (#145)
nikhilprasanth May 6, 2026
0da0d2f
fix: path traversal in finetune Gradio handlers (closes #1293)
AAtomical May 12, 2026
80b19f1
Merge pull request #1294 from AAtomical/fix/path-traversal-finetune-g…
SWivid May 13, 2026
a663e28
update AMD ROCm install to 7.2 for RDNA 3.5/4 support
Kaihui-AMD May 18, 2026
ff24bdb
Merge pull request #1298 from Kaihui-AMD/docs/update-amd-rocm-install
SWivid May 18, 2026
c1667b3
Add Intel XPU (Intel Arc GPU) inference support (#55)
smashingtags May 19, 2026
6037217
Add FunAudioLLM ecosystem section with links to ASR projects
LauraGPT May 25, 2026
055a851
Add FunAudioLLM ecosystem section with links to ASR projects
LauraGPT May 25, 2026
67e1954
Add FunAudioLLM ecosystem section with links to ASR projects
LauraGPT May 25, 2026
071ce7e
Add OmniVoice-MLX to community projects (#165)
ailuntx May 28, 2026
8b40215
Add tts-audiobook-tool project to community list (#167)
zeropointnine May 28, 2026
3ce000a
Update docs
zhu-han May 28, 2026
1330a67
docs: point README to Higgs Audio v3 (HF + API), archive v2/v2.5 guid…
SilinMeng0510 Jun 5, 2026
734d99a
Fix END_PUNCTUATION dropping members due to unescaped triple quote (#…
MinDongRyul Jun 11, 2026
451f23d
Add LA Studio to community projects (#195)
dduongtrandai Jun 24, 2026
f77501c
Expose pad_duration and fade_duration to let users control fade-in/ou…
zhu-han Jul 3, 2026
9ec2124
v1.1.21 gradio>=6.15.0
SWivid Jul 5, 2026
4ab2c75
Add CI workflows (PyPI publish, pre-commit) and format codebase with …
zhu-han Jul 6, 2026
996f73b
Bump version to 0.2.0
zhu-han Jul 6, 2026
c33fd18
Add audio.cpp to community projects (#209)
0xShug0 Jul 9, 2026
d037d48
Update formatting rules
zhu-han Jul 9, 2026
2876c6e
Fix flex_attention silently running fp32 under bf16 mixed precision (…
shuaills Jul 11, 2026
8102b02
Honor asr_model_name when the ASR model is loaded lazily (#221)
HenryVarro666 Jul 13, 2026
404783f
Add VoiceClonePrompt.save()/load() for cross-session voice reuse (#223)
HenryVarro666 Jul 15, 2026
30cafb3
Allow placing the ASR model on a chosen device (#224)
HenryVarro666 Jul 15, 2026
426fd2a
Avoid decoding the full reference audio just to get its duration (#220)
aradix85 Jul 15, 2026
bd1c9c9
Add opt-in text normalization (numbers, dates, currency) via generate…
HenryVarro666 Jul 15, 2026
0bbb883
Add stale issue workflow (#228)
zhu-han Jul 16, 2026
6cb3544
docs: add LocalText2Voice to community projects (#230)
estebanstifli Jul 18, 2026
d1f810e
fix: allocate fix_duration across chunks to avoid N× duration blow-up
Jul 23, 2026
2606bcc
Merge pull request #1309 from Wing900/fix/fix-duration-chunk-allocation
SWivid Jul 23, 2026
b71ccc5
v1.1.22
SWivid Jul 23, 2026
17988dc
feat: FlashInfer-accelerated inference (2-2.9x lossless speedup) (#239)
yuekaizhang Jul 30, 2026
f94ba6a
feat: add optional LoRA finetuning support via PEFT (#243)
kawshikbuet17 Aug 5, 2026
e1570e8
Add voicestudio/models/__init__.py registry, refresh status tracking …
b-re-w Aug 19, 2026
f432c86
Graft upstream git history for dia
b-re-w Aug 19, 2026
22e81ab
Graft upstream git history for higgs_audio_v2
b-re-w Aug 19, 2026
cabef6e
Graft upstream git history for parler_tts
b-re-w Aug 19, 2026
0494cf0
Graft upstream git history for qwen3_tts
b-re-w Aug 19, 2026
240e6fb
Graft upstream git history for f5_tts
b-re-w Aug 19, 2026
55014de
Graft upstream git history for cosyvoice_v1
b-re-w Aug 19, 2026
a5b5cfa
Graft upstream git history for cosyvoice_v2
b-re-w Aug 19, 2026
129bd89
Graft upstream git history for cosyvoice_v3
b-re-w Aug 19, 2026
2ec0592
Graft upstream git history for omnivoice
b-re-w Aug 19, 2026
2916707
Graft upstream git history for spark_tts
b-re-w Aug 19, 2026
bb6aa27
Graft upstream git history for chroma
b-re-w Aug 19, 2026
d426a5c
Graft upstream git history for prompt_tts_pp
b-re-w Aug 19, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
52 changes: 34 additions & 18 deletions PROJECT.md
Original file line number Diff line number Diff line change
Expand Up @@ -178,21 +178,37 @@ VoiceStudio to restore it — don't fork the whole processor.

## Status tracking

Update this table as each model's migration lands.

| Model | Status | Notes |
|---|---|---|
| Qwen3-TTS | In progress | Model/config are import relays to transformers-tts. Processor subclass adds `encode`/`encode_voice_design`/`encode_custom_voice` task dispatch with `RuntimeError` on task mismatch. `encode_voice_clone` raises `NotImplementedError`: transformers-tts's `Qwen3TTSProcessor` has no reference-audio input path yet. |
| Parler-TTS | Not started | Second in migration order. Target: use HF-registered `dac`, closest transformers lineage TBD |
| Higgs-Audio v2 | Done | Import relay to transformers-tts. Reference processor pattern for other audio-tokenizer models |
| Higgs-Audio v3 | In progress | `bosonai/higgs-tts-3-4b`. Weights-only checkpoint (no upstream v3 code); backbone is a plain `Qwen3Model` (no dual-FFN). Reuses transformers-tts's `Qwen3Model`, `HiggsAudioV2Embeddings`, `HiggsAudioV2PreTrainedModel`, and `HiggsAudioV2TokenizerModel`; `HiggsAudioV3Model`/`ForConditionalGeneration` are new since v2's dual-FFN decoder layer does not fit. No checkpoint conversion or real-weights test run yet; the reference-audio prompt layout in `HiggsAudioV3Processor` is unverified against the real checkpoint. |
| Chroma | In progress | Backbone/decoder/generation loop reimplemented against transformers-tts's Llama, Qwen2.5-Omni thinker, and Mimi codec classes (no full Chroma architecture exists upstream in transformers-tts to relay to). Processor subclasses `Qwen2_5OmniProcessor` to add the reference-audio voice-cloning prompt. |
| Spark-TTS | Not started | Already vendored in `dep/Spark-TTS` |
| Dia | Not started | Already marked "fully tested (by HF)" in old README |
| CosyVoice (v1/v2/v3) | In progress | Three separate model folders (`cosyvoice_v1`, `cosyvoice_v2`, `cosyvoice_v3`), matching the upstream repo's own class inheritance: v2's Qwen2-backbone LLM and causal flow decoder subclass v1's flow-matching classes, v3's LLM and DiT flow decoder subclass v2's. `cosyvoice_v1` reuses `Wav2Vec2ConformerEncoder` for the text encoder/LLM backbone; `cosyvoice_v2`/`v3` reuse `Qwen2Model`. Flow-matching U-Net/DiT estimators and the HiFTNet vocoder are new code (no transformers-tts lineage). `dep/CosyVoice` deleted. See each folder's README for known gaps (JIT/TRT/vLLM/streaming/bistream not ported; speech tokenizer and speaker encoder are external ONNX artifacts). |
| F5-TTS | In progress | Full reimplementation (DiT flow-matching, no existing transformers-tts lineage). `RMSNorm` reuses `LlamaRMSNorm`. Model predicts mel spectrograms only; `F5TTSProcessor.decode` requires an external vocoder (e.g. `vocos`) to render audio. `dep/F5-TTS` deleted. |
| PromptTTS++ (`prompt_tts_pp`) | In progress | Acoustic model and vocoder are `FastSpeech2Conformer`/`FastSpeech2ConformerHifiGan` from transformers-tts, conditioned via `speaker_embedding`. The BERT-based prompt encoder is newly implemented since it has no equivalent in transformers-tts. No pretrained checkpoint conversion done yet (upstream ships no public checkpoint). |
| audiotools dependency removal | Not started | |
| vocos dependency removal | Not started | |
| speechbrain fork removal | Not started | Check if upstream now supports the required torch version |
| UTMOSv2 decoupling | Not started | Route through `evaluate`, may need `voicestudio/metrics/` |
Update this table as each model's migration lands. "Code landed" means the folder exists,
imports, and passed a static review pass (architecture, CE-training requirement, CLAUDE.md
conventions) with known issues fixed. It does NOT mean the model has been run against a real
pretrained checkpoint yet; see "Runtime-verified" for that.

| Model | Code landed | Runtime-verified (real checkpoint) | Upstream git history preserved | Notes |
|---|---|---|---|---|
| Qwen3-TTS | Yes | No | No | Import relays to transformers-tts. Processor subclass adds `encode`/`encode_voice_design`/`encode_custom_voice` task dispatch with `RuntimeError` on task mismatch. `encode_voice_clone` raises `NotImplementedError`. |
| Parler-TTS | Yes | No | No | Import bug (`isin_mps_friendly`) fixed. Still a ~3300-line near-verbatim port with scattered `# Copied from` comments rather than a `modular_parler_tts.py`; revisit for a proper modular conversion. |
| Higgs-Audio v2 | Yes | No | No | Import relay to transformers-tts. Reviewed clean, no issues found. |
| Higgs-Audio v3 | Yes | No | No | `bosonai/higgs-tts-3-4b`, weights-only checkpoint (no upstream v3 code). Reuses transformers-tts's `Qwen3Model`/`HiggsAudioV2*` classes; `HiggsAudioV3Model`/`ForConditionalGeneration` are new. `audio_labels`-only crash and default-off text CE loss fixed. |
| Chroma | Yes | No | No | Backbone/decoder reimplemented against transformers-tts's Llama, Qwen2.5-Omni thinker, and Mimi codec classes. Processor kwargs-merging and a labels-dereference guard fixed. |
| Spark-TTS | Yes | No | No | Config previously had no LLM sub-config, so plain construction silently produced an untrained model; fixed. License header consolidated onto `modeling_spark_tts.py` only; dead code and processor/model duplication cleaned up. |
| Dia | Yes | No | No | Import relay to transformers-tts's native Dia. Reviewed clean. |
| CosyVoice v1 | Yes | No | No | `Wav2Vec2ConformerEncoder`-based text encoder/LLM backbone; own flow-matching/vocoder code. |
| CosyVoice v2 | Yes | No | No | Subclasses v1; `Qwen2Model` LLM backbone. |
| CosyVoice v3 | Yes | No | No | Subclasses v2. DiT attention-mask shape bug (crashed on padded batches) fixed. |
| F5-TTS | Yes | No | No | Full reimplementation, DiT flow-matching. Predicts mel spectrograms only; `F5TTSProcessor.decode` needs an external vocoder. `forward()` previously had no `labels`/loss path at all; added. |
| PromptTTS++ (`prompt_tts_pp`) | Yes | No | No | `FastSpeech2Conformer`/`FastSpeech2ConformerHifiGan` from transformers-tts, conditioned via a newly-implemented BERT-based prompt encoder. `return_dict=False` tuple-indexing bug fixed. No public upstream checkpoint exists. |
| OmniVoice | Yes | No | No | No transformers-tts lineage (closest in spirit to CSM/Moshi); modeling code is new, audio tokenizer reused from transformers-tts's `HiggsAudioV2TokenizerModel`. Training-time sample masking and reference-audio auto-transcription (ASR) not ported. |
| audiotools dependency removal | Not started | | | |
| vocos dependency removal | Not started | | | |
| speechbrain fork removal | Not started | | | Check if upstream now supports the required torch version |
| UTMOSv2 decoupling | Not started | | | Route through `evaluate`, may need `voicestudio/metrics/` |

### Known gap: upstream git history

The per-model migration procedure above calls for preserving each upstream repo's git
history via `git subtree`/`filter-repo`. In practice none of the models above got this: every
folder's earliest commit in this repo is a fresh "Add X" commit with no upstream authorship in
its ancestry. Retrofitting this means, per model: clone upstream fresh, `git filter-repo
--to-subdirectory-filter` to rewrite its history under that model's folder path, then graft it
onto `develop`'s history (e.g. `git merge --allow-unrelated-histories -X ours`, since the
grafted paths land in a non-colliding subfolder) and force-push. In progress.
32 changes: 32 additions & 0 deletions voicestudio/models/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
from . import (
chroma,
cosyvoice_v1,
cosyvoice_v2,
cosyvoice_v3,
dia,
f5_tts,
higgs_audio_v2,
higgs_audio_v3,
omnivoice,
parler_tts,
prompt_tts_pp,
qwen3_tts,
spark_tts,
)


__all__ = [
"chroma",
"cosyvoice_v1",
"cosyvoice_v2",
"cosyvoice_v3",
"dia",
"f5_tts",
"higgs_audio_v2",
"higgs_audio_v3",
"omnivoice",
"parler_tts",
"prompt_tts_pp",
"qwen3_tts",
"spark_tts",
]
207 changes: 207 additions & 0 deletions voicestudio/models/chroma/_history/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,207 @@
# Byte-compiled / optimized / DLL files
__pycache__/
*.py[codz]
*$py.class

# C extensions
*.so

# Distribution / packaging
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib/
lib64/
parts/
sdist/
var/
wheels/
share/python-wheels/
*.egg-info/
.installed.cfg
*.egg
MANIFEST

# PyInstaller
# Usually these files are written by a python script from a template
# before PyInstaller builds the exe, so as to inject date/other infos into it.
*.manifest
*.spec

# Installer logs
pip-log.txt
pip-delete-this-directory.txt

# Unit test / coverage reports
htmlcov/
.tox/
.nox/
.coverage
.coverage.*
.cache
nosetests.xml
coverage.xml
*.cover
*.py.cover
.hypothesis/
.pytest_cache/
cover/

# Translations
*.mo
*.pot

# Django stuff:
*.log
local_settings.py
db.sqlite3
db.sqlite3-journal

# Flask stuff:
instance/
.webassets-cache

# Scrapy stuff:
.scrapy

# Sphinx documentation
docs/_build/

# PyBuilder
.pybuilder/
target/

# Jupyter Notebook
.ipynb_checkpoints

# IPython
profile_default/
ipython_config.py

# pyenv
# For a library or package, you might want to ignore these files since the code is
# intended to run in multiple environments; otherwise, check them in:
# .python-version

# pipenv
# According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
# However, in case of collaboration, if having platform-specific dependencies or dependencies
# having no cross-platform support, pipenv may install dependencies that don't work, or not
# install all needed dependencies.
#Pipfile.lock

# UV
# Similar to Pipfile.lock, it is generally recommended to include uv.lock in version control.
# This is especially recommended for binary packages to ensure reproducibility, and is more
# commonly ignored for libraries.
#uv.lock

# poetry
# Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
# This is especially recommended for binary packages to ensure reproducibility, and is more
# commonly ignored for libraries.
# https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
#poetry.lock
#poetry.toml

# pdm
# Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
# pdm recommends including project-wide configuration in pdm.toml, but excluding .pdm-python.
# https://pdm-project.org/en/latest/usage/project/#working-with-version-control
#pdm.lock
#pdm.toml
.pdm-python
.pdm-build/

# pixi
# Similar to Pipfile.lock, it is generally recommended to include pixi.lock in version control.
#pixi.lock
# Pixi creates a virtual environment in the .pixi directory, just like venv module creates one
# in the .venv directory. It is recommended not to include this directory in version control.
.pixi

# PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
__pypackages__/

# Celery stuff
celerybeat-schedule
celerybeat.pid

# SageMath parsed files
*.sage.py

# Environments
.env
.envrc
.venv
env/
venv/
ENV/
env.bak/
venv.bak/

# Spyder project settings
.spyderproject
.spyproject

# Rope project settings
.ropeproject

# mkdocs documentation
/site

# mypy
.mypy_cache/
.dmypy.json
dmypy.json

# Pyre type checker
.pyre/

# pytype static type analyzer
.pytype/

# Cython debug symbols
cython_debug/

# PyCharm
# JetBrains specific template is maintained in a separate JetBrains.gitignore that can
# be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
# and can be added to the global gitignore or merged into this file. For a more nuclear
# option (not recommended) you can uncomment the following to ignore the entire idea folder.
#.idea/

# Abstra
# Abstra is an AI-powered process automation framework.
# Ignore directories containing user credentials, local state, and settings.
# Learn more at https://abstra.io/docs
.abstra/

# Visual Studio Code
# Visual Studio Code specific template is maintained in a separate VisualStudioCode.gitignore
# that can be found at https://github.com/github/gitignore/blob/main/Global/VisualStudioCode.gitignore
# and can be added to the global gitignore or merged into this file. However, if you prefer,
# you could uncomment the following to ignore the entire vscode folder
.vscode/

# Ruff stuff:
.ruff_cache/

# PyPI configuration file
.pypirc

# Cursor
# Cursor is an AI-powered code editor. `.cursorignore` specifies files/directories to
# exclude from AI features like autocomplete and code analysis. Recommended for sensitive data
# refer to https://docs.cursor.com/context/ignore-files
.cursorignore
.cursorindexingignore

# Marimo
marimo/_static/
marimo/_lsp/
__marimo__/
Loading
Loading