Skip to content

mtmd: add KaniTTS-2 support - #28609

Draft
hans00 wants to merge 4 commits into
ggml-org:masterfrom
hans00:mtmd-kani-tts2
Draft

mtmd: add KaniTTS-2 support#28609
hans00 wants to merge 4 commits into
ggml-org:masterfrom
hans00:mtmd-kani-tts2

Conversation

@hans00

@hans00 hans00 commented Sep 8, 2026

Copy link
Copy Markdown

Overview

Support https://huggingface.co/nineninesix/kani-tts-2-en and https://huggingface.co/nineninesix/kani-tts-2-pt
It based on LFM2 backbone and NeMo Nano Codec 22 kHz / 0.6 kbps decoder.

Additional information

Requirements

@github-actions github-actions Bot added documentation Improvements or additions to documentation model Model specific mtmd Related to multimodal functionality (video/image/audio) conversion labels Sep 8, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 8, 2026

Copy link
Copy Markdown

Hi @hans00, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

Convert the WavLM speaker encoder and backbone speaker projection into the NeMo mmproj, encode reference audio through MTMD, and insert the projected embedding into the Kani prompt.

Assisted-by: Codex
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion documentation Improvements or additions to documentation model Model specific mtmd Related to multimodal functionality (video/image/audio)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant