Skip to content

Repository files navigation

EasyFunASR — Chinese Voice Input (FunASR + GNOME)

English | 中文

A resident-GPU Chinese voice-input tool for Linux. Double-tap Ctrl (or Super+Alt+V) to start/stop recording — recognized text is auto-pasted at your cursor. The ASR model stays loaded in VRAM, so there's no cold start; recognition takes ~2 s.

Features

  • 🎙️ Model resident in GPU VRAM — instant record-and-paste (no 30 s cold start)
  • ⌨️ Trigger by double-tap Ctrl / Super+Alt+V / tray menu
  • 📊 Live status icon in the top bar (idle / recording / stopped / starting)
  • 🔧 One-line installer (Ubuntu/Debian/Fedora/Arch; auto-detects GPU / desktop / driver)
  • 🔒 Fully local, runs on open-source FunASR models — nothing leaves your machine

Screenshots

Tray menu

Requirements

  • Linux with an X11 session (GNOME recommended; Wayland is not supported for paste — see Known limitations)
  • NVIDIA GPU (optional; falls back to slow CPU mode without one)
  • Python ≥ 3.10
  • Desktop: GNOME (other desktops need manual hotkey binding)

Quick install

git clone https://github.com/huchi996/easyfunasr.git
cd easyfunasr
./install.sh

install.sh will: detect distro → install system packages → create .venv → install torch (cu128 if GPU, else CPU) + funasr → generate autostart entries → register the GNOME hotkey. It needs sudo for the system packages.

easyfunasr start      # start everything (first run loads the model, ~30 s)

Usage

  • Double-tap Ctrl (two presses within 0.35 s) → start recording; double-tap again → stop, text is pasted at the cursor
  • Or Super+Alt+V
  • Or click the microphone icon in the top bar → menu "Toggle / On / Off"
  • Key chords (Ctrl+C / Ctrl+V, etc.) won't trigger it

easyfunasr CLI

easyfunasr start [all|dictation|double-ctrl|tray]
easyfunasr stop
easyfunasr restart
easyfunasr status              # per-component + socket status
easyfunasr toggle              # same as the hotkey
easyfunasr reload              # restart the recognition service
easyfunasr logs [dictation|double-ctrl|tray]
easyfunasr install-autostart | uninstall-autostart

Configuration (env overrides)

Paths are computed dynamically by easyfunasr/config.py; tweak via environment variables:

Variable Default Description
EASYFUNASR_DEVICE auto (cuda:0/cpu) Force device
EASYFUNASR_MODEL / _VAD / _PUNC paraformer-zh / fsmn-vad / ct-punc Models
EASYFUNASR_DOUBLE_CTRL_THRESHOLD 0.35 Double-Ctrl interval (seconds)
EASYFUNASR_KEYBINDING <Super><Alt>v GNOME hotkey
EASYFUNASR_LOG_VIEWER gnome-text-editor Tray "View log" command

Directory layout

easyfunasr/
├── install.sh / uninstall.sh
├── requirements.txt
├── easyfunasr/          # Python package (config + dictation/toggle/double_ctrl/tray/recognize/cli)
├── scripts/             # autostart launchers (self-contained, resolve project root at runtime)
├── assets/autostart/    # .desktop templates (@PROJECT_ROOT@ placeholder)
├── logs/                # placeholder (actual logs live in ~/.local/state/easyfunasr/)
└── .venv/               # created by install.sh (gitignored)

Runtime artifacts: socket/PID in $XDG_RUNTIME_DIR/, logs in ~/.local/state/easyfunasr/, models in ~/.cache/modelscope/.

Uninstall

./uninstall.sh            # keeps .venv and models
./uninstall.sh --purge    # also removes .venv (~7.5 GB) and the model cache

Manual install (without install.sh)

  1. Install system packages: xdotool xclip ffmpeg libportaudio2 python3-venv python3-pip (GNOME also needs gir1.2-ayatanaappindicator3-0.1)
  2. python3 -m venv .venv && .venv/bin/pip install -U pip
  3. .venv/bin/pip install --extra-index-url https://download.pytorch.org/whl/cu128 torch torchaudio (drop --extra-index-url for CPU-only)
  4. .venv/bin/pip install -r requirements.txt
  5. ./easyfunasr/cli.py install-autostart

Troubleshooting

  • Hotkey says "service not running": the service is still loading (first ~30 s); check easyfunasr status or tail ~/.local/state/easyfunasr/dictation.log
  • Empty / garbage recognition: check the mic (arecord -l) and mute state; very short recordings may come back empty
  • Double-Ctrl doesn't trigger: requires an X11 session (not Wayland); chords like Ctrl+C won't trigger it by design
  • GPU OOM: 6 GB cards are fine; if running other large models at the same time, force CPU with EASYFUNASR_DEVICE=cpu
  • Tray icon missing: install gnome-shell-extension-appindicator and enable the ubuntu-appindicators extension

Known limitations

  • No Wayland support: paste (xdotool) and key listening (XRECORD) depend on X11. Wayland would need wtype / wl-clipboard (not implemented yet)
  • Paste uses "clipboard + Ctrl+V", so it briefly occupies the clipboard (auto-restored after ~0.4 s)
  • Recording stops manually via double-tap (no automatic silence detection yet)

How it works

A single UNIX socket decouples the pieces: a resident service (dictation.py) holds the model in VRAM and accepts toggle/status/quit; two triggers (double_ctrl.py via XRECORD, toggle.py via the GNOME hotkey) send toggle; a tray app (tray.py, system Python + AppIndicator) orchestrates start/stop. A unified easyfunasr CLI replaces the launch scripts. All paths are centralized in easyfunasr/config.py — zero hardcoded absolute paths.

Credits

Built on Alibaba DAMO Academy's FunASR and the Paraformer model.

License

MIT (see LICENSE). FunASR follows its own open-source license.

About

常驻 GPU 的中文语音输入工具(FunASR + GNOME) | 双击 Ctrl 触发 | 任务栏托盘 | 一键 install.sh

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages