Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
Python speech recognition toolkit for agent serving
FunASR gives you reusable speech models and pipelines for transcription, streaming ASR, punctuation, VAD, speaker-aware output, and emotion tagging. You can use it from Python, the CLI, OpenAI-compatible endpoints, or the MCP server examples, depending on how you want your agent or app to call it.
Builders who need speech-to-text pipelines, speaker-aware transcription, or agent-facing audio services.
You can turn audio into structured text and expose it to agents or apps through Python, CLI, or API serving.
What it does
Training and inference toolkit
Provides Python modules for models, datasets, losses, metrics, optimizers, schedulers, and training utilities under `funasr/`.
Offline and streaming ASR
Supports batch transcription and chunked streaming models such as Paraformer streaming and Fun-ASR-Nano.
Speech pipeline components
Combines ASR with VAD, punctuation, speaker embeddings, and emotion recognition in one pipeline.
CLI for audio jobs
Includes the `funasr` command for transcription, JSON output, SRT subtitles, and speaker-aware runs.
Serving options
Includes OpenAI-compatible API examples, MCP server examples, and deployment guides for local and edge setups.
Model and deployment guides
Documents model selection, deployment matrix, migration from Whisper, benchmarking, and runtime options like vLLM and llama.cpp.
How to get it
- 1Found FunASR useful? Star the project so more builders can find it.
# CPU-only installs can use the default PyPI wheels. pip install torch torchaudio pip install funasr
- 2For GPU quickstarts, install the PyTorch and torchaudio wheels that match your NVIDIA…
python - <<'PY' import torch print(torch.cuda.is_available()) PY
- 3Run
pip install funasr
- 4Run
git clone https://github.com/modelscope/FunASR.git && cd FunASR pip install -e ./
README
Industrial speech recognition toolkit for offline, streaming, and edge deployment.
ASR · VAD · punctuation · speaker pipelines · emotion and audio-event models · OpenAI-compatible serving
Quick Start · Model selection · Models · Deployment matrix · Deployment hub · Docs · Benchmark · Contribute
Quick Start
Native Transformers
For Fun-ASR-Nano transcription with the Hugging Face API, start with the Transformers 5.17.0 CPU quickstart. No FunASR toolkit or remote Python code is needed.
Space · Notebook · Python / batch examples
FunASR toolkit and pipelines
No local setup? Open the Colab quickstart to transcribe a public sample or upload your own audio in a browser.
Found FunASR useful? Star the project so more builders can find it.
# CPU-only installs can use the default PyPI wheels.
pip install torch torchaudio
pip install funasr
For GPU quickstarts, install the PyTorch and torchaudio wheels that match your NVIDIA driver from pytorch.org before installing FunASR. After installation, confirm the GPU is visible:
python - <<'PY'
import torch
print(torch.cuda.is_available())
PY
Only use device="cuda" when this prints True; otherwise use device="cpu"
or reinstall PyTorch with the correct CUDA wheel.
FunASR toolkit GPU example: Fun-ASR-Nano (Chinese, English, Japanese, and Chinese dialect groups and regional accents; the separate native Transformers CPU path is linked above):
from funasr import AutoModel
model = AutoModel(model="FunAudioLLM/Fun-ASR-Nano-2512", device="cuda")
result = model.generate(input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav")
print(result[0]["text"])
For the separate 31-language checkpoint, use Fun-ASR-MLT-Nano-2512. Language coverage is checkpoint-specific, so Nano and MLT-Nano should be treated as distinct model choices.
For a CPU-first example with five-language ASR plus emotion and audio-event tags, use SenseVoiceSmall. The pipeline below combines it with FSMN-VAD and CAM++ for speaker-aware VAD segments; these are not native speaker outputs of the SenseVoiceSmall checkpoint. See the SenseVoice paper, Hugging Face checkpoint, and GGUF edge checkpoint.
from funasr import AutoModel
from funasr.utils.postprocess_utils import rich_transcription_postprocess
model = AutoModel(model="iic/SenseVoiceSmall", vad_model="fsmn-vad", spk_model="cam++", device="cpu")
result = model.generate(
input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav",
batch_size_s=300,
)
# The AutoModel pipeline returns VAD segments with speaker ids and timestamps:
for seg in result[0]["sentence_info"]:
print(f"[{seg['start']/1000:.1f}s] Speaker {seg['spk']}: {rich_transcription_postprocess(seg['sentence'])}")
This prints each returned segment's start time in seconds, anonymous speaker index, and text with SenseVoice tags removed. Text and segment boundaries depend on the audio and checkpoint; no fixed transcript is asserted here.
CAM++ extracts spk_embedding vectors. AutoModel clusters those embeddings
and assigns speaker indices to VAD segments. Indices are local to a recording,
not known-person identities. See the SDK contract for
the component and result boundaries. Change to device="cuda" only after
verifying a compatible GPU environment as described above.
Scale & deploy the flagship
At scale, accelerate Fun-ASR-Nano with vLLM (batch processing):
from funasr.auto.auto_model_vllm import AutoModelVLLM
model = AutoModelVLLM(model="FunAudioLLM/Fun-ASR-Nano-2512", tensor_parallel_size=1)
results = model.generate(["audio1.wav", "audio2.wav"], language="auto")
Deploy as API server: Local SenseVoice CPU recipe · Nano GPU serving and pinned vLLM setup
Use with AI agents: MCP Server for Claude/Cursor · OpenAI API for LangChain/Dify/AutoGen
Use with voice agents: OpenClaw realtime plugin for self-hosted Talk and Voice Call transcription
Why FunASR?
FunASR is a toolkit: choose the task, checkpoint, and runtime separately. Support in one model or adapter does not imply support in every serving backend.
| Task | Checkpoint or pipeline | Runtime entrypoint | Important limitation |
|---|---|---|---|
| File transcription with emotion/event tags | SenseVoiceSmall | Python AutoModel, CPU or GPU | Five-language checkpoint; tags do not identify speakers. |
| LLM-based file transcription | Fun-ASR-Nano | AutoModel; split-engine AutoModelVLLM for the documented GPU path | Base Nano covers zh/en/ja and Chinese dialects/accents; timestamp support depends on checkpoint and path. |
| Broader multilingual transcription | Fun-ASR-MLT-Nano | Python AutoModel | Separate 31-language checkpoint; do not transfer its coverage to base Nano. |
| Chunked live transcription | Paraformer-zh-streaming | Streaming SDK or runtime WebSocket service | Use the streaming checkpoint and per-session cache, not an offline checkpoint. |
| Speaker-aware file transcription | SenseVoiceSmall + FSMN-VAD + CAM++ | AutoModel with VAD and embedding clustering | Anonymous indices within a recording, not enrolled-speaker identification. |
| Joint text, timestamps, and speakers | MOSS-Transcribe-Diarize, third-party OpenMOSS | FunASR adapter or upstream backend in the MOSS guide | Offline, recording-local anonymous labels; no external VAD/speaker pipeline for its unified path. |
| Native CPU/edge transcription | Fun-ASR-Nano or SenseVoiceSmall GGUF | llama.cpp runtime | Requires matching converted weights; GGUF is not a Python AutoModel checkpoint. |
See the Model Zoo and deployment matrix for checkpoint, interface, and licensing boundaries. Benchmark on your own audio and hardware before choosing a runtime.
Trying FunASR for the first time? Use the Colab quickstart before setting up a local environment. Choosing a first model? Start with the model selection guide. Planning a switch from Whisper or a cloud ASR provider? Use the migration guide and benchmark example to test representative audio, map features, and roll out safely.
Installation
pip install funasr
From source / Requirements
git clone https://github.com/modelscope/FunASR.git && cd FunASR
pip install -e ./
Requirements: Python ≥ 3.8. Install PyTorch + torchaudio first (pytorch.org), then pip install funasr.
Model Zoo
This list includes third-party models. OpenMOSS publishes MOSS-Transcribe-Diarize; FunASR provides an adapter, not ownership of its weights. Its unified path is offline, with anonymous labels scoped to each recording, not realtime or known-person identification. Model licenses are separate from the toolkit's MIT license.
| Model | Task | Languages | Params | Links |
|---|---|---|---|---|
| Fun-ASR-Nano | ASR | zh/en/ja + Chinese dialects and accents | 800M | ⭐ HF / Transformers · HF / FunASR GGUF |
| Fun-ASR-MLT-Nano | ASR | 31 languages | 800M | ⭐ 🤗 |
| SenseVoiceSmall | ASR + emotion + events | zh/en/ja/ko/yue | 234M | ⭐ 🤗 GGUF paper |
| MOSS-Transcribe-Diarize | Third-party OpenMOSS: offline ASR + timestamps + anonymous speakers | See official card | See official card | 🤗 guide |
| Paraformer-zh | ASR + timestamps | zh/en | 220M | ⭐ 🤗 |
| Paraformer-zh-streaming | Streaming ASR | zh/en | 220M | ⭐ 🤗 |
| Qwen3-ASR | ASR, 52 languages | multilingual | 1.7B | usage |
| GLM-ASR-Nano | ASR, 17 languages | multilingual | 1.5B | usage |
| Whisper-large-v3 | ASR + translation | multilingual | 1550M | usage |
| Whisper-large-v3-turbo | ASR + translation | multilingual | 809M | usage |
| ct-punc | Punctuation | zh/en | 290M | ⭐ 🤗 |
| fsmn-vad | VAD | zh/en | 0.4M | ⭐ 🤗 |
| cam++ | Speaker embeddings (pipeline component) | — | 7.2M | ⭐ 🤗 |
| emotion2vec+large | Emotion recognition | — | 300M | ⭐ 🤗 |
Usage
Python tutorial · SDK parameters and outputs · Training · Model registration
from funasr import AutoModel
# Chinese production (VAD + ASR + punctuation + speaker)
model = AutoModel(model="paraformer-zh", vad_model="fsmn-vad", punc_model="ct-punc", spk_model="cam++", device="cuda")
result = model.generate(input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav", hotword="关键词 20")
# Optional Silero VAD (install first: python -m pip install "funasr[silero]")
model = AutoModel(
model="paraformer-zh", vad_model="silero-vad", device="cuda",
vad_kwargs={"silero_threshold": 0.5, "silero_min_silence_duration_ms": 100},
)
result = model.generate(input="audio.wav")
# Streaming real-time (feed audio chunk by chunk)
import soundfile as sf
model = AutoModel(model="paraformer-zh-streaming", device="cuda")
audio, sr = sf.read("speech.wav", dtype="float32") # 16 kHz mono
chunk_size = [0, 10, 5] # 600 ms chunks
chunk_stride = chunk_size[1] * 960
cache = {}
n_chunks = (len(audio) - 1) // chunk_stride + 1
for i in range(n_chunks):
chunk = audio[i * chunk_stride : (i + 1) * chunk_stride]
res = model.generate(input=chunk, cache=cache, is_final=(i == n_chunks - 1),
chunk_size=chunk_size, encoder_chunk_look_back=4, decoder_chunk_look_back=1)
if res[0]["text"]:
print(res[0]["text"], end="", flush=True)
# Emotion recognition
model = AutoModel(model="emotion2vec_plus_large", device="cuda")
result = model.generate(input="audio.wav", granularity="utterance")
CLI (Agent-Friendly)
# Transcribe audio (simplest)
funasr audio.wav
# JSON output (for AI agents)
funasr audio.wav --output-format json
# SRT subtitles
funasr audio.wav --output-format srt --output-dir ./subs
# Speaker diarization + timestamps
funasr audio.wav --spk --timestamps -f json
# Choose model and language
funasr audio.wav --model paraformer --language zh
# Batch transcribe
funasr *.wav --output-format srt --output-dir ./output
Available models: sensevoice (default), paraformer, paraformer-en, fun-asr-nano
Deploy
Start a local SenseVoice CPU service from a fresh directory in a POSIX shell with Python 3.11. This installs the PyPI release into a separate environment, not this source checkout. Keep the unauthenticated service on loopback; use the security guide before exposing it to other clients.
python3.11 -m venv .venv-funasr-http
. .venv-funasr-http/bin/activate
python -m pip install torch torchaudio
python -m pip install funasr fastapi uvicorn python-multipart
python -m pip check
funasr-server --host 127.0.0.1 --port 8000 --model sensevoice --device cpu
Wait for model download and server startup. In a second terminal, use the same directory and curl 7.76+ to download a public Chinese audio sample and transcribe it. The request uses the preloaded model; no fixed transcript or speaker labels are promised.
curl --fail --location https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/BAC009S0764W0121.wav -o sample.wav && \
curl --fail-with-body http://127.0.0.1:8000/v1/audio/transcriptions \
-F file=@sample.wav \
-F model=sensevoice \
-F response_format=verbose_json
For offline joint ASR and anonymous speaker labels (moss-transcribe-diarize), prepare the separate environment
in the MOSS service, Docker, Kubernetes, vLLM, SGLang, LocalAI, and FunClip guide →.
It is an alternative service, not another command in the CPU environment. Stop the
CPU service before reusing port 8000. For Nano GPU serving, follow the
pinned split-engine guide and inspect the actual backend logs;
selecting a model does not by itself prove that vLLM was loaded.
# Docker streaming service
docker pull registry.cn-hangzhou.aliyuncs.com/funasr_repo/funasr:funasr-runtime-sdk-online-cpu-0.1.12
CPU / Edge — llama.cpp / GGUF (no GPU, no Python)
Run SenseVoice / Paraformer / Fun-ASR-Nano as a single self-contained binary on CPU and edge devices — this is to FunASR what whisper.cpp is to Whisper, but with ~3× lower CER than whisper.cpp on Chinese. Built-in FSMN-VAD, no Python at runtime.
# Linux / macOS: run from the extracted release directory
bash download-funasr-model.sh sensevoice ./gguf # or: paraformer | nano
./llama-funasr-sensevoice -m ./gguf/sensevoice-small-q8.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav
# → 欢迎大家来体验达摩院推出的语音识别模型
# Windows PowerShell: run from the extracted archive root (with the `hf` CLI installed)
hf download FunAudioLLM/SenseVoiceSmall-GGUF sensevoice-small-q8.gguf --local-dir .\gguf
hf download FunAudioLLM/fsmn-vad-GGUF fsmn-vad.gguf --local-dir .\gguf
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav
# Use the windows-x64-vulkan package with a current AMD, Intel, or NVIDIA Vulkan driver:
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav --backend vulkan
# Use the windows-x64-cuda package on RTX 30-class GPUs:
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav --backend cuda
Use funasr-llamacpp-linux-x64-vulkan.tar.gz on Linux GPU systems with a
working Vulkan driver/ICD:
./llama-funasr-sensevoice -m ./gguf/sensevoice-small-q8.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav --backend vulkan
The Windows Vulkan ZIP uses the system Vulkan loader supplied by the GPU driver; installing the Vulkan SDK is only necessary when building from source. Both Vulkan packages currently accelerate SenseVoiceSmall.
Tagged releases provide two Windows CUDA packages. The standard
windows-x64-cuda ZIP targets CUDA architecture 86, while
windows-x64-cuda-blackwell targets architecture 120 (sm_120) for RTX 50 /
Blackwell GPUs. Both ZIPs bundle the required cuBLAS DLLs and use the static MSVC
runtime, so users need a compatible NVIDIA driver but not a separate CUDA Toolkit
installation. CI verifies the architecture and package boundary; it does not prove
inference on physical Blackwell hardware.
Prebuilt binaries: Releases · v0.2.6 · Linux Vulkan tarball · Windows Vulkan zip · Windows CUDA zip · Windows Blackwell CUDA zip · Download & quickstart: funasr.com/deploy/llama-cpp · GGUF models: Hugging Face · Docs & benchmarks: runtime/llama.cpp/
OpenAI API example → · Gradio demo → · Client recipes → · JavaScript/TypeScript recipes → · Kubernetes template → · Workflow recipes → · Postman collection → · OpenAPI spec → · Security guide → · Deployment matrix → · Deployment docs → · Agent integration →
Benchmark
The historical benchmark report and split-engine measurements retain their original results. They are separate records, not universal speed rankings or production capacity guarantees.
Use the RTFx and reproducibility notes to compare checkpoint/revision, audio set, hardware, batching, warmup, timing scope, and CER/WER. Offline throughput is not streaming latency. The migration benchmark example helps measure your own recordings with the same evaluation scope.
What's new
- MOSS-Transcribe-Diarize brings long-form ASR, timestamps, and anonymous speaker labels to FunASR services, Docker, Kubernetes, vLLM/SGLang workflows, and FunClip. Deploy MOSS ->
- FunASR 1.4.15 adds tested NumPy 2 compatibility and fixes streaming KWS/VAD boundaries and checkpoint ranking. Install with
python -m pip install -U "funasr==1.4.15". Release and verification scope -> - Native Transformers: Released 5.17.0 supports Fun-ASR-Nano with the official
-hfcheckpoint, CPU examples and a notebook. Get started ->
See GitHub Releases for the complete changelog and downloadable assets.
Community
Start with troubleshooting before reporting a problem. Include your exact model, runtime, environment and a minimal reproduction.
| 📖 Documentation | 🐛 Issues |
| 💬 Discussions | 🤗 HuggingFace |
| 🤝 Contributing | 🌐 funasr.com |
| 🗺️ Repository roles & roadmap | 📈 Growth plan |
| 🧩 Community projects | 💡 Use-case showcase |
Star History
License
- FunASR toolkit source code in this repository: MIT License.
- Pretrained model weights are licensed separately. Check the license shown on each model card; when a model card links to the FunASR Model Open Source License Agreement, those terms apply.
Citations
@inproceedings{gao2023funasr,
author={Zhifu Gao and others},
title={FunASR: A Fundamental End-to-End Speech Recognition Toolkit},
booktitle={INTERSPEECH},
year={2023}
}
Files in the repo
- .github
- benchmarks
- data
- docs
- examples
- fun_text_processing
- funasr
- gh-pages-output
- integrations
- model_zoo
- runtime
- scripts
- tests
- tests_models
- web-pages
- .gitignore
- .pre-commit-config.yaml
- Acknowledge.md
- benchmark_vllm.py
- CONTRIBUTING.md
- Contribution.md
- glama.json
- LICENSE
- MinMo_gitlab
- MODEL_LICENSE
- pyproject.toml
- README_ja.md
- README_ko.md
- README_zh.md
- README.md
- SECURITY.md
- setup.py
- training.html
Discussion (0)
Ask about usage, or say what you built with itSign in to join the discussion.
No comments yet. Be the first to say what this is good for.
More frameworks & sdks
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

Open-source Agent Operating System
JavaScript in-page GUI agent. Control web interfaces with natural language.