- name
- qwen-asr
- description
- >-
- metadata
qwen-asr
Local, CPU-only speech-to-text powered by Qwen3-ASR. No API key or cloud needed.
- Source code: huanglizhuo/QwenASR
- Based on: antirez/qwen-asr (original C implementation)
Install
Run the install script to download the pre-built binary and model:
bash {baseDir}/scripts/install.shThis will:
- Download the
qwen-asrbinary for your platform from GitHub Releases - Download the
qwen3-asr-0.6bmodel (~1.5 GB) from HuggingFace
Usage
Transcribe an audio file
bash {baseDir}/scripts/transcribe.sh <audio-file>Supports any audio format: wav, mp3, m4a, ogg, flac, opus, webm, aac, etc. Non-WAV files are automatically converted via ffmpeg (must be installed).
Or call qwen-asr directly (WAV only):
qwen-asr -d ~/.openclaw/tools/qwen-asr/qwen3-asr-0.6b -i <audio-file> --silentFrom stdin
cat audio.wav | qwen-asr -d ~/.openclaw/tools/qwen-asr/qwen3-asr-0.6b --stdin --silentCommon parameters
| Flag | Description |
|---|---|
--silent | Print only transcription text (no progress) |
--language <lang> | Force language (e.g., zh, en) |
-S <seconds> | Segmented mode — split audio into chunks |
--stream | Streaming mode — process audio in real time |
--stdin | Read audio from stdin |
Model path
Default model directory: ~/.openclaw/tools/qwen-asr/qwen3-asr-0.6b