🔊 Runs whisper large-v3 locally via MLX (Apple Silicon GPU). Zero network egress — audio stays on disk, no keys to provision, no metered billing. Throughput: ~2.7x realtime (8 min audio ≈ 3 min wall clock).
• audio and video in — mp3, wav, m4a, flac, aiff, mp4, mov — anything ffmpeg can decode
• segment-level timestamps — start/end to the centisecond
• subtitle export — plaintext and .srt, no post-processing step
• VAD gating — silero VAD isolates speech first, so music, silence, and room tone never reach the decoder
• confidence filtering — segments scored on log-probability, compression ratio, and lexical diversity; anything below threshold gets dropped
• domain priming — pass in names, acronyms, or jargon and the decoder biases toward them; large win on proper nouns
• offline after first run — one model download, then zero network calls
| machine | apple silicon mac (m1 or later) |
| python | 3.10+ |
| ffmpeg | brew install ffmpeg |
| disk | ~3 gb for model weights |
git clone https://github.com/oppo123451/mlx-transcribe.git
cd mlx-transcribe
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtmodel weights download automatically the first time you run it and cache to
~/.cache/huggingface, so it only happens once.
command line:
python pipeline.py interview.mp3prints timestamped segments and writes transcript.txt and transcript.srt.
as a library:
from pipeline import transcribe
result = transcribe("interview.mp3")
print(result["text"]) # full transcript
print(result["duration"]) # length in seconds
for seg in result["segments"]:
print(seg["start"], seg["end"], seg["text"])with domain context — this is the single biggest accuracy lever, use it:
result = transcribe(
"earnings_call.mp3",
context="EBITDA, basis points, year-over-year, NASDAQ, Q3 guidance",
)pick a different model size to trade accuracy for speed:
result = transcribe("podcast.mp3", model="mlx-community/whisper-small-mlx")audio file
| ffmpeg decode any codec -> 16 khz mono float32
| silero vad find speech regions, skip everything else
| whisper large-v3 (mlx/metal) transcribe each region
| confidence filter drop low-scoring segments
| timestamp offset + clamp map back to the real timeline
transcript + srt
the stack
- mlx — apple's array framework. runs the model on the mac gpu through metal, using unified memory so there is no cpu/gpu copying
- whisper large-v3 — openai's encoder-decoder asr model, 99 languages, the full size version
- silero vad — tiny, fast voice activity detection model
- ffmpeg — decodes literally any audio or video format you throw at it
mit — do whatever you want with it :)