Skip to content

About

Local audio transcription pipeline with voice-activity detection and hallucination filtering. Whisper large-v3 on Apple Silicon via MLX.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

11 Commits

Folders and files

Repository files navigation

mlx-transcribe

🔊 Runs whisper large-v3 locally via MLX (Apple Silicon GPU). Zero network egress — audio stays on disk, no keys to provision, no metered billing. Throughput: ~2.7x realtime (8 min audio ≈ 3 min wall clock).

what it does

•	audio and video in — mp3, wav, m4a, flac, aiff, mp4, mov — anything ffmpeg can decode
•	segment-level timestamps — start/end to the centisecond
•	subtitle export — plaintext and .srt, no post-processing step
•	VAD gating — silero VAD isolates speech first, so music, silence, and room tone never reach the decoder
•	confidence filtering — segments scored on log-probability, compression ratio, and lexical diversity; anything below threshold gets dropped
•	domain priming — pass in names, acronyms, or jargon and the decoder biases toward them; large win on proper nouns
•	offline after first run — one model download, then zero network calls

requirements

machine apple silicon mac (m1 or later)
python 3.10+
ffmpeg brew install ffmpeg
disk ~3 gb for model weights

install

git clone https://github.com/oppo123451/mlx-transcribe.git
cd mlx-transcribe
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

model weights download automatically the first time you run it and cache to ~/.cache/huggingface, so it only happens once.

usage

command line:

python pipeline.py interview.mp3

prints timestamped segments and writes transcript.txt and transcript.srt.

as a library:

from pipeline import transcribe

result = transcribe("interview.mp3")

print(result["text"])          # full transcript
print(result["duration"])      # length in seconds

for seg in result["segments"]:
    print(seg["start"], seg["end"], seg["text"])

with domain context — this is the single biggest accuracy lever, use it:

result = transcribe(
    "earnings_call.mp3",
    context="EBITDA, basis points, year-over-year, NASDAQ, Q3 guidance",
)

pick a different model size to trade accuracy for speed:

result = transcribe("podcast.mp3", model="mlx-community/whisper-small-mlx")

how it works

audio file
    |  ffmpeg                        decode any codec -> 16 khz mono float32
    |  silero vad                    find speech regions, skip everything else
    |  whisper large-v3 (mlx/metal)  transcribe each region
    |  confidence filter             drop low-scoring segments
    |  timestamp offset + clamp      map back to the real timeline
transcript + srt

the stack

  • mlx — apple's array framework. runs the model on the mac gpu through metal, using unified memory so there is no cpu/gpu copying
  • whisper large-v3 — openai's encoder-decoder asr model, 99 languages, the full size version
  • silero vad — tiny, fast voice activity detection model
  • ffmpeg — decodes literally any audio or video format you throw at it

license

mit — do whatever you want with it :)

About

Local audio transcription pipeline with voice-activity detection and hallucination filtering. Whisper large-v3 on Apple Silicon via MLX.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages