Speech-to-math demo using Whisper for transcription and a fine-tuned ByT5 decoder for math notation.
- Model:
vibhuiitj/byt5-base-whispermath-a100-checkpoint-10724 - Training data:
vibhuiitj/whispermath-input-output - Web demo Space: https://huggingface.co/spaces/vibhuiitj/whispermath-webdemo
- Direct demo: https://vibhuiitj-whispermath-webdemo.hf.space
audio
-> faster-whisper
-> spoken transcript
-> fine-tuned ByT5
-> LaTeX-like math text
-> KaTeX render
Example:
audio/transcript: integral from zero to pi of sine x dx
model output: \int_0^\pi \sin x dx
Whisper is not fine-tuned here. The fine-tuned model is ByT5, trained for:
spoken math text -> math / LaTeX-like text
phase-1/ Whisper transcription + early rule parser
phase-2/ Dataset construction
phase-3-decoder/ ByT5 training, eval, prediction scripts
webdemo/ FastAPI + browser demo, deployed to HF Spaces
Final dataset:
vibhuiitj/whispermath-input-output
Schema:
{
"input_text": "x squared minus y squared equals four",
"output_text": "x^2-y^2=4",
"type": "latex"
}Types:
latex formula-heavy rows
mixed natural language with math
normal normal text copied through as a control task
Split used during training/eval:
train: 146,981
validation: 1,500
test: 1,500
Base:
google/byt5-base
Fine-tuned checkpoint:
vibhuiitj/byt5-base-whispermath-a100-checkpoint-10724
ByT5 was used because it is byte-level and handles LaTeX characters like:
\ { } _ ^
without adding custom tokenizer tokens.
Config:
phase-3-decoder/configs/byt5_base_a100_80gb.yaml
Main settings:
model_name: google/byt5-base
dataset_id: vibhuiitj/whispermath-input-output
num_train_epochs: 3
learning_rate: 5e-5
max_source_length: 512
max_target_length: 512
per_device_train_batch_size: 8
gradient_accumulation_steps: 4
bf16: trueRun:
cd phase-3-decoder
python src/train_byt5.py --config configs/byt5_base_a100_80gb.yamlText-only:
cd phase-3-decoder
python src/predict.py \
--model vibhuiitj/byt5-base-whispermath-a100-checkpoint-10724 \
"x squared minus y squared equals four"Expected:
x^2-y^2=4
Web demo:
cd webdemo
WHISPERMATH_WHISPER_MODEL=small.en \
python -m uvicorn app:app --host 127.0.0.1 --port 8766Open:
http://127.0.0.1:8766
Use medium.en locally for better Whisper transcription:
WHISPERMATH_WHISPER_MODEL=medium.en python -m uvicorn app:app --host 127.0.0.1 --port 8766Evaluation:
cd phase-3-decoder
python src/evaluate.py \
--model vibhuiitj/byt5-base-whispermath-a100-checkpoint-10724 \
--max-samples 150 \
--num-beams 1Results on 150 held-out rows:
| Type | Count | Exact | CER |
|---|---|---|---|
| normal | 55 | 18.2% | 0.468 |
| latex | 53 | 3.8% | 0.260 |
| mixed | 42 | 0.0% | 0.559 |
Comparable 30-row sample vs earlier smoke checkpoint:
| Type | Old CER | New CER |
|---|---|---|
| normal | 0.654 | 0.563 |
| latex | 0.656 | 0.222 |
| mixed | 0.660 | 0.553 |
The biggest improvement is on pure LaTeX/formula rows.
Good:
x squared minus y squared equals four
-> x^2-y^2=4
integral from zero to pi of sine x dx
-> \int_0^\pi \sin x dx
Known weak cases:
limit as x tends to zero of sine x by x
-> \lim_{x\to 0}\sin x\by x
the derivative of x cubed plus five x squared minus seven
-> \frac{d}{x^3+5x^2-7}
tiny.enis too weak for spoken math.small.enis the best free-CPU default.medium.engives better transcripts locally but is slower.- The decoder still needs more data for fractions, derivatives, limits, grouped expressions, and phrases like
whole square. - The web UI includes an editable transcript box so Whisper errors can be corrected before re-decoding with ByT5.