Skip to content

About

solve mathematics via speech

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

WhisperMath

Speech-to-math demo using Whisper for transcription and a fine-tuned ByT5 decoder for math notation.

Links

What It Does

audio
  -> faster-whisper
  -> spoken transcript
  -> fine-tuned ByT5
  -> LaTeX-like math text
  -> KaTeX render

Example:

audio/transcript: integral from zero to pi of sine x dx
model output:     \int_0^\pi \sin x dx

Whisper is not fine-tuned here. The fine-tuned model is ByT5, trained for:

spoken math text -> math / LaTeX-like text

Repo Layout

phase-1/           Whisper transcription + early rule parser
phase-2/           Dataset construction
phase-3-decoder/   ByT5 training, eval, prediction scripts
webdemo/           FastAPI + browser demo, deployed to HF Spaces

Dataset

Final dataset:

vibhuiitj/whispermath-input-output

Schema:

{
  "input_text": "x squared minus y squared equals four",
  "output_text": "x^2-y^2=4",
  "type": "latex"
}

Types:

latex   formula-heavy rows
mixed   natural language with math
normal  normal text copied through as a control task

Split used during training/eval:

train:      146,981
validation: 1,500
test:       1,500

Model

Base:

google/byt5-base

Fine-tuned checkpoint:

vibhuiitj/byt5-base-whispermath-a100-checkpoint-10724

ByT5 was used because it is byte-level and handles LaTeX characters like:

\ { } _ ^

without adding custom tokenizer tokens.

Training

Config:

phase-3-decoder/configs/byt5_base_a100_80gb.yaml

Main settings:

model_name: google/byt5-base
dataset_id: vibhuiitj/whispermath-input-output
num_train_epochs: 3
learning_rate: 5e-5
max_source_length: 512
max_target_length: 512
per_device_train_batch_size: 8
gradient_accumulation_steps: 4
bf16: true

Run:

cd phase-3-decoder
python src/train_byt5.py --config configs/byt5_base_a100_80gb.yaml

Inference

Text-only:

cd phase-3-decoder
python src/predict.py \
  --model vibhuiitj/byt5-base-whispermath-a100-checkpoint-10724 \
  "x squared minus y squared equals four"

Expected:

x^2-y^2=4

Web demo:

cd webdemo
WHISPERMATH_WHISPER_MODEL=small.en \
python -m uvicorn app:app --host 127.0.0.1 --port 8766

Open:

http://127.0.0.1:8766

Use medium.en locally for better Whisper transcription:

WHISPERMATH_WHISPER_MODEL=medium.en python -m uvicorn app:app --host 127.0.0.1 --port 8766

Results

Evaluation:

cd phase-3-decoder
python src/evaluate.py \
  --model vibhuiitj/byt5-base-whispermath-a100-checkpoint-10724 \
  --max-samples 150 \
  --num-beams 1

Results on 150 held-out rows:

Type Count Exact CER
normal 55 18.2% 0.468
latex 53 3.8% 0.260
mixed 42 0.0% 0.559

Comparable 30-row sample vs earlier smoke checkpoint:

Type Old CER New CER
normal 0.654 0.563
latex 0.656 0.222
mixed 0.660 0.553

The biggest improvement is on pure LaTeX/formula rows.

Examples

Good:

x squared minus y squared equals four
-> x^2-y^2=4
integral from zero to pi of sine x dx
-> \int_0^\pi \sin x dx

Known weak cases:

limit as x tends to zero of sine x by x
-> \lim_{x\to 0}\sin x\by x
the derivative of x cubed plus five x squared minus seven
-> \frac{d}{x^3+5x^2-7}

Notes

  • tiny.en is too weak for spoken math.
  • small.en is the best free-CPU default.
  • medium.en gives better transcripts locally but is slower.
  • The decoder still needs more data for fractions, derivatives, limits, grouped expressions, and phrases like whole square.
  • The web UI includes an editable transcript box so Whisper errors can be corrected before re-decoding with ByT5.

About

solve mathematics via speech

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages