Speech, sound by sound.
Turn a short recording into IPA. Listen back and explore what the model hears.
Your audio
20 sec maxDrop an audio clip here
or
WAV, MP3 & other audio formats · up to 25 MB
Choose a clip to get started.
IPA transcription
Select a sound to highlight its audio span. Times are approximate.
Click or drag to seek · arrow keys to fine-tune
Confidence shows how strongly the model predicts each sound.
Model detailsFrame timeline & raw output
Colors distinguish sounds; gray marks CTC blanks, which can occur inside speech. Each frame advances 20 ms. The slider seeks audio; emission times are approximate sound boundaries.