See your voice.
Record your voice, and this tool automatically generates your speech production profile by analyzing your vowel space, estimating your vocal-tract length, and capturing key features of your lip shape, tongue contour, and voice — all directly in your browser.
What this tool does
Acoustic Vowel Profile
Automatic F0, F1, F2 and F3 extraction via LPC, plotted on a linked 2D/3D vowel chart.
Open Vowel Profile →Vocal-Tract Length
An estimated acoustic vocal-tract length from your formants, with an animated 3D tract.
Open Vocal Tract →Classroom Mode
Click through vowels to demonstrate source–filter theory live, with or without recorded data.
Open Classroom Mode →Speech Profile Report
A clinical-style, exportable report summarising every measurement in one document.
Open Speech Profile →Record your target vowels
Read each vowel clearly for about a second in a quiet room, close to the microphone — or click the ⇪ upload button on any vowel card to use a short existing audio clip instead. Either way, the system automatically finds the steady portion of the recording and extracts F0, F1, F2 and F3 — you never need to enter formant values by hand.
Optional: record additional IPA vowels (full chart, 22 more symbols) ▾
These aren't required — the 6 core vowels above are enough for a profile. Anything you record here is analyzed too, and shows up in the vowel chart, tract table and report alongside the core set.
Tongue Contour
A small neural network predicts a 20-point tongue contour directly from your microphone, frame by frame, entirely in your browser. Enable the microphone below, hit start, and speak — watch it track in real time as you move through different vowels. Details on how the model works are on the How It Works page.
The microphone is shared across the whole app — enabling it here also enables it for recording vowels, and vice versa.
Lip Capture
Reads lip shape (open / closed / rounded / spread) live from your webcam using on-device face landmarks — video never leaves your browser. Enable the camera below before recording vowels on the Record tab, and each vowel recording will also capture a lip-shape reading alongside the audio.
Runs entirely on-device (WASM face landmarks) — video is never uploaded or stored anywhere, only the two numbers (mouth openness, mouth roundedness) derived from it. When enabled, each vowel recording on the Record tab also captures a lip-shape reading alongside the audio, shown in the vowel cards, Classroom Mode, and the Speech Profile report.
Calibrating your resting mouth width (once the camera is on) improves the roundedness reading, since it's measured relative to your own neutral position rather than a generic average.
Acoustic vowel chart & articulatory tongue contour
Click any vowel to link its acoustic measurements to an illustrative, ultrasound-style tongue contour. The contour is a teaching aid blended from cardinal-vowel articulatory targets using your formant values — it is not a direct reconstruction of your actual tongue movement.
Each vowel you've recorded is run once through the same causal-LSTM model, directly on that vowel's own audio (not live) — this overlays every resulting contour on one chart, colour-coded and labelled, so you can compare tongue posture across your vowels at a glance. This is a separate pipeline from the hand-authored illustrative contour above.
Estimated vocal-tract length
Each formant behaves like a resonance of a tube closed at the glottis and open at the lips. Working the resonance formula backwards from your measured formants gives an estimated acoustic vocal-tract length — not a direct anatomical measurement. Click a vowel below (or a row in the table) to see its own estimate; click "Average" to go back to the session-wide figure.
| Vowel | F1 | F2 | F3 | Est. VTL (cm) |
|---|
Classroom demonstration
Select a vowel to simultaneously animate tongue configuration, lip shape, vocal-tract resonance, and the acoustic vowel space. Falls back to typical adult formant and lip-rounding values for any vowel you haven't recorded, so this works even with no speaker data loaded.
Different tongue configurations → different vocal-tract resonances → different formant frequencies → different positions in acoustic space. Lip rounding mainly shifts F3 (and, for close vowels, F2) without changing tongue height.
How it works
The mathematics behind the two outputs, in plain terms.
Source–filter theory
Speech is modelled as a source — the buzzing sound made by the vocal folds — passed through a filter, the changing shape of the vocal tract (throat, mouth, lips). The filter boosts some frequencies and damps others, and it is those boosted frequencies — formants — that a listener hears as different vowels.
In the frequency domain, convolution becomes multiplication: Speech(f) = Source(f) × Filter(f).
Fourier transform
To see which frequencies are present, the app converts short windows of the recording from the time domain into the frequency domain using an FFT (Fast Fourier Transform), an efficient way of computing:
x(t) is the microphone signal over time; X(f) tells you how much energy is present at each frequency f. Sliding this window across the recording and stacking the results produces the spectrogram shown on the Vowel Profile tab.
Formants (F1, F2, F3)
Formants are peaks in the spectrum caused by resonances of the vocal tract — the same way a bottle resonates when you blow across it. This app finds them with linear predictive coding (LPC): it fits a small set of resonant "poles" to a short window of the signal (via the Levinson–Durbin recursion), then converts each pole's angle to a frequency and its distance from the unit circle to a bandwidth. Poles with a plausible frequency and a narrow bandwidth are kept as formant candidates.
- F1 tracks tongue height (low F1 → high tongue, close to the palate).
- F2 tracks tongue frontness/backness (high F2 → tongue fronted).
- F3 is influenced by lip rounding, tongue-tip shape and pharyngeal width.
Vocal-tract length from formants
Approximating the vocal tract as a uniform tube, closed at the glottis and open at the lips, its resonances form a quarter-wavelength series:
Fn is the n-th formant, c is the speed of sound (≈35,000 cm/s in warm, moist air), and L is the effective vocal-tract length. Rearranged for L and applied separately to F1 (n=1), F2 (n=2) and F3 (n=3), then averaged:
This is a simplified single-tube model. Real vocal tracts are non-uniform, so this is reported as an estimated acoustic length, distinct from an anatomical measurement.
Clinical acoustic measures (brief)
Beyond F0/F1/F2/F3, the Speech Profile report includes measures drawn from clinical voice-and-speech acoustics. Each reflects a different aspect of how the voice is produced, and each is associated with — not diagnostic of — the conditions listed. The same acoustic pattern can arise from several different causes, including a noisy recording.
| Measure | What it reflects | Commonly associated with |
|---|---|---|
| Jitter | Cycle-to-cycle pitch-period variation | Dysphonia, vocal-fold lesions, irregular vibration |
| Shimmer | Cycle-to-cycle amplitude variation | Dysphonia, glottic insufficiency |
| F0 SD | Pitch stability over the vowel | Tremor, Parkinsonian and other neurological voice disorders |
| HNR | Harmonic energy relative to noise | Breathy or rough voice quality |
| CPP | Strength/regularity of the harmonic structure | Overall dysphonia severity |
| H1–H2 | Relative strength of the first two harmonics | Breathy vs. pressed voice quality |
| Vowel space area | Size of the /i–a–u/ articulatory triangle | Dysarthria, reduced articulatory movement |
| Formant centralization ratio | How centralized vowels are in the vowel space | Parkinson's disease, hypokinetic dysarthria |
Jitter, shimmer, HNR and CPP here are simplified, frame-based approximations of the pitch-synchronous measures used in clinical tools like Praat or MDVP — suitable for relative, within-session comparison and teaching, not for clinical measurement. This tool also does not include prosody/timing measures (speech rate, pauses) or tremor analysis, since those need connected speech and several seconds of sustained phonation respectively, rather than short isolated vowels.
The Tongue Contour tab: acoustic-to-articulatory inversion
Everywhere else in this app, the tongue shape you see is illustrative — a hand-authored shape per cardinal vowel, blended by your formant readings (see "Formants" above). The Tongue Contour tab is different: it's the direct output of a small neural network trained to predict tongue shape from audio, a task called acoustic-to-articulatory inversion (AAI).
The model is a 2-layer causal LSTM (a recurrent network that only looks at current and past audio, never future audio — a requirement for anything running live rather than analyzing a finished recording). At each ~10ms step it takes 39 acoustic features — 13 MFCCs plus their first and second derivatives (Δ and ΔΔ), the same feature representation used throughout the acoustic-to-articulatory inversion literature — and outputs 20 (x, y) points tracing the tongue's surface from root to tip.
It runs entirely client-side: the trained weights are embedded in aai-weights.js, and the feature extraction plus LSTM forward pass are reimplemented from scratch in plain JavaScript (aai-model.js) — no server, no API call, nothing leaves your browser.
Two practical details worth knowing: the model normalizes its inputs against a running estimate of their mean and spread rather than a fixed constant, recalibrating to your microphone and voice within the first ~200ms of speaking — treat the first fraction of a second of any recording as a warm-up. And the delta features need a few frames of audio just after the moment being predicted, so live predictions lag the actual sound by roughly 80ms — imperceptible while speaking continuously, but the reason the very last instant of a cutoff recording won't be reflected in the final displayed frame.
Known limitations & validation
This is a browser-based teaching and screening tool, not a diagnostic instrument. Formant tracking on noisy recordings, overlapping consonant transitions, or very short vowels can be unreliable — the app flags low-confidence measurements. For research use, results should be spot-checked against dedicated acoustic analysis software such as Praat on the same recordings, and F0/F1/F2/F3/duration/VTL differences documented before drawing conclusions. Individual tongue shape is illustrative only; genuine articulatory reconstruction requires imaging techniques such as ultrasound tongue imaging, EMA, or MRI.
Your Speech Profile
A complete, exportable report combining the acoustic vowel profile and estimated vocal-tract profile.