Speech Profile — SpeechLab IITD logo
My Speech Profile
SpeechLab IITD · Speech Analysis Class
SpeechLab · IIT Delhi

See your voice.

Acoustic & articulatory speech science

Record your voice, and this tool automatically generates your speech production profile by analyzing your vowel space, estimating your vocal-tract length, and capturing key features of your lip shape, tongue contour, and voice — all directly in your browser.

Developed by SpeechLab, IIT Delhi — for a speech analysis class
F1 · F2 · F3
/i e a o u ə/

What this tool does

Four automated outputs, generated from your own recordings
🎙️

Acoustic Vowel Profile

Automatic F0, F1, F2 and F3 extraction via LPC, plotted on a linked 2D/3D vowel chart.

Open Vowel Profile →
📐

Vocal-Tract Length

An estimated acoustic vocal-tract length from your formants, with an animated 3D tract.

Open Vocal Tract →
🎓

Classroom Mode

Click through vowels to demonstrate source–filter theory live, with or without recorded data.

Open Classroom Mode →
📋

Speech Profile Report

A clinical-style, exportable report summarising every measurement in one document.

Open Speech Profile →
6Target vowels /i e a o u ə/
F0·F1·F2·F3Extracted automatically
100%Runs in your browser
0Data ever uploaded
Step 1

Record your target vowels

Read each vowel clearly for about a second in a quiet room, close to the microphone — or click the ⇪ upload button on any vowel card to use a short existing audio clip instead. Either way, the system automatically finds the steady portion of the recording and extracts F0, F1, F2 and F3 — you never need to enter formant values by hand.

Input
Want to watch your tongue move while you record? Open Tongue Contour → Want lip shape captured alongside each recording? Open Lip Capture →
Optional: record additional IPA vowels (full chart, 22 more symbols) ▾

These aren't required — the 6 core vowels above are enough for a profile. Anything you record here is analyzed too, and shows up in the vowel chart, tract table and report alongside the core set.

Record at least 3 vowels to continue.
Live · Experimental

Tongue Contour

A small neural network predicts a 20-point tongue contour directly from your microphone, frame by frame, entirely in your browser. Enable the microphone below, hit start, and speak — watch it track in real time as you move through different vowels. Details on how the model works are on the How It Works page.

Live AI-predicted contour
Direct neural-network output, updated in real time — not the blended illustrative contour used elsewhere in this app.
Microphone & status
Inference time
–
Frames predicted
Off

The microphone is shared across the whole app — enabling it here also enables it for recording vowels, and vice versa.

This model was trained entirely on synthetic formant-resonance audio paired with a stylised contour shape — not on real measured articulography (EMA / ultrasound / rt-MRI) or real recorded voices. Treat this as a live demonstration that an acoustic-to-articulatory inversion pipeline can run end to end in a browser — not as a validated measurement of your actual tongue movement. See the How It Works page, or README-AAI.md, for full details.
Optional

Lip Capture

Reads lip shape (open / closed / rounded / spread) live from your webcam using on-device face landmarks — video never leaves your browser. Enable the camera below before recording vowels on the Record tab, and each vowel recording will also capture a lip-shape reading alongside the audio.

Camera
Lips: —
Camera off — lip shape won't be captured
About this feature

Runs entirely on-device (WASM face landmarks) — video is never uploaded or stored anywhere, only the two numbers (mouth openness, mouth roundedness) derived from it. When enabled, each vowel recording on the Record tab also captures a lip-shape reading alongside the audio, shown in the vowel cards, Classroom Mode, and the Speech Profile report.

Calibrating your resting mouth width (once the camera is on) improves the roundedness reading, since it's measured relative to your own neutral position rather than a generic average.

Purely illustrative: thresholds for open / closed / rounded / spread are heuristic, not a validated articulatory measurement.
Output 1

Acoustic vowel chart & articulatory tongue contour

Click any vowel to link its acoustic measurements to an illustrative, ultrasound-style tongue contour. The contour is a teaching aid blended from cardinal-vowel articulatory targets using your formant values — it is not a direct reconstruction of your actual tongue movement.

F1–F2 acoustic vowel space view in 3D (F1·F2·F3) →
Illustrative tongue contour
Educational model, styled after ultrasound tongue imaging — blended from cardinal-vowel articulatory targets by F1 (height) / F2 (frontness) / F3 (rounding). Not a measured articulatory image.
AI-predicted tongue contours — all recorded vowels Open live Tongue Contour →

Each vowel you've recorded is run once through the same causal-LSTM model, directly on that vowel's own audio (not live) — this overlays every resulting contour on one chart, colour-coded and labelled, so you can compare tongue posture across your vowels at a glance. This is a separate pipeline from the hand-authored illustrative contour above.

Direct neural-network output per vowel, averaged over each recording's stable frames — not the blended illustrative contour used elsewhere in this app.
This model was trained entirely on synthetic formant-resonance audio, not on real measured articulography or real recorded voices — treat every line here as a demonstration that the inversion pipeline runs on each vowel, not as a validated measurement of your actual tongue shape. Full caveats on the How It Works page.
Signal for selected vowel
Waveform
Spectrogram (0–5 kHz) with F1 / F2 / F3 tracks
Output 2

Estimated vocal-tract length

Each formant behaves like a resonance of a tube closed at the glottis and open at the lips. Working the resonance formula backwards from your measured formants gives an estimated acoustic vocal-tract length — not a direct anatomical measurement. Click a vowel below (or a row in the table) to see its own estimate; click "Average" to go back to the session-wide figure.

3D vocal tract
Select a vowel to estimate length
Per-vowel estimates
VowelF1F2F3Est. VTL (cm)
Averaged estimate
—
Formant-based n
F1·F2·F3
Labelled explicitly as estimated acoustic vocal-tract length. It can differ from a person's actual anatomical vocal-tract length, which would require imaging (MRI, ultrasound, or X-ray) to measure directly.
Teaching mode

Classroom demonstration

Select a vowel to simultaneously animate tongue configuration, lip shape, vocal-tract resonance, and the acoustic vowel space. Falls back to typical adult formant and lip-rounding values for any vowel you haven't recorded, so this works even with no speaker data loaded.

Tongue configuration
Lip shape
–
Acoustic vowel space position
Vowel
–
F1
–
F2
–
F3
–
Est. VTL
–

Different tongue configurations → different vocal-tract resonances → different formant frequencies → different positions in acoustic space. Lip rounding mainly shifts F3 (and, for close vowels, F2) without changing tongue height.

Reference

How it works

The mathematics behind the two outputs, in plain terms.

Source–filter theory

Speech is modelled as a source — the buzzing sound made by the vocal folds — passed through a filter, the changing shape of the vocal tract (throat, mouth, lips). The filter boosts some frequencies and damps others, and it is those boosted frequencies — formants — that a listener hears as different vowels.

speech signal(t) = source(t) ✳ vocal-tract filter(t)

In the frequency domain, convolution becomes multiplication: Speech(f) = Source(f) × Filter(f).

Fourier transform

To see which frequencies are present, the app converts short windows of the recording from the time domain into the frequency domain using an FFT (Fast Fourier Transform), an efficient way of computing:

X(f) = ∫ x(t) · e−2πift dt

x(t) is the microphone signal over time; X(f) tells you how much energy is present at each frequency f. Sliding this window across the recording and stacking the results produces the spectrogram shown on the Vowel Profile tab.

Formants (F1, F2, F3)

Formants are peaks in the spectrum caused by resonances of the vocal tract — the same way a bottle resonates when you blow across it. This app finds them with linear predictive coding (LPC): it fits a small set of resonant "poles" to a short window of the signal (via the Levinson–Durbin recursion), then converts each pole's angle to a frequency and its distance from the unit circle to a bandwidth. Poles with a plausible frequency and a narrow bandwidth are kept as formant candidates.

  • F1 tracks tongue height (low F1 → high tongue, close to the palate).
  • F2 tracks tongue frontness/backness (high F2 → tongue fronted).
  • F3 is influenced by lip rounding, tongue-tip shape and pharyngeal width.

Vocal-tract length from formants

Approximating the vocal tract as a uniform tube, closed at the glottis and open at the lips, its resonances form a quarter-wavelength series:

Fn ≈ (2n − 1)·c / (4L)

Fn is the n-th formant, c is the speed of sound (≈35,000 cm/s in warm, moist air), and L is the effective vocal-tract length. Rearranged for L and applied separately to F1 (n=1), F2 (n=2) and F3 (n=3), then averaged:

L ≈ (2n − 1)·c / (4·Fn)

This is a simplified single-tube model. Real vocal tracts are non-uniform, so this is reported as an estimated acoustic length, distinct from an anatomical measurement.

Clinical acoustic measures (brief)

Beyond F0/F1/F2/F3, the Speech Profile report includes measures drawn from clinical voice-and-speech acoustics. Each reflects a different aspect of how the voice is produced, and each is associated with — not diagnostic of — the conditions listed. The same acoustic pattern can arise from several different causes, including a noisy recording.

MeasureWhat it reflectsCommonly associated with
JitterCycle-to-cycle pitch-period variationDysphonia, vocal-fold lesions, irregular vibration
ShimmerCycle-to-cycle amplitude variationDysphonia, glottic insufficiency
F0 SDPitch stability over the vowelTremor, Parkinsonian and other neurological voice disorders
HNRHarmonic energy relative to noiseBreathy or rough voice quality
CPPStrength/regularity of the harmonic structureOverall dysphonia severity
H1–H2Relative strength of the first two harmonicsBreathy vs. pressed voice quality
Vowel space areaSize of the /i–a–u/ articulatory triangleDysarthria, reduced articulatory movement
Formant centralization ratioHow centralized vowels are in the vowel spaceParkinson's disease, hypokinetic dysarthria

Jitter, shimmer, HNR and CPP here are simplified, frame-based approximations of the pitch-synchronous measures used in clinical tools like Praat or MDVP — suitable for relative, within-session comparison and teaching, not for clinical measurement. This tool also does not include prosody/timing measures (speech rate, pauses) or tremor analysis, since those need connected speech and several seconds of sustained phonation respectively, rather than short isolated vowels.

Acoustic features indicate patterns associated with a disorder; they do not diagnose a disorder by themselves. Interpretation of these measures should be made by a qualified speech-language pathologist or physician as part of a full clinical evaluation.

The Tongue Contour tab: acoustic-to-articulatory inversion

Everywhere else in this app, the tongue shape you see is illustrative — a hand-authored shape per cardinal vowel, blended by your formant readings (see "Formants" above). The Tongue Contour tab is different: it's the direct output of a small neural network trained to predict tongue shape from audio, a task called acoustic-to-articulatory inversion (AAI).

The model is a 2-layer causal LSTM (a recurrent network that only looks at current and past audio, never future audio — a requirement for anything running live rather than analyzing a finished recording). At each ~10ms step it takes 39 acoustic features — 13 MFCCs plus their first and second derivatives (Δ and ΔΔ), the same feature representation used throughout the acoustic-to-articulatory inversion literature — and outputs 20 (x, y) points tracing the tongue's surface from root to tip.

audio → MFCC + Δ + ΔΔ (39-dim) → 2-layer causal LSTM → 20-point tongue contour

It runs entirely client-side: the trained weights are embedded in aai-weights.js, and the feature extraction plus LSTM forward pass are reimplemented from scratch in plain JavaScript (aai-model.js) — no server, no API call, nothing leaves your browser.

Two practical details worth knowing: the model normalizes its inputs against a running estimate of their mean and spread rather than a fixed constant, recalibrating to your microphone and voice within the first ~200ms of speaking — treat the first fraction of a second of any recording as a warm-up. And the delta features need a few frames of audio just after the moment being predicted, so live predictions lag the actual sound by roughly 80ms — imperceptible while speaking continuously, but the reason the very last instant of a cutoff recording won't be reflected in the final displayed frame.

Trained entirely on synthetic audio (formant-resonance filters standing in for vowels), not on real measured articulography (EMA, ultrasound, or rt-MRI) or real recorded voices. This tab demonstrates that the acoustic-to-articulatory inversion pipeline runs live, end to end, in a browser — it is not a validated measurement of your actual tongue movement. Full technical details, including what retraining on real lab data would involve, are in the project's README-AAI.md.

Known limitations & validation

This is a browser-based teaching and screening tool, not a diagnostic instrument. Formant tracking on noisy recordings, overlapping consonant transitions, or very short vowels can be unreliable — the app flags low-confidence measurements. For research use, results should be spot-checked against dedicated acoustic analysis software such as Praat on the same recordings, and F0/F1/F2/F3/duration/VTL differences documented before drawing conclusions. Individual tongue shape is illustrative only; genuine articulatory reconstruction requires imaging techniques such as ultrasound tongue imaging, EMA, or MRI.

Summary

Your Speech Profile

A complete, exportable report combining the acoustic vowel profile and estimated vocal-tract profile.