Cascade formant speech synthesizer

A synthesizer for languages that don't exist

No samples, no recordings. A Klatt synthesizer written from scratch and rendered a sample at a time: a glottal pulse through five cascaded resonators and a nasal pole–zero pair. In front of it sits a pronouncing front-end — lexicon, stress, vowel reduction, consonant loci, flapping — so it speaks English properly first. Behind it, a language layer rewrites the syllables on the way out until it stops being English at all.

Clearest
Vol 85

Spectrogram · 0–5 kHz
Touch first · pitch, tract size, speed, consonant
PITCH SIZE
Presets
Sound design · 48 controls behind five tabs

Syllable trace · bold marks stress
— press Speak —

Rendered a sample at a time

A browser's filters are musical EQ: gain and Q set independently, describing no vocal tract that exists. So the voice is built sample by sample from Klatt's own resonator, y[n] = A·x[n] + B·y[n−1] + C·y[n−2], with coefficients recomputed as the mouth moves.

Why it can be understood

A lexicon of common words, then context rules for the rest. English spelling is a poor guide to sound — island read literally comes out as ISS-land.

Length carries meaning

You hear bad before its final consonant arrives, because the vowel runs half again as long as the one in bat. Tense vowels stretch, lax ones clip, and everything slows into a pause.

Consonant loci

The bursts of b d g are nearly identical. You hear the difference in whether F2 rises or falls on the way into the vowel, so each stop aims the formants at its own locus before releasing.

The nose subtracts

Every other filter here adds a resonance. A nasal opens the velum and the sealed mouth becomes a side branch that cancels a band — near 750 Hz for m, 1600 for n. That notch is the whole difference between them.

Nothing repeats exactly

Intelligibility comes from the formants; sounding human comes from the excitation. No two glottal cycles in a real voice are the same length or the same loudness, so no two here are either.