No samples, no recordings. A Klatt synthesizer written from scratch and rendered a sample at a time: a glottal pulse through five cascaded resonators and a nasal pole–zero pair. In front of it sits a pronouncing front-end — lexicon, stress, vowel reduction, consonant loci, flapping — so it speaks English properly first. Behind it, a language layer rewrites the syllables on the way out until it stops being English at all.
A browser's filters are musical EQ: gain and Q set independently, describing no vocal
tract that exists. So the voice is built sample by sample from Klatt's own resonator,
y[n] = A·x[n] + B·y[n−1] + C·y[n−2], with coefficients recomputed as the
mouth moves.
A lexicon of common words, then context rules for the rest. English spelling is a poor
guide to sound — island read literally comes out as ISS-land.
You hear bad before its final consonant arrives, because the vowel runs half again as long as the one in bat. Tense vowels stretch, lax ones clip, and everything slows into a pause.
The bursts of b d g are nearly identical. You hear the difference in whether
F2 rises or falls on the way into the vowel, so each stop aims the formants at its own
locus before releasing.
Every other filter here adds a resonance. A nasal opens the velum and the sealed mouth
becomes a side branch that cancels a band — near 750 Hz for m, 1600
for n. That notch is the whole difference between them.
Intelligibility comes from the formants; sounding human comes from the excitation. No two glottal cycles in a real voice are the same length or the same loudness, so no two here are either.