Every take passed the check and the voice still changed between lines
A generated narration changed speaker every few lines. The quality gate that existed to catch exactly that passed all 17 lines on the first take, word perfect. The gate scored transcript similarity, and a transcriber cannot hear pitch, so the defect sat in full view of a check that was structurally incapable of seeing it.
TL;DR · THE FIX
A check that passes on the wrong axis is the same failure as a health endpoint returning 200 from a layer that never touched the database. My take-selection gate scored generated speech by transcribing it and comparing the text, so it measured intelligibility and nothing else while the fundamental frequency wandered 92 Hz across one video. The fix is to score the axis the defect lives on (median F0 per take), anchor to the first accepted output rather than the first intelligible one, and measure the baseline before changing anything so you know which half is a regression.
The symptom
A generated voiceover: seventeen lines, one script, one voice setting. Played back, it sounds like three people taking turns. The voices are close enough that a transcript would never show it, but a listener stops following the content and starts wondering what is wrong with the audio.
The pipeline that produced it has a quality gate whose job is to reject bad takes and re-roll them, and it had run:
line 01 take 1 ACCEPTED
line 02 take 1 ACCEPTED
line 03 take 1 ACCEPTED
...
line 17 take 1 ACCEPTED
17/17 accepted on first take. 100% word accuracy.
Seventeen out of seventeen, word perfect, no retries. By its own report this was the cleanest run the tool had ever produced.
What I tried first
The instinct is to distrust the generator and start tuning it: a different seed, a different sampling temperature, a harder pin on the voice reference. That is how you spend an afternoon.
I did something duller first and measured the previous video, the one that had already shipped and sounded fine. Without a baseline you cannot tell a regression from a house style, and you will “fix” things that were never broken. Both videos through the same measurement:
F0 spread (max - min of per-line median pitch)
already shipped 25 Hz
new video 92 Hz
articulation rate
already shipped 222 wpm
new video 224 wpm
The pitch spread is nearly four times wider, a real regression that matches what I was hearing. The articulation rate is effectively identical, so the “too fast” complaint also sitting in my notes was the voice the channel had always used. Without that baseline I would have spent hours slowing down a deliberate choice.
What was happening
The gate generated a take, transcribed it with Whisper, compared the transcript to the script, accepted if the words matched closely enough, and otherwise re-rolled. The comparison is text against text. Whisper hands you words, with no pitch, timbre, or speaker identity attached, so two takes of the same line in two different voices produce byte-identical transcripts and both score 100%.
The gate was measuring intelligibility, correctly and reliably, on every line, and intelligibility was never the problem. It could not observe the thing it had been built to prevent, so its 17/17 said nothing about voice.
The second half of the mechanism is what made the drift possible. The model samples a voice latent per generation, so a re-roll produces a different voice rather than a cleaner version of the same one. A retry loop whose exit condition ignores an attribute will therefore drift on that attribute, every time, by construction. The line that ended up at 214 Hz against a body sitting around 140 Hz was almost certainly one that got re-rolled once for a genuine reason, came back intelligible, and was accepted as it was.
The same thing happens outside audio. Retry an image generation “until the text is legible” and the composition drifts. Retry a structured extraction “until it parses” and the field semantics drift. The exit condition is the only thing holding the output still, and everything it does not name is free to move.
The fix
Score the axis the defect lives on. Here, pitch:
import numpy as np
import librosa
def median_f0(path, fmin=70, fmax=350):
y, sr = librosa.load(path, sr=None)
f0, voiced, _ = librosa.pyin(y, fmin=fmin, fmax=fmax, sr=sr)
voiced_f0 = f0[voiced & ~np.isnan(f0)]
if voiced_f0.size == 0:
return None
return float(np.median(voiced_f0))
Then anchor, and the anchor is neither a constant you configure nor the running mean:
anchor = None # set by the FIRST ACCEPTED line, not the first generated one
TOLERANCE_HZ = 12
for line in script:
for attempt in range(MAX_TAKES):
wav = generate(line)
if not transcript_matches(wav, line):
continue # unintelligible, re-roll as before
pitch = median_f0(wav)
if anchor is None:
anchor = pitch # first line clearing BOTH bars defines the voice
accept(wav)
break
if abs(pitch - anchor) <= TOLERANCE_HZ:
accept(wav)
break
# intelligible but wrong voice: re-roll
else:
raise RuntimeError("no take within tolerance of anchor for: " + line)
Anchoring on the first accepted output rather than the first generated one matters because take one of line one is as likely to be an outlier as any other take. Anchor on a bad take and the whole video matches it.
Two smaller repairs came out of the same pass. The script says two hundred and the transcriber writes 200; both are correct, and a naive string comparison rejects a perfect take and sends you into a re-roll loop that can only end in drift, so numerals get normalised out of both sides before scoring. And short inputs turned out to be the hard case: a one-word line, "Wrong.", came back as "So it's a caged sickle". With almost no context for the model to condition on, a single word is more likely to be mangled than a full sentence, so verification cannot be skipped on inputs that look trivial.
The result
before after
F0 spread 92 Hz 25 Hz
articulation 224 wpm 222 wpm
Spread back to the reference video’s level with articulation untouched, which confirms the change did what it claimed and nothing else.
The lesson
I have written this failure up before in a completely different layer. A free-tier database keepalive pinged an auth health endpoint and got a clean 200 every day for weeks while the database it was meant to keep awake was never touched by that request. Here it is again in a generative pipeline: a check that cannot observe the failure mode will report success during the failure, and its green carries no information.
For any gate you write, name the defect you are trying to catch in one sentence before writing the check (“the voice changes between lines”), then ask whether the thing you are about to measure could physically detect that sentence being true. Transcript similarity cannot detect a voice change, and that is answerable in ten seconds. Measure the baseline before you change anything, or you cannot separate a regression from a deliberate choice and you will burn time on the half that was never broken. And anchor retries to the first accepted output, because anything the exit condition does not name will drift unless you hold it still on purpose.
The 17/17 made this harder to find. A tool with no check at all would have sent me straight to listening. A tool with a confident, precise, wrong check sent me looking everywhere except at the check itself.
Discussion
Powered by GitHub. Sign in to leave a comment.