Play the audio and follow the subtitles. Each word is colored by the speaker that method assigned; a red underline means it disagrees with ground truth.
Source: Harper Valley Bank (CC BY 4.0) — no patient audio.
What each method knows:AWS and W+ECAPA self-enrolled are fully blind — profiles
come from the recording's own clusters, with no outside reference audio. W+ECAPA enrolled and FT v2
were given about 20 s of reference audio from each person's clean channel — available for care partners
in the archive, rarely for patients. Compare the first two for a fair test, the last two to see what enrollment adds.
Scores
method
words written
speaker attribution
wrong words
Subtitles
FT v2 — extracted audio
The other methods only label; they never split the audio. FT v2 actually extracts it — this is the per-person output.