We show some examples of audios regenerated with different representations.
"Complete" representations almost exactly reproduce copy-synthesis, whereas
disentangled ones tend to sample different speakers and conditions, which vary
with the chosen seed. Using guidance increases audio quality and reduces variance
between seeds.
Reference
Vocos
Representation
Seeds
Notes
Guidance is classifier-free guidance on the flow-matching model. Off
(w = 1) samples the conditional field as trained; on
(w = 2) extrapolates towards the distribution mode.
Utterances are from LibriTTS test-clean, unseen in
training.