Draw from the archive
Five random clips from each collection, sampled from jobs with all three token outputs committed to R2. The seed and selection ranks are recorded.
FIELD NOTES / 001AUDIO TOKENIZER SANITY CHECK
The same moment. Four versions.
Listen for what survives the trip.
Decoded from previously saved R2 tokens.
No fresh encoding. Exact input audio verified.
Chatterbox S3 uses the source clip’s speaker and mel reference, plus a prompt prefix from its saved tokens. Qwen and Fish reconstruct from saved tokens alone. Fish’s original encoding normalized each clip to −16 LUFS, so its reconstruction may sound louder or quieter than the source.
HOW THIS ROOM WAS MADE
A listening check, not a leaderboard.
No hand-picked best takes.
Five random clips from each collection, sampled from jobs with all three token outputs committed to R2. The seed and selection ranks are recorded.
Each source is decoded and cropped using its saved sample boundaries. Its PCM hash must match the audio originally passed to the tokenizers. Enhanced source inputs are labeled.
The existing R2 shards are downloaded and hash-verified. Matching decoders reconstruct all saved codebooks. The clips are never encoded again.
Raw decoder output is preserved: no gain matching, denoising, or time warping. Switching follows the original clip’s timeline; downloads retain the full output, including decoder tail padding.