FIELD NOTES / 001AUDIO TOKENIZER SANITY CHECK

Listen to the
round trip.●

The same moment. Four versions.
Listen for what survives the trip.

Enter the listening room
THE SIGNAL PATH STORED TOKENS
AMPLITUDE
IN01: ENCODE02: STORE03: DECODEOUT
One source.
Three return paths.
QCF
QWEN · CHATTERBOX S3 · FISH DAC
15random clips
03source collections
45token reconstructions

Decoded from previously saved R2 tokens.
No fresh encoding. Exact input audio verified.

15 recordings · 60 listening lanes RAW LEVELS · NO GAIN MATCHING
Loading the listening room…

Chatterbox S3 uses the source clip’s speaker and mel reference, plus a prompt prefix from its saved tokens. Qwen and Fish reconstruct from saved tokens alone. Fish’s original encoding normalized each clip to −16 LUFS, so its reconstruction may sound louder or quieter than the source.

HOW THIS ROOM WAS MADE

An honest
round trip.

A listening check, not a leaderboard.
No hand-picked best takes.

Download the provenance manifest ↗
01

Draw from the archive

Five random clips from each collection, sampled from jobs with all three token outputs committed to R2. The seed and selection ranks are recorded.

02

Recover the exact input

Each source is decoded and cropped using its saved sample boundaries. Its PCM hash must match the audio originally passed to the tokenizers. Enhanced source inputs are labeled.

03

Decode the stored codes

The existing R2 shards are downloaded and hash-verified. Matching decoders reconstruct all saved codebooks. The clips are never encoded again.

04

Leave the differences audible

Raw decoder output is preserved: no gain matching, denoising, or time warping. Switching follows the original clip’s timeline; downloads retain the full output, including decoder tail padding.