A listening atlas for emotion, voice and vocal bursts
Hear what the Humaneness Ears Base and Medium encoders predict on real audio. Explore vocal bursts in time, compare Orange speaker embeddings, and listen across emotion, voice-style and quality rankings. They use OpenAI Whisper Base and Whisper Small encoder backbones respectively; every player opens over HTTPS.
Start listening
Choose a listening path. The audio streams separately, so these pages open in the browser.
Find laughter, cries and other bursts in time
50 clips with start and end bars, three predicted burst labels and all 192 Base/Medium scores.
Listen to the burst atlas ↗ 02 · Full annotationOne recording, every annotation explained
Hear the source and inspect its transcript, captions, emotions, voice dimensions and timed events.
Open the complete example ↗ 03 · Orange timbre · 128DDo the nearest voices sound alike?
20 queries compare the Orange teacher’s three neighbors with each model’s three predicted neighbors.
Compare timbre neighbors ↗ 04 · Orange identity · 250DHear the identity-embedding neighborhood
The same 20-query study for Orange identity vectors, with cosine values beside every player.
Compare identity neighbors ↗Explore what the scores mean
Each page ranks the same 1,000-clip pool separately for Humaneness Ears Base and Medium.
Five listening points along each score
Genuineness, burst blend, AudioBox, DNSMOS, background noise and aesthetics, from low to high.
Browse 11 quality axes ↗ Emotion and voice40 EmoNet emotions + 57 VoiceNet dimensions
Hear the top ten Base and Medium predictions for anger, warmth, tension and every other dimension.
Browse 97 dimensions ↗Data, models and measured accuracy
Listening examples help interpret predictions; independent benchmarks measure agreement.
Independent benchmark report
Compare Whisper and CLAP on human emotion ratings, speaker vectors and burst timing.
Read the full results ↗ ReproduceModel weights, inference code and methods
Check checkpoints, training details, target definitions and their source repositories.
Open the model repository ↗ DatasetGemini annotation dataset
Download validated WebDataset shards with audio or annotation sidecars and recorded source rights.
Browse the dataset ↗How to interpret the audio
These listening pages are qualitative examples, not an accuracy estimate.
The 1,000 Emolia examples were excluded from the two-epoch Gemini fine-tune, but exact overlap with the earlier S1–S10 curriculum is possible. The 50 burst clips were selected for varied Gemini-labeled events and model confidence. For independent numbers, use the benchmark report above.
The Whisper audio encoders have no text decoder and do not transcribe speech. Orange vectors are teacher/model embeddings, not named speaker identities. Class softmax values are uncalibrated. Every event interval has a start and end; score z values describe position in the training reference distribution, not probability.
Audio comes from Emolia or individually CC BY 4.0 marked records in the Scaling Ladder. Page manifests retain source ID, SHA256 and split. Taxonomies and teacher packages are linked in the model README.