Interactive listening evidence · Base & Medium

A listening atlas for emotion, voice and vocal bursts

Hear what the Humaneness Ears Base and Medium encoders predict on real audio. Explore vocal bursts in time, compare Orange speaker embeddings, and listen across emotion, voice-style and quality rankings. They use OpenAI Whisper Base and Whisper Small encoder backbones respectively; every player opens over HTTPS.

1,000Emolia listening clips · 499 DE / 501 EN
50timed vocal-burst examples · 139 Gemini events
192continuous scores from each model
108emotion, voice and quality listening pages

Start listening

Choose a listening path. The audio streams separately, so these pages open in the browser.

Explore what the scores mean

Each page ranks the same 1,000-clip pool separately for Humaneness Ears Base and Medium.

Data, models and measured accuracy

Listening examples help interpret predictions; independent benchmarks measure agreement.

How to interpret the audio

These listening pages are qualitative examples, not an accuracy estimate.

The 1,000 Emolia examples were excluded from the two-epoch Gemini fine-tune, but exact overlap with the earlier S1–S10 curriculum is possible. The 50 burst clips were selected for varied Gemini-labeled events and model confidence. For independent numbers, use the benchmark report above.

The Whisper audio encoders have no text decoder and do not transcribe speech. Orange vectors are teacher/model embeddings, not named speaker identities. Class softmax values are uncalibrated. Every event interval has a start and end; score z values describe position in the training reference distribution, not probability.

Audio comes from Emolia or individually CC BY 4.0 marked records in the Scaling Ladder. Page manifests retain source ID, SHA256 and split. Taxonomies and teacher packages are linked in the model README.