I’ve just finished publishing preview material for all 17 categories of Human Vocality Primitives, a purpose-recorded dataset of non-lexical human vocal sound.
The premise is straightforward: most vocal audio in existing corpora captures these sounds incidentally, buried inside speech or performance, labeled as whatever the recording was actually for. HVP captures them as the subject rather than the byproduct. Breath, phonation modes, resonance behavior, and vocal gesture, each isolated and recorded deliberately.
The taxonomy is organized around control axes rather than sound labels. Each category isolates a single dimension of vocal production and captures systematic variation along it, with most categories including light modal anchor conditions as a within-category baseline.
Four subcategory groupings:
- Airflow & Airstream Primitives (4 categories) — sustained breath states, whisper and aspiration, breathing cycles, plosive and consonant bursts
- Phonation Mode Primitives (5 categories) — modal, breathy and semi-modal, vocal fry and subharmonic, falsetto/loft, pressed and constricted
- Formant & Resonance Primitives (3 categories) — vowel morphing and formant shifts, nasalization and velum control, overtone and harmonic formant interaction
- Gestural & Expressive Primitives (5 categories) — non-lexical vocal gestures, emotional vocalizations, mouth clicks and pops, inhales and reverse phonation, effort and exertion
Technical details: 96 kHz / 24-bit WAV, mono, captured in a single controlled room with one mic position and one signal chain per category. No mastering, compression, or normalization. Full metadata schema covering source, articulation labeling, capture conditions, and provenance.
All material is purpose-recorded and rights-cleared, with no scraped or third-party content.
Preview collection: Human Vocality Primitives (Previews) - a Harmonic-Frontier-Audio Collection
Happy to answer questions about the taxonomy, the capture approach, or anything else.