Datasets:
Dataset Viewer
The dataset viewer is not available for this split.
The number of columns (7698) exceeds the maximum supported number of columns (1000). This is a current limitation of the datasets viewer. You can reduce the number of columns if you want the viewer to work.
Error code: TooManyColumnsError
Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
Katachi Benchmark Corpora (juxtnorth/katachi-benchmarks)
High-precision phonetic & acoustic benchmark suites for Japanese speech evaluation, pitch accent contour analysis, and micro-stutter detection.
Dataset Structure & Download Archives
This repository hosts pre-processed, downsampled (16kHz mono), and loudness-normalized (-23 LUFS) benchmark datasets for the Katachi evaluation engine. Assets are packaged as modular domain bundles for fast, zero-dependency streaming:
corpora/jvs-certification-16k.tar.gz: JVS (Japanese Versatile Speech) Corpus — 100 native speakers, 16kHz Sinc downsampled, normalized to -23 LUFS (15,000 utterances).corpora/jsut-certification-16k.tar.gz: JSUT basic5000 Corpus — Single speaker phonetically balanced sentences, 16kHz Sinc downsampled, -23 LUFS (7,696 utterances).corpora/synthetic-adversarial.tar.gz: 14 adversarial stress test fixtures across Categories A-D (glottal blocks, pitch drops, flutter, noise).database/kanjium-accents.sqlite: Pre-compiled Kanjium pitch accent SQLite database.models/: Wav2Vec2 Japanese phoneme CTC and Silero VAD ONNX models.manifest.lock.json: Cryptographic SHA-256 manifest of all assets.
Acoustic & Processing Standards
- Sample Rate: 16,000 Hz (Blackman-Harris windowed Sinc anti-aliased resampling)
- Loudness: -23.0 LUFS integrated loudness (ITU-R BS.1770 / EBU R128)
- Format: 16-bit Linear PCM Mono WAV
- Companion Descriptors: Full mora alignments, LPC vowel formants (F1-F3), and YIN/pYIN F0 pitch contours.
Licensing & Attribution Register
This dataset is partitioned under multiple licenses per upstream authors:
| Asset / Partition | Provenance / Authors | Upstream License | Commercial Use |
|---|---|---|---|
JVS Corpus (corpora/jvs-certification-16k.tar.gz) |
Takaki, Kameoka, Yamagishi (NII) | CC-BY-SA-4.0 |
Permitted (with attribution) |
JSUT Corpus (corpora/jsut-certification-16k.tar.gz) |
Sonobe, Takamichi, Saruwatari (UTokyo) | CC-BY-SA-4.0 |
Permitted (with attribution) |
Kanjium SQLite (database/) |
Lars Yencken, mifunetoshiro, EDRDG | CC-BY-SA-4.0 |
Permitted (with attribution) |
Synthetic Fixtures (corpora/synthetic-adversarial.tar.gz) |
Voicevox / No.7 Project (小岩井ことり) | Custom Non-Commercial | Restricted (Requires paid contract) |
Mandatory Attribution Notices
- JVS: Derived from the JVS Corpus (© 2019 National Institute of Informatics, Japan) by Shinnosuke Takaki, Hirokazu Kameoka, and Junichi Yamagishi, licensed under CC BY-SA 4.0.
- JSUT: Derived from the JSUT Corpus (© 2017 Saruwatari Lab, The University of Tokyo) by Ryosuke Sonobe, Shinnosuke Takamichi, and Hiroshi Saruwatari, licensed under CC BY-SA 4.0.
- Kanjium: Pitch accent data derived from the Kanjium project (© Lars Yencken, mifunetoshiro), built upon EDICT/JMdict dictionary files (© Electronic Dictionary Research and Development Group), licensed under CC BY-SA 4.0.
- VOICEVOX No.7: Synthetic test fixtures generated using VOICEVOX and Voice Library No.7. Credit: VOICEVOX:No.7 (Voice Actor: Kotori Koiwai). Official Terms of Use: https://voiceseven.com/ and https://voicevox.hiroshiba.jp/. Non-commercial research and benchmarking use only. Commercial use requires a separate license from the No.7 Project.
- Downloads last month
- 21