Dataset Viewer
Duplicate
The dataset viewer is not available for this split.
The number of columns (7698) exceeds the maximum supported number of columns (1000). This is a current limitation of the datasets viewer. You can reduce the number of columns if you want the viewer to work.
Error code:   TooManyColumnsError

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

Katachi Benchmark Corpora (juxtnorth/katachi-benchmarks)

High-precision phonetic & acoustic benchmark suites for Japanese speech evaluation, pitch accent contour analysis, and micro-stutter detection.

Dataset Structure & Download Archives

This repository hosts pre-processed, downsampled (16kHz mono), and loudness-normalized (-23 LUFS) benchmark datasets for the Katachi evaluation engine. Assets are packaged as modular domain bundles for fast, zero-dependency streaming:

  • corpora/jvs-certification-16k.tar.gz: JVS (Japanese Versatile Speech) Corpus — 100 native speakers, 16kHz Sinc downsampled, normalized to -23 LUFS (15,000 utterances).
  • corpora/jsut-certification-16k.tar.gz: JSUT basic5000 Corpus — Single speaker phonetically balanced sentences, 16kHz Sinc downsampled, -23 LUFS (7,696 utterances).
  • corpora/synthetic-adversarial.tar.gz: 14 adversarial stress test fixtures across Categories A-D (glottal blocks, pitch drops, flutter, noise).
  • database/kanjium-accents.sqlite: Pre-compiled Kanjium pitch accent SQLite database.
  • models/: Wav2Vec2 Japanese phoneme CTC and Silero VAD ONNX models.
  • manifest.lock.json: Cryptographic SHA-256 manifest of all assets.

Acoustic & Processing Standards

  • Sample Rate: 16,000 Hz (Blackman-Harris windowed Sinc anti-aliased resampling)
  • Loudness: -23.0 LUFS integrated loudness (ITU-R BS.1770 / EBU R128)
  • Format: 16-bit Linear PCM Mono WAV
  • Companion Descriptors: Full mora alignments, LPC vowel formants (F1-F3), and YIN/pYIN F0 pitch contours.

Licensing & Attribution Register

This dataset is partitioned under multiple licenses per upstream authors:

Asset / Partition Provenance / Authors Upstream License Commercial Use
JVS Corpus (corpora/jvs-certification-16k.tar.gz) Takaki, Kameoka, Yamagishi (NII) CC-BY-SA-4.0 Permitted (with attribution)
JSUT Corpus (corpora/jsut-certification-16k.tar.gz) Sonobe, Takamichi, Saruwatari (UTokyo) CC-BY-SA-4.0 Permitted (with attribution)
Kanjium SQLite (database/) Lars Yencken, mifunetoshiro, EDRDG CC-BY-SA-4.0 Permitted (with attribution)
Synthetic Fixtures (corpora/synthetic-adversarial.tar.gz) Voicevox / No.7 Project (小岩井ことり) Custom Non-Commercial Restricted (Requires paid contract)

Mandatory Attribution Notices

  • JVS: Derived from the JVS Corpus (© 2019 National Institute of Informatics, Japan) by Shinnosuke Takaki, Hirokazu Kameoka, and Junichi Yamagishi, licensed under CC BY-SA 4.0.
  • JSUT: Derived from the JSUT Corpus (© 2017 Saruwatari Lab, The University of Tokyo) by Ryosuke Sonobe, Shinnosuke Takamichi, and Hiroshi Saruwatari, licensed under CC BY-SA 4.0.
  • Kanjium: Pitch accent data derived from the Kanjium project (© Lars Yencken, mifunetoshiro), built upon EDICT/JMdict dictionary files (© Electronic Dictionary Research and Development Group), licensed under CC BY-SA 4.0.
  • VOICEVOX No.7: Synthetic test fixtures generated using VOICEVOX and Voice Library No.7. Credit: VOICEVOX:No.7 (Voice Actor: Kotori Koiwai). Official Terms of Use: https://voiceseven.com/ and https://voicevox.hiroshiba.jp/. Non-commercial research and benchmarking use only. Commercial use requires a separate license from the No.7 Project.
Downloads last month
21