Instructions to use cstr/ecapa-lid-107-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- speechbrain
How to use cstr/ecapa-lid-107-GGUF with speechbrain:
# interface not specified in config.json
- Notebooks
- Google Colab
- Kaggle
ECAPA-TDNN Language Identification (GGUF)
GGUF conversion of speechbrain/lang-id-voxlingua107-ecapa for use with CrispASR.
Model Details
- Architecture: ECAPA-TDNN (SE-Res2Net + Attentive Statistical Pooling)
- Parameters: 21M
- Size: 43 MB (F16)
- Languages: 107 (VoxLingua107 dataset)
- License: Apache 2.0
- Training: SpeechBrain on VoxLingua107 (6,628 hours YouTube speech)
Usage with CrispASR
# As LID pre-step for any ASR backend
crispasr -m whisper-large-v3.gguf --lid-backend ecapa -l auto -f audio.wav
# Model auto-downloads on first use, or specify path:
crispasr -m model.gguf --lid-backend ecapa --lid-model ecapa-lid-107-f16.gguf -l auto -f audio.wav
Accuracy
Tested on 12-language edge-TTS benchmark (3 samples per language):
| Language | Accuracy | Confidence |
|---|---|---|
| English | 3/3 | pβ₯0.99 |
| German | 3/3 | pβ₯0.99 |
| French | 3/3 | pβ₯0.99 |
| Spanish | 3/3 | pβ₯0.96 |
| Japanese | 3/3 | pβ₯0.99 |
| Chinese | 3/3 | pβ₯0.99 |
| Korean | 3/3 | pβ₯0.99 |
| Russian | 3/3 | pβ₯0.99 |
| Arabic | 3/3 | pβ₯0.99 |
| Hindi | 3/3 | pβ₯0.99 |
| Portuguese | 3/3 | pβ₯0.99 |
| Italian | 3/3 | pβ₯0.99 |
Files
| File | Size | Description |
|---|---|---|
ecapa-lid-107-f16.gguf |
43 MB | F16 weights (recommended) |
Conversion
python models/convert-ecapa-tdnn-lid-to-gguf.py \
--input speechbrain/lang-id-voxlingua107-ecapa \
--output ecapa-lid-107-f16.gguf
Architecture
Input: 16kHz PCM β 60-dim mel fbank (SpeechBrain STFT, n_fft=400)
β Sentence-level mean normalization
β Block0: Conv1d(60β1024, k=5) + ReLU + BN
β Block1-3: SE-Res2Net (1024 channels, 8 sub-bands, dilations 2/3/4)
β MFA: concatenate block1-3 outputs β Conv1d(3072β3072, k=1) + ReLU + BN
β ASP: Attentive Statistical Pooling β [6144]
β BN + FC(6144β256) β embedding
β Classifier: BN β Linear(256β512) + BN + LeakyReLU β Linear(512β107)
Citation
@inproceedings{ravanelli2021speechbrain,
title={SpeechBrain: A General-Purpose Speech Toolkit},
author={Ravanelli, Mirco and others},
booktitle={Proceedings of the 22nd Annual Conference of the International Speech Communication Association (INTERSPEECH)},
year={2021}
}
Provenance and EU AI Act Art. 53 note
- Upstream model: speechbrain/lang-id-voxlingua107-ecapa β published by
speechbrain. - Upstream licence:
apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented β where it is documented at all β by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
- Downloads last month
- 780
Hardware compatibility
Log In to add your hardware
16-bit
Model tree for cstr/ecapa-lid-107-GGUF
Base model
speechbrain/lang-id-voxlingua107-ecapa