Voice Activity Detection
ONNX
vadonnx
vad

ten-vad (ONNX, repackaged for vadonnx)

ONNX voice activity detection model packaged for use with vadonnx.

  • Upstream source: TEN-framework/ten-vad
  • License: Apache-2.0 with additional conditions set by Agora (the upstream LICENSE, which GitHub reads as NOASSERTION). Not a permissive license: see below.
  • Sample rate: 16000 Hz
  • Frame size: 256 samples
  • Stateful: True

This repository redistributes the model in ONNX form together with a signature.json describing its input/output wiring. All rights and the original license belong to the upstream authors.

License terms

The upstream LICENSE is Apache-2.0 plus these conditions, quoted in substance:

  1. You may not deploy ten-vad in a way that competes with Agora's offerings, or that allows others to compete with them, including enabling any third party to develop or deploy applications.
  2. You may deploy ten-vad solely to create and enable deployment of your own applications, for your benefit and that of your direct end users. You may note "Powered by ten-vad" in your documentation.
  3. Derivative works remain subject to the same license.

The full text is at the license_link above. Read it before you build a product on this model. This mirror exists so that vadonnx can fetch the graph for your own application. It is not an invitation to redistribute the graph as a component.

Usage

from vadonnx import load_vad
vad = load_vad("ten")
segments = vad.get_speech_segments(audio, sample_rate=16000)

Note: feature extraction uses TEN's native library; the pure-ONNX path in vadonnx is experimental.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support