Sub-modality profiles
Input / output
Input:audio from Audio Track (any supported encoding; adapted to int16 @ 16 kHz mono).
Outputs (every chunk while listening):
Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
Silero VAD profiles for real-time and batch speech segmentation.
| Profile | Use case |
|---|---|
| Real-Time Voice Assistant | Low-latency conversational / command endpoints |
| High-Noise / Industrial | Factory floors, vehicles, call centres |
| Batch Transcription / STT | Offline Whisper-style segmentation (~28 s max chunk) |
| Quiet Studio / Whisperer | Distant mic, soft speech, long pauses |
| Custom | Manual Silero parameters |
audio from Audio Track (any supported encoding; adapted to int16 @ 16 kHz mono).
Outputs (every chunk while listening):
{
"speech_probability": 0.87,
"is_speaking": true,
"sample_rate_hz": 16000,
"channels": 1
}
{
"audio": "<numpy int16>",
"speech_probability": 0.92,
"is_speaking": false,
"start_timestamp_sec": 1.2,
"end_timestamp_sec": 3.8,
"sample_rate_hz": 16000,
"channels": 1
}
| Preset | Duration | Samples @ 16 kHz |
|---|---|---|
| None | Full segment | — |
| Wake Word Engine | 80 ms | 1280 |
| Speech-To-Text | 4 s | 64000 |
| Custom | User-defined | — |