Skip to main content
Voice Assistant uses Silero VAD to detect speech start/end, extract utterances, and optionally re-chunk output for downstream nodes.

Sub-modality profiles

Input / output

Input: audio from Audio Track (any supported encoding; adapted to int16 @ 16 kHz mono). Outputs (every chunk while listening):
Outputs (when speech ends):

Output buffer preset (optional)

After a full utterance is detected, audio can be re-chunked for downstream consumers: Set the upstream Audio Track buffer preset to Voice Assistant (32 ms) so chunks are 512 samples—the native Silero frame size. Larger chunks still work (the node reframes internally) but add latency.