Skip to main content
Sound Security Guard monitors audio for a single active scenario at a time using the pretrained Audio Spectrogram Transformer (MIT/ast-finetuned-audioset-10-10-0.4593). The model outputs 527 AudioSet class probabilities (multi-label sigmoid). No fine-tuning is required—scenarios map to existing AudioSet label strings.

Strict scenario isolation

Only labels belonging to the selected scenario are evaluated. For example, with Glass Break active, a passing police siren is ignored even if the model detects it strongly.

Scenarios (sub-modalities)

Parameters

Input / output

Input: Same as VA—audio adapted to PCM S16LE int16 @ 16 kHz mono. Interim output (window filling / no alert):
Alert output:
Wire event_detected to a send_alert node or conditional gate (see Alerts).

Analysis window

Audio accumulates until the analysis window is full (default 4 s), then AST runs inference. The buffer advances by 50% hop (half-window overlap) so events near window boundaries are not missed. After an alert, a cooldown suppresses duplicate notifications.

Example workflow

Edge install

First inference downloads the AST weights from Hugging Face.