Extracting Speech Segments with Silero VAD and ONNX Runtime

Chronological Source Flow
Back

AI Fusion Summary

Two labs utilize ONNX Runtime to process a 14-second conversation decoded via FFmpeg into 16 kHz mono waveforms. The first lab employs Silero VAD to detect speech and extract segments as WAV files using a CPU execution provider. The second lab uses Pyannote Segmentation 3.0 to identify speaker changes and split the recording into individual utterances. Both tests verify the ability to isolate speech and separate alternating speakers into contiguous segments for downstream processing.
Community Comments
Loading updates...
0