A deep learning system identifying speakers from short audio clips. Built a 1D residual CNN trained on the full FFT spectrum of raw audio with noise augmentation, reaching 95.3% test accuracy at ~110ms inference latency on CPU. Suitable for real-time voice authentication and security applications. Deployed as a live, interactive demo via Streamlit.
-
Updated
Jul 30, 2026 - Python