Repository navigation
Audio Feature Extraction
Raw audio is a time-domain signal. Machine learning algorithms generally perform better when the signal is represented through informative numerical features.
VoxShield uses Librosa to transform audio into a fixed-length feature vector.
The proposed expanded feature set contains 191 features.
The final feature count should be confirmed against the implementation before documenting it as the deployed configuration.
Mel-Frequency Cepstral Coefficients represent the spectral characteristics of audio using a representation related to human auditory perception.
They are widely used in speech and audio processing.
How they work
- Audio is divided into short frames.
- A frequency spectrum is calculated.
- Frequencies are mapped onto the Mel scale.
- Logarithmic filter-bank energies are transformed using a cepstral operation.
- The resulting coefficients describe spectral characteristics.
Implementation
mfcc = librosa.feature.mfcc(
y=y,
sr=sr,
n_mfcc=40
)The exact parameters should match your final implementation.
The model aggregates MFCC values using their mean and standard deviation.
Proposed feature count:
40 MFCC coefficients × 2 statistics = 80 featuresMFCCs provide a compact representation of spectral characteristics that can be useful for distinguishing different types of speech audio.
Important: MFCCs alone do not reliably establish whether a voice is genuine or cloned. They are input features for the classifier, not proof of authenticity.
Reference: Librosa feature extraction documentation
A Mel Spectrogram represents how energy is distributed across Mel-scaled frequency bands over time.
It provides a time-frequency representation of the audio signal.
Audio Signal
↓
Short-Time Fourier Transform
↓
Mel Filter Bank
↓
Mel SpectrogramThe model calculates statistical summaries of the Mel Spectrogram.
Proposed feature count:
40 Mel bands × 2 statistics = 80 featuresSynthetic speech can contain differences in spectral patterns, but their usefulness depends on the dataset, recording conditions, and the model.
Mel Spectrograms complement MFCCs by retaining a different representation of spectral energy.
Reference: Librosa feature extraction documentation
Spectral features describe the distribution of energy across frequencies.
VoxShield includes:
- Spectral Centroid
- Spectral Bandwidth
- Spectral Rolloff
- Spectral Contrast
Measures the weighted center of the frequency spectrum.
It is often interpreted as a measure related to the brightness of a sound.
Higher centroid → More energy concentrated toward higher frequencies
Lower centroid → More energy concentrated toward lower frequencies
This is a descriptive interpretation, not a direct indicator of cloned speech.
Measures the spread of the spectrum around its centroid.
It helps describe how widely distributed the frequency energy is.
Represents the frequency below which a specified proportion of spectral energy is concentrated.
It can provide information about the distribution of high-frequency energy.
Describes differences between spectral peaks and valleys across frequency bands.
This may capture structural characteristics of the audio spectrum.
Proposed Feature Count
| Feature | Statistics | Count |
|---|---|---|
| Spectral Centroid | Mean + Standard Deviation | 2 |
| Spectral Bandwidth | Mean + Standard Deviation | 2 |
| Spectral Rolloff | Mean + Standard Deviation | 2 |
| Spectral Contrast | Mean | 7 |
| Total | 13 |
Pitch describes the perceived fundamental frequency of a sound.
It is closely associated with the fundamental frequency of voiced speech.
VoxShield Usage
The proposed implementation extracts pitch-related statistics:
Pitch mean
Pitch standard deviation
Speech generated or modified by different systems may exhibit differences in pitch behavior.
However, pitch varies naturally between speakers, genders, languages, emotions, and recording environments.
Therefore, pitch should be treated as a supporting feature rather than an independent detection mechanism.
Limitation
Pitch estimation can be unreliable in noisy audio or unvoiced regions. The implementation should handle invalid or missing pitch estimates consistently.
Zero Crossing Rate measures how frequently an audio waveform crosses the zero-amplitude axis.
It provides a time-domain description of signal variation.
VoxShield Usage
ZCR mean
ZCR standard deviation
ZCR may help characterize differences in signal properties, including aspects of noisiness and spectral behavior.
It is not specific to synthetic speech and should not be interpreted independently.
Chroma features represent the distribution of energy across the twelve pitch classes.
They are more commonly associated with musical and harmonic analysis but can also describe certain characteristics of audio signals.
VoxShield Usage The model calculates the mean of the 12 chroma dimensions.
12 chroma features × 1 statistic = 12 features
Chroma is not inherently a voice-cloning detector. Its contribution to speech classification should be validated experimentally.
Root Mean Square (RMS) energy describes the magnitude of an audio signal over time.
VoxShield Usage
RMS mean
RMS standard deviation
It can provide information about amplitude variation and recording characteristics.
However, microphone distance, background noise, gain, and volume normalization can significantly affect RMS values.