and their implications on human behavior.
-
Pitch: This is perceived as the 'highness' or 'lowness' of a voice, and is related to the frequency of the vocal folds' vibrations. It's often used to convey emotional states. For instance, people tend to speak in higher pitches when they're excited and in lower pitches when they're sad or bored. Pitch can also be a marker of sociolinguistic variables like age, gender, and cultural background.
-
Intensity: This is perceived as the 'loudness' or 'softness' of a voice, and is related to the amplitude of the sound wave. It can also convey emotional states, as well as emphasize certain parts of a message. For instance, someone might speak more loudly when they're angry or excited, or when they're trying to stress a particular point.
-
Formants: These are the resonant frequencies of the vocal tract, and they play a crucial role in distinguishing different vowels and consonants. By studying someone's formant frequencies, you might be able to glean information about their physiological characteristics (like the length and shape of their vocal tract) and their linguistic background.
[Control methods used in a study of the vowels" - Peterson and Barney (1952)]
-
Jitter and Shimmer: These are measures of the variation in pitch and amplitude, respectively, over time. Higher levels of jitter and shimmer are often associated with vocal pathologies, like hoarseness or dysphonia, but can also be a sign of nervousness or stress.
-
Speech rate and pauses: The speed at which someone talks, and the frequency and duration of their pauses, can provide information about their cognitive processes (like how quickly they're thinking or how certain they are of what they're saying), as well as their emotional state and cultural background.
Goldman-Eisler, F. (1968). Psycholinguistics: Experiments in spontaneous speech. Academic Press.
Speech rate and personality" by A. R. Bradlow et al. (1999)
-
Prosody: This encompasses all the 'melodic' features of speech, like pitch, rhythm, and intonation. Prosody is often used to convey emotional states, to differentiate between statements and questions, and to provide structure to speech.
Acoustic correlates of emotional change in speech by Scherer (2003)
["Analysis of prosodic variation in speech for clinical psychology" by M. P. Aylett, B. R. Cowan, and S. Watkins (2013)]
- Suggests that variations in prosody, which are the rhythmic and intonational aspects of speech, could be useful in identifying a range of psychological conditions, potentially including those associated with emotional detachment or dissociation.
Analysis of prosodic variation in speech for clinical depression
-
Harmonics-to-noise ratio: Can provide information about vocal quality and health.
Ferrand CT. Harmonics-to-noise ratio: an index of vocal aging. J Voice. 2002 Dec
-
Spectral tilt: Can provide information about vocal effort.
-
Voice Quality:
Listener responses to different speaker attributes in naturalness and segmental intelligibility"; E. Hanson (1997)
-
Lombard Effect:
. Various spectral parameters: Can help with identifying individual speakers or speech sounds.
Raise in pitch, increased amplitude, increased rate of speech: excitement.
-
Spectrogram:
-
Mel Spectrogram: This is a type of spectrogram that emphasizes the mel scale, a perceptual scale of pitches judged by listeners to be equal in distance from one another. It's often used in speech and speaker recognition tasks. There's evidence that it can capture certain vocal characteristics related to emotional states and mental health conditions, such as depression or stress.
-
Chromagram:
Harmonic/Percussive Separation Using Median Filtering" by Fitzgerald (2010).
-
Mel Frequency Cepstral Coefficients (MFCC)-gram: These are a type of spectral representation that mimics the human auditory system's response. They are often used in speech and speaker recognition tasks. Some research suggests that they can be used to detect certain types of speech pathologies and mental health conditions. For instance, speech under stress or associated with depression may show different MFCC patterns.
Speech Emotion Recognition using Mel Frequency Cepstral Coefficient and SVM Classifier
-
Glottal Flow Features: These features relate to the airflow from the lungs through the vocal folds. Abnormal glottal flow can be an indicator of vocal fold pathologies. There's also some evidence that these features might change under different emotional states, though more research is needed in this area.
-
Fundamental Frequency (F0): Often perceived as pitch, F0 is the rate of vibration of the vocal folds. It is a key parameter in conveying prosodic information (intonation, stress, etc.). Variations in F0 are linked to emotional state, and changes in F0 might reflect psychological stress or health issues, like neurological disorders affecting vocal fold control.
[Vocal Frequency and Personality: A Cross-cultural Study" by J. Laver (1968]
-
Gammatone Frequency Cepstral Coefficients (GFCCs)
- The effectiveness of MFCC and GFCC representations are compared and evaluated over emotion and intensity classification tasks with fully connected and recurrent neural network architectures. The results provide evidence that GFCCs outperform MFCCs in speech emotion recognition.
-
Linear Prediction Cepstrum Coefficients (LPC)
- Information obtained from linear prediction analysis has increased the ability of discriminating different languages. It also shows that language identification performance may be increased by encompassing temporal information by including delta and acceleration features. Besides, the performance of our test system has proved the feasibility of the modeling language by a single Gaussian Mixture Model instead of using complex system such as phonetic recogniser followed by language modelling or large vocabulary continuous speech recognition system.
-
Line Spectral Frequency (LSF)
-
Discrete Wavelet Transform (DWT): Permits high-frequency events identification with an enhanced temporal resolution [45, 46, 47]. A wavelet is a waveform of effectively limited duration that has an average value of zero. Many wavelets also display orthogonality, an ideal feature of compact signal representation [46]. WT is a signal processing technique that can be used to represent real-life non-stationary signals with high efficiency [33, 46]. It has the ability to mine information from the transient signals concurrently in both time and frequency domains [33, 45, 48]. Daubechies and others have developed an orthogonal DWT specially designed for analyzing a finite set of observations over the set of scales (dyadic discretization) [47].
-
Perceptual Linear Prediction (PLP): Combines the critical bands, intensity-to-loudness compression and equal loudness pre-emphasis in the extraction of relevant information from speech. It is rooted in the nonlinear bark scale and was initially intended for use in speech recognition tasks by eliminating the speaker dependent features [11]. PLP gives a representation conforming to a smoothed short-term spectrum that has been equalized and compressed similar to the human hearing making it similar to the MFCC. In the PLP approach, several prominent features of hearing are replicated and the consequent auditory like spectrum of speech is approximated by an autoregressive all–pole model [52]. PLP gives minimized resolution at high frequencies that signifies auditory filter bank based approach, yet gives the orthogonal outputs that are similar to the cepstral analysis. It uses linear predictions for spectral smoothing, hence, the name is perceptual linear prediction [28]. PLP is a combination of both spectral analysis and linear prediction analysis.
Voice disorders can manifest in various ways, including hoarseness, breathiness, strain, pitch problems, or complete voice loss. These disorders can stem from several causes, such as vocal cord damage, neurological conditions, or lifestyle factors like smoking or alcohol use.
Drug use, including both legal substances like alcohol and tobacco and illegal substances like cocaine or opioids, can significantly impact vocal health in several ways:
-
Alcohol: Chronic alcohol consumption can cause dehydration, which may affect vocal fold lubrication, leading to vocal fatigue or strain. Moreover, alcohol can lead to acid reflux or gastroesophageal reflux disease (GERD), where stomach acid flows back into the esophagus. This reflux can irritate and potentially damage the vocal cords, leading to hoarseness or other voice problems.
-
Tobacco: Smoking is one of the leading causes of laryngeal cancer, which can severely impact vocal function. Even in the absence of cancer, smoking can cause inflammation and irritation of the vocal cords (also known as smoker's laryngitis), leading to a raspy, hoarse voice. Additionally, smoking can impair the function of the cilia, tiny hair-like structures in the lungs and respiratory tract that help filter out harmful substances, making the voice box more vulnerable to infections and disease.
-
Cocaine: Cocaine use can cause a range of problems related to the voice, particularly when it's snorted. These may include hoarseness, a decreased sense of smell, and a chronically runny or stuffy nose. Long-term use can even lead to perforation of the nasal septum.
-
Opioids: These can potentially lead to hormonal changes that affect the voice. Chronic opioid use can cause a decrease in several hormones, including those that are crucial for vocal fold lubrication and voice pitch. Furthermore, opioids can lead to a slowed respiratory rate, which can potentially impact voice quality.
-
Methamphetamine: Also known as "meth," this stimulant can lead to severe dental issues, causing what's known as "meth mouth." The resulting dental problems can affect speech production, leading to impaired articulation or other voice changes.
-
Inhalants: These are various substances (often household chemicals) that people inhale to get high. Chronic use can lead to damage to the vocal folds and other parts of the respiratory tract, leading to voice changes.
The impact of these substances on the voice can also be compounded by other effects of drug use, like poor nutrition, lack of sleep, and increased risk of infections, all of which can further compromise vocal health.
The DCIEM Map Task Corpus: Spontaneous dialogue under sleep deprivation and drug treatment
Causes and paralinguistic correlates of interpersonal equivocation
Introversion/Extraversion
The acoustics of extraversion" by F. Nolan, V. Oh, and A. Lee (2011
- Found correlations between certain vocal features (such as average pitch, pitch variability, and intensity) and the personality trait of extraversion.
"The sound of extraversion" by C. G. Jones, M. C. Hall, and P. G. Schmid (2020)
- Uses machine learning to identify associations between acoustic features of speech and the personality trait of extraversion.
Extraversion as a Moderator of the Efficacy of Self-Esteem Maintenance Strategies
Acoustic Correlates of Personality: A Meta-Analysis" by S. McAleer, P. Todorov, and V. Belin (2014)
- Found that while there are some acoustic cues related to perceived personality traits, the evidence is somewhat inconsistent, and further research is needed.
"Cues to Deception" by P. DePaulo et al. (2003)
- Comprehensive review on the nonverbal and verbal cues of deceptive behavior, including voice-related aspects. It's noted that vocal stress and specific speech patterns can be indicative of lying, but the effectiveness of these cues can be variable.
Empirical Interpretation of the Relationship Between Speech Acoustic Context and Emotion Recognition
Emotion recognition and confidence ratings predicted by vocal stimulus type and prosodic parameters
Vocal biomarkers of depression based on motor incoordination - E. J. Godoy and J. Rose (2019)
- Explores how the vocal tracts of depressed individuals can be affected, which results in a distinctive acoustic signature.
- Examines the vocal and speech characteristics that may emerge during the recounting of traumatic memories, which is often associated with dissociation.
Profiling Elements of Prosody in Speech-Communication (PEPS-C)
- PEPS-C comprises 12 tasks, addressing receptive and expressive skills in parallel. The tasks are at two levels, examining prosodic function and prosodic form, respectively. PEPS-C defines four main linguistic functions conveyed by prosody, with a receptive and an expressive task for each
[Expressed emotion and voice acoustic measures of parents of children with autism spectrum disorder" by S. Wong, D. Kuhl, and P. Knaus (2019)]
- How vocal characteristics might differ in parents of children with autism, who may experience higher levels of stress and emotional detachment.
The hypothesis of apraxia of speech in children with autism spectrum disorder
Relations of sex and dialect to reduction
Towards a definition and working model of stress and its effects on speech