Signal Typing and Physical Measures
The acoustic analysis of voice requires not only appropriate measurement techniques but also recognition that different voices exhibit fundamentally different signal characteristics. A severely dysphonic voice may be so aperiodic that traditional perturbation measures fail entirely. This section explores classification schemes for categorizing vocal signals based on their periodicity and examines the physical measurement approaches used to quantify fluctuations and perturbations. Understanding both signal types and measurement methods is essential for selecting appropriate analysis techniques and interpreting results accurately.
Signal Type Classification
Voice signals range from nearly periodic (healthy modal phonation) to completely aperiodic (severe dysphonia or whisper). Classification schemes help clinicians and researchers identify which signals can be reliably analyzed with period-based perturbation measures.
Titze’s Four-Type System
One widely used classification divides voice signals into four types based on the degree and nature of aperiodicity:
Type 1: Nearly Periodic Signals
Characteristics:
- Predominantly periodic with small cycle-to-cycle variations
- Clear fundamental frequency identifiable in all cycles
- Harmonic structure visible in spectrum
- Jitter and shimmer at or near normal limits (about <1% and <3%); somewhat higher values still count as Type 1 as long as every cycle remains identifiable
Clinical examples:
- Normal healthy voices
- Mild dysphonia with small vocal fold lesions
- Well-compensated mild vocal fold asymmetries
Measurement implications:
- All perturbation measures (jitter, shimmer, HNR) are reliable
- Autocorrelation and waveform matching work well
- Traditional acoustic analysis is appropriate
Type 1 signals represent the ideal case for acoustic analysis. Period detection algorithms perform well, and perturbation measures provide meaningful quantification of voice quality deviations.
Type 2: Signals with Subharmonics or Frequency Jumps
Characteristics:
- Periodic regions interrupted by subharmonics, biphonation, or frequency jumps
- Spectrum shows subharmonic components (F0/2, F0/3) alongside the fundamental
- Alternating glottal cycles with different characteristics
- May exhibit diplophonia (perception of two simultaneous pitches)
Clinical examples:
- Moderate dysphonia with significant vocal fold asymmetry
- Unilateral vocal fold paralysis
- Large unilateral lesions (cysts, polyps)
- Spasmodic dysphonia during voice breaks
Measurement implications:
- Period-to-period measures become unreliable
- Traditional jitter/shimmer measures may fail or give artifactually high values
- Requires specialized analysis recognizing subharmonic structure
- HNR still provides useful information
Figure 11.1: Example of Type 2 signal showing subharmonic component at F0/2, creating alternating long and short periods in the waveform.
The challenge with Type 2 signals is that “period” becomes ambiguous. Should analysis track the shorter oscillation period or the longer period encompassing the subharmonic pattern? Different algorithms make different choices, leading to inconsistent results across analysis systems.
Type 3: Chaotic or Aperiodic Signals
Characteristics:
- Little or no obvious periodicity
- Fundamental frequency difficult or impossible to determine
- Spectrum shows broadband noise with weak or absent harmonic structure
- No consistent pattern from cycle to cycle
Clinical examples:
- Severe dysphonia
- Complete vocal fold paralysis with compensatory false fold phonation
- Extensive scarring or tissue loss
- Severe neurological voice disorders
Measurement implications:
- Traditional perturbation measures fail completely
- Cannot reliably identify “periods” for jitter/shimmer calculation
- Spectral measures (spectral slope, cepstral peak prominence) may provide useful information
- HNR typically very low (<5 dB) but calculation methods vary
Type 3 signals challenge the fundamental assumption of period-based analysis: that individual periods can be identified. Attempts to apply traditional jitter and shimmer measures to Type 3 signals produce meaningless results.
Type 4: Aphonic Signals
Characteristics:
- No vocal fold vibration
- Pure noise signal
- No identifiable fundamental frequency
- Spectrum lacks harmonic structure entirely
Clinical examples:
- Whisper
- Severe vocal fold paralysis without compensation
- Complete aphonia following laryngectomy
- Severe abductor spasmodic dysphonia during episodes
Measurement implications:
- No F0 to measure; perturbation analysis impossible
- Analysis limited to noise characteristics
- Spectral analysis shows only turbulence noise characteristics
Type 4 signals represent the endpoint of the periodicity continuum. These signals cannot be analyzed with any F0-based or period-based measures.
Practical Classification Challenges
Real clinical voices often exhibit mixed characteristics:
Time-varying types: A single voice sample may shift between Type 1 and Type 2, or Type 2 and Type 3, during sustained phonation
Task dependence: A voice might be Type 1 at comfortable pitch and loudness but Type 2 or 3 at very high or very low pitch
Boundary ambiguity: The boundary between Type 1 with high perturbation and Type 2 is not always clear
These challenges mean that signal typing requires both algorithmic detection and expert perceptual judgment.
Period Detection Methods
Accurate measurement of jitter and shimmer depends on reliably identifying the boundaries of individual glottal periods. Several algorithms approach this problem differently.
Waveform Matching
Principle: Find the time shift that maximizes the similarity between adjacent waveform segments.
Algorithm:
- Select a reference waveform segment (e.g., 2-3 periods)
- Find the best match in the following portion of signal
- The time shift giving maximum correlation indicates period length
- Update reference and repeat
Advantages:
- Works well for Type 1 signals with consistent waveform shape
- Relatively robust to noise
Disadvantages:
- Can fail with Type 2 signals (subharmonics create ambiguous matches)
- Computationally intensive
- May “drift” over long analysis windows
Autocorrelation
Principle: Calculate the correlation between the signal and time-shifted versions of itself.
Algorithm:
R(τ) = Σ x(t) × x(t + τ)
The time lag τ at which R(τ) shows its first major peak (excluding τ=0) indicates the period.
Advantages:
- Computationally efficient via FFT
- Well-established theoretical foundation
- Works reasonably well for Type 1 signals
Disadvantages:
- Sensitive to noise in Type 2 signals
- May identify subharmonic period instead of fundamental period
- Assumes stationarity over analysis window
The autocorrelation method forms the basis for many acoustic analysis programs, including Praat’s pitch detection algorithm.
Cepstral Analysis
Principle: The cepstrum (inverse Fourier transform of log power spectrum) shows a peak at the fundamental period.
Algorithm:
- Calculate power spectrum: |FFT(x)|²
- Take logarithm: log(|FFT(x)|²)
- Inverse FFT: cepstrum = IFFT(log|FFT(x)|²)
- Peak location in cepstrum indicates period
Advantages:
- Particularly robust for harmonic signals
- Less sensitive to formant structure than time-domain methods
- The cepstral peak prominence (CPP) provides a useful voice quality measure
Disadvantages:
- Requires longer analysis windows than autocorrelation
- Resolution limited by window length
- Can fail with Type 3 signals
Cepstral peak prominence (CPP) has emerged as a particularly valuable measure. It quantifies the prominence of the cepstral peak and correlates well with overall voice quality, remaining more robust than jitter/shimmer for Type 2 signals.
Zero-Crossing Analysis
Principle: Count zero-crossings in the filtered signal to estimate period.
Algorithm:
- High-pass filter to remove DC offset
- Identify points where signal crosses zero
- Period ≈ time between alternate zero-crossings (or between peaks)
Advantages:
- Computationally simple
- Works in real-time
Disadvantages:
- Very sensitive to noise
- Fails with complex waveforms
- Not suitable for clinical acoustic analysis
Zero-crossing methods are rarely used in modern voice analysis due to their limitations.
Electroglottography (EGG) Period Detection
Principle: Use the EGG signal, which reflects vocal fold contact area, to identify periods more reliably than the acoustic signal.
Advantages:
- Less affected by formant structure and radiation characteristics
- Often clearer period markers than acoustic signal
- Particularly useful for Type 2 signals
Disadvantages:
- Requires specialized EGG equipment
- Less clinically accessible than acoustic recording
- EGG “period” may differ slightly from acoustic period
EGG can serve as a valuable adjunct to acoustic analysis, particularly for research purposes or challenging clinical cases.
Common Software Tools
Multiple commercial and open-source software packages provide perturbation analysis. Each uses different algorithms and provides different measures.
Praat
Developer: Paul Boersma and David Weenink, University of Amsterdam Platform: Windows, Mac, Linux (free, open-source)
Strengths:
- Highly flexible scripting capability
- Excellent documentation
- Active user community
- Voice report provides jitter, shimmer, HNR
Perturbation measures:
- Jitter (local), Jitter (local, absolute), Jitter (rap), Jitter (ppq5)
- Shimmer (local), Shimmer (local, dB), Shimmer (apq3), Shimmer (apq5), Shimmer (apq11)
- HNR (harmonics-to-noise ratio)
Algorithms: Autocorrelation-based pitch detection with cross-correlation verification
Considerations: Very sensitive to analysis settings (pitch range, voicing threshold). Results may differ from commercial systems.
Multi-Dimensional Voice Program (MDVP)
Developer: Kay Pentax (now part of Pentax Medical) Platform: Windows (commercial)
Strengths:
- Industry standard in clinical voice labs
- Extensive normative database
- Comprehensive measure set (>30 parameters)
- Standardized protocols
Perturbation measures:
- Jitter (Jitt%), Relative Average Perturbation (RAP), Pitch Perturbation Quotient (PPQ)
- Shimmer (Shim%), Amplitude Perturbation Quotient (APQ)
- Noise-to-Harmonics Ratio (NHR), Voice Turbulence Index (VTI)
- Signal-to-Noise Ratio (SNR)
Algorithms: Proprietary waveform-matching algorithm
Considerations: Expensive; normative data specific to MDVP may not apply to other systems.
Dr. Speech (Tiger Electronics)
Developer: Tiger DRS, Inc. Platform: Windows (commercial)
Strengths:
- User-friendly interface
- Real-time visual biofeedback
- Combined acoustic and aerodynamic analysis (with hardware)
Perturbation measures:
- Jitter, Shimmer, HNR
- Various smoothing levels
Considerations: Less commonly used in research; limited published normative data.
TF32
Developer: Paul Milenkovic, University of Wisconsin Platform: Windows (free for research/clinical use)
Strengths:
- Specialized for perturbation analysis
- Very sophisticated handling of Type 2 signals
- Detailed signal typing
Perturbation measures:
- Multiple jitter and shimmer variants
- Automatic signal type classification
- Harmonic component analysis
Algorithms: Advanced autocorrelation with subharmonic detection
Considerations: Less user-friendly interface; steeper learning curve.
CSL (Computerized Speech Lab) / KayPENTAX Systems
Developer: Kay Pentax Platform: Windows (commercial)
Strengths:
- Integrated hardware and software
- High-quality recording system
- Industry standard
Perturbation measures:
- Includes MDVP as module
- Additional visual analysis tools (spectrography, pitch tracking)
Considerations: Expensive hardware system; primarily for dedicated voice labs.
Reliability and Validity Considerations
The reliability and validity of perturbation measures depend on multiple factors beyond the choice of algorithm.
Test-Retest Reliability
Within-session reliability: Repeated measures from the same recording session typically show high correlation (r > 0.90 for jitter and shimmer in Type 1 signals).
Between-session reliability: Measures from different recording sessions show more variability (r = 0.70-0.90) due to:
- True day-to-day voice variation
- Different vocal effort or pitch level
- Recording condition variations
- Microphone positioning differences
Clinical implication: For tracking treatment progress, use consistent recording protocols and consider the standard error of measurement when interpreting changes.
Inter-System Agreement
Different analysis systems produce different absolute values for the same voice sample:
Jitter: Agreement is moderate (r = 0.80-0.95) but absolute values may differ by 10-30%
Shimmer: Agreement is lower (r = 0.70-0.90) with larger absolute differences
HNR: Calculation methods vary substantially; values may differ by several dB
Clinical implication: Compare results only to normative data obtained with the same analysis system. Do not mix systems when tracking an individual over time.
Recording Quality Effects
Perturbation measures are highly sensitive to recording conditions:
Signal-to-Noise Ratio
- SNR < 25 dB: Shimmer becomes artifactually elevated; HNR decreased
- SNR 25-35 dB: Acceptable for clinical use with caution
- SNR > 35 dB: Preferred for research; minimal noise contamination
Source of noise: Background environmental noise, microphone self-noise, electronic system noise
Microphone Type and Placement
- Head-mounted microphones: Provide most consistent mouth-to-microphone distance; preferred for clinical work
- Hand-held microphones: Convenient but distance variation affects amplitude measures
- Desk-mounted microphones: Distance variation affects both amplitude and noise floor
Mouth-to-microphone distance: Standard protocols use 5-10 cm for head-mounted, 15-30 cm for desk-mounted microphones
Analog-to-Digital Conversion
- Sampling rate: Minimum 20 kHz; 44.1 or 48 kHz recommended
- Bit depth: 16-bit adequate for clinical use; 24-bit preferred for research
- Anti-aliasing: Essential to prevent high-frequency noise folding into analysis band
Analysis Settings
Analysis results depend critically on user-selected parameters:
Pitch Range Settings
Setting too narrow a pitch range excludes valid F0 candidates; setting too wide a range allows pitch-doubling or pitch-halving errors.
Recommendation: Adjust pitch range based on preliminary pitch tracking visualization.
Voicing Threshold
Higher thresholds exclude weak periods; lower thresholds include noise segments as “periods.”
Impact: Can dramatically affect jitter/shimmer in borderline Type 1-2 signals.
Window Length
- Too short: Inadequate frequency resolution; unstable estimates
- Too long: Assumes stationarity; inappropriate for varying voice quality
- Typical: 0.5-1.0 second for sustained vowel analysis
Task and Phonetic Context
Perturbation values vary with:
Vowel Type
- High vowels (/i/, /u/): Often show lower perturbation than low vowels (/a/)
- Possible mechanism: Different glottal configurations for different vowels
Recommendation: Use sustained /a/ as standard vowel for consistency with normative data.
Pitch Level
- Comfortable pitch: Lowest perturbation
- Very high or very low pitch: Increased perturbation even in healthy voices
Recommendation: Use comfortable pitch level; specify pitch when reporting results.
Loudness Level
- Comfortable loudness: Lowest perturbation
- Very soft phonation: Increased perturbation due to reduced vocal fold contact
- Very loud phonation: Increased perturbation due to nonlinear tissue effects
Recommendation: Use comfortable loudness; specify intensity when reporting results.
Advanced Measures
Beyond traditional jitter and shimmer, several advanced measures provide additional voice quality information.
Cepstral Peak Prominence (CPP)
CPP quantifies how prominent the cepstral peak is relative to the overall cepstrum:
Advantages:
- More robust than jitter/shimmer for Type 2 signals
- Correlates strongly with overall dysphonia severity
- Less affected by analysis parameter choices
Normal values: CPP > 10-12 dB (values vary with window length and averaging method)
Phonatory Frequency Range
While not a perturbation measure, the phonatory frequency range (lowest to highest sustainable F0) provides important functional information:
- Normal: >2 octaves for trained singers; 1.5-2 octaves for non-singers
- Reduced range: Suggests vocal fold stiffness, mass lesions, or neural control problems
Maximum Phonation Time (MPT)
Maximum phonation time measures the longest sustainable vowel on a single breath:
- Normal: >15-20 seconds for adults
- Reduced MPT: May indicate respiratory insufficiency, glottal incompetence, or inefficient vocal technique
While MPT is not an acoustic measure of perturbation, it complements acoustic analysis in comprehensive voice assessment.
Electroglottographic Measures
When EGG is available:
Contact Quotient (CQ): Percentage of glottal cycle with vocal fold contact Speed Quotient (SQ): Ratio of opening to closing phase durations
These measures provide information about glottal configuration complementing acoustic perturbation analysis.
Key Takeaways
- ✅ Voice signals are classified as Type 1 (nearly periodic), Type 2 (subharmonics/biphonation), Type 3 (chaotic), or Type 4 (aphonic)
- ✅ Traditional perturbation measures (jitter, shimmer) are reliable only for Type 1 signals; Type 2-4 require specialized analysis approaches
- ✅ Period detection methods include waveform matching, autocorrelation, and cepstral analysis, each with distinct strengths and limitations
- ✅ Common software tools (Praat, MDVP, TF32) use different algorithms and produce different absolute values for the same voice
- ✅ Recording quality significantly affects perturbation measures; SNR >35 dB is preferred for research, >25 dB acceptable for clinical use
- ✅ Analysis results depend on user-selected parameters including pitch range, voicing threshold, and window length
- ✅ Test-retest reliability is high within sessions but moderate between sessions; always use the same system and protocol for longitudinal tracking
- ✅ Advanced measures like cepstral peak prominence provide more robust assessment for moderately disordered voices
Related Topics
- Some Definitions
- Sources of Fluctuation and Perturbation
- Clinical and Pedagogical Issues
- Acoustic Theory of Sound Production
- Source-Filter Theory
Further Reading
- Titze, I. R. (1995). Workshop on acoustic voice analysis: Summary statement. National Center for Voice and Speech, Iowa City, IA.
- Deliyski, D. D., Shaw, H. S., & Evans, M. K. (2005). Adverse effects of environmental noise on acoustic voice quality measurements. Journal of Voice, 19(1), 15-28.
- Brockmann, M., Drinnan, M. J., Storck, C., & Carding, P. N. (2011). Reliable jitter and shimmer measurements in voice clinics: The relevance of vowel, gender, vocal intensity, and fundamental frequency effects in a typical clinical task. Journal of Voice, 25(1), 44-53.
- Maryn, Y., Corthals, P., Van Cauwenberge, P., Roy, N., & De Bodt, M. (2010). Toward improved ecological validity in the acoustic measurement of overall voice quality: Combining continuous speech and sustained vowels. Journal of Voice, 24(5), 540-555.
- Heman-Ackah, Y. D., Michael, D. D., & Goding Jr, G. S. (2002). The relationship between cepstral peak prominence and selected parameters of dysphonia. Journal of Voice, 16(1), 20-27.
- Deliyski, D. D., & Shaw, H. S. (2008). Studying vocal fold vibrations with videokymography. In Voice and speech quality after treatment of early glottic carcinoma (pp. 19-39). Karger Publishers.