Signal Typing and Physical Measures

signal-analysis measurement acoustic-analysis jitter-measurement shimmer-measurement signal-types
Last updated: 2026-01-28

Signal Typing and Physical Measures

The acoustic analysis of voice requires not only appropriate measurement techniques but also recognition that different voices exhibit fundamentally different signal characteristics. A severely dysphonic voice may be so aperiodic that traditional perturbation measures fail entirely. This section explores classification schemes for categorizing vocal signals based on their periodicity and examines the physical measurement approaches used to quantify fluctuations and perturbations. Understanding both signal types and measurement methods is essential for selecting appropriate analysis techniques and interpreting results accurately.

Signal Type Classification

Voice signals range from nearly periodic (healthy modal phonation) to completely aperiodic (severe dysphonia or whisper). Classification schemes help clinicians and researchers identify which signals can be reliably analyzed with period-based perturbation measures.

Titze’s Four-Type System

One widely used classification divides voice signals into four types based on the degree and nature of aperiodicity:

Type 1: Nearly Periodic Signals

Characteristics:

  • Predominantly periodic with small cycle-to-cycle variations
  • Clear fundamental frequency identifiable in all cycles
  • Harmonic structure visible in spectrum
  • Jitter and shimmer at or near normal limits (about <1% and <3%); somewhat higher values still count as Type 1 as long as every cycle remains identifiable

Clinical examples:

  • Normal healthy voices
  • Mild dysphonia with small vocal fold lesions
  • Well-compensated mild vocal fold asymmetries

Measurement implications:

  • All perturbation measures (jitter, shimmer, HNR) are reliable
  • Autocorrelation and waveform matching work well
  • Traditional acoustic analysis is appropriate

Type 1 signals represent the ideal case for acoustic analysis. Period detection algorithms perform well, and perturbation measures provide meaningful quantification of voice quality deviations.

Type 2: Signals with Subharmonics or Frequency Jumps

Characteristics:

  • Periodic regions interrupted by subharmonics, biphonation, or frequency jumps
  • Spectrum shows subharmonic components (F0/2, F0/3) alongside the fundamental
  • Alternating glottal cycles with different characteristics
  • May exhibit diplophonia (perception of two simultaneous pitches)

Clinical examples:

  • Moderate dysphonia with significant vocal fold asymmetry
  • Unilateral vocal fold paralysis
  • Large unilateral lesions (cysts, polyps)
  • Spasmodic dysphonia during voice breaks

Measurement implications:

  • Period-to-period measures become unreliable
  • Traditional jitter/shimmer measures may fail or give artifactually high values
  • Requires specialized analysis recognizing subharmonic structure
  • HNR still provides useful information

Figure 11.1: Example of Type 2 signal showing subharmonic component at F0/2, creating alternating long and short periods in the waveform.

The challenge with Type 2 signals is that “period” becomes ambiguous. Should analysis track the shorter oscillation period or the longer period encompassing the subharmonic pattern? Different algorithms make different choices, leading to inconsistent results across analysis systems.

Type 3: Chaotic or Aperiodic Signals

Characteristics:

  • Little or no obvious periodicity
  • Fundamental frequency difficult or impossible to determine
  • Spectrum shows broadband noise with weak or absent harmonic structure
  • No consistent pattern from cycle to cycle

Clinical examples:

  • Severe dysphonia
  • Complete vocal fold paralysis with compensatory false fold phonation
  • Extensive scarring or tissue loss
  • Severe neurological voice disorders

Measurement implications:

  • Traditional perturbation measures fail completely
  • Cannot reliably identify “periods” for jitter/shimmer calculation
  • Spectral measures (spectral slope, cepstral peak prominence) may provide useful information
  • HNR typically very low (<5 dB) but calculation methods vary

Type 3 signals challenge the fundamental assumption of period-based analysis: that individual periods can be identified. Attempts to apply traditional jitter and shimmer measures to Type 3 signals produce meaningless results.

Type 4: Aphonic Signals

Characteristics:

  • No vocal fold vibration
  • Pure noise signal
  • No identifiable fundamental frequency
  • Spectrum lacks harmonic structure entirely

Clinical examples:

  • Whisper
  • Severe vocal fold paralysis without compensation
  • Complete aphonia following laryngectomy
  • Severe abductor spasmodic dysphonia during episodes

Measurement implications:

  • No F0 to measure; perturbation analysis impossible
  • Analysis limited to noise characteristics
  • Spectral analysis shows only turbulence noise characteristics

Type 4 signals represent the endpoint of the periodicity continuum. These signals cannot be analyzed with any F0-based or period-based measures.

Practical Classification Challenges

Real clinical voices often exhibit mixed characteristics:

Time-varying types: A single voice sample may shift between Type 1 and Type 2, or Type 2 and Type 3, during sustained phonation

Task dependence: A voice might be Type 1 at comfortable pitch and loudness but Type 2 or 3 at very high or very low pitch

Boundary ambiguity: The boundary between Type 1 with high perturbation and Type 2 is not always clear

These challenges mean that signal typing requires both algorithmic detection and expert perceptual judgment.

Period Detection Methods

Accurate measurement of jitter and shimmer depends on reliably identifying the boundaries of individual glottal periods. Several algorithms approach this problem differently.

Waveform Matching

Principle: Find the time shift that maximizes the similarity between adjacent waveform segments.

Algorithm:

  1. Select a reference waveform segment (e.g., 2-3 periods)
  2. Find the best match in the following portion of signal
  3. The time shift giving maximum correlation indicates period length
  4. Update reference and repeat

Advantages:

  • Works well for Type 1 signals with consistent waveform shape
  • Relatively robust to noise

Disadvantages:

  • Can fail with Type 2 signals (subharmonics create ambiguous matches)
  • Computationally intensive
  • May “drift” over long analysis windows

Autocorrelation

Principle: Calculate the correlation between the signal and time-shifted versions of itself.

Algorithm:

R(τ) = Σ x(t) × x(t + τ)

The time lag τ at which R(τ) shows its first major peak (excluding τ=0) indicates the period.

Advantages:

  • Computationally efficient via FFT
  • Well-established theoretical foundation
  • Works reasonably well for Type 1 signals

Disadvantages:

  • Sensitive to noise in Type 2 signals
  • May identify subharmonic period instead of fundamental period
  • Assumes stationarity over analysis window

The autocorrelation method forms the basis for many acoustic analysis programs, including Praat’s pitch detection algorithm.

Cepstral Analysis

Principle: The cepstrum (inverse Fourier transform of log power spectrum) shows a peak at the fundamental period.

Algorithm:

  1. Calculate power spectrum: |FFT(x)|²
  2. Take logarithm: log(|FFT(x)|²)
  3. Inverse FFT: cepstrum = IFFT(log|FFT(x)|²)
  4. Peak location in cepstrum indicates period

Advantages:

  • Particularly robust for harmonic signals
  • Less sensitive to formant structure than time-domain methods
  • The cepstral peak prominence (CPP) provides a useful voice quality measure

Disadvantages:

  • Requires longer analysis windows than autocorrelation
  • Resolution limited by window length
  • Can fail with Type 3 signals

Cepstral peak prominence (CPP) has emerged as a particularly valuable measure. It quantifies the prominence of the cepstral peak and correlates well with overall voice quality, remaining more robust than jitter/shimmer for Type 2 signals.

Zero-Crossing Analysis

Principle: Count zero-crossings in the filtered signal to estimate period.

Algorithm:

  1. High-pass filter to remove DC offset
  2. Identify points where signal crosses zero
  3. Period ≈ time between alternate zero-crossings (or between peaks)

Advantages:

  • Computationally simple
  • Works in real-time

Disadvantages:

  • Very sensitive to noise
  • Fails with complex waveforms
  • Not suitable for clinical acoustic analysis

Zero-crossing methods are rarely used in modern voice analysis due to their limitations.

Electroglottography (EGG) Period Detection

Principle: Use the EGG signal, which reflects vocal fold contact area, to identify periods more reliably than the acoustic signal.

Advantages:

  • Less affected by formant structure and radiation characteristics
  • Often clearer period markers than acoustic signal
  • Particularly useful for Type 2 signals

Disadvantages:

  • Requires specialized EGG equipment
  • Less clinically accessible than acoustic recording
  • EGG “period” may differ slightly from acoustic period

EGG can serve as a valuable adjunct to acoustic analysis, particularly for research purposes or challenging clinical cases.

Common Software Tools

Multiple commercial and open-source software packages provide perturbation analysis. Each uses different algorithms and provides different measures.

Praat

Developer: Paul Boersma and David Weenink, University of Amsterdam Platform: Windows, Mac, Linux (free, open-source)

Strengths:

  • Highly flexible scripting capability
  • Excellent documentation
  • Active user community
  • Voice report provides jitter, shimmer, HNR

Perturbation measures:

  • Jitter (local), Jitter (local, absolute), Jitter (rap), Jitter (ppq5)
  • Shimmer (local), Shimmer (local, dB), Shimmer (apq3), Shimmer (apq5), Shimmer (apq11)
  • HNR (harmonics-to-noise ratio)

Algorithms: Autocorrelation-based pitch detection with cross-correlation verification

Considerations: Very sensitive to analysis settings (pitch range, voicing threshold). Results may differ from commercial systems.

Multi-Dimensional Voice Program (MDVP)

Developer: Kay Pentax (now part of Pentax Medical) Platform: Windows (commercial)

Strengths:

  • Industry standard in clinical voice labs
  • Extensive normative database
  • Comprehensive measure set (>30 parameters)
  • Standardized protocols

Perturbation measures:

  • Jitter (Jitt%), Relative Average Perturbation (RAP), Pitch Perturbation Quotient (PPQ)
  • Shimmer (Shim%), Amplitude Perturbation Quotient (APQ)
  • Noise-to-Harmonics Ratio (NHR), Voice Turbulence Index (VTI)
  • Signal-to-Noise Ratio (SNR)

Algorithms: Proprietary waveform-matching algorithm

Considerations: Expensive; normative data specific to MDVP may not apply to other systems.

Dr. Speech (Tiger Electronics)

Developer: Tiger DRS, Inc. Platform: Windows (commercial)

Strengths:

  • User-friendly interface
  • Real-time visual biofeedback
  • Combined acoustic and aerodynamic analysis (with hardware)

Perturbation measures:

  • Jitter, Shimmer, HNR
  • Various smoothing levels

Considerations: Less commonly used in research; limited published normative data.

TF32

Developer: Paul Milenkovic, University of Wisconsin Platform: Windows (free for research/clinical use)

Strengths:

  • Specialized for perturbation analysis
  • Very sophisticated handling of Type 2 signals
  • Detailed signal typing

Perturbation measures:

  • Multiple jitter and shimmer variants
  • Automatic signal type classification
  • Harmonic component analysis

Algorithms: Advanced autocorrelation with subharmonic detection

Considerations: Less user-friendly interface; steeper learning curve.

CSL (Computerized Speech Lab) / KayPENTAX Systems

Developer: Kay Pentax Platform: Windows (commercial)

Strengths:

  • Integrated hardware and software
  • High-quality recording system
  • Industry standard

Perturbation measures:

  • Includes MDVP as module
  • Additional visual analysis tools (spectrography, pitch tracking)

Considerations: Expensive hardware system; primarily for dedicated voice labs.

Reliability and Validity Considerations

The reliability and validity of perturbation measures depend on multiple factors beyond the choice of algorithm.

Test-Retest Reliability

Within-session reliability: Repeated measures from the same recording session typically show high correlation (r > 0.90 for jitter and shimmer in Type 1 signals).

Between-session reliability: Measures from different recording sessions show more variability (r = 0.70-0.90) due to:

  • True day-to-day voice variation
  • Different vocal effort or pitch level
  • Recording condition variations
  • Microphone positioning differences

Clinical implication: For tracking treatment progress, use consistent recording protocols and consider the standard error of measurement when interpreting changes.

Inter-System Agreement

Different analysis systems produce different absolute values for the same voice sample:

Jitter: Agreement is moderate (r = 0.80-0.95) but absolute values may differ by 10-30%

Shimmer: Agreement is lower (r = 0.70-0.90) with larger absolute differences

HNR: Calculation methods vary substantially; values may differ by several dB

Clinical implication: Compare results only to normative data obtained with the same analysis system. Do not mix systems when tracking an individual over time.

Recording Quality Effects

Perturbation measures are highly sensitive to recording conditions:

Signal-to-Noise Ratio

  • SNR < 25 dB: Shimmer becomes artifactually elevated; HNR decreased
  • SNR 25-35 dB: Acceptable for clinical use with caution
  • SNR > 35 dB: Preferred for research; minimal noise contamination

Source of noise: Background environmental noise, microphone self-noise, electronic system noise

Microphone Type and Placement

  • Head-mounted microphones: Provide most consistent mouth-to-microphone distance; preferred for clinical work
  • Hand-held microphones: Convenient but distance variation affects amplitude measures
  • Desk-mounted microphones: Distance variation affects both amplitude and noise floor

Mouth-to-microphone distance: Standard protocols use 5-10 cm for head-mounted, 15-30 cm for desk-mounted microphones

Analog-to-Digital Conversion

  • Sampling rate: Minimum 20 kHz; 44.1 or 48 kHz recommended
  • Bit depth: 16-bit adequate for clinical use; 24-bit preferred for research
  • Anti-aliasing: Essential to prevent high-frequency noise folding into analysis band

Analysis Settings

Analysis results depend critically on user-selected parameters:

Pitch Range Settings

Setting too narrow a pitch range excludes valid F0 candidates; setting too wide a range allows pitch-doubling or pitch-halving errors.

Recommendation: Adjust pitch range based on preliminary pitch tracking visualization.

Voicing Threshold

Higher thresholds exclude weak periods; lower thresholds include noise segments as “periods.”

Impact: Can dramatically affect jitter/shimmer in borderline Type 1-2 signals.

Window Length

  • Too short: Inadequate frequency resolution; unstable estimates
  • Too long: Assumes stationarity; inappropriate for varying voice quality
  • Typical: 0.5-1.0 second for sustained vowel analysis

Task and Phonetic Context

Perturbation values vary with:

Vowel Type

  • High vowels (/i/, /u/): Often show lower perturbation than low vowels (/a/)
  • Possible mechanism: Different glottal configurations for different vowels

Recommendation: Use sustained /a/ as standard vowel for consistency with normative data.

Pitch Level

  • Comfortable pitch: Lowest perturbation
  • Very high or very low pitch: Increased perturbation even in healthy voices

Recommendation: Use comfortable pitch level; specify pitch when reporting results.

Loudness Level

  • Comfortable loudness: Lowest perturbation
  • Very soft phonation: Increased perturbation due to reduced vocal fold contact
  • Very loud phonation: Increased perturbation due to nonlinear tissue effects

Recommendation: Use comfortable loudness; specify intensity when reporting results.

Advanced Measures

Beyond traditional jitter and shimmer, several advanced measures provide additional voice quality information.

Cepstral Peak Prominence (CPP)

CPP quantifies how prominent the cepstral peak is relative to the overall cepstrum:

Advantages:

  • More robust than jitter/shimmer for Type 2 signals
  • Correlates strongly with overall dysphonia severity
  • Less affected by analysis parameter choices

Normal values: CPP > 10-12 dB (values vary with window length and averaging method)

Phonatory Frequency Range

While not a perturbation measure, the phonatory frequency range (lowest to highest sustainable F0) provides important functional information:

  • Normal: >2 octaves for trained singers; 1.5-2 octaves for non-singers
  • Reduced range: Suggests vocal fold stiffness, mass lesions, or neural control problems

Maximum Phonation Time (MPT)

Maximum phonation time measures the longest sustainable vowel on a single breath:

  • Normal: >15-20 seconds for adults
  • Reduced MPT: May indicate respiratory insufficiency, glottal incompetence, or inefficient vocal technique

While MPT is not an acoustic measure of perturbation, it complements acoustic analysis in comprehensive voice assessment.

Electroglottographic Measures

When EGG is available:

Contact Quotient (CQ): Percentage of glottal cycle with vocal fold contact Speed Quotient (SQ): Ratio of opening to closing phase durations

These measures provide information about glottal configuration complementing acoustic perturbation analysis.


Key Takeaways

  • ✅ Voice signals are classified as Type 1 (nearly periodic), Type 2 (subharmonics/biphonation), Type 3 (chaotic), or Type 4 (aphonic)
  • ✅ Traditional perturbation measures (jitter, shimmer) are reliable only for Type 1 signals; Type 2-4 require specialized analysis approaches
  • ✅ Period detection methods include waveform matching, autocorrelation, and cepstral analysis, each with distinct strengths and limitations
  • ✅ Common software tools (Praat, MDVP, TF32) use different algorithms and produce different absolute values for the same voice
  • ✅ Recording quality significantly affects perturbation measures; SNR >35 dB is preferred for research, >25 dB acceptable for clinical use
  • ✅ Analysis results depend on user-selected parameters including pitch range, voicing threshold, and window length
  • ✅ Test-retest reliability is high within sessions but moderate between sessions; always use the same system and protocol for longitudinal tracking
  • ✅ Advanced measures like cepstral peak prominence provide more robust assessment for moderately disordered voices

Further Reading

  1. Titze, I. R. (1995). Workshop on acoustic voice analysis: Summary statement. National Center for Voice and Speech, Iowa City, IA.
  2. Deliyski, D. D., Shaw, H. S., & Evans, M. K. (2005). Adverse effects of environmental noise on acoustic voice quality measurements. Journal of Voice, 19(1), 15-28.
  3. Brockmann, M., Drinnan, M. J., Storck, C., & Carding, P. N. (2011). Reliable jitter and shimmer measurements in voice clinics: The relevance of vowel, gender, vocal intensity, and fundamental frequency effects in a typical clinical task. Journal of Voice, 25(1), 44-53.
  4. Maryn, Y., Corthals, P., Van Cauwenberge, P., Roy, N., & De Bodt, M. (2010). Toward improved ecological validity in the acoustic measurement of overall voice quality: Combining continuous speech and sustained vowels. Journal of Voice, 24(5), 540-555.
  5. Heman-Ackah, Y. D., Michael, D. D., & Goding Jr, G. S. (2002). The relationship between cepstral peak prominence and selected parameters of dysphonia. Journal of Voice, 16(1), 20-27.
  6. Deliyski, D. D., & Shaw, H. S. (2008). Studying vocal fold vibrations with videokymography. In Voice and speech quality after treatment of early glottic carcinoma (pp. 19-39). Karger Publishers.