Some Definitions

terminology jitter shimmer perturbation vibrato harmonics-to-noise-ratio
Last updated: 2026-01-28

Some Definitions

The study of vocal fluctuations and perturbations requires precise terminology to distinguish different types of variation, their time scales, and their measurement methods. This section establishes the foundational definitions necessary for understanding acoustic analysis of voice and interpreting clinical and research reports. While some terms (like “jitter” and “shimmer”) have become standard in voice science, others (like “vibrato” vs. “tremor”) carry both technical and cultural meanings that require careful clarification.

Perturbation vs. Fluctuation

The fundamental distinction in vocal variation separates perturbations from fluctuations based on their temporal characteristics and perceptual salience.

Perturbation

Perturbation refers to cycle-to-cycle irregularity in the fundamental frequency or amplitude of vocal fold vibration. These variations occur on the time scale of individual glottal cycles (typically 2-20 ms for normal speaking fundamental frequencies). Perturbations arise from inherent biological noise in neuromuscular control, minor asymmetries in vocal fold structure, and turbulent airflow. They are generally not perceived as distinct modulations but rather contribute to overall voice quality, particularly sensations of roughness or hoarseness when excessive.

Fluctuation

Fluctuation describes slower, more periodic modulations of fundamental frequency or amplitude, typically occurring at rates between 3 and 15 Hz. These variations span multiple glottal cycles and are often perceptible as distinct oscillations in pitch or loudness. Fluctuations may be intentional (as in artistic vibrato) or unintentional (as in pathological tremor). Unlike perturbations, which represent noise or irregularity, fluctuations often exhibit relatively regular periodicity.

Jitter: Frequency Perturbation

Jitter quantifies the cycle-to-cycle variability in fundamental period (or fundamental frequency). Multiple jitter measures exist, each with different sensitivity characteristics and clinical utility.

Absolute Jitter (Jitter Absolute, Jita)

Absolute jitter represents the average absolute difference between consecutive periods:

Jita = (1 / N-1) × Σ |Ti - Ti+1|

Where:

  • Ti = duration of period i
  • N = number of periods analyzed

Expressed in microseconds (μs) or milliseconds (ms), absolute jitter provides a dimension-specific measure but depends on the fundamental frequency of the voice sample. Typical values for healthy voices range from 20-50 μs.

Jitter Percent (Jitter %, Jitt)

Jitter percent normalizes absolute jitter by the average period, making it independent of fundamental frequency:

Jitt = (Jita / Tavg) × 100%

Where Tavg is the average period. This normalization allows comparison across different pitch levels and between male and female voices. Healthy voices typically exhibit jitter percent values below 1.0%, with values above 1.0% suggesting potential pathology.

Relative Average Perturbation (RAP)

RAP compares each period to the average of itself and its two neighbors, reducing sensitivity to slow pitch changes or vibrato:

RAP = (1 / N-2) × Σ |Ti - [(Ti-1 + Ti + Ti+1) / 3]|

Divided by average period and expressed as a percentage, RAP provides better resistance to slow fluctuations than simple jitter percent. Normal values are typically below 0.68%.

Pitch Perturbation Quotient (PPQ)

PPQ extends the averaging concept by comparing each period to the average of itself and four neighbors (two on each side):

PPQ = (1 / N-4) × Σ |Ti - [(Ti-2 + Ti-1 + Ti + Ti+1 + Ti+2) / 5]|

This five-period smoothing provides even greater resistance to slow fluctuations. Normal values are typically below 0.84%. PPQ is particularly useful when analyzing voices with vibrato, as it better isolates cycle-to-cycle perturbations from the slower vibrato modulation.

Shimmer: Amplitude Perturbation

Shimmer quantifies cycle-to-cycle variability in amplitude, analogous to how jitter quantifies frequency variability.

Shimmer in dB (Shimmer dB, ShdB)

Shimmer in dB expresses the average absolute difference between the amplitudes of consecutive periods on a logarithmic scale:

ShdB = (1 / N-1) × Σ |20 log10(Ai / Ai+1)|

Where Ai represents the peak-to-peak amplitude of period i. The logarithmic scale better represents human perception of intensity changes. Healthy voices typically show shimmer dB values below 0.35 dB.

Shimmer Percent (Shimmer %, Shim)

Shimmer percent expresses amplitude variability on a linear scale:

Shim = (1 / N-1) × Σ |Ai - Ai+1| / Aavg × 100%

Normal values are typically below 3.0%, though this measure is more sensitive to recording conditions than shimmer dB.

Amplitude Perturbation Quotient (APQ)

APQ applies the same smoothing principle to amplitude as PPQ does to frequency:

APQ = (1 / N-10) × Σ |Ai - [(Ai-5 + ... + Ai + ... + Ai+5) / 11]|

This eleven-period smoothing (period i and five neighbors on each side) helps distinguish true amplitude perturbation from slower amplitude fluctuations like tremolo. Normal values are typically below 3.07%.

Harmonics-to-Noise Ratio (HNR)

The Harmonics-to-Noise Ratio quantifies the proportion of periodic (harmonic) energy to aperiodic (noise) energy in the voice signal. Unlike jitter and shimmer, which analyze individual periods, HNR evaluates the overall spectral quality of the voice.

Definition and Calculation

HNR is typically calculated using autocorrelation methods:

HNR (dB) = 10 log10(ACmax / (1 - ACmax))

Where ACmax is the maximum value of the autocorrelation function (excluding the value at zero lag). Higher HNR values indicate more periodic signals with less noise. Healthy modal phonation typically yields HNR values above 10-13 dB, while breathy or rough voices show reduced HNR.

Noise-to-Harmonics Ratio (NHR)

NHR represents the inverse concept:

NHR = 1 / (10^(HNR/10))

Some analysis programs report NHR rather than HNR. Healthy voices show NHR values below 0.19, with higher values indicating increased noise energy relative to harmonic content.

Vibrato and Tremor

Distinguishing vibrato from tremor requires considering both acoustic characteristics and functional context.

Vibrato

Vibrato describes a quasi-periodic modulation of fundamental frequency and amplitude occurring at rates typically between 5-7 Hz with extents of 50-100 cents (half to whole semitone). Vibrato is generally considered:

  • Intentional and controlled: Singers can initiate, sustain, and terminate vibrato voluntarily
  • Regular in rate and extent: Shows relatively consistent periodicity and depth
  • Aesthetically valued: Considered a desirable quality in most Western classical singing traditions
  • Combined modulation: Involves both F0 and amplitude changes in coordinated fashion

Tremor

Tremor refers to involuntary rhythmic oscillation, typically:

  • Rate: typically 3-8 Hz depending on the disorder (essential voice tremor most often 4-7 Hz, Parkinsonian tremor somewhat slower), so it overlaps the vibrato range
  • Irregular: Less consistent in rate and extent than vibrato
  • Involuntary: Cannot be reliably controlled or eliminated by the speaker/singer
  • Variable origin: May originate from neurological conditions (essential tremor, Parkinson’s disease), respiratory instability, or laryngeal muscle dysfunction

The boundary between wide vibrato and tremor can be ambiguous, particularly in aging singers whose voluntary vibrato may become less regular and partially involuntary.

Coefficient of Variation

The coefficient of variation (CV) provides a normalized measure of variability for any acoustic parameter:

CV = (standard deviation / mean) × 100%

Applied to fundamental frequency or amplitude, CV quantifies the overall variability relative to the mean value. This measure proves particularly useful for characterizing fluctuations (like vibrato or tremor) rather than cycle-to-cycle perturbations. A vibrato with 6% frequency CV, for example, has a standard deviation of 6% of the mean frequency.

Directional Perturbation Factor (DPF)

Some perturbation measures distinguish between consecutive periods that increase vs. decrease in duration or amplitude. The Directional Perturbation Factor identifies periods where the direction of change (longer/shorter, louder/softer) differs from the average trend. Higher DPF values suggest more irregular, chaotic variability rather than smooth trends.

Signal-to-Noise Ratio (SNR)

While not specific to voice, Signal-to-Noise Ratio measures the acoustic signal level relative to background noise in the recording:

SNR (dB) = 20 log10(Asignal / Anoise)

Adequate SNR (typically >30 dB for clinical voice analysis) is essential for reliable perturbation measurement, as background noise can artificially inflate shimmer and reduce HNR.


Key Takeaways

  • ✅ Perturbations are cycle-to-cycle irregularities while fluctuations are slower, periodic modulations
  • ✅ Jitter quantifies frequency perturbation with measures including absolute jitter, jitter percent, RAP, and PPQ
  • ✅ Shimmer quantifies amplitude perturbation using shimmer dB, shimmer percent, and APQ
  • ✅ Smoothed measures (RAP, PPQ, APQ) better isolate perturbations from slow fluctuations like vibrato
  • ✅ HNR measures the ratio of periodic to aperiodic energy, with higher values indicating better voice quality
  • ✅ Vibrato is intentional, regular modulation (5-7 Hz) while tremor is involuntary and often irregular (about 3-8 Hz, overlapping the vibrato range)
  • ✅ Normative values provide clinical reference points: jitter <1%, shimmer <3%, HNR >10 dB for healthy voices

Further Reading

  1. Titze, I. R. (1995). Workshop on acoustic voice analysis: Summary statement. National Center for Voice and Speech. Iowa City, IA.
  2. Baken, R. J., & Orlikoff, R. F. (2000). Clinical measurement of speech and voice (2nd ed.). San Diego: Singular Publishing Group. [Chapters 8-9]
  3. Karnell, M. P. (1991). Laryngeal perturbation analysis: Minimum length of analysis window. Journal of Speech and Hearing Research, 34, 544-548.
  4. Pinto, N. B., & Titze, I. R. (1990). Unification of perturbation measures in speech signals. Journal of the Acoustical Society of America, 87, 1278-1289.
  5. Kreiman, J., & Gerratt, B. R. (2005). Perception of aperiodicity in pathological voice. Journal of the Acoustical Society of America, 117, 2201-2211.
  6. Dejonckere, P. H., Bradley, P., Clemente, P., et al. (2001). A basic protocol for functional assessment of voice pathology. European Archives of Oto-Rhino-Laryngology, 258, 77-82.