Test Utterances

clinical assessment measurement protocols standardization
Last updated: 2025-02-07

Test Utterances

Clinical voice assessment requires eliciting voice samples suitable for perceptual evaluation, acoustic analysis, and documentation of vocal function. The choice of test utterances—specific phonatory tasks requested from patients—profoundly affects the resulting measurements and clinical interpretations. Standardized protocols enable comparison across patients, time points, and clinical sites while balancing the need for controlled conditions against ecological validity. Understanding the rationale, implementation, and limitations of various test utterances informs appropriate clinical decision-making and meaningful interpretation of assessment results.

Sustained Vowels

Sustained vowels represent the most common test utterance for acoustic voice analysis, involving phonation of a single vowel at comfortable pitch and loudness for several seconds.

Rationale and Advantages

Sustained vowels provide:

Phonatory isolation: Eliminates articulatory movements, isolating laryngeal function Stationarity: Promotes relatively stable F₀ and amplitude facilitating perturbation analysis Standardization: Simple instructions with minimal room for interpretation variation Cross-linguistic applicability: Basic vowels exist across languages Reduced cognitive load: Minimal memory or linguistic processing demands Extended analysis windows: Several seconds of continuous phonation enable reliable statistics

These advantages make sustained vowels the standard for perturbation analysis, enabling jitter, shimmer, and harmonics-to-noise ratio measurement.

Vowel Selection

Standard protocols typically employ:

/a/ (as in “father”): Most common choice due to:

  • Open vocal tract minimizing formant-source interaction
  • Comfortable articulation position
  • Relatively stable acoustic characteristics
  • Cross-cultural familiarity

/i/ (as in “see”): Sometimes included to assess:

  • Higher formants affecting harmonics
  • Increased vocal tract impedance effects
  • Vowel-dependent stability differences

/u/ (as in “moon”): Occasionally used for:

  • Lip rounding effects
  • Different formant configuration
  • Comparison with other vowels

Research shows vowel choice affects measured perturbation values by 10-30%, with /a/ typically producing lowest jitter and shimmer. Standardizing on one vowel (usually /a/) facilitates comparison while recognizing that complete voice assessment may include multiple vowels.

Duration

Typical protocols request:

  • Minimum: 3 seconds (providing 200-400 cycles at normal pitch)
  • Standard: 5 seconds (balancing adequacy and fatigue avoidance)
  • Maximum phonation time: As long as possible (separate measure of respiratory-phonatory coordination)

Analysis typically excludes the first and last 0.5-1.0 seconds (onset and offset transients), analyzing the stable middle portion. Very long sustained vowels may show increasing perturbation from fatigue, though this pattern itself provides diagnostic information.

Pitch and Loudness Specification

Instructions typically specify:

  • Comfortable pitch: The speaker’s habitual conversational pitch
  • Comfortable loudness: Typical conversational loudness (approximately 70 dB SPL at 30 cm)

Some protocols include additional conditions:

  • High pitch (upper comfortable range)
  • Low pitch (lower comfortable range)
  • Soft phonation (piano)
  • Loud phonation (forte)

Extreme conditions often increase perturbation even in normal voices, while pathology may show exaggerated effects or inability to achieve the targets.

Reading Passages

Reading passages provide standardized connected speech samples enabling assessment of voice use in more natural communicative contexts.

Standard Reading Passages

Several passages achieve widespread clinical use:

Rainbow Passage (Fairbanks, 1960):

  • 330 words, requiring approximately 60 seconds
  • Phonetically balanced (representative distribution of English phonemes)
  • Relatively emotionally neutral content
  • Widely used in English-speaking contexts

Grandfather Passage:

  • Shorter alternative (approximately 40 seconds)
  • Simpler vocabulary suitable for diverse populations
  • Clear, declarative sentences

CAPE-V Sentences (Consensus Auditory-Perceptual Evaluation of Voice):

  • Six brief sentences designed for perceptual evaluation
  • Includes sustained vowels and connected speech
  • Balance vowels and consonants
  • Brief duration reduces fatigue

Language-Specific Passages:

  • Analogous passages developed for other languages
  • Maintain similar phonetic balance and emotional neutrality

Advantages of Reading Passages

Reading passages provide:

Ecological validity: Closer approximation to real communication than sustained vowels Prosodic variation: Natural intonation, stress, and rhythm patterns Articulation-phonation interaction: Reveals coordination between systems Speaking fundamental frequency sampling: Documents typical pitch use Intensity variation assessment: Captures natural loudness modulation Vocal endurance: Multiple sentences reveal fatigue or consistency

These characteristics make reading passages valuable complements to sustained vowel analysis, revealing voice problems that may not appear in sustained phonation.

Analysis Challenges

Reading passages present analytical challenges:

Nonstationarity: Continuous pitch and loudness changes violate perturbation analysis assumptions Pause segmentation: Automated analysis must identify speech versus silence Vowel identification: Extracting specific vowels from running speech requires sophisticated algorithms Spectral variability: Rapidly changing formants affect acoustic measures Individual reading styles: Rate, prosody, and expression vary across individuals

Consequently, reading passage analysis focuses primarily on:

  • Perceptual evaluation using rating scales
  • Cepstral measures (CPP) tolerant of nonstationarity
  • Fundamental frequency statistics (mean, range, variability)
  • Intensity statistics
  • Speaking rate and pause patterns

Syllable Repetition Tasks

Diadochokinesis (DDK) or syllable repetition tasks involve rapid repetition of simple syllables, assessing motor control and coordination.

Standard DDK Tasks

/pa-pa-pa…: Rapid bilabial plosive repetition assessing lip movement speed and precision

/ta-ta-ta…: Rapid alveolar plosive repetition assessing tongue tip movement

/ka-ka-ka…: Rapid velar plosive repetition assessing tongue dorsum movement

/pa-ta-ka…: Alternating motion rate (AMR) assessing sequencing and coordination across articulators

Typical instructions request “as fast and steady as possible” for 5-10 seconds per syllable sequence.

While primarily articulatory tasks, DDK provides voice-relevant information:

Voicing consistency: Maintaining voicing across rapid cycles reveals laryngeal control Amplitude stability: Regular amplitude indicates steady respiratory support Rate capabilities: Maximum repetition rate reflects motor system integrity Rhythm regularity: Consistent inter-syllable intervals indicate good control Vocal quality maintenance: Ability to sustain quality during rapid movements

Neurological disorders affecting voice often show DDK abnormalities, making these tasks valuable in differential diagnosis.

Pitch Glides

Pitch glides (sirens) involve smooth transitions across the vocal range from lowest comfortable pitch to highest and vice versa.

Protocol

Standard glide tasks:

  • Begin at comfortable pitch
  • Glide smoothly upward to highest comfortable pitch
  • Glide smoothly downward to lowest comfortable pitch
  • Maintain steady loudness throughout
  • Duration typically 3-5 seconds each direction

Information Provided

Pitch glides reveal:

Frequency range: Maximum and minimum achievable F₀ Register transitions: Smooth versus abrupt shift locations between registers Pitch control: Smoothness versus jerky transitions Vocal quality changes: Shifts in resonance or strain across range Break points: Frequencies where phonation becomes unstable Effort patterns: Locations requiring excessive effort or tension

These features provide diagnostic information about:

  • Vocal fold mass and tension capabilities
  • Cricothyroid function (pitch raising)
  • Register development and blending
  • Presence of nodules or other lesions affecting specific pitch ranges
  • Neurological control of laryngeal muscles

Maximum Phonation Time

Maximum phonation time (MPT) measures the longest duration a patient can sustain phonation on a single breath.

Protocol

Standard MPT protocol:

  1. Instruct patient to take a deep breath
  2. Phonate /a/ at comfortable pitch and loudness
  3. Continue until air is exhausted
  4. Measure duration in seconds
  5. Repeat 3 times, recording maximum value

Normative Values

Adult males: 20-35 seconds (mean ~25) Adult females: 15-25 seconds (mean ~20) Children (age 6-10): 10-15 seconds

Values below 10 seconds suggest respiratory-phonatory dysfunction.

Clinical Interpretation

MPT reflects:

Respiratory capacity: Lung volume and capacity for sustained exhalation Glottal closure efficiency: Incomplete closure wastes air, reducing MPT Respiratory-laryngeal coordination: Ability to manage airflow for phonation Pulmonary function: May correlate with vital capacity in respiratory disease

Reduced MPT may indicate:

  • Incomplete glottal closure (paralysis, atrophy, bowing)
  • Respiratory weakness
  • Poor breath management technique
  • Pulmonary disease
  • Neurological disorders affecting respiratory control

However, many factors influence MPT including body size, physical fitness, and effort. Some highly trained singers achieve 40-60+ seconds through exceptional technique rather than lung capacity alone.

Conversational Speech Sampling

Conversational speech samples involve spontaneous speaking during interview or discussion, providing maximally ecological assessment.

Elicitation Methods

Structured interview: Standard questions prompting extended responses

  • “Tell me about your job/family/hobbies”
  • “Describe what you did yesterday”
  • “Explain your voice problem and how it affects you”

Picture description: Describing complex scenes requiring sustained speech

Monologue topics: Speaking about familiar topics without interruption

Interactive conversation: Natural dialogue with clinician

Advantages

Conversational samples provide:

Maximum ecological validity: True representation of real-world voice use Natural prosody and dynamics: Authentic pitch and loudness variation Emotional content: Real emotional expression affecting voice Individual style: Reveals personal communication patterns Extended duration: Several minutes of speech for analysis Functional impact assessment: Direct observation of communicative limitations

Analysis Challenges

Conversational speech presents maximal analytical challenges:

Extreme nonstationarity: Constant pitch, loudness, articulation changes Content variability: Each sample differs, limiting comparability Environmental factors: Background noise, room acoustics affect recording quality Interaction effects: Partner speech, turn-taking affect production Cognitive-linguistic demands: Content generation affects voice production

Consequently, conversational speech analysis focuses primarily on:

  • Perceptual rating scales (CAPE-V, GRBAS)
  • Cepstral measures
  • Statistical summaries of F₀ and intensity
  • Voice quality consistency assessment
  • Communication effectiveness evaluation

Standardization and Protocol Considerations

Achieving reliable, comparable measurements requires attention to standardization.

Instruction Standardization

Provide:

  • Consistent wording: Same instructions across patients and sessions
  • Demonstration: Model the task when appropriate
  • Practice trials: Allow familiarization before recorded sample
  • Clear expectations: Specify pitch, loudness, duration requirements
  • Feedback: Confirm understanding before recording

Recording Standards

Maintain:

  • Consistent equipment: Same microphone, recorder, settings
  • Fixed distance: Standard mouth-to-microphone distance (typically 10 cm)
  • Quiet environment: Minimize background noise (<50 dB SPL ambient)
  • Calibration: Regular equipment calibration and checks
  • Documentation: Record all relevant parameters (equipment, settings, room)

Order Effects

Consider:

  • Fatigue: Later tasks may show degraded performance
  • Warm-up: Early tasks may not represent stable phonation
  • Learning: Practice effects may improve later tasks
  • Counterbalancing: Varying task order across patients when research requires

Ecological Validity Versus Control

Test utterance selection involves tradeoffs between experimental control and ecological validity.

Control-Focused Approach

Sustained vowels maximize:

  • Stationarity enabling perturbation analysis
  • Standardization across patients, sessions, sites
  • Isolation of phonatory function from articulation
  • Reliable quantitative measures

But minimize:

  • Resemblance to real communication
  • Prosodic variation
  • Articulation-phonation interaction
  • Functional relevance

Ecological-Focused Approach

Conversational speech maximizes:

  • Real-world relevance
  • Functional communication assessment
  • Natural prosody and variation
  • Communicative impact evaluation

But minimizes:

  • Measurement reliability
  • Quantitative analysis feasibility
  • Cross-patient comparability
  • Controlled variable isolation

Balanced Protocols

Comprehensive voice assessment typically includes:

  1. Sustained vowels for quantitative acoustic analysis
  2. Reading passage for semi-controlled connected speech
  3. Conversation for ecological functional assessment
  4. Special tasks (pitch glides, MPT, DDK) as clinically indicated

This multi-task approach balances control and validity, providing complementary information addressing different assessment goals.

Summary

Test utterances for clinical voice assessment include sustained vowels (providing phonatory isolation and analysis stability), reading passages (balancing standardization and ecological validity), syllable repetition tasks (assessing motor control), pitch glides (revealing frequency range and transitions), maximum phonation time (measuring respiratory-phonatory coordination), and conversational speech (maximizing functional relevance). Each utterance type offers distinct advantages and limitations regarding standardization, analytical feasibility, and real-world relevance.

Sustained vowels enable reliable perturbation analysis but sacrifice ecological validity. Conversational speech provides functional assessment but challenges quantitative measurement. Comprehensive protocols combine multiple utterances addressing different assessment goals while maintaining standardization through consistent instructions, recording conditions, and documentation. The choice among utterances depends on assessment objectives, available analysis resources, patient capabilities, and the need to balance experimental control against functional relevance. Understanding test utterance characteristics enables appropriate protocol design and meaningful interpretation of clinical voice assessment results.


Key Takeaways

  • ✅ Sustained vowels provide phonatory isolation and stationarity enabling reliable perturbation analysis with /a/ as the most common choice
  • ✅ Reading passages (Rainbow Passage, CAPE-V sentences) balance standardization with ecological validity for connected speech assessment
  • ✅ Syllable repetition tasks (DDK) assess motor control and coordination relevant to voice production
  • ✅ Pitch glides reveal frequency range, register transitions, and vocal control across the phonational range
  • ✅ Maximum phonation time measures respiratory-phonatory coordination with normative values of 20-35 seconds for adult males, 15-25 for females
  • ✅ Conversational speech maximizes ecological validity but presents analytical challenges requiring perceptual rating or robust measures like cepstral analysis
  • ✅ Standardization requires consistent instructions, recording conditions, equipment, and documentation to enable reliable comparison
  • ✅ Comprehensive assessment protocols include multiple utterance types balancing experimental control and functional relevance

Further Reading

  1. Fairbanks, G. (1960). Voice and articulation drillbook (2nd ed.). New York: Harper & Row.
  2. Kempster, G. B., Gerratt, B. R., Abbott, K. V., Barkmeier-Kraemer, J., & Hillman, R. E. (2009). Consensus auditory-perceptual evaluation of voice: Development of a standardized clinical protocol. American Journal of Speech-Language Pathology, 18(2), 124-132.
  3. Baken, R. J., & Orlikoff, R. F. (2000). Clinical measurement of speech and voice (2nd ed.). San Diego, CA: Singular Publishing Group.
  4. Zraick, R. I., Kempster, G. B., Connor, N. P., Thibeault, S., Klaben, B. K., Bursac, Z., & Glaze, L. E. (2011). Establishing validity of the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V). American Journal of Speech-Language Pathology, 20(1), 14-22.