Test Utterances
Clinical voice assessment requires eliciting voice samples suitable for perceptual evaluation, acoustic analysis, and documentation of vocal function. The choice of test utterances—specific phonatory tasks requested from patients—profoundly affects the resulting measurements and clinical interpretations. Standardized protocols enable comparison across patients, time points, and clinical sites while balancing the need for controlled conditions against ecological validity. Understanding the rationale, implementation, and limitations of various test utterances informs appropriate clinical decision-making and meaningful interpretation of assessment results.
Sustained Vowels
Sustained vowels represent the most common test utterance for acoustic voice analysis, involving phonation of a single vowel at comfortable pitch and loudness for several seconds.
Rationale and Advantages
Sustained vowels provide:
Phonatory isolation: Eliminates articulatory movements, isolating laryngeal function Stationarity: Promotes relatively stable F₀ and amplitude facilitating perturbation analysis Standardization: Simple instructions with minimal room for interpretation variation Cross-linguistic applicability: Basic vowels exist across languages Reduced cognitive load: Minimal memory or linguistic processing demands Extended analysis windows: Several seconds of continuous phonation enable reliable statistics
These advantages make sustained vowels the standard for perturbation analysis, enabling jitter, shimmer, and harmonics-to-noise ratio measurement.
Vowel Selection
Standard protocols typically employ:
/a/ (as in “father”): Most common choice due to:
- Open vocal tract minimizing formant-source interaction
- Comfortable articulation position
- Relatively stable acoustic characteristics
- Cross-cultural familiarity
/i/ (as in “see”): Sometimes included to assess:
- Higher formants affecting harmonics
- Increased vocal tract impedance effects
- Vowel-dependent stability differences
/u/ (as in “moon”): Occasionally used for:
- Lip rounding effects
- Different formant configuration
- Comparison with other vowels
Research shows vowel choice affects measured perturbation values by 10-30%, with /a/ typically producing lowest jitter and shimmer. Standardizing on one vowel (usually /a/) facilitates comparison while recognizing that complete voice assessment may include multiple vowels.
Duration
Typical protocols request:
- Minimum: 3 seconds (providing 200-400 cycles at normal pitch)
- Standard: 5 seconds (balancing adequacy and fatigue avoidance)
- Maximum phonation time: As long as possible (separate measure of respiratory-phonatory coordination)
Analysis typically excludes the first and last 0.5-1.0 seconds (onset and offset transients), analyzing the stable middle portion. Very long sustained vowels may show increasing perturbation from fatigue, though this pattern itself provides diagnostic information.
Pitch and Loudness Specification
Instructions typically specify:
- Comfortable pitch: The speaker’s habitual conversational pitch
- Comfortable loudness: Typical conversational loudness (approximately 70 dB SPL at 30 cm)
Some protocols include additional conditions:
- High pitch (upper comfortable range)
- Low pitch (lower comfortable range)
- Soft phonation (piano)
- Loud phonation (forte)
Extreme conditions often increase perturbation even in normal voices, while pathology may show exaggerated effects or inability to achieve the targets.
Reading Passages
Reading passages provide standardized connected speech samples enabling assessment of voice use in more natural communicative contexts.
Standard Reading Passages
Several passages achieve widespread clinical use:
Rainbow Passage (Fairbanks, 1960):
- 330 words, requiring approximately 60 seconds
- Phonetically balanced (representative distribution of English phonemes)
- Relatively emotionally neutral content
- Widely used in English-speaking contexts
Grandfather Passage:
- Shorter alternative (approximately 40 seconds)
- Simpler vocabulary suitable for diverse populations
- Clear, declarative sentences
CAPE-V Sentences (Consensus Auditory-Perceptual Evaluation of Voice):
- Six brief sentences designed for perceptual evaluation
- Includes sustained vowels and connected speech
- Balance vowels and consonants
- Brief duration reduces fatigue
Language-Specific Passages:
- Analogous passages developed for other languages
- Maintain similar phonetic balance and emotional neutrality
Advantages of Reading Passages
Reading passages provide:
Ecological validity: Closer approximation to real communication than sustained vowels Prosodic variation: Natural intonation, stress, and rhythm patterns Articulation-phonation interaction: Reveals coordination between systems Speaking fundamental frequency sampling: Documents typical pitch use Intensity variation assessment: Captures natural loudness modulation Vocal endurance: Multiple sentences reveal fatigue or consistency
These characteristics make reading passages valuable complements to sustained vowel analysis, revealing voice problems that may not appear in sustained phonation.
Analysis Challenges
Reading passages present analytical challenges:
Nonstationarity: Continuous pitch and loudness changes violate perturbation analysis assumptions Pause segmentation: Automated analysis must identify speech versus silence Vowel identification: Extracting specific vowels from running speech requires sophisticated algorithms Spectral variability: Rapidly changing formants affect acoustic measures Individual reading styles: Rate, prosody, and expression vary across individuals
Consequently, reading passage analysis focuses primarily on:
- Perceptual evaluation using rating scales
- Cepstral measures (CPP) tolerant of nonstationarity
- Fundamental frequency statistics (mean, range, variability)
- Intensity statistics
- Speaking rate and pause patterns
Syllable Repetition Tasks
Diadochokinesis (DDK) or syllable repetition tasks involve rapid repetition of simple syllables, assessing motor control and coordination.
Standard DDK Tasks
/pa-pa-pa…: Rapid bilabial plosive repetition assessing lip movement speed and precision
/ta-ta-ta…: Rapid alveolar plosive repetition assessing tongue tip movement
/ka-ka-ka…: Rapid velar plosive repetition assessing tongue dorsum movement
/pa-ta-ka…: Alternating motion rate (AMR) assessing sequencing and coordination across articulators
Typical instructions request “as fast and steady as possible” for 5-10 seconds per syllable sequence.
Voice-Related Information
While primarily articulatory tasks, DDK provides voice-relevant information:
Voicing consistency: Maintaining voicing across rapid cycles reveals laryngeal control Amplitude stability: Regular amplitude indicates steady respiratory support Rate capabilities: Maximum repetition rate reflects motor system integrity Rhythm regularity: Consistent inter-syllable intervals indicate good control Vocal quality maintenance: Ability to sustain quality during rapid movements
Neurological disorders affecting voice often show DDK abnormalities, making these tasks valuable in differential diagnosis.
Pitch Glides
Pitch glides (sirens) involve smooth transitions across the vocal range from lowest comfortable pitch to highest and vice versa.
Protocol
Standard glide tasks:
- Begin at comfortable pitch
- Glide smoothly upward to highest comfortable pitch
- Glide smoothly downward to lowest comfortable pitch
- Maintain steady loudness throughout
- Duration typically 3-5 seconds each direction
Information Provided
Pitch glides reveal:
Frequency range: Maximum and minimum achievable F₀ Register transitions: Smooth versus abrupt shift locations between registers Pitch control: Smoothness versus jerky transitions Vocal quality changes: Shifts in resonance or strain across range Break points: Frequencies where phonation becomes unstable Effort patterns: Locations requiring excessive effort or tension
These features provide diagnostic information about:
- Vocal fold mass and tension capabilities
- Cricothyroid function (pitch raising)
- Register development and blending
- Presence of nodules or other lesions affecting specific pitch ranges
- Neurological control of laryngeal muscles
Maximum Phonation Time
Maximum phonation time (MPT) measures the longest duration a patient can sustain phonation on a single breath.
Protocol
Standard MPT protocol:
- Instruct patient to take a deep breath
- Phonate /a/ at comfortable pitch and loudness
- Continue until air is exhausted
- Measure duration in seconds
- Repeat 3 times, recording maximum value
Normative Values
Adult males: 20-35 seconds (mean ~25) Adult females: 15-25 seconds (mean ~20) Children (age 6-10): 10-15 seconds
Values below 10 seconds suggest respiratory-phonatory dysfunction.
Clinical Interpretation
MPT reflects:
Respiratory capacity: Lung volume and capacity for sustained exhalation Glottal closure efficiency: Incomplete closure wastes air, reducing MPT Respiratory-laryngeal coordination: Ability to manage airflow for phonation Pulmonary function: May correlate with vital capacity in respiratory disease
Reduced MPT may indicate:
- Incomplete glottal closure (paralysis, atrophy, bowing)
- Respiratory weakness
- Poor breath management technique
- Pulmonary disease
- Neurological disorders affecting respiratory control
However, many factors influence MPT including body size, physical fitness, and effort. Some highly trained singers achieve 40-60+ seconds through exceptional technique rather than lung capacity alone.
Conversational Speech Sampling
Conversational speech samples involve spontaneous speaking during interview or discussion, providing maximally ecological assessment.
Elicitation Methods
Structured interview: Standard questions prompting extended responses
- “Tell me about your job/family/hobbies”
- “Describe what you did yesterday”
- “Explain your voice problem and how it affects you”
Picture description: Describing complex scenes requiring sustained speech
Monologue topics: Speaking about familiar topics without interruption
Interactive conversation: Natural dialogue with clinician
Advantages
Conversational samples provide:
Maximum ecological validity: True representation of real-world voice use Natural prosody and dynamics: Authentic pitch and loudness variation Emotional content: Real emotional expression affecting voice Individual style: Reveals personal communication patterns Extended duration: Several minutes of speech for analysis Functional impact assessment: Direct observation of communicative limitations
Analysis Challenges
Conversational speech presents maximal analytical challenges:
Extreme nonstationarity: Constant pitch, loudness, articulation changes Content variability: Each sample differs, limiting comparability Environmental factors: Background noise, room acoustics affect recording quality Interaction effects: Partner speech, turn-taking affect production Cognitive-linguistic demands: Content generation affects voice production
Consequently, conversational speech analysis focuses primarily on:
- Perceptual rating scales (CAPE-V, GRBAS)
- Cepstral measures
- Statistical summaries of F₀ and intensity
- Voice quality consistency assessment
- Communication effectiveness evaluation
Standardization and Protocol Considerations
Achieving reliable, comparable measurements requires attention to standardization.
Instruction Standardization
Provide:
- Consistent wording: Same instructions across patients and sessions
- Demonstration: Model the task when appropriate
- Practice trials: Allow familiarization before recorded sample
- Clear expectations: Specify pitch, loudness, duration requirements
- Feedback: Confirm understanding before recording
Recording Standards
Maintain:
- Consistent equipment: Same microphone, recorder, settings
- Fixed distance: Standard mouth-to-microphone distance (typically 10 cm)
- Quiet environment: Minimize background noise (<50 dB SPL ambient)
- Calibration: Regular equipment calibration and checks
- Documentation: Record all relevant parameters (equipment, settings, room)
Order Effects
Consider:
- Fatigue: Later tasks may show degraded performance
- Warm-up: Early tasks may not represent stable phonation
- Learning: Practice effects may improve later tasks
- Counterbalancing: Varying task order across patients when research requires
Ecological Validity Versus Control
Test utterance selection involves tradeoffs between experimental control and ecological validity.
Control-Focused Approach
Sustained vowels maximize:
- Stationarity enabling perturbation analysis
- Standardization across patients, sessions, sites
- Isolation of phonatory function from articulation
- Reliable quantitative measures
But minimize:
- Resemblance to real communication
- Prosodic variation
- Articulation-phonation interaction
- Functional relevance
Ecological-Focused Approach
Conversational speech maximizes:
- Real-world relevance
- Functional communication assessment
- Natural prosody and variation
- Communicative impact evaluation
But minimizes:
- Measurement reliability
- Quantitative analysis feasibility
- Cross-patient comparability
- Controlled variable isolation
Balanced Protocols
Comprehensive voice assessment typically includes:
- Sustained vowels for quantitative acoustic analysis
- Reading passage for semi-controlled connected speech
- Conversation for ecological functional assessment
- Special tasks (pitch glides, MPT, DDK) as clinically indicated
This multi-task approach balances control and validity, providing complementary information addressing different assessment goals.
Summary
Test utterances for clinical voice assessment include sustained vowels (providing phonatory isolation and analysis stability), reading passages (balancing standardization and ecological validity), syllable repetition tasks (assessing motor control), pitch glides (revealing frequency range and transitions), maximum phonation time (measuring respiratory-phonatory coordination), and conversational speech (maximizing functional relevance). Each utterance type offers distinct advantages and limitations regarding standardization, analytical feasibility, and real-world relevance.
Sustained vowels enable reliable perturbation analysis but sacrifice ecological validity. Conversational speech provides functional assessment but challenges quantitative measurement. Comprehensive protocols combine multiple utterances addressing different assessment goals while maintaining standardization through consistent instructions, recording conditions, and documentation. The choice among utterances depends on assessment objectives, available analysis resources, patient capabilities, and the need to balance experimental control against functional relevance. Understanding test utterance characteristics enables appropriate protocol design and meaningful interpretation of clinical voice assessment results.
Key Takeaways
- ✅ Sustained vowels provide phonatory isolation and stationarity enabling reliable perturbation analysis with /a/ as the most common choice
- ✅ Reading passages (Rainbow Passage, CAPE-V sentences) balance standardization with ecological validity for connected speech assessment
- ✅ Syllable repetition tasks (DDK) assess motor control and coordination relevant to voice production
- ✅ Pitch glides reveal frequency range, register transitions, and vocal control across the phonational range
- ✅ Maximum phonation time measures respiratory-phonatory coordination with normative values of 20-35 seconds for adult males, 15-25 for females
- ✅ Conversational speech maximizes ecological validity but presents analytical challenges requiring perceptual rating or robust measures like cepstral analysis
- ✅ Standardization requires consistent instructions, recording conditions, equipment, and documentation to enable reliable comparison
- ✅ Comprehensive assessment protocols include multiple utterance types balancing experimental control and functional relevance
Related Topics
- Jitter and Shimmer
- Signals with Small Perturbations
- Fundamental Frequency Profile
- Nonstationarity and Trends
Further Reading
- Fairbanks, G. (1960). Voice and articulation drillbook (2nd ed.). New York: Harper & Row.
- Kempster, G. B., Gerratt, B. R., Abbott, K. V., Barkmeier-Kraemer, J., & Hillman, R. E. (2009). Consensus auditory-perceptual evaluation of voice: Development of a standardized clinical protocol. American Journal of Speech-Language Pathology, 18(2), 124-132.
- Baken, R. J., & Orlikoff, R. F. (2000). Clinical measurement of speech and voice (2nd ed.). San Diego, CA: Singular Publishing Group.
- Zraick, R. I., Kempster, G. B., Connor, N. P., Thibeault, S., Klaben, B. K., Bursac, Z., & Glaze, L. E. (2011). Establishing validity of the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V). American Journal of Speech-Language Pathology, 20(1), 14-22.