Frequency Spectrum of the Resonating Tube
When a tube resonates, it does not respond equally to all frequencies. Instead, it exhibits a frequency spectrum—a characteristic pattern of resonant peaks at specific frequencies (formants) with varying sharpness and amplitude. Understanding this frequency response is essential for comprehending how the vocal tract shapes the glottal source spectrum into recognizable speech sounds. The spectrum of a resonating tube reveals not only which frequencies are emphasized but also how sharply tuned those resonances are, which profoundly affects vowel quality and voice timbre.
The Concept of Frequency Response
The frequency response describes how a system responds to inputs at different frequencies.
Transfer Function
The transfer function H(f) relates input to output as a function of frequency:
H(f) = Output(f) / Input(f)
For the vocal tract:
- Input: Volume velocity at the glottis (glottal airflow)
- Output: Sound pressure at the lips (radiated sound)
- Transfer function: Represents filtering effect of vocal tract
Magnitude Response: |H(f)| shows how much each frequency is amplified or attenuated.
Phase Response: arg[H(f)] shows phase shift introduced at each frequency (less perceptually important for speech).
Resonance Peaks
The magnitude spectrum |H(f)| of a resonating tube shows prominent peaks at resonant frequencies:
Formants: The peaks in |H(f)| correspond to natural frequencies where the tube responds most strongly. These are the formants F1, F2, F3, etc.
Peak Height: The amplitude of each peak depends on damping—less damping produces taller, sharper peaks.
Peak Width: The bandwidth of each peak depends on energy losses—more losses create broader peaks.
Valleys: Between formant peaks, the magnitude response drops, showing frequency regions where the tube responds weakly.
Figure 6.5: Frequency response magnitude showing formant peaks at F1, F2, and F3 with characteristic bandwidths and amplitudes.
Resonance and Standing Waves
Resonances occur at frequencies where standing wave patterns satisfy boundary conditions.
Standing Wave Formation
At resonant frequencies, incident and reflected waves interfere constructively:
Pressure Nodes: Locations where pressure variations cancel (minimum pressure amplitude). Occur at open ends (lip opening).
Pressure Antinodes: Locations where pressure variations reinforce (maximum pressure amplitude). Occur at closed ends (glottis) and at internal reflection points.
Particle Velocity: Pattern complementary to pressure—velocity nodes at pressure antinodes and vice versa.
Resonance Condition
For a quarter-wave tube (closed-open), resonances occur when:
Fₙ = (2n - 1)(c/4L)
where n = 1, 2, 3, …
First Formant (n = 1): F₁ = c/4L
- Standing wave has quarter-wavelength fitting in tube length
- Pressure maximum at closed end, minimum at open end
- Simplest resonance mode
Second Formant (n = 2): F₂ = 3c/4L
- Three-quarter wavelength fits in tube
- One additional pressure node inside tube
- More complex standing wave pattern
Higher Formants: Each successive formant adds another half-wavelength, increasing pattern complexity.
Multiple Formants
A single tube typically has an infinite series of resonances:
Formant Spacing: For uniform tube, formants are evenly spaced by c/2L ≈ 1000 Hz (for L = 17.5 cm).
Simultaneous Resonances: All formants exist simultaneously. The tube responds to components of the input spectrum at all formant frequencies.
Spectral Shape: The overall frequency response reflects all formants together, creating the characteristic filtering effect that defines vowel quality.
Formant Bandwidth and Quality Factor
Real resonances are not infinitely sharp but have finite bandwidth due to energy dissipation.
Sources of Damping
Energy losses broaden resonance peaks:
Viscous Losses: Friction at tube walls converts acoustic energy to heat. Proportional to surface area and frequency. Higher frequencies experience greater viscous loss.
Thermal Losses: Temperature oscillations at walls dissipate energy through thermal conduction. Also frequency-dependent.
Radiation Losses: Acoustic energy radiating from the mouth opening is lost from the system. Increases with frequency.
Yielding Walls: Vocal tract walls are not perfectly rigid. Wall vibration absorbs acoustic energy, particularly at lower frequencies where wall compliance matters more.
Formant Bandwidth Definition
Bandwidth (Bₙ): The width of the formant peak, measured at the half-power point (3 dB down from peak).
Half-Power Point: Frequency at which |H(f)|² falls to half its maximum value, corresponding to |H(f)| falling to 1/√2 ≈ 0.707 of peak value.
Typical Values:
B₁ ≈ 50-100 Hz (first formant)
B₂ ≈ 70-150 Hz (second formant)
B₃ ≈ 100-200 Hz (third formant)
B₄ ≈ 150-300 Hz (fourth formant)
Trend: Higher formants generally have larger bandwidths due to increased losses at higher frequencies.
Quality Factor (Q)
The quality factor Q quantifies resonance sharpness:
Q = Fₙ / Bₙ
Interpretation:
- High Q: Sharp, narrow resonance; little damping; long ringing time
- Low Q: Broad, rounded resonance; substantial damping; short decay
Typical Vocal Tract Values:
Q₁ = F₁/B₁ ≈ 500/70 ≈ 7-10
Q₂ = F₂/B₂ ≈ 1500/120 ≈ 10-15
Q₃ = F₃/B₃ ≈ 2500/150 ≈ 15-20
Comparison:
- Musical instruments (flute, violin): Q ≈ 50-100 (sharp resonances)
- Vocal tract: Q ≈ 10-20 (moderate resonances)
- Megaphone: Q ≈ 2-5 (broad resonances)
Physical Significance of Bandwidth
Time-Domain Correspondence: Bandwidth relates to resonance decay time:
τ = 1/(πB)
where τ is the time constant for exponential decay.
Narrow Bandwidth (small B): Long decay time—resonance persists after excitation stops. System has “memory.”
Wide Bandwidth (large B): Short decay time—resonance dies quickly. System responds rapidly to changes.
Speech Implications: Moderate Q (moderate bandwidth) allows rapid formant transitions during connected speech while maintaining sufficient spectral definition for vowel identification.
The Complete Spectral Response
The vocal tract transfer function combines all formants into a complete frequency response.
Mathematical Representation
The transfer function can be approximated as a product of resonances:
H(f) = A · ∏[1 / (1 - (f/Fₙ)² + j(f·Bₙ)/(Fₙ²))]
where the product is over all formants n = 1, 2, 3, …
Components:
- A: Overall amplitude constant
- Resonance terms: Each formant contributes a peak centered at Fₙ with bandwidth Bₙ
- Complex denominator: Creates resonance peak shape
Spectral Shape Characteristics
The magnitude spectrum |H(f)| has characteristic features:
Formant Peaks: Prominent maxima at F1, F2, F3, etc., with heights depending on Q factors and source-formant frequency relationships.
Anti-resonances: Valleys between formants where response is minimal. Frequency components in these regions are attenuated.
Overall Slope: Generally decreasing at high frequencies due to radiation characteristics and increased losses.
Low-Frequency Roll-off: Below F1, response drops. Very low frequencies are not efficiently radiated.
Source-Filter Interaction
The radiated speech spectrum combines source and filter:
Speech Spectrum = Source Spectrum × Transfer Function
or in decibels:
Speech(dB) = Source(dB) + Transfer(dB)
Source Spectrum: Glottal airflow has energy at F0 and harmonics, with amplitude decreasing approximately 12 dB/octave.
Filter Spectrum: Vocal tract transfer function emphasizes formant regions, suppresses intervening frequencies.
Output Spectrum: Product shows peaks at harmonics nearest to formants. These harmonics are strongest in radiated sound.
Figure 6.14: Illustration of source-filter theory showing glottal source spectrum, vocal tract transfer function, and resulting speech output spectrum.
Frequency Spacing and Formant Interactions
The relationship between F0 and formant frequencies affects spectral fine structure.
Harmonic-Formant Alignment
Harmonically Excited Resonances: The glottal source produces harmonics at integer multiples of F0. Formants amplify whichever harmonics are nearest.
Sparse Spectrum (low F0, male voices):
- F0 ≈ 100 Hz means harmonics at 100, 200, 300, 400, … Hz
- Multiple harmonics fall within each formant bandwidth
- Formant peaks clearly visible in spectrum
- F1 might amplify 4th-6th harmonics, F2 might amplify 15th-20th harmonics
Dense Spectrum (high F0, female/child voices):
- F0 ≈ 250 Hz means harmonics at 250, 500, 750, 1000, … Hz
- Fewer harmonics per formant bandwidth
- F1 might amplify only 2nd-3rd harmonics
- Individual harmonics more prominent than formant envelope
Very High F0 (soprano, above 500 Hz):
- F0 may exceed F1 for closed vowels
- Only one or two harmonics below F2
- Formant structure less apparent; individual harmonics dominate
- Vowel identification relies more on F1-F2 spacing
Formant Tuning
Singers sometimes adjust formant frequencies to align with harmonics:
Principle: Maximum acoustic output when formant frequency matches a harmonic of F0.
F1 Tuning: On high pitches, singers raise F1 (by lowering jaw/tongue) to match F0 or second harmonic. Creates louder, more resonant tone.
Singer’s Formant: Male classical singers develop formant clustering around 2800-3200 Hz, creating a strong spectral peak that projects through orchestral accompaniment.
Acoustic Consequences: Aligned formant-harmonic produces maximum amplitude at that frequency, enhancing loudness and carrying power.
Factors Affecting Spectral Shape
Various anatomical and physiological factors influence the frequency response.
Vocal Tract Length
Shorter Tract (children, women):
- Higher formant frequencies: Fₙ ∝ 1/L
- Formants more widely spaced in Bark or mel scale (perceptual spacing)
- Same vowel shapes produce proportionally higher formants
Longer Tract (men):
- Lower formant frequencies
- Average adult male F1 ≈ 500 Hz vs. average adult female F1 ≈ 550 Hz for /a/
Length Modification: Lip protrusion lengthens tract, lowering all formants. Larynx lowering also lengthens tract.
Constriction Degree
Narrow Constriction (high vowels):
- F1 lowered
- Sharper formant peaks due to stronger coupling effects
- Greater acoustic separation between oral and pharyngeal cavities
Wide Constriction (low vowels):
- F1 raised
- Formants closer to uniform tube values
- Less defined cavity separation
Tissue Properties
Rigid Walls (idealized):
- Higher Q (sharper resonances)
- Neglects yielding wall losses
Compliant Walls (realistic):
- Wall vibration absorbs energy, particularly at lower frequencies
- Increases formant bandwidth, especially B1
- More realistic modeling includes wall impedance
Coupling and Side Branches
Nasal Coupling: Opening velopharyngeal port introduces:
- Additional resonances (nasal formants)
- Anti-resonances (zeros) that cancel energy
- Increased damping (larger bandwidths)
Piriform Sinuses: Small cavities lateral to laryngopharynx:
- Create anti-resonances around 4-5 kHz
- Contribute to individual voice quality differences
Sublingual Cavity: Space under tongue, variable with tongue position:
- Affects F3 and higher formants
- More prominent for some vowels
Measurement and Analysis
Quantifying the frequency response requires specialized techniques.
Impulse Response Method
Principle: Apply brief acoustic impulse at glottis, record output at lips.
Processing: Fourier transform of output reveals transfer function magnitude and phase.
Advantages: Directly measures complete frequency response.
Challenges: Difficult to apply controlled impulse at glottis in vivo. More suitable for physical models or computational simulations.
Acoustic Impedance Measurement
Principle: Measure impedance looking into tract from lips or glottis.
Relationship: Formants correspond to impedance minima (anti-resonances of impedance).
Application: Acoustic reflection techniques can measure input impedance, from which transfer function is derived.
Inverse Filtering
Principle: Record speech output, estimate glottal flow waveform, compute transfer function.
Process:
- Record radiated speech
- Remove lip radiation effect (6 dB/octave boost)
- Estimate formant frequencies and bandwidths
- Construct inverse filter to recover glottal flow
- Transfer function is inverse of the inverse filter
Advantage: Can be applied to actual speech production.
Challenge: Assumptions about source-filter independence and accurate formant estimation required.
Linear Predictive Coding (LPC)
Principle: Model vocal tract as all-pole filter, estimate parameters from speech signal.
Output: Formant frequencies and bandwidths directly from model parameters.
Advantages: Computationally efficient, widely used in speech processing.
Limitations: Assumes all-pole model (no zeros); may introduce spurious formants; sensitive to analysis parameters.
Clinical and Practical Applications
Understanding the frequency spectrum of resonating tubes has diverse applications.
Voice Quality Assessment
Spectral Slope: Overall spectral tilt indicates vocal effort and glottal configuration:
- Steep negative slope: Breathy voice, incomplete closure
- Shallow slope: Pressed voice, hyperadduction
Formant Clarity: Well-defined formant peaks indicate:
- Healthy vocal tract resonance
- Appropriate formant bandwidths
- Efficient vocal tract configuration
Spectral Noise: Aperiodic energy between harmonics indicates:
- Turbulence (breathiness, roughness)
- Incomplete vocal fold closure
- Pathological tissue changes
Vowel Recognition Systems
Feature Extraction: Automatic speech recognition systems extract formant frequencies from spectral peaks.
Formant Tracking: Algorithms identify and track F1, F2, F3 over time, providing features for vowel classification.
Robustness: Systems must handle variation in formant amplitudes, bandwidths, and harmonic spacing across speakers and conditions.
Acoustic Design
Concert Hall Acoustics: Understanding resonance and frequency response guides room design for optimal speech intelligibility.
Hearing Aid Design: Frequency shaping in hearing aids accounts for speech spectrum characteristics, emphasizing formant regions.
Telephone Systems: Bandwidth limitations (300-3400 Hz) designed to preserve F1 and F2 for most vowels, sacrificing F3 and higher frequencies.
Singing Training
Spectral Analysis: Real-time spectral displays help singers visualize:
- Formant frequency adjustments
- Harmonic-formant alignment
- Singer’s formant development
Resonance Strategies: Understanding formant behavior guides pedagogical instructions for vowel modification and resonance tuning.
Summary
The frequency spectrum of a resonating tube describes how the tube responds to different input frequencies, with the vocal tract transfer function showing prominent resonant peaks (formants) at natural frequencies where standing waves satisfy boundary conditions. For a uniform quarter-wave tube, formants occur at Fₙ = (2n-1)(c/4L), typically spaced at approximately 1000 Hz intervals for an adult male vocal tract. Each formant peak has a finite bandwidth (typically 50-300 Hz) due to energy losses from viscosity, thermal conduction, radiation, and wall compliance, with bandwidth increasing for higher formants.
The quality factor Q = Fₙ/Bₙ quantifies resonance sharpness, with typical vocal tract values ranging from 7-20, representing moderate damping that balances spectral definition against rapid transient response needed for speech. The complete transfer function combines all formants into a characteristic spectral shape that filters the glottal source spectrum, amplifying harmonics near formant frequencies while attenuating components between formants. The interaction between harmonic spacing (determined by F0) and formant spacing affects spectral fine structure, with low F0 producing sparse spectra where formant envelope is prominent and high F0 producing dense spectra where individual harmonics dominate.
Formant frequencies and bandwidths depend on vocal tract dimensions, constriction degree, and tissue properties, with shorter tracts producing higher formants and narrow constrictions producing lower F1 with sharper resonances. Measurement techniques including inverse filtering, acoustic impedance measurements, and linear predictive coding allow quantification of formant parameters from speech signals. Understanding frequency spectra informs clinical voice assessment, automatic speech recognition, acoustic design, and singing pedagogy, providing quantitative links between vocal tract configuration and acoustic output.
Key Takeaways
- ✅ The vocal tract transfer function shows resonant peaks (formants) at natural frequencies where standing waves satisfy boundary conditions
- ✅ Formant bandwidth quantifies peak width at half-power points, typically ranging from 50-100 Hz for F1 to 150-300 Hz for F4
- ✅ Quality factor Q = Fₙ/Bₙ describes resonance sharpness, with vocal tract Q typically 7-20, balancing definition and transient response
- ✅ Energy losses from viscosity, thermal conduction, radiation, and wall compliance broaden formant peaks, with losses increasing at higher frequencies
- ✅ The speech spectrum is the product of glottal source spectrum and vocal tract transfer function, amplifying harmonics near formants
- ✅ Harmonic-formant alignment depends on F0, with low F0 showing clear formant envelope and high F0 showing dominant individual harmonics
- ✅ Formant tuning in singing aligns formants with harmonics to maximize acoustic output at specific frequencies
- ✅ Measurement techniques including inverse filtering and LPC extract formant frequencies and bandwidths for voice analysis and synthesis
Related Topics
- Quarter-Wave Resonance
- Formant Bandwidth
- Reflection Coefficients
- The F1-F2 Vowel Chart
- The Glottal Source Function
- The Spectral Slope
Further Reading
- Fant, G. (1960). Acoustic Theory of Speech Production. The Hague: Mouton.
- Stevens, K. N. (1998). Acoustic Phonetics. Cambridge, MA: MIT Press.
- Titze, I. R. (2000). Principles of Voice Production (2nd ed.). Iowa City: National Center for Voice and Speech.
- Flanagan, J. L. (1972). Speech Analysis Synthesis and Perception (2nd ed.). New York: Springer-Verlag.
- Sundberg, J. (1987). The Science of the Singing Voice. DeKalb, IL: Northern Illinois University Press.
- Kent, R. D., & Read, C. (2002). Acoustic Analysis of Speech (2nd ed.). Albany, NY: Singular Publishing Group.