A Three-Tube Approximation

vowels acoustics formants vocal-tract modeling tube-resonance
Last updated: 2025-02-07

A Three-Tube Approximation

The three-tube model represents a significant advance in accurately predicting formant frequencies for vowels. By dividing the vocal tract into three cylindrical sections—pharyngeal, oral, and lip—this approximation captures the essential acoustic effects of tongue position, constriction location, and lip configuration. The three-tube model provides intuitive insights into how articulatory adjustments affect specific formants, making it valuable for both theoretical understanding and practical applications in voice pedagogy and clinical intervention.

From Uniform Tube to Three-Tube Model

The progression from simple to more complex models illustrates the relationship between anatomical realism and acoustic accuracy.

The Uniform Tube Model

The simplest vocal tract model treats it as a uniform cylindrical tube:

Configuration: Single tube of constant cross-sectional area, closed at one end (glottis), open at the other (lips).

Formant Prediction: Fₙ = (2n-1)(c/4L), where L is total tract length (approximately 17-17.5 cm for adult males).

Results: Formants evenly spaced at approximately 1000 Hz intervals (500, 1500, 2500, 3500 Hz…).

Limitations: Predicts only the neutral schwa-like vowel. Cannot account for different vowel qualities because it ignores vocal tract shape variations.

The Two-Tube Model

Adding one constriction creates a two-tube model:

Configuration: Two tubes of different areas joined at the primary tongue constriction point. Back cavity (pharynx) and front cavity (oral cavity anterior to constriction).

Acoustic Principle: Relative lengths and areas of the two sections determine formant frequencies. Moving the constriction forward or backward shifts formants systematically.

Improvements: Can predict how F1 and F2 vary with tongue height and advancement, explaining major vowel category distinctions.

Remaining Limitations: Ignores lip effects, which are particularly important for rounded vowels and higher formants.

The Three-Tube Model

Adding the lip section completes the basic model:

Configuration: Three cylindrical sections representing:

  1. Pharyngeal tube (back cavity): From glottis to tongue constriction
  2. Oral tube (front cavity): From constriction to lips
  3. Lip tube: Short section representing lip protrusion or constriction

Advantages: Accounts for lip rounding effects, provides better predictions for F3 and higher formants, maintains interpretability while improving accuracy.

Three-tube model diagram Figure 6.11: Three-tube approximation showing pharyngeal, oral, and lip sections for different vowel configurations.

The Three-Tube Configuration

Each of the three sections has specific acoustic properties determined by its dimensions.

Pharyngeal Section (Back Cavity)

The pharyngeal tube extends from the glottis to the location of the primary tongue constriction.

Length (Lp): Varies with tongue position:

  • Front vowels: Long pharyngeal section (10-12 cm)
  • Back vowels: Short pharyngeal section (3-6 cm)
  • Central vowels: Intermediate length (7-9 cm)

Cross-sectional Area (Ap): Varies with tongue root position and pharyngeal width:

  • Expanded pharynx (ATR): Ap = 3-5 cm²
  • Neutral pharynx: Ap = 2-3 cm²
  • Constricted pharynx (RTR): Ap = 1-2 cm²

Boundary Conditions: Closed at glottal end (approximately), open to oral section at the constriction.

Acoustic Role: Pharyngeal length and area strongly influence F1. Long pharynx (front vowels) lowers F1. Expanded pharynx (large Ap) also lowers F1.

Oral Section (Front Cavity)

The oral tube extends from the tongue constriction to the lips.

Length (Lo): Complementary to pharyngeal length (Lp + Lo + Ll ≈ L total):

  • Front vowels: Short oral section (4-6 cm)
  • Back vowels: Long oral section (10-13 cm)
  • Central vowels: Intermediate length (7-9 cm)

Cross-sectional Area (Ao): Determined by the degree of tongue constriction:

  • High vowels (/i/, /u/): Small Ao at constriction (0.3-0.5 cm²)
  • Mid vowels (/e/, /o/): Moderate Ao (0.5-1.5 cm²)
  • Low vowels (/a/): Large Ao (2-4 cm²), minimal constriction

Acoustic Role: Oral length primarily affects F2. Short oral cavity (front vowels) raises F2. Long oral cavity (back vowels) lowers F2. Constriction area affects multiple formants.

Lip Section

The lip tube represents the acoustic effects of lip configuration.

Length (Ll): Determined by lip protrusion:

  • Rounded/protruded lips: Ll = 1-2 cm
  • Neutral lips: Ll = 0.5 cm
  • Spread lips: Ll ≈ 0 cm (can be modeled as zero)

Cross-sectional Area (Al): Lip aperture area:

  • Rounded vowels: Small Al (0.5-2 cm²)
  • Unrounded vowels: Larger Al (2-5 cm²)

Acoustic Role: Lip section primarily lowers formants through tract lengthening. Protruded lips lower F2 and F3 particularly. Small lip aperture increases radiation impedance, affecting higher formants more than lower ones.

Acoustic Theory of the Three-Tube Model

The formant frequencies of the three-tube system depend on the acoustic properties of each section and how they couple.

Wave Propagation in Concatenated Tubes

Sound waves propagate through the three-tube system, experiencing reflections at each area discontinuity:

Glottis-Pharynx Junction: Expansion from small glottal opening to larger pharynx. Partial reflection of compression waves (converted to rarefaction) and rarefaction waves (converted to compression).

Pharynx-Oral Junction (Constriction): Area change (usually a decrease for vowels). Reflection coefficient depends on area ratio. Significant reflection for high vowels with narrow constriction.

Oral-Lip Junction: Area change depending on lip configuration. Reflection affects higher formants particularly.

Lip-Air Interface: Large expansion to infinite area. Strong reflection (compression inverts to rarefaction), crucial for establishing standing waves.

Formant Determination

Formants occur at frequencies where constructive interference produces standing wave resonances:

Boundary Conditions:

  • Pressure maximum (velocity minimum) at closed glottal end
  • Pressure minimum (velocity maximum) at open lip end
  • Continuity of pressure and volume velocity at tube junctions

Resonance Condition: The combination of tube lengths and reflection coefficients determines which frequencies satisfy standing wave conditions. Unlike the uniform tube (analytical solution), the three-tube system typically requires numerical solution.

Perturbation Theory

A powerful approach for understanding formant shifts uses perturbation theory:

Principle: Small changes in tube dimensions produce predictable formant frequency changes based on the standing wave pattern at each formant.

Key Insight: The effect of a local area change on formant frequency depends on whether that location has high or low particle velocity for that particular formant mode.

Formant Raising Rule: Decreasing area at a location of high particle velocity raises that formant’s frequency.

Formant Lowering Rule: Increasing area at a location of high particle velocity lowers that formant’s frequency.

This explains why specific articulatory maneuvers affect some formants more than others.

Relating Three-Tube Parameters to Vowel Categories

Specific vowel categories correspond to characteristic three-tube configurations.

High Front Vowel /i/

Articulatory Configuration: Tongue raised and fronted, creating narrow constriction near hard palate. Lips spread.

Three-Tube Parameters:

  • Lp = 10-12 cm (long pharynx)
  • Lo = 4-5 cm (short oral cavity)
  • Ll = 0-0.5 cm (minimal lip protrusion)
  • Ap = 2-4 cm² (moderate to large pharyngeal area)
  • Ao = 0.3-0.5 cm² (narrow constriction)
  • Al = 3-5 cm² (large lip aperture)

Resulting Formants:

  • F1 ≈ 250-350 Hz (long pharynx, narrow constriction)
  • F2 ≈ 2200-2800 Hz (short front cavity)
  • F3 ≈ 3000-3500 Hz

High Back Vowel /u/

Articulatory Configuration: Tongue raised and retracted, creating constriction near velum. Lips rounded and protruded.

Three-Tube Parameters:

  • Lp = 4-6 cm (short pharynx)
  • Lo = 10-12 cm (long oral cavity)
  • Ll = 1.5-2 cm (significant protrusion)
  • Ap = 2-3 cm² (moderate pharyngeal area)
  • Ao = 0.3-0.5 cm² (narrow constriction)
  • Al = 0.5-1 cm² (small rounded aperture)

Resulting Formants:

  • F1 ≈ 250-350 Hz (narrow constriction, though shorter pharynx raises F1 somewhat)
  • F2 ≈ 600-900 Hz (long front cavity, lip rounding lowers F2)
  • F3 ≈ 2100-2500 Hz (lip rounding strongly lowers F3)

Low Front Vowel /a/

Articulatory Configuration: Tongue lowered and somewhat fronted. Jaw maximally open. Lips neutral to slightly spread.

Three-Tube Parameters:

  • Lp = 8-10 cm (moderately long pharynx)
  • Lo = 7-9 cm (moderate oral cavity)
  • Ll = 0.5-1 cm (minimal protrusion)
  • Ap = 2-3 cm² (moderate pharyngeal area)
  • Ao = 2-4 cm² (large area, minimal constriction)
  • Al = 3-5 cm² (large aperture)

Resulting Formants:

  • F1 ≈ 700-1000 Hz (large constriction area, smaller pharynx)
  • F2 ≈ 1100-1500 Hz (intermediate front/back cavity balance)
  • F3 ≈ 2400-2800 Hz

Mid Central Vowel /ə/ (Schwa)

Articulatory Configuration: Neutral tongue position, approximating uniform tube. Minimal constriction.

Three-Tube Parameters:

  • Lp = 8-9 cm (moderate pharynx)
  • Lo = 8-9 cm (moderate oral cavity)
  • Ll = 0.5 cm (minimal lip configuration)
  • Ap = 2-3 cm² (moderate areas throughout)
  • Ao = 2-3 cm²
  • Al = 3-4 cm²

Resulting Formants:

  • F1 ≈ 500-600 Hz
  • F2 ≈ 1400-1600 Hz
  • F3 ≈ 2400-2600 Hz

Close to uniform tube predictions, representing the neutral vocal tract configuration.

Effects of Individual Parameters

Systematically varying each parameter reveals its acoustic consequences.

Constriction Location (Pharyngeal vs. Oral Length)

Moving Constriction Forward (increasing Lp, decreasing Lo):

  • F1: Decreases (longer pharynx acts as larger Helmholtz resonator)
  • F2: Increases (shorter front cavity raises its resonance)
  • F3: Variable depending on detailed configuration

Moving Constriction Backward (decreasing Lp, increasing Lo):

  • F1: Increases (shorter pharynx)
  • F2: Decreases (longer front cavity)
  • F3: Variable

Conclusion: Front-back tongue position (constriction location) primarily controls F2, with secondary effects on F1.

Constriction Degree (Area at Constriction)

Narrowing Constriction (decreasing Ao):

  • F1: Decreases (smaller constriction area increases coupling between cavities)
  • F2: May increase or decrease depending on formant-specific patterns
  • Overall: Greater separation between front and back cavity resonances

Widening Constriction (increasing Ao):

  • F1: Increases (approaches uniform tube behavior)
  • F2: Moves toward neutral values
  • Overall: Formants converge toward uniform tube values

Conclusion: Constriction degree (related to tongue height) primarily controls F1.

Pharyngeal Expansion

Expanding Pharynx (increasing Ap):

  • F1: Decreases (larger back cavity resonator)
  • F2: Smaller effect
  • Overall: Creates “darker” vowel quality

Constricting Pharynx (decreasing Ap):

  • F1: Increases
  • F2: Smaller effect
  • Overall: Creates “brighter,” tenser quality

Conclusion: Advanced Tongue Root (expanding pharynx) systematically lowers F1.

Lip Protrusion and Rounding

Increasing Protrusion (increasing Ll, decreasing Al):

  • F1: Decreases slightly
  • F2: Decreases substantially (200-400 Hz)
  • F3: Decreases substantially (300-500 Hz)
  • Overall: All formants lower, tract lengthening effect

Lip Spreading (Ll → 0, increasing Al):

  • F1: Increases slightly
  • F2: Increases (50-200 Hz)
  • F3: Increases
  • Overall: Opposite of rounding effect

Conclusion: Lip configuration affects all formants, with strongest effects on F2 and F3.

Computational Implementation

Modern analysis uses numerical methods to compute formant frequencies from three-tube parameters.

Transfer Function Approach

The vocal tract transfer function relates glottal volume velocity to radiated sound pressure:

Calculation: For given tube dimensions, compute acoustic impedances of each section, apply junction boundary conditions, determine overall transfer function H(f).

Formants: Appear as peaks (resonances) in |H(f)|, where impedance looking back from lips is minimum.

Advantages: Rigorous, handles any number of tube sections, includes losses and radiation impedance.

Area Function to Formant Mapping

Input: Specify Lp, Lo, Ll, Ap, Ao, Al (six parameters).

Processing:

  1. Compute impedances for each tube section
  2. Apply junction conditions (pressure and volume velocity continuity)
  3. Solve for frequencies where resonance condition is met
  4. Extract first several formant frequencies

Output: F1, F2, F3, F4, … and potentially formant bandwidths.

Inverse Problem: Formants to Area Function

Given target formant frequencies, optimization can find tube parameters:

Objective: Minimize difference between predicted and target formants.

Search Space: Six-dimensional space of tube parameters (or more for more sections).

Result: Estimates of articulatory configuration that would produce target formants.

Application: Helps identify what articulatory adjustment is needed to achieve desired acoustic output.

Clinical and Pedagogical Applications

The three-tube model provides practical guidance for voice training and therapy.

Formant-Targeted Voice Modification

Goal: Modify specific formant to achieve desired acoustic effect.

Strategy: Identify which tube parameter most affects that formant, adjust articulation accordingly.

Examples:

  • Raising F1 (for high notes): Widen constriction (lower tongue), expand pharynx, or open mouth
  • Lowering F2 (for darker tone): Retract tongue or round lips
  • Raising F2 (for brighter tone): Front tongue or spread lips

Understanding Compensatory Articulation

Concept: Multiple articulatory configurations can produce similar formant patterns.

Example: F2 lowering can be achieved by:

  • Retracting tongue (increasing Lo)
  • Rounding lips (increasing Ll, decreasing Al)
  • Some combination

Clinical Relevance: Patients with anatomical or motor limitations can potentially compensate using alternative articulatory strategies to achieve target acoustics.

Vowel Intelligibility in Dysarthria

Problem: Reduced articulatory range in dysarthria compresses vowel space.

Analysis: Three-tube model reveals which articulatory dimensions are most affected:

  • Limited tongue range: Reduced Lp variation, compressed F2 range
  • Limited jaw opening: Reduced Ao variation, compressed F1 range
  • Limited lip rounding: Reduced Ll and Al variation, affects F2/F3

Intervention: Target the most functionally recoverable dimensions to maximize vowel differentiation.

Singing Pedagogy

Vowel Modification Rules: Three-tube model explains traditional pedagogical guidance:

“Modify /i/ toward /e/ on high notes”:

  • Rationale: Increase Ao slightly (widen constriction) to raise F1
  • Acoustic result: F1 rises to match rising F0, improving efficiency

“Drop your jaw”:

  • Increases Ao and overall tract opening
  • Raises F1, useful for high notes
  • May also expand pharynx (lower F1 component), requiring careful balance

“Cover the tone”:

  • Retract tongue slightly (increase Lo), lower larynx (increase all lengths), round lips
  • Lowers F2 and F3, creating darker timbre

Extensions and Limitations

The three-tube model balances simplicity and accuracy, but more complex models exist.

Four and Five-Tube Models

Additional sections can represent:

  • Sublingual cavity: Space under tongue
  • Piriform sinuses: Lateral pharyngeal recesses
  • Nasal cavity: For nasal vowels
  • Separate larynx tube: Below pharynx

Improved Accuracy: Better predictions, particularly for F3 and higher formants.

Cost: More parameters, less intuitive interpretation.

Continuous Area Functions

Instead of Discrete Sections: Specify A(x) as continuous function of distance x along tract.

Methods: Acoustic wave propagation solved numerically along continuously-varying tract.

Data Source: MRI or other imaging provides detailed tract shape.

Advantages: Maximum accuracy, no assumption of cylindrical sections.

Applications: Research, detailed vowel production studies, synthesis of natural-sounding speech.

Non-Cylindrical Sections

Real vocal tract sections are not perfect cylinders:

Elliptical Cross-Sections: Width and height differ, particularly in pharynx and oral cavity.

Effect: Primarily affects higher formants (F4 and above). Lower formants relatively robust to cross-sectional shape.

Modeling: Can use equivalent circular cross-sections with same area for reasonable approximation.

Losses and Damping

Basic three-tube model often assumes lossless propagation:

Reality: Viscous and thermal losses at walls, yielding and radiation losses.

Effect: Formant bandwidths (not just frequencies). Losses create finite Q factor for each formant.

Enhancement: Adding losses improves realism, allows bandwidth prediction, better matches natural speech spectra.

Summary

The three-tube approximation models the vocal tract as three concatenated cylindrical sections: pharyngeal (from glottis to tongue constriction), oral (from constriction to lips), and lip (representing lip protrusion). This model captures the essential acoustic effects of tongue position, constriction location and degree, and lip configuration, providing accurate formant frequency predictions for most vowels. The pharyngeal length and area primarily affect F1, with long pharynx and expanded pharynx lowering F1. The oral cavity length primarily affects F2, with short front cavity (front vowels) raising F2 and long front cavity (back vowels) lowering F2. The lip section affects all formants but especially F2 and F3, with protrusion and rounding lowering these formants substantially.

Specific vowel categories correspond to characteristic three-tube configurations: /i/ has long pharynx, short oral cavity, and minimal lip protrusion, yielding low F1 and high F2; /u/ has short pharynx, long oral cavity, and significant protrusion, yielding low F1 and low F2; /a/ has moderate dimensions with large constriction area, yielding high F1 and intermediate F2. Perturbation theory explains how local area changes affect formants based on the particle velocity distribution for each formant mode, providing intuitive understanding of formant tuning.

The three-tube model informs clinical and pedagogical applications by identifying which articulatory adjustment affects which formant, enabling targeted voice modification strategies. It explains vowel modification rules in singing, compensatory articulation in dysarthria, and formant tuning for acoustic efficiency. Extensions including four or more tube sections, continuous area functions, and inclusion of acoustic losses provide greater accuracy at the cost of increased complexity, but the three-tube model remains valuable for its balance of predictive power and interpretability.


Key Takeaways

  • ✅ The three-tube model divides the vocal tract into pharyngeal, oral, and lip sections, each characterized by length and cross-sectional area
  • ✅ Pharyngeal length and area primarily determine F1, with long/expanded pharynx lowering F1 and short/constricted pharynx raising F1
  • ✅ Oral cavity length primarily determines F2, with short front cavity raising F2 (front vowels) and long front cavity lowering F2 (back vowels)
  • ✅ Lip protrusion and rounding lower all formants but affect F2 and F3 most strongly, typically lowering F2 by 200-400 Hz
  • ✅ Constriction degree (related to tongue height) affects F1, with narrow constriction lowering F1 and wide constriction raising F1
  • ✅ Different vowels correspond to characteristic three-tube configurations: /i/ (long pharynx, short oral), /u/ (short pharynx, long oral, protruded lips), /a/ (moderate dimensions, large constriction area)
  • ✅ Perturbation theory explains that decreasing area at high particle velocity locations raises formant frequency, providing intuitive formant tuning rules
  • ✅ The three-tube model enables targeted voice modification strategies by identifying which articulatory parameter affects which formant

Further Reading

  1. Fant, G. (1960). Acoustic Theory of Speech Production. The Hague: Mouton.
  2. Stevens, K. N. (1998). Acoustic Phonetics. Cambridge, MA: MIT Press.
  3. Story, B. H. (2005). A parametric model of the vocal tract area function for vowel and consonant simulation. Journal of the Acoustical Society of America, 117(5), 3231-3254.
  4. Titze, I. R. (2000). Principles of Voice Production (2nd ed.). Iowa City: National Center for Voice and Speech.
  5. Flanagan, J. L. (1972). Speech Analysis Synthesis and Perception (2nd ed.). New York: Springer-Verlag.
  6. Chiba, T., & Kajiyama, M. (1941). The Vowel: Its Nature and Structure. Tokyo: Kaiseikan.