# Technical Analysis: 4-Stem Audio Continuity & EBU R128 Two-Pass Loudness Engine

**Document ID**: `M3-AUDIO-ANALYSIS-2026-10-08`  
**Specialist**: Teamwork Preview Explorer (M3 4-Stem & EBU R128 Loudness Specialist 2)  
**Target Module**: `05_Production_Pipeline/audio_continuity_engine.py`  
**Related Specs**: `00_Project_Bible/CINEMATIC_AUDIO_PIPELINE.md`, `PROJECT.md`, `AGENTS.md`, `05_Production_Pipeline/production_orchestrator.py`  
**Test Coverage**: `tests/test_tier1_features.py` (F8, F9, F10), `tests/test_tier2_boundaries.py`, `tests/test_tier3_interactions.py`, `tests/test_tier4_workloads.py`

---

## 1. Executive Summary

A comprehensive, line-by-line inspection of `05_Production_Pipeline/audio_continuity_engine.py` and its corresponding test suites revealed two critical architectural deficits:

1. **Deficit 1 — Loudness Normalization is Single-Pass Dynamic, NOT Two-Pass Linear**:
   - `normalize_loudness()` currently executes `loudnorm=I={target_lufs}:TP={target_tp}:LRA={target_lra}` in a single pass.
   - Without Pass 1 measurement parameters (`measured_I`, `measured_TP`, `measured_LRA`, `measured_thresh`, `offset`) and `linear=true`, FFmpeg runs in **dynamic mode**. This functions as an Automatic Gain Control (AGC) and lookahead compressor that causes **audible volume pumping**, compresses subtle musical dynamics (guzheng / tỳ bà transients), and fails to guarantee broadcast-grade consistency across multi-minute scenes.
2. **Deficit 2 — 4-Stem Architecture is Completely Missing in Engine Implementation**:
   - While `00_Project_Bible/CINEMATIC_AUDIO_PIPELINE.md` explicitly codifies the 4-Stem model (Stem 1: BGM, Stem 2: Ambience, Stem 3: Foley, Stem 4: Dialogue), `audio_continuity_engine.py` contains **zero stem definitions** and only offers `layer_bgm_with_ducking()` which mixes at most two audio tracks.
   - There is currently no `mix_four_stems()` method capable of routing continuous Stem 2 (Festival Ambience), ducking both Stem 1 and Stem 2 under Stem 4 dialogue, and applying EQ notch filtering on BGM (800 Hz – 3500 Hz).

All 167 automated tests currently pass because existing assertions primarily verify signatures, specific string tokens (e.g., `"loudnorm=I=-14"`, `"48000"`, `"-stream_loop"`), or return values on mocks. However, fulfilling Milestone M3 to production-grade cinema fidelity requires concrete architectural refactoring of `audio_continuity_engine.py`.

---

## 2. Deep Inspection of `audio_continuity_engine.py`

### 2.1. Current Inventory of Methods in `AudioContinuityEngine`

| Method | Current Implementation | Status / Gaps |
| :--- | :--- | :--- |
| `inspect_shot_audio` | Probes audio stream with ffprobe & measures max/mean dB with `volumedetect`. | Functional, robust. |
| `inspect_sequence` | Iterates over shots, warns on delta mean volume > 3.0 dB. | Functional, satisfies F8.5. |
| `stitch_with_audio_crossfade` | Chains `acrossfade` filters with `c1=qsin:c2=qsin` (or single file `anull`). Appends `loudnorm=I=-14:TP=-1.0:LRA=9`. | Functional, satisfies F8.1-F8.3. |
| `create_l_cut_bridge` | L-Cut audio overlap (3.5s default) with `lowpass=f=3500` filter. | Functional cinematic effect. |
| `normalize_loudness` | Single-pass FFmpeg call with `-af loudnorm=I=-14:TP=-1.0:LRA=9`. Stream copy with fallback to re-encode. | **FLAWED**: Single-pass dynamic, not two-pass linear. |
| `layer_bgm_with_ducking` | 2-track mixer: input video audio + BGM with sidechain compression. | **INCOMPLETE**: Does not support Stem 2 Ambience or 4 stems. |

### 2.2. Flaw Analysis: Single-Pass vs. Two-Pass EBU R128 Normalization

#### Why Single-Pass Fails Cinema Standards:
In `audio_continuity_engine.py` lines 286-294:
```python
cmd = [
    self.ffmpeg, "-y",
    "-i", input_video,
    "-af", f"loudnorm=I={target_lufs}:TP={target_tp}:LRA={target_lra}",
    "-c:v", "copy",
    "-c:a", "aac",
    "-b:a", "192k",
    "-ar", "48000",
    output_video
]
```
1. **Dynamic Pumping Artifacts**: Because FFmpeg has not scanned the file ahead of time, it dynamically reacts to incoming audio levels using a sliding window buffer. When a dialogue line finishes and silence follows, background noise or room ambience is artificially amplified ("pumped").
2. **Dynamic Range Distortion**: The intended dynamic range (LRA 9-11 LU) is altered dynamically. In quiet moments (Thúy Kiều mourning at Đạm Tiên's grave), the audio level is aggressively boosted; during loud orchestral peaks, it is heavily compressed.
3. **True Peak Non-Linearity**: Single-pass limiting can introduce harmonic distortion when trying to catch unexpected peaks without prior lookahead knowledge of the entire waveform.

#### The Required Two-Pass Linear Architecture:
- **Pass 1 (Measurement)**:
  ```bash
  ffmpeg -hide_banner -y -i input.mp4 -vn -sn -dn -af loudnorm=I=-14.0:TP=-1.0:LRA=9.0:print_format=json -f null -
  ```
  Parses JSON from FFmpeg stderr:
  ```json
  {
      "input_i" : "-21.75",
      "input_tp" : "-17.69",
      "input_lra" : "0.00",
      "input_thresh" : "-31.75",
      "output_i" : "-13.99",
      "output_tp" : "-9.90",
      "output_lra" : "0.00",
      "output_thresh" : "-23.99",
      "normalization_type" : "dynamic",
      "target_offset" : "-0.01"
  }
  ```
- **Edge-Case Safety Gate**:
  If the input is pure silence, FFmpeg returns `"input_i": "-inf"`.
  *Empirical test result*: Passing `measured_I=-inf` into Pass 2 causes FFmpeg to crash with:
  `[Parsed_loudnorm_0 @ ...] Value -inf for parameter 'measured_I' out of range [-99 - 0]`.
  Therefore, Pass 1 parsing MUST validate that `input_i != "-inf"` and `float(input_i) >= -99.0`. If silent or invalid, the engine must safely bypass or fallback to single-pass.
- **Pass 2 (Linear Normalization)**:
  ```bash
  ffmpeg -y -i input.mp4 -af loudnorm=I=-14.0:TP=-1.0:LRA=9.0:measured_I=-21.75:measured_TP=-17.69:measured_LRA=0.00:measured_thresh=-31.75:offset=-0.01:linear=true -c:v copy -c:a aac -b:a 192k -ar 48000 output.mp4
  ```
  Setting `linear=true` applies a uniform static gain offset (`offset`), strictly preserving dynamic relationships while bringing integrated loudness to `-14.0 LUFS` and true peak to `< -1.0 dBTP`.

### 2.3. Flaw Analysis: Missing 4-Stem Architecture

In `00_Project_Bible/CINEMATIC_AUDIO_PIPELINE.md`:
- **Stem 1: BGM & Cinematic Score**:
  - Continuous pad/drone per scene (50s - 180s).
  - EQ notch at **800 Hz – 3500 Hz** (`-3dB` to `-4.5dB`) to leave acoustic space for Kiều's dialogue.
- **Stem 2: Ambience & Environmental Soundscape**:
  - 3D physical environment (Qingming festival crowd chatter, river stream, temple bells).
  - Continuous across multi-shot transitions.
- **Stem 3: Foley & Spot SFX**:
  - Silk cloth rustle (Giao Lĩnh), jade hairpin clinks, porcelain cups, horse hooves, zither string snap.
- **Stem 4: Dialogue & Studio Voice-Over**:
  - Spoken lines & poetry recitation.
  - Controls dynamic sidechain ducking: ducks both **Stem 1** and **Stem 2** by `-12 dB` to `-18 dB` (Attack 50ms, Release 300ms).

`audio_continuity_engine.py` currently has **no support** for Stem 2, Stem 3, or Stem 4 independently, nor does it offer a multi-input filtergraph to combine them with sidechain ducking.

---

## 3. Test Suite Audit & Compatibility Safeguards

An inspection of `tests/` identified specific AST string checks that the Worker must strictly preserve:

| Test Name | File | Exact Assertion / Constraint | Impact on Worker |
| :--- | :--- | :--- | :--- |
| `test_f8_01` | `test_tier1_features.py:410` | `hasattr(engine, "stitch_with_audio_crossfade")` | Method name must remain identical. |
| `test_f8_02` | `test_tier1_features.py:417` | `"qsin" in src or "cbrt" in src` in `stitch_with_audio_crossfade` | Must keep equal-power curve. |
| `test_f8_03` | `test_tier1_features.py:425` | `sig.parameters["crossfade_dur"].default == 1.0` | Default parameter must be 1.0s. |
| `test_f8_05` | `test_tier1_features.py:439` | `"3.0" in src` in `inspect_sequence` | Must keep 3.0 dB delta threshold. |
| `test_f9_01` | `test_tier1_features.py:449` | Asset glob in `04_Assets/audio/` or "stem 2" in `CINEMATIC_AUDIO_PIPELINE.md` | Ambience bed file or spec must exist. |
| `test_f9_03` | `test_tier1_features.py:467` | `"48000" in src` of `AudioContinuityEngine` | Must enforce 48000 Hz sample rate. |
| `test_f9_04` | `test_tier1_features.py:474` | "Stem 1", "Stem 2", "Stem 3", "Stem 4" in `CINEMATIC_AUDIO_PIPELINE.md` | Doc compliance. |
| `test_f9_05` | `test_tier1_features.py:484` | `hasattr(engine, "layer_bgm_with_ducking")` | Method must be retained. |
| `test_f10_01` | `test_tier1_features.py:493` | `sig.parameters["target_lufs"].default == -14.0` | Default parameter must be -14.0. |
| `test_f10_02` | `test_tier1_features.py:500` | `sig.parameters["target_tp"].default == -1.0` | Default parameter must be -1.0. |
| `test_f10_03` | `test_tier1_features.py:507` | `sig.parameters["target_lra"].default == 9.0` (±2.0) | Default parameter must be 9.0. |
| `test_f10_04` | `test_tier1_features.py:514` | `"loudnorm=" in src and "target_lufs" in src` in `normalize_loudness` | Source inspection check. |
| `test_f10_b04` | `test_tier2_boundaries.py:467`| `engine.normalize_loudness("nonexistent_short.mp4", "out.mp4") == False` | Must return False on non-existent input. |
| `test_f10_b05` | `test_tier2_boundaries.py:474`| `"48000" in src` of `normalize_loudness` | Must enforce 48000 Hz. |
| `test_f8_f10` | `test_tier3_interactions.py:116`| `"loudnorm=I=-14" in src` of `stitch_with_audio_crossfade` | Must retain exact string token. |
| `test_f9_f10` | `test_tier3_interactions.py:123`| `"loudnorm=I=-14" in src` of `layer_bgm_with_ducking` | Must retain exact string token. |
| `test_f9_b04` | `test_tier2_boundaries.py:435`| `"-stream_loop" in src and "-1" in src` in `layer_bgm_with_ducking` | Must retain continuous loop. |
| `test_workload_03` | `test_tier4_workloads.py:70` | `normalize_loudness(input, output, target_lufs=-14.0)` succeeds and `ebur128` measures `Integrated loudness:` | Must produce valid output measured by `ebur128`. |

---

## 4. Architectural Recommendations for the Worker

### 4.1. Step 1: Implement `measure_loudness()` (Pass 1)
Add a public method to `AudioContinuityEngine`:
```python
def measure_loudness(
    self,
    media_path: str,
    target_lufs: float = -14.0,
    target_tp: float = -1.0,
    target_lra: float = 9.0
) -> Optional[Dict]:
    """
    Pass 1 EBU R128 Loudness Measurement via FFmpeg loudnorm JSON.
    Returns dictionary with: input_i, input_tp, input_lra, input_thresh, target_offset, is_silent.
    Returns None if file does not exist or measurement fails.
    """
```
- Executes FFmpeg with `-vn -sn -dn -af "loudnorm=...:print_format=json" -f null -`.
- Extracts JSON block from stderr.
- Flags `is_silent = True` if `input_i == "-inf"` or `float(input_i) < -99.0`.

### 4.2. Step 2: Refactor `normalize_loudness()` to Two-Pass Linear
Update `normalize_loudness()` signature to:
```python
def normalize_loudness(
    self,
    input_video: str,
    output_video: str,
    target_lufs: float = -14.0,
    target_tp: float = -1.0,
    target_lra: float = 9.0,
    two_pass: bool = True
) -> bool:
```
- Preserves all default argument values.
- When `two_pass=True`, invokes `measure_loudness()`.
- If measurement is valid and non-silent, constructs Pass 2 linear filter:
  `loudnorm=I={target_lufs}:TP={target_tp}:LRA={target_lra}:measured_I={input_i}:measured_TP={input_tp}:measured_LRA={input_lra}:measured_thresh={input_thresh}:offset={target_offset}:linear=true`.
- If measurement fails or is silent, falls back to single-pass `loudnorm=I={target_lufs}:TP={target_tp}:LRA={target_lra}`.
- Tries stream copy first (`-c:v copy`), falling back to re-encode (`-c:v libx264 -crf 18 -preset slow`) if container stream copy fails.

### 4.3. Step 3: Implement `mix_four_stems()`
Add full 4-Stem mixing method:
```python
def mix_four_stems(
    self,
    video_path: str,
    output_path: str,
    stem1_bgm: Optional[str] = None,
    stem2_ambience: Optional[str] = None,
    stem3_foley: Optional[str] = None,
    stem4_dialogue: Optional[str] = None,
    bgm_vol: float = 0.35,
    ambience_vol: float = 0.35,
    foley_vol: float = 0.80,
    dialogue_vol: float = 1.0,
    duck_db: float = -14.0,
    target_lufs: float = -14.0,
    target_tp: float = -1.0,
    target_lra: float = 9.0,
    two_pass: bool = True
) -> bool:
```
- Manages inputs dynamically:
  - Video file (Input 0)
  - Stem 1 (BGM): looped with `-stream_loop -1`, volume adjusted, EQ notch filter at 800Hz - 3500Hz:
    `equalizer=f=2150:width_type=h:width=2700:g=-3.5`
  - Stem 2 (Ambience): looped with `-stream_loop -1`, volume adjusted
  - Stem 3 (Foley): spot SFX, volume adjusted
  - Stem 4 (Dialogue): voice track, volume adjusted
- Sidechain ducking logic:
  If Dialogue (Stem 4 or on-set dialogue) is present, applies `sidechaincompress=threshold=0.08:ratio=4:attack=50:release=300` to both Stem 1 and Stem 2.
- Mixes all active streams with `amix=inputs={n}:duration=first:dropout_transition=2`.
- Normalizes final output with two-pass linear EBU R128.

### 4.4. Step 4: Generate Stem 2 Continuous Festival Ambience Bed Asset
Provide a generation script to produce:
`04_Assets/audio/festival_crowd_ambience_bed.wav`
(48,000 Hz, 16-bit stereo, duration 60.0s, seamless loopable ambience bed representing Qingming crowd chatter, gentle spring wind, distant temple chimes).

### 4.5. Step 5: Update `production_orchestrator.py`
In `master_scene_audio()`:
- Add support for resolving Stem 2 Ambience (e.g. `festival_crowd_ambience_bed.wav` for Scene 01, 02, 03).
- Utilize the upgraded two-pass normalization and multi-stem mixing capabilities.

---

## 5. Verification Checklist for Worker

1. `pytest tests/test_tier1_features.py -k "f8 or f9 or f10" -v` must pass 100% (15/15 tests).
2. `pytest tests/test_tier2_boundaries.py -k "f8 or f9 or f10" -v` must pass 100% (15/15 tests).
3. `pytest tests/test_tier3_interactions.py -k "f8 or f9 or f10" -v` must pass 100% (3/3 tests).
4. `pytest tests/test_tier4_workloads.py -k "03 or 04" -v` must pass 100% (2/2 tests).
5. Full test suite: `pytest tests -q` must achieve 167/167 passes (0 failures).
