# TECHNICAL INVESTIGATION REPORT: M3 AUDIO ENGINE & CROSSFADE CONCATENATION

**Project**: Thập Ngũ Niên (The Fifteen Springs) — EP01 Production & Universal Standards  
**Target File**: `05_Production_Pipeline/audio_continuity_engine.py`  
**Related Files**: `05_Production_Pipeline/production_orchestrator.py`, `05_Production_Pipeline/assemble_ep01_feature.py`, `tests/test_tier1_features.py`, `tests/test_tier2_boundaries.py`, `tests/test_tier3_interactions.py`, `tests/test_tier4_workloads.py`  
**Date**: 2026-10-08  
**Investigator**: teamwork_preview_explorer (M3 Audio Engine Specialist 1)

---

## 1. Executive Summary

Milestone M3 establishes the 4-Stem Audio Continuity Engine and EBU R128 broadcast mastering for the 180-minute AI cinema production of "Thập Ngũ Niên". The core engine implemented in `05_Production_Pipeline/audio_continuity_engine.py` provides equal-power crossfading (`qsin`), EBU R128 loudness normalization (-14 LUFS), dynamic sidechain ducking, and J-Cut/L-Cut acoustic bridges.

However, an exhaustive empirical inspection revealed **four critical vulnerabilities** in the current crossfade and concatenation pipeline:
1. **Cumulative Audio/Video Duration Desynchronization**: Chaining FFmpeg's `acrossfade=d=D` filter across $N$ shots shortens total audio duration by $(N - 1) \cdot D$ seconds relative to the video duration. In an 8-shot scene with $D = 1.0$s, audio is truncated by 7.0 seconds; by shot 8, audio begins 7.0 seconds ahead of the picture, causing severe dialogue desynchronization and a dead silence dropout at the scene tail.
2. **Missing Audio Stream Fatal Crash**: If any input video lacks an audio track (0 audio streams), `stitch_with_audio_crossfade` crashes with `Stream specifier ':a' in filtergraph description matches no streams. Error binding filtergraph inputs/outputs: Invalid argument`.
3. **Channel Downgrade on Mixed Streams**: If Shot 1 is mono and Shot 2 is stereo, FFmpeg's `concat` filter silently downgrades the entire scene master to mono.
4. **Passive RMS Warning Without Active Gain Staging**: While `inspect_sequence` detects volume deltas > 3.0 dB, `stitch_with_audio_crossfade` performs zero gain adjustment before crossfading. Clips with large volume disparities suffer severe compressor pumping and loudness jumps.
5. **Orchestrator Disconnect**: `production_orchestrator.py::concat_scene_shots` does not utilize `AudioContinuityEngine.stitch_with_audio_crossfade`, falling back to raw hard cuts. Furthermore, neither orchestrator nor feature assembler follows the mandatory `_v*` file versioning standard.

---

## 2. Deep Dive: `audio_continuity_engine.py` Implementation

### 2.1 Component Breakdown

| Method | Role | Current Status | Issues Found |
| :--- | :--- | :--- | :--- |
| `inspect_shot_audio` | Probes stream codec, channels, sample rate, max/mean dB | Functional via `ffprobe` & `volumedetect` | None; handles missing file and 0 streams correctly. |
| `inspect_sequence` | Audits adjacent shot volume deltas | Functional (warns if $\Delta > 3.0$ dB) | Advisory only; does not return actionable gain staging offsets. |
| `stitch_with_audio_crossfade` | Chains `acrossfade` + video `concat` + `loudnorm` | Passes basic tests | Truncates audio by $(N-1)D$, crashes on 0 audio streams, ignores mono/stereo mismatches. |
| `normalize_loudness` | EBU R128 broadcast normalization | Functional (-14 LUFS, -1.0 dBTP) | Uses single-pass dynamic `loudnorm` rather than two-pass linear normalization. |
| `layer_bgm_with_ducking` | Dynamic sidechain ducking of Stem 1 under Stem 3/4 | Functional (-14dB ducking) | Crashes if input video has 0 audio streams (`[0:a]` specifier missing). |
| `create_l_cut_bridge` | Acoustic lowpass filter (3500Hz) for spatial bridge | Functional | Hardcodes 10s shot duration assumption (`atrim=start=6.5:end=10.0`). |

---

## 3. Audio Crossfade Concatenation & The Duration Shrinkage Phenomenon

### 3.1 Mathematical & Physical Root Cause
In FFmpeg, the `acrossfade=d=D:c1=qsin:c2=qsin` filter overlaps the tail of stream $A$ with the head of stream $B$ by duration $D$. Consequently:
$$\text{Duration}(A \oplus_{\text{acrossfade}} B) = \text{Duration}(A) + \text{Duration}(B) - D$$

When concatenating $N$ shots of duration $T$ each:
- Video duration (via `concat=n=N:v=1:a=0`):
  $$T_{\text{video}} = \sum_{i=1}^N T_i = N \times T$$
- Audio duration (via chained `acrossfade`):
  $$T_{\text{audio}} = \sum_{i=1}^N T_i - (N - 1) \cdot D = N \times T - (N - 1) \cdot D$$

### 3.2 Empirical Verification
We executed an empirical test using two synthetic 5.0-second video clips (`v1.mp4`, `v2.mp4`) through `stitch_with_audio_crossfade(crossfade_dur=1.0)`:
```json
{
  "streams": [
    {"codec_type": "video", "duration": "10.000000"},
    {"codec_type": "audio", "duration": "9.000000"}
  ]
}
```
**Result**:
- Video duration = 10.000s
- Audio duration = 9.000s (exactly 1.000s missing at the tail)
- In Shot 2, the audio begins at $t = 4.0$s instead of $t = 5.0$s (1.0s before the visual cut).
- Across a full scene of 10 shots ($10 \times 10s = 100s$), the audio ends at $91.0s$, and Shot 10 audio plays during Shot 9!

### 3.3 Architectural Solution
In cinema post-production, shot-to-shot cutting requires two distinct modes:
1. **Mode A: Sample-Accurate Boundary Micro-Crossfade (Straight Cuts / Default)**
   - To eliminate click/pop and DC offset thumps at shot cuts without any timeline shrinkage:
   - Apply micro equal-power fades: `afade=t=out:st={T_i - \delta}:d={\delta}:curve=qsin` at the tail and `afade=t=in:st=0:d={\delta}:curve=qsin` at the head ($\delta = 30$ms).
   - Concatenate via `concat=n=N:v=1:a=1`.
   - **Empirical result**: Video = 10.000000s, Audio = 10.000000s. Zero drift, zero lip-sync mismatch, zero clicks/pops!
2. **Mode B: True Transition Crossfade with Video Alignment or Tail Padding**
   - When large crossfades ($D \ge 0.5$s) are desired:
   - If visual transitions are used: use `xfade` on video with `offset = T_1 - D` so both video and audio overlap identically.
   - If visual is hard-cut: pad the audio tail using `apad` to guarantee $T_{\text{audio}} = T_{\text{video}}$, avoiding premature audio freeze.

---

## 4. Active RMS Volume Staging (Gain Staging)

### 4.1 The Flaw in Current Passive Thresholding
`inspect_sequence` performs `volumedetect` and logs:
```python
delta_mean = abs(r["mean_volume_db"] - prev["mean_volume_db"])
if delta_mean > 3.0:
    print(f"⚠️ CẢNH BÁO LỆCH ÂM LƯỢNG: Chênh {delta_mean:.1f} dB (Cần Audio Normalize)")
```
However, `stitch_with_audio_crossfade` takes no action on this warning. When Shot 1 (mean volume -32 dB) meets Shot 2 (mean volume -14 dB), the sudden 18 dB step change enters `acrossfade`. Equal-power curves (`qsin`) cannot compensate for an 18 dB power disparity. The subsequent `loudnorm` filter then applies sudden pumping attenuation.

### 4.2 Proposed Active RMS Gain Staging Engine
Before crossfading:
1. Measure `mean_volume_db` for each shot $i$.
2. Compute reference target $R_{\text{target}} = -24.0$ dBFS (or sequence median of non-silent shots).
3. Compute gain adjustment: $\Delta G_i = R_{\text{target}} - \text{mean\_volume\_db}_i$.
4. Safety clamping: Clamp $\Delta G_i \in [-12.0, +6.0]$ dB, ensuring $\text{max\_volume\_db}_i + \Delta G_i \le -1.0$ dBFS (headroom protection).
5. Apply per-stream gain: `[a{i}]volume={gain_adj}dB[a{i}_staged]` before crossfading.

---

## 5. Sample Rate, Channel Layout, and Missing Audio Resilience

### 5.1 Missing Audio Stream Vulnerability
When a clip without audio is passed to `stitch_with_audio_crossfade`:
- FFmpeg outputs: `Stream specifier ':a' in filtergraph description matches no streams.`
- The entire concat process aborts with returncode 1.
- **Fix**: Pre-inspect each input with `ffprobe`. For clips with `has_audio == False`:
  Inject `aevalsrc=0:d={dur}:s=48000:c=stereo[a{i}_norm]` into the filtergraph.

### 5.2 Mixed Channel Layout (Mono vs. Stereo) Vulnerability
When Shot 1 is mono and Shot 2 is stereo:
- FFmpeg's `concat` adopts the channel layout of stream 0 (`mono`).
- The entire output is converted to mono, destroying the stereo panning of Stem 2 and Stem 3.
- **Fix**: Pre-condition every audio stream:
  `[{i}:a]aresample=48000,aformat=sample_rates=48000:channel_layouts=stereo[a{i}_norm]`

---

## 6. EBU R128 Normalization: Single-Pass vs. Two-Pass Linear

### 6.1 Single-Pass Limitations
Single-pass `loudnorm=I=-14:TP=-1.0:LRA=9` employs real-time dynamic gain compression. High-energy transients (thunder, sword clashes) trigger abrupt ducking, followed by slow release breathing.

### 6.2 Verified Two-Pass Linear Implementation
In two-pass normalization:
1. **Pass 1 (Analysis)**:
   ```bash
   ffmpeg -y -i input.mp4 -af loudnorm=I=-14:TP=-1.0:LRA=9:print_format=json -f null -
   ```
   Extract `input_i`, `input_tp`, `input_lra`, `input_thresh`, `target_offset`.
2. **Pass 2 (Linear Normalization)**:
   ```bash
   ffmpeg -y -i input.mp4 -af loudnorm=I=-14:TP=-1.0:LRA=9:measured_I={input_i}:measured_TP={input_tp}:measured_LRA={input_lra}:measured_thresh={input_thresh}:offset={target_offset}:linear=true -c:v copy -c:a aac -b:a 192k -ar 48000 output.mp4
   ```
Empirical testing confirmed: Exit code 0, linear application, zero dynamic compression artifacts.

---

## 7. Pipeline Integration & Versioning Compliance

1. **`production_orchestrator.py`**:
   - `concat_scene_shots` builds hard-cut FFmpeg strings manually.
   - It should import and call `AudioContinuityEngine.stitch_with_audio_crossfade`.
   - Default output paths should automatically resolve version suffixes (`_v1.mp4`, `_v2.mp4`).
2. **`assemble_ep01_feature.py`**:
   - Assembles grand feature master using manual concat.
   - Output filename must be `thap_ngu_nien_ep01_grand_master_v1.mp4`.

---

## 8. Test Suite Analysis (`test_f8_*` across Tiers 1-4)

### 8.1 Current Test Coverage Matrix
- **Tier 1 (Features)**:
  - `test_f8_01`: Verifies method availability and callability. (PASS)
  - `test_f8_02`: Verifies presence of `"qsin"` or `"cbrt"` curve in source. (PASS)
  - `test_f8_03`: Verifies default `crossfade_dur == 1.0`. (PASS)
  - `test_f8_04`: Verifies `callable(concat_scene_shots)`. (PASS)
  - `test_f8_05`: Verifies warning threshold at `"3.0"` dB. (PASS)
- **Tier 2 (Boundaries)**:
  - `test_f8_b01`: Empty list returns False. (PASS)
  - `test_f8_b02`: Single file passthrough. (PASS)
  - `test_f8_b03`: Nonexistent file returns `has_audio=False`. (PASS)
  - `test_f8_b04`: Zero stream file returns -99.0 dB. (PASS)
  - `test_f8_b05`: Duration $\le 0$ rejected. (PASS)
- **Tier 3 (Interactions)**:
  - `test_f8_f9`: Verifies crossfade and ducking coexistence. (PASS)
  - `test_f8_f10`: Verifies `loudnorm=I=-14` in crossfade source. (PASS)
- **Tier 4 (Workloads)**:
  - `test_workload_03`: Synthetic 2-pass loudness normalization. (PASS)
  - `test_workload_04`: Synthetic 2-shot crossfade assembly ($D=0.5$s). (PASS)

### 8.2 Testing Blind Spots (Gaps to Harden)
The existing test suite does NOT test:
1. Concat with a mute video (0 audio streams) -> Currently crashes in production.
2. Channel layout mismatch (mono + stereo) -> Currently silently degrades to mono.
3. Sample rate mismatch (44.1kHz + 48kHz) -> Currently risks resampling drift.
4. Audio/video duration divergence test ($|T_{\text{video}} - T_{\text{audio}}| < 0.05$s).
5. Active gain staging verification when adjacent shot delta exceeds 3.0 dB.
