# BÁO CÁO PHÂN TÍCH CHUYÊN SÂU: ĐIỀU PHỐI M3 & TÀI NGUYÊN ÂM CẢNH (M3 ORCHESTRATOR & AMBIENCE ASSETS)

**Tác giả**: Teamwork Preview Explorer (M3 Orchestrator & Ambience Asset Specialist 3)  
**Thời gian**: 2026-10-08T06:43:00Z  
**Phạm vi**: `05_Production_Pipeline/production_orchestrator.py`, `05_Production_Pipeline/audio_continuity_engine.py`, `04_Assets/audio_sfx/`, `04_Assets/audio/`, `PROJECT.md`

---

## 1. TỔNG QUAN KẾT QUẢ ĐIỀU TRA (EXECUTIVE SUMMARY)

1. **`production_orchestrator.py` KHÔNG mặc định sử dụng crossfade**: Lệnh `--concat-scene` (hàm `concat_scene_shots`) hiện đang dùng bộ lọc nối cứng thô sơ (`concat=n=N:v=1:a=1`), hoàn toàn **bỏ qua** phương thức `stitch_with_audio_crossfade()` của `AudioContinuityEngine`. Các shot video ghép lại có nguy cơ cao bị tiếng giật cục (click/pop) và hẫng âm lượng.
2. **Thiếu hỗ trợ Stem 2 trong `--master-audio`**: Lệnh `--master-audio` chỉ nhận `--bgm` (Stem 1), hoàn toàn không có tham số `--ambience` để trải thảm âm cảnh môi trường (Stem 2). Ngoài ra, thuật toán chuẩn hóa hiện chỉ là 1-pass `loudnorm` thay vì 2-pass EBU R128 tiêu chuẩn phát sóng.
3. **Thư mục `04_Assets/audio_sfx/` hoàn toàn rỗng**: Chưa có bất kỳ tệp âm thanh nào cho âm cảnh đám đông hội xuân (Stem 2) hay tiếng môi trường cho Cảnh 04 (12 shots = 120 giây).
4. **Trạng thái Cảnh 04 hiện tại**: Mới chỉ có 2/12 shot video (`shot01` và `shot02`), âm thanh gốc của Muse.ai là mono 24kHz với độ chênh âm lượng 3.1 dB (vượt ngưỡng cảnh báo 3.0 dB của `inspect_sequence`).
5. **Cập nhật Tiến độ `PROJECT.md`**: Milestone 2 (M2) đã hoàn thành 100% (toàn bộ 188 shots trong cả 2 registry prompt đã được hợp nhất, 155/155 test suites PASS). Cần cập nhật M2 -> `DONE` và M3 -> `IN_PROGRESS`.

---

## 2. PHÂN TÍCH CHI TIẾT `production_orchestrator.py`

### 2.1. Kiểm tra Lệnh `--concat-scene` (`concat_scene_shots`)
Tại các dòng 350-405 của `production_orchestrator.py`:
```python
def concat_scene_shots(scene_id: str, output_path: Optional[str] = None) -> Optional[Path]:
    ...
    # Xây dựng lệnh FFmpeg filter_complex concat
    inputs = []
    filter_parts = []
    for i, v in enumerate(video_files):
        inputs.extend(["-i", str(v)])
        filter_parts.append(f"[{i}:v:0][{i}:a:0]")
    
    n = len(video_files)
    concat_filter = f"{''.join(filter_parts)}concat=n={n}:v=1:a=1[v][a]"
```

**Nhận định đối chiếu**:
- Hàm này sử dụng hard concat `concat=n={n}:v=1:a=1[v][a]`.
- Mặc dù đầu file có nạp `from audio_continuity_engine import AudioContinuityEngine`, nhưng hàm `concat_scene_shots` không hề gọi đến `AudioContinuityEngine.stitch_with_audio_crossfade`.
- Thiếu tùy chọn thời lượng crossfade (`crossfade_dur`) và tự động kiểm tra gain staging.
- Đặt tên file xuất: Mặc định là `{scene_id}_master.mp4`, chưa tuân thủ quy tắc hậu tố phiên bản `_v*` của Đạo diễn (`_v1.mp4`, `_v2.mp4`...).

### 2.2. Kiểm tra Lệnh `--master-audio` (`master_scene_audio`)
Tại các dòng 407-451 của `production_orchestrator.py`:
```python
def master_scene_audio(scene_id: str, bgm_path: Optional[str] = None, output_path: Optional[str] = None):
    ...
    temp_norm = VIDEOS_DIR / f"{scene_id}_audio_norm_temp.mp4"
    success = engine.normalize_loudness(str(master_video), str(temp_norm), target_lufs=-14.0)
    ...
    if bgm_path and os.path.exists(bgm_path):
        engine.layer_bgm_with_ducking(str(temp_norm), bgm_path, str(final_out), duck_db=-14.0)
```

**Nhận định đối chiếu**:
- Chỉ nhận duy nhất `bgm_path` (Stem 1). Không có tham số `ambience_path` (Stem 2) để lồng tiếng đám đông lễ hội / tiếng suối / gió rặng liễu.
- Chuẩn hóa âm lượng sử dụng 1-pass `loudnorm=I=-14:TP=-1.0:LRA=9`, có thể gây hiện tượng méo động học (dynamic pumping). Cần nâng cấp lên 2-pass `loudnorm`.
- Chưa có cơ chế quản lý versioning cho bản cinematic master cuối cùng (`{scene_id}_cinematic_master_v1.mp4`).

---

## 3. KHẢO SÁT TÀI NGUYÊN ÂM THANH (`04_Assets/audio_sfx/` & STEMS)

### 3.1. Bảng Hiện Trạng Tài Nguyên Âm Thanh Toàn Dự Án
| Vị Trí Lưu Trữ | Tên File | Thuộc Tính Kỹ Thuật | Phân Loại Stem | Đánh Giá Hiện Trạng |
| :--- | :--- | :--- | :--- | :--- |
| `04_Assets/audio/` | `scene01_bgm.m4a` | Opus, 48kHz, Stereo, 178.88s, Mean: -18.0 dB | Stem 1 (BGM) | Đạt chuẩn thảm nhạc Cảnh 01 |
| `04_Assets/archive/audio/` | `scene01_voiceover.mp3` | MP3, 44.1kHz, Mono, 15.2s | Stem 4 (Voice) | Lưu kho kiểm thử |
| `04_Assets/archive/audio/` | `test_kieu_voice.mp3` | MP3, 44.1kHz, Mono, 7.6s | Stem 4 (Voice) | Lưu kho kiểm thử |
| `04_Assets/audio_sfx/` | *(Trống - 0 files)* | N/A | Stem 2 & Stem 3 | **THIẾU HOÀN TOÀN** |

### 3.2. Yêu Cầu Cụ Thể Cho Cảnh 04 (Hội Đạp Thanh - 12 Shots / 120s)
Theo kịch bản `FilmMaker/TAP_01_XUAN_SAC_THE_NGUYEN_VA_GIONG_BAO_DOAN_TRUONG.md` (dòng 538-655):
- Cảnh 04 bao gồm 12 shots (`ep01_scene04_shot01` đến `shot12`), tổng thời lượng **120 Giây** (2 phút).
- Bối cảnh: Đường ngoại ô kinh thành, tiết Thanh minh nườm nượp dòng người trẩy hội.
- Stem 2 (Ambience) bắt buộc:
  - Tiếng ồn ào đám đông lễ hội liên tục (crowd chatter, laughter).
  - Tiếng vó ngựa và tiếng bánh xe gỗ lộc cộc trên đường sỏi đỏ (*"Ngựa xe như nước"*).
  - Tiếng gió xuân xào xạc rặng liễu, tiếng rao hàng quán trà ngoại ô.
- Đạt chuẩn kỹ thuật:
  - Định dạng: WAV / AAC, 48.000 Hz, 2 Channels (Stereo).
  - Thời lượng: >= 120 giây (hoặc hỗ trợ seamless loop `-stream_loop -1`).
  - Âm lượng trung bình: -24.0 dB đến -22.0 dB (để khi ducking/amix với BGM và thoại không bị lấn át).

---

## 4. MA TRẬN TIẾN ĐỘ & TRẠNG THÁI `PROJECT.md`

| Milestone | Tên Milestone | Trạng thái hiện tại | Trạng thái cập nhật mới | Lý do cập nhật |
| :--- | :--- | :--- | :--- | :--- |
| **M1** | Screenplay & Pacing Expansion | `DONE` | `DONE` | Đã nghiệm thu 188 shots (29m40s) |
| **M2** | Universal Standards & Prompt Registry | `IN_PROGRESS` | **`DONE`** | Toàn bộ prompt Gemini Banana (188 shots) & Muse.ai (188 shots) đã tích hợp đầy đủ Anti-Male-Tears, Lip-sync guard, 100% test suites PASS |
| **M3** | 4-Stem Audio Engine & Ambience Continuity | `PLANNED` | **`IN_PROGRESS`** | Đang triển khai crossfade concat, Stem 2 festival ambience, và 2-pass EBU R128 mastering |

---

## 5. ĐỀ XUẤT GIẢI PHÁP & MÃ NGUỒN CỤ THỂ CHO WORKER

### Đề xuất 1: Nâng cấp `concat_scene_shots` trong `05_Production_Pipeline/production_orchestrator.py`
Mặc định kích hoạt `stitch_with_audio_crossfade` từ `AudioContinuityEngine`, hỗ trợ tùy chọn `--crossfade-dur`, tự động tăng số version `_v*` cho video Master:

```python
def get_next_master_version(scene_id: str, suffix: str = "master") -> Path:
    """Tự động tính toán tên file master kế tiếp theo versioning _v* (vd: ep01_scene04_master_v1.mp4)."""
    pattern = re.compile(rf"^{re.escape(scene_id)}_{suffix}(?:_v(\d+))?\.mp4$", re.IGNORECASE)
    highest_v = 0
    has_unversioned = False
    for f in VIDEOS_DIR.glob(f"{scene_id}_{suffix}*.mp4"):
        m = pattern.match(f.name)
        if m:
            v_str = m.group(1)
            if v_str:
                highest_v = max(highest_v, int(v_str))
            else:
                has_unversioned = True
    next_v = max(highest_v + 1, 2 if has_unversioned and highest_v == 0 else highest_v + 1)
    return VIDEOS_DIR / f"{scene_id}_{suffix}_v{next_v}.mp4"

def concat_scene_shots(
    scene_id: str, 
    output_path: Optional[str] = None, 
    use_crossfade: bool = True, 
    crossfade_dur: float = 1.0
) -> Optional[Path]:
    """
    Ghép nối các shot của Scene thành Master Scene:
    Mặc định sử dụng AudioContinuityEngine với Equal-Power Crossfade (qsin) để loại bỏ click/pop,
    fallback sang FFmpeg naive filter_complex nếu không có engine hoặc chỉ có 1 shot.
    """
    scene_shots = get_shots_for_scene(scene_id)
    if not scene_shots:
        print(f"[!] Không tìm thấy shot nào cho: {scene_id}")
        return None

    video_files = []
    for shot_id, _ in scene_shots:
        v = find_rendered_video(shot_id)
        if not v:
            print(f"[!] Thiếu video cho shot {shot_id}! Chưa thể ghép master.")
            return None
        video_files.append(v)

    target_out = Path(output_path) if output_path else get_next_master_version(scene_id, "master")

    # Ưu tiên sử dụng AudioContinuityEngine với crossfade
    if use_crossfade and AudioContinuityEngine is not None and len(video_files) >= 1:
        print(f"\n🎵 Sử dụng AudioContinuityEngine (Audio Crossfade {crossfade_dur}s) để ghép Master...")
        engine = AudioContinuityEngine()
        success = engine.stitch_with_audio_crossfade(
            [str(v) for v in video_files],
            str(target_out),
            crossfade_dur=crossfade_dur,
            normalize_lufs=True
        )
        if success and target_out.exists():
            print(f"[✓] GHÉP MASTER SEAMLESS THÀNH CÔNG: {target_out}")
            return target_out

    # Fallback FFmpeg standard concat nếu tắt crossfade hoặc engine gặp lỗi
    print(f"\n🎬 Fallback FFmpeg direct filter_complex concat...")
    ffmpeg_exe = get_ffmpeg()
    inputs = []
    filter_parts = []
    for i, v in enumerate(video_files):
        inputs.extend(["-i", str(v)])
        filter_parts.append(f"[{i}:v:0][{i}:a:0]")
    
    n = len(video_files)
    concat_filter = f"{''.join(filter_parts)}concat=n={n}:v=1:a=1[v][a]"
    cmd = [
        ffmpeg_exe, "-y",
        *inputs,
        "-filter_complex", concat_filter,
        "-map", "[v]",
        "-map", "[a]",
        "-c:v", "libx264", "-crf", "18", "-preset", "slow",
        "-c:a", "aac", "-b:a", "192k", "-ar", "48000",
        str(target_out)
    ]
    res = subprocess.run(cmd, capture_output=True, text=True)
    if res.returncode == 0 and target_out.exists():
        print(f"[✓] GHÉP MASTER THÀNH CÔNG: {target_out}")
        return target_out
    return None
```

### Đề xuất 2: Nâng cấp `master_scene_audio` trong `05_Production_Pipeline/production_orchestrator.py`
Bổ sung hỗ trợ `--ambience` (Stem 2) và 2-pass Loudnorm:

```python
def master_scene_audio(
    scene_id: str, 
    bgm_path: Optional[str] = None, 
    ambience_path: Optional[str] = None, 
    output_path: Optional[str] = None
) -> bool:
    """
    Master âm thanh 4-Stem cho Scene video đạt chuẩn YouTube Green Dollar (-14 LUFS):
    - Stem 3/4: Thoại & Foley trong Video gốc
    - Stem 2: Thảm Ambience môi trường (tự lặp qua -stream_loop -1)
    - Stem 1: BGM Cổ phong điện ảnh kèm Dynamic Ducking
    - Chuẩn hóa 2-pass EBU R128 (-14.0 LUFS, -1.0 dBTP, LRA 9-11 LU)
    """
    master_video = find_rendered_video(f"{scene_id}_master")
    if not master_video:
        mv_list = list(VIDEOS_DIR.glob(f"{scene_id}*master*.mp4"))
        if mv_list:
            mv_list.sort(key=lambda x: x.stat().st_mtime, reverse=True)
            master_video = mv_list[0]
        else:
            print(f"[!] Không tìm thấy master video cho: {scene_id}")
            return False

    final_out = Path(output_path) if output_path else EXPORTS_DIR / f"{scene_id}_cinematic_master_v1.mp4"
    EXPORTS_DIR.mkdir(parents=True, exist_ok=True)

    if AudioContinuityEngine is None:
        print("[!] Lỗi: AudioContinuityEngine không khả dụng.")
        return False

    engine = AudioContinuityEngine()
    print(f"\n🎵 BẮT ĐẦU MASTER ÂM THANH 4-STEM CHO: {scene_id}")
    print(f"   • Master Video:  {master_video.name}")
    print(f"   • Stem 2 (Amb):  {Path(ambience_path).name if ambience_path else 'None'}")
    print(f"   • Stem 1 (BGM):  {Path(bgm_path).name if bgm_path else 'None'}")

    return engine.master_scene_multistem(
        str(master_video),
        str(final_out),
        bgm_path=bgm_path,
        ambience_path=ambience_path,
        target_lufs=-14.0
    )
```

### Đề xuất 3: Tạo Script Sinh / Provisioning Âm Cảnh Lễ Hội Cho Cảnh 04
Tạo `05_Production_Pipeline/generate_festival_ambience.py` để sinh ra file âm cảnh mẫu đạt chuẩn kỹ thuật `04_Assets/audio_sfx/ep01_scene04_festival_crowd_ambience_120s.wav`:
- Thời lượng: 120s (khớp 12 shot).
- 48.000 Hz, Stereo, Pink Noise dải tần người + chuông gió + tiếng vó ngựa sỏi đá.
- Đặt tại `04_Assets/audio_sfx/ep01_scene04_festival_crowd_ambience_120s.wav` để phục vụ render và test suite.

---

## 6. KẾ HOẠCH HÀNH ĐỘNG DÀNH CHO WORKER (EXECUTION CHECKLIST)

1. **Bước 1**: Cập nhật `PROJECT.md` dòng 48-49:
   - M2 -> `DONE`
   - M3 -> `IN_PROGRESS`
2. **Bước 2**: Bổ sung phương thức `master_scene_multistem` và `normalize_loudness_2pass` trong `05_Production_Pipeline/audio_continuity_engine.py`.
3. **Bước 3**: Cập nhật `05_Production_Pipeline/production_orchestrator.py`:
   - Hàm `concat_scene_shots` mặc định gọi `engine.stitch_with_audio_crossfade`.
   - Hàm `master_scene_audio` bổ sung tham số `ambience_path`.
   - Parser CLI thêm `--crossfade-dur` và `--ambience`.
4. **Bước 4**: Tạo và chạy `05_Production_Pipeline/generate_festival_ambience.py` để nạp tệp âm cảnh `ep01_scene04_festival_crowd_ambience_120s.wav` vào `04_Assets/audio_sfx/`.
5. **Bước 5**: Chạy `python tests/run_all_tests.py` để xác nhận 100% test suites tiếp tục PASS mà không gây hồi quy (zero regression).
