# Muse2API - AI Agent Integration & API Reference Guide

Tài liệu đặc tả kỹ thuật này được thiết kế chuẩn hóa để **AI Agent** (OpenAI Assistants, LangChain, AutoGen, CrewAI, Antigravity, LlamaIndex, v.v.) có thể đọc hiểu, gọi công cụ (Tool Calling / Function Calling) và tự động tương tác với cổng dịch vụ **Muse2API**.

---

## 1. Tổng quan hệ thống (System Overview)

- **Tên dịch vụ**: Muse2API Gateway
- **Kiến trúc**: Cổng API bất đồng bộ (FastAPI) đóng gói giao diện web muse.ai thành chuẩn tương thích 100% với **OpenAI API**.
- **Địa chỉ dịch vụ (Base URL)**: `http://127.0.0.1:18610` (hoặc `http://localhost:18610`)
- **OpenAI Compatible Endpoint**: `http://127.0.0.1:18610/v1`
- **Phương thức xác thực (Authentication)**:
  - Header: `Authorization: Bearer <API_KEY>`
  - File chứa API Key cục bộ: `c:\Projects\Muse2API\data\api_key`
  - Giá trị mẫu hiện tại: `m2a--nqA-RT1mZ4AWalrdDVzSYkhI9Cksd8O`
- **Cơ chế tải lên & Quản lý Account**: Hệ thống tự động phân bổ request vào bể tài khoản (Account Pool) qua thuật toán xoay vòng (Round-Robin / LRU), tự động xử lý anti-bot Turnstile qua driver ẩn danh `invisible_playwright`.

---

## 2. Danh mục Models & Bí danh (Model Catalog & Aliases)

Hệ thống tự động map tên model từ các chuẩn phổ biến về model nội bộ của Muse.ai:

| Loại hình tác vụ | Model ID nội bộ | Các bí danh hỗ trợ (Aliases) | Mô tả tính năng |
| :--- | :--- | :--- | :--- |
| **Video Generation** | `muse-video` | `sora` | Text-to-Video & Image-to-Video (First-frame animation) |
| **Image Generation** | `muse-image` | `dall-e-3`, `gpt-image-1` | Text-to-Image & Image Editing/Cutout background |
| **Chat / Reasoning** | `muse-chat` | `gpt-4o`, `gpt-4o-mini`, `gpt-4.1`, `gpt-5`, `claude-sonnet-4`, `deepseek-chat` | Hội thoại đa lượt, hỗ trợ đính kèm hình ảnh |

---

## 3. Quy trình làm việc của AI Agent (Agent Workflows)

### 3.1. Tạo Video (Video Generation Workflow - Asynchronous)

Do việc sinh video AI mất từ 30 giây đến 3 phút, API video vận hành theo mô hình **Bất đồng bộ (Async Task Pattern)**:

```
[Agent] ─── POST /v1/videos ───► [Muse2API] (Trả về task_id)
   │
   ▼ (Vòng lặp Polling: mỗi 3-5s)
[Agent] ─── GET /v1/videos/{task_id} ───► [Muse2API]
   │
   ├── status == "pending" / "running" (Tiếp tục đợi)
   ├── status == "succeeded" (Lấy video URL trong result.url)
   └── status == "failed" (Xử lý lỗi / Thử lại)
```

#### Bước 1: Gửi tác vụ tạo video
- **Endpoint**: `POST /v1/videos`
- **Headers**:
  - `Authorization: Bearer <API_KEY>`
  - `Content-Type: application/json`
- **Body JSON**:
  ```json
  {
    "model": "sora",
    "prompt": "A cinematic shot of a futuristic cyberpunk city at rainy night, neon signs reflecting on wet asphalt, 4k resolution",
    "size": "16:9",
    "duration": 5,
    "image": "https://example.com/first_frame.png"
  }
  ```
- **Mô tả tham số**:
  - `prompt` *(string, bắt buộc)*: Mô tả chi tiết cảnh quay bằng tiếng Anh.
  - `model` *(string, tùy chọn)*: `"sora"` hoặc `"muse-video"` (mặc định: `muse-video`).
  - `size` *(string, tùy chọn)*: Tỷ lệ khung hình (`"16:9"`, `"9:16"`, `"1:1"`).
  - `duration` *(integer, tùy chọn)*: Thời lượng từ `1` đến `10` giây (mặc định: `5`).
  - `image` *(string, tùy chọn)*: URL ảnh công khai hoặc base64 data URI để làm khung hình bắt đầu (First-frame / Image-to-Video).
- **Phản hồi mẫu**:
  ```json
  {
    "id": "vid_a1b2c3d4e5",
    "object": "video",
    "status": "pending",
    "progress": 0,
    "created_at": 1728557400,
    "model": "muse-video",
    "result": null,
    "error": null
  }
  ```

#### Bước 2: Kiểm tra tiến độ và nhận kết quả
- **Endpoint**: `GET /v1/videos/{task_id}`
- **Headers**:
  - `Authorization: Bearer <API_KEY>`
- **Phản hồi khi đang xử lý (`status`: "running")**:
  ```json
  {
    "id": "vid_a1b2c3d4e5",
    "object": "video",
    "status": "running",
    "progress": 45,
    "created_at": 1728557400,
    "model": "muse-video",
    "result": null,
    "error": null
  }
  ```
- **Phản hồi khi hoàn tất (`status`: "succeeded")**:
  ```json
  {
    "id": "vid_a1b2c3d4e5",
    "object": "video",
    "status": "succeeded",
    "progress": 100,
    "created_at": 1728557400,
    "model": "muse-video",
    "result": {
      "url": "http://127.0.0.1:18610/v1/media/vid_a1b2c3d4e5.mp4",
      "mime": "video/mp4",
      "width": 1280,
      "height": 720
    },
    "error": null
  }
  ```

---

### 3.2. Tạo Hình Ảnh (Image Generation & Edits)

#### Tạo ảnh mới (Text-to-Image)
- **Endpoint**: `POST /v1/images/generations`
- **Body JSON**:
  ```json
  {
    "model": "dall-e-3",
    "prompt": "An oil painting of a cottage by the lake in autumn, golden hour lighting",
    "n": 1,
    "size": "1024x1024",
    "response_format": "url",
    "background": "auto"
  }
  ```
- **Tùy chọn background**:
  - `"auto"` hoặc `"opaque"`: Giữ nguyên nền gốc do AI tạo ra.
  - `"transparent"`: Tự động tách chủ thể khỏi nền và trả về file PNG trong suốt (RGBA).
- **Phản hồi chuẩn OpenAI**:
  ```json
  {
    "created": 1728557400,
    "data": [
      {
        "url": "http://127.0.0.1:18610/v1/media/img_f6e5d4c3.png",
        "revised_prompt": "An oil painting of a serene wooden cottage nestled beside a calm lake surrounded by vibrant autumn trees..."
      }
    ]
  }
  ```

#### Chỉnh sửa / Tạo ảnh biến thể theo ảnh tham chiếu (Image-to-Image)
- **Endpoint**: `POST /v1/images/edits`
- **Body JSON**:
  ```json
  {
    "model": "dall-e-3",
    "prompt": "Transform this character into anime style with glowing eyes",
    "image": "https://example.com/character.png",
    "size": "1024x1024"
  }
  ```
*(Hỗ trợ tối đa 4 ảnh tham chiếu trong danh sách `image`)*.

---

### 3.3. Hội thoại Trợ lý AI (Chat Completions)

Tương thích hoàn toàn với thư viện `openai` hoặc REST client chuẩn:
- **Endpoint**: `POST /v1/chat/completions`
- **Body JSON**:
  ```json
  {
    "model": "gpt-4o",
    "messages": [
      {"role": "system", "content": "You are a creative cinema director AI assistant."},
      {"role": "user", "content": "Write a 5-second cinematic prompt for an intro scene."}
    ],
    "stream": false
  }
  ```
*(Có thể bật `"stream": true` để nhận Server-Sent Events (SSE) theo chuẩn OpenAI)*.

---

### 3.4. Kiểm tra sức khỏe dịch vụ (Health Check)

Dành cho AI Agent kiểm tra trước khi thực thi chuỗi tác vụ:
- **Endpoint**: `GET /readyz`
- **Phản hồi**:
  ```json
  {
    "ready": true,
    "driver": {
      "driver": "invisible",
      "ok": true,
      "tabs": 2,
      "busy_tabs": 0
    },
    "accounts": {
      "total": 3,
      "available": 3,
      "invalid": 0,
      "inflight": 0
    }
  }
  ```

---

## 4. Định nghĩa Function Calling / Tool Definitions (Dành cho AI Agent)

Dưới đây là schema JSON chuẩn theo định dạng **OpenAI Tools / JSON Schema** để các AI Agent nạp trực tiếp vào danh sách công cụ:

```json
[
  {
    "type": "function",
    "function": {
      "name": "generate_muse_video",
      "description": "Tạo video AI từ văn bản (Text-to-Video) hoặc từ hình ảnh tham chiếu (Image-to-Video) thông qua Muse2API. Trả về task_id để theo dõi tiến trình.",
      "parameters": {
        "type": "object",
        "properties": {
          "prompt": {
            "type": "string",
            "description": "Mô tả chi tiết phân cảnh video cần tạo bằng tiếng Anh (ví dụ: cinematic camera movement, lighting, subject action)."
          },
          "size": {
            "type": "string",
            "enum": ["16:9", "9:16", "1:1"],
            "default": "16:9",
            "description": "Tỷ lệ khung hình: 16:9 cho màn hình ngang/điện ảnh, 9:16 cho video dọc TikTok/Shorts, 1:1 cho video vuông."
          },
          "duration": {
            "type": "integer",
            "minimum": 1,
            "maximum": 10,
            "default": 5,
            "description": "Thời lượng video tính bằng giây (khuyên dùng 5s)."
          },
          "image": {
            "type": "string",
            "description": "URL ảnh hợp lệ (hoặc data URI) để làm khung hình bắt đầu cho video (Image-to-Video)."
          }
        },
        "required": ["prompt"]
      }
    }
  },
  {
    "type": "function",
    "function": {
      "name": "get_muse_video_status",
      "description": "Kiểm tra tiến độ render video và lấy URL kết quả tải video khi hoàn tất.",
      "parameters": {
        "type": "object",
        "properties": {
          "task_id": {
            "type": "string",
            "description": "ID của tác vụ video nhận được từ lệnh generate_muse_video (ví dụ: vid_abc123)."
          }
        },
        "required": ["task_id"]
      }
    }
  },
  {
    "type": "function",
    "function": {
      "name": "generate_muse_image",
      "description": "Tạo hình ảnh nghệ thuật AI hoặc tách nền tự động từ văn bản mô tả.",
      "parameters": {
        "type": "object",
        "properties": {
          "prompt": {
            "type": "string",
            "description": "Mô tả chi tiết hình ảnh cần tạo bằng tiếng Anh."
          },
          "size": {
            "type": "string",
            "enum": ["1024x1024", "1792x1024", "1024x1792"],
            "default": "1024x1024",
            "description": "Kích thước ảnh độ phân giải cao."
          },
          "background": {
            "type": "string",
            "enum": ["auto", "opaque", "transparent"],
            "default": "auto",
            "description": "Chọn 'transparent' nếu muốn tự động tách nền trả về file PNG trong suốt."
          }
        },
        "required": ["prompt"]
      }
    }
  }
]
```

---

## 5. Thư viện & Mã nguồn tích hợp mẫu (SDK & Code Examples)

### 5.1. Tích hợp trực tiếp bằng thư viện chính thức `openai-python`

```python
import os
from openai import OpenAI

# Khởi tạo client trỏ trực tiếp tới Muse2API
client = OpenAI(
    base_url="http://127.0.0.1:18610/v1",
    api_key="m2a--nqA-RT1mZ4AWalrdDVzSYkhI9Cksd8O"  # Hoặc đọc từ c:\Projects\Muse2API\data\api_key
)

# 1. Chat Completion
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "user", "content": "Gợi ý 3 kịch bản video ngắn 5 giây về chủ đề thiên nhiên kỳ vĩ."}
    ]
)
print("Chat Response:", response.choices[0].message.content)

# 2. Image Generation
image_res = client.images.generate(
    model="dall-e-3",
    prompt="A surreal floating island with crystal waterfalls, fantasy style, 8k",
    size="1024x1024"
)
print("Image URL:", image_res.data[0].url)
```

---

### 5.2. Hàm Python tự hành hoàn chỉnh để tạo Video (Tự động Polling)

```python
import time
import requests

class MuseVideoAgentClient:
    def __init__(self, base_url: str = "http://127.0.0.1:18610", api_key: str = "m2a--nqA-RT1mZ4AWalrdDVzSYkhI9Cksd8O"):
        self.base_url = base_url.rstrip("/")
        self.headers = {
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json"
        }

    def create_video_and_wait(self, prompt: str, size: str = "16:9", duration: int = 5, image: str = None, timeout: int = 600) -> str:
        """
        Gửi yêu cầu tạo video và tự động theo dõi đến khi có URL video tải về.
        """
        payload = {
            "model": "sora",
            "prompt": prompt,
            "size": size,
            "duration": duration
        }
        if image:
            payload["image"] = image

        # 1. Khởi tạo task
        resp = requests.post(f"{self.base_url}/v1/videos", headers=self.headers, json=payload)
        resp.raise_for_status()
        task_data = resp.json()
        task_id = task_data["id"]
        print(f"[Agent] Đã tạo task: {task_id}, bắt đầu polling...")

        # 2. Polling loop
        started = time.time()
        while time.time() - started < timeout:
            time.sleep(4)  # Nghỉ 4 giây giữa các lần kiểm tra
            status_resp = requests.get(f"{self.base_url}/v1/videos/{task_id}", headers=self.headers)
            status_resp.raise_for_status()
            data = status_resp.json()

            status = data.get("status")
            progress = data.get("progress", 0)
            print(f"[Agent] Trạng thái task: {status} ({progress}%)")

            if status == "succeeded":
                video_url = data["result"]["url"]
                print(f"[Agent] Hoàn thành! Video URL: {video_url}")
                return video_url
            elif status == "failed":
                err = data.get("error")
                raise RuntimeError(f"Tạo video thất bại: {err}")

        raise TimeoutError("Hết thời gian chờ tạo video (>10 phút).")

# Sử dụng:
# client = MuseVideoAgentClient()
# url = client.create_video_and_wait(prompt="Cyberpunk drone flying through neon-lit alleys")
```

---

## 6. Xử lý lỗi và Chiến lược Retry cho Agent (Error Handling)

| Mã HTTP | Tên lỗi | Nguyên nhân gốc rễ | Hành vi đề xuất cho AI Agent |
| :--- | :--- | :--- | :--- |
| `400` | `invalid_request_error` | Sai định dạng JSON, prompt trống, size không hỗ trợ | Kiểm tra lại schema, không retry giữ nguyên tham số |
| `401` | `authentication_error` | API Key không chính xác hoặc thiếu header `Bearer` | Đọc lại key từ `data/api_key` |
| `404` | `not_found` | Sai `task_id` hoặc media đã bị dọn dẹp | Dừng polling, thông báo task không tồn tại |
| `429` | `rate_limit_exceeded` | Tất cả tài khoản trong pool đang bận (`busy_tabs`) | Tạm dừng (Exponential backoff) từ 5-15 giây rồi thử lại |
| `502` / `504` | `upstream_error` / `timeout` | Mạng upstream muse.ai phản hồi chậm hoặc bị chặn | Chờ 10 giây và retry tối đa 3 lần |

---

## 7. Quy tắc tối ưu Prompt cho Agent (Best Practices)

1. **Ngôn ngữ**: Luôn dịch prompt sang **tiếng Anh** trước khi gửi tới Muse2API để mô hình sinh hình ảnh/video đạt độ chi tiết cao nhất.
2. **Cấu trúc Prompt Video khuyến nghị**:
   - `[Subject]` (Chủ thể chính) + `[Action/Movement]` (Hành động, chuyển động) + `[Environment/Background]` (Bối cảnh không gian) + `[Camera Angle]` (Góc máy: pan, zoom in, drone shot, fpv) + `[Lighting & Style]` (Ánh sáng: golden hour, volumetric light, cyberpunk, cinematic 4k).
3. **Kích thước phù hợp ngữ cảnh**:
   - TikTok, Instagram Reels, YouTube Shorts: `"size": "9:16"`
   - YouTube tiêu chuẩn, Phim ảnh, Web: `"size": "16:9"`
   - Bài đăng Instagram, Ảnh đại diện: `"size": "1:1"`
