# Hướng Dẫn Phát Triển Song Song DSCons Trên Windows (GTX 1660 Super) & macOS (Apple Silicon)

Tài liệu này hướng dẫn thiết lập môi trường phát triển song song cho dự án **DSCons** trên cả 2 nền tảng: **macOS (Apple Silicon)** và **Windows (NVIDIA GTX 1660 Super)**.

---

## I. MÔ HÌNH CHỌN TỐI ƯU: QWEN 2.5 3B INSTRUCT

Sau khi đánh giá các dòng mô hình SLM kích thước 2.5B - 3B, dự án DSCons chọn **Qwen 2.5 3B Instruct** làm mô hình local mặc định tối ưu nhất nhờ các ưu điểm:

- **Khả năng Tiếng Việt xuất sắc**: Hiểu ngữ cảnh xây dựng, thuật ngữ chuyên ngành và xưng hô theo đúng Persona.
- **Hỗ trợ Structured Output / JSON mode**: Tương thích hoàn hảo với Pydantic schemas của DSCons khi trích xuất dữ liệu rủi ro, mốc dự án hay danh mục hồ sơ thiếu.
- **Dung lượng siêu nhẹ**: Chỉ chiếm ~2.1 GB VRAM / Unified Memory.
- **Tốc độ phản hồi cực nhanh**:
  - Trên **macOS (M1/M2/M3/M4)**: 60 - 90 tokens/sec (via `mlx-lm`).
  - Trên **Windows (GTX 1660 Super)**: 45 - 70 tokens/sec (via CUDA / `Ollama` / `llama-cpp-python`).

---

## II. BẢNG MÔ HÌNH THEO NỀN TẢNG

| Nền tảng | GPU / Phần cứng | Thư viện LLM | Tên Model Cấu hình | Cấu hình `.env` |
| --- | --- | --- | --- | --- |
| **macOS** | Apple Silicon M-Series | `mlx-lm` | `mlx-community/Qwen2.5-3B-Instruct-4bit` | `LLM_PROVIDER=mlx`<br>`LLM_MODEL_NAME=mlx-community/Qwen2.5-3B-Instruct-4bit` |
| **Windows** | NVIDIA GTX 1660 Super (CUDA) | `Ollama` / `vLLM` / `LM Studio` | `qwen2.5:3b-instruct` | `LLM_PROVIDER=openai_compatible`<br>`LLM_API_BASE_URL=http://127.0.0.1:11434/v1`<br>`LLM_MODEL_NAME=qwen2.5:3b-instruct` |

---

## III. QUY TRÌNH THIẾT LẬP TRÊN WINDOWS (GTX 1660 SUPER)

### 1. Chuẩn bị LLM Engine với Ollama (Khuyên dùng trên Windows)
1. Tải và cài đặt **Ollama cho Windows** tại: [ollama.com](https://ollama.com) (Ollama sẽ tự nhận diện NVIDIA GTX 1660 Super và bật CUDA acceleration).
2. Tải model Qwen 2.5 3B Instruct:
   ```bash
   ollama pull qwen2.5:3b-instruct
   ```
3. Ollama sẽ mở một server chuẩn OpenAI API tại: `http://localhost:11434/v1`.

### 2. Khởi tạo Services Phụ trợ (PostgreSQL & Qdrant)
Chạy Docker Desktop trên Windows (bật WSL2 backend):
```bash
docker-compose up -d postgres qdrant
```

### 3. Cấu hình file `.env` trên Windows
Tạo hoặc sửa file `.env` trên máy Windows:
```env
APP_ENV=development
LLM_PROVIDER=openai_compatible
LLM_API_BASE_URL=http://127.0.0.1:11434/v1
LLM_MODEL_NAME=qwen2.5:3b-instruct
QDRANT_URL=http://127.0.0.1:6333
DATABASE_URL=postgresql://dscons_user:dscons_password@127.0.0.1:5432/dscons_db
```

### 4. Chạy Backend FastAPI trên Windows
```bash
python -m venv .venv
.venv\Scripts\activate
pip install -e .
python -m uvicorn app.main:app --host 127.0.0.1 --port 8000 --reload
```

---

## IV. QUY TRÌNH THIẾT LẬP TRÊN MACOS (APPLE SILICON)

Cấu hình file `.env` trên MacBook:
```env
APP_ENV=development
LLM_PROVIDER=mlx
LLM_MODEL_NAME=mlx-community/Qwen2.5-3B-Instruct-4bit
QDRANT_URL=http://127.0.0.1:6333
DATABASE_URL=postgresql://dscons_user:dscons_password@127.0.0.1:5432/dscons_db
```

Khởi chạy backend:
```bash
source .venv/bin/activate
python3 -m uvicorn app.main:app --host 127.0.0.1 --port 8000 --reload
```

---

## V. ĐẶC ĐIỂM KIẾN TRÚC MÃ NGUỒN ĐÃ ĐƯỢC CHUẨN HÓA

Service `LLMClient` trong [`app/services/llm_client.py`](file:///Users/pain/Projects/DSCons/app/services/llm_client.py) đã được nâng cấp cơ chế **Auto-Fallback / Multi-Backend**:
- Nếu phát hiện thư viện `mlx_lm` (macOS), client sẽ ưu tiên dùng **MLX** để sinh văn bản mượt mà trực tiếp trên Unified Memory.
- Nếu chạy trên **Windows** (không có `mlx_lm`), client tự động chuyển sang mode `openai_compatible` để gọi đến Ollama/vLLM local thông qua CUDA của card GTX 1660 Super.
- Mọi router, agent personas và API contract hoàn toàn giống nhau 100% trên cả 2 hệ điều hành.
