# Kế hoạch xử lý triệt để fail-soft/fallback che lỗi thật trong DSCons

## 1. Phạm vi và mục tiêu

Tài liệu này tập trung vào vấn đề #3 đã được chẩn đoán trước đó: nhiều service aggregate/report của DSCons đang ưu tiên trả response schema-compatible và fallback rỗng, nhưng chưa có contract observability đủ rõ để caller, dashboard, health check và vận hành phân biệt được:

- dữ liệu thật đầy đủ
- dữ liệu thật nhưng suy giảm một phần
- dữ liệu đang fallback
- dữ liệu không đáng tin để tiếp tục workflow phụ thuộc
- lỗi nội bộ cần fail-closed thay vì tiếp tục che đi

Phạm vi bám trực tiếp vào code hiện tại:

- `app/services/company_operational_state_service.py`
- `app/services/remediation_planning_service.py`
- `app/services/dossier_report_service.py`
- `app/api/routes.py`
- `tests/test_company_operational_state_service.py`
- `tests/test_remediation_planning_service.py`
- `README.md`

Mục tiêu của kế hoạch không phải loại bỏ toàn bộ fail-soft, mà là chuẩn hóa **degraded-state observability** để fail-soft có kiểm soát, có thể đo, có thể cảnh báo và không che mất root cause.

---

## 2. Root-cause sâu

## 2.1. Fail-soft hiện tại là behavior hợp lý nhưng chưa có semantic contract

Các service hiện tại đã chọn hướng “đừng làm vỡ API”:

- `CompanyOperationalStateService` bọc nhiều nguồn bằng `_safe_*`
- `RemediationPlanningService` trả danh sách rỗng hoặc fallback plan khi PostgreSQL/manifest lỗi
- `DossierReportService` trả report rỗng khi Qdrant lỗi
- `/health/readiness` chỉ nhìn vài cờ ready cơ bản
- README còn mô tả fallback như một tính năng tích cực, nhưng chưa nói rõ mức độ tin cậy dữ liệu

Điều này giúp UI và client không crash, nhưng vì contract quá mỏng, downstream khó biết khi nào response còn dùng được cho quyết định nghiệp vụ.

## 2.2. `used_fallback: bool` đang quá nghèo thông tin

Hiện response thường chỉ có:

- `source`
- `used_fallback`
- `message`

Ba trường này chưa đủ để trả lời các câu hỏi vận hành quan trọng:

- fallback ở **nguồn chính** hay chỉ ở enrichment phụ?
- data đang **stale**, **partial**, hay **empty because no data**?
- caller có thể tiếp tục automation hay phải chặn thao tác?
- lỗi là do config, dependency outage, schema drift hay bug transform?
- mức độ ảnh hưởng là informational, degraded hay blocking?

Hậu quả là một response “items=[]” có thể mang 3 nghĩa hoàn toàn khác nhau:

- thật sự không có dữ liệu
- dependency chết nên không đọc được dữ liệu
- transform lỗi nên bị nuốt và trả rỗng

## 2.3. Aggregate service đang trộn lẫn critical path và optional enrichment

Trong `CompanyOperationalStateService`, project data, employee logs, dossier review, operations cases/actions/blockers, remediation plans đều được gom về cùng một response. Nhưng fail policy chưa được phân tầng rõ:

- có nguồn nên coi là **core** cho digital twin
- có nguồn chỉ là **enrichment**
- có phần là **presentation-level summary** có thể rơi rụng mà không đổi bản chất dữ liệu

Hiện tại code xử lý gần giống nhau: bắt exception, trả empty payload, nối message. Điều này làm cho lỗi ở tầng summary/presentation có hình thức giống lỗi mất dữ liệu nguồn.

## 2.4. Message text đang gánh quá nhiều nghĩa máy phải suy ra

Hiện lỗi thường nằm trong chuỗi tiếng Việt như:

- “đã trả fallback rỗng an toàn”
- “Không thể nạp ... Chi tiết: ...”
- “coverage manifest chưa khả dụng...”

Chuỗi này hữu ích cho người đọc log, nhưng không phải machine-readable contract. Dashboard, API client, alerting, test automation không nên phải parse text để suy luận tình trạng hệ thống.

## 2.5. Health/readiness hiện tách rời với mức suy giảm của API nghiệp vụ

`/health/readiness` hiện mới kiểm tra:

- application
- postgres
- static assets

Nhưng các API nghiệp vụ quan trọng còn phụ thuộc vào:

- Qdrant cho dossier readiness report
- coverage manifest/file system cho remediation planning
- Postgres dữ liệu workflow/reviews
- khả năng build aggregate projections
- SSE publisher ổn định

Vì vậy có thể xảy ra tình huống:

- `/health/readiness` báo ready
- nhưng `/v1/dossiers/{project_code}/readiness` đang fallback rỗng do Qdrant lỗi
- hoặc `/v1/remediation/plans` đang degrade vì manifest hỏng
- hoặc `/v1/company/operational-state` đang trả thiếu cả operations/remediation mà vẫn 200 như bình thường

## 2.6. Logging/metrics chưa có cấu trúc cho degraded state

Hiện root cause chủ yếu chỉ lộ ra qua:

- exception string gắn vào `message`
- có thể xuất hiện trong server logs nếu framework in trace

Chưa có mô hình thống nhất cho:

- error code
- dependency name
- endpoint/service bị ảnh hưởng
- fail policy đã áp dụng
- severity
- degraded count/rate
- fallback reason cardinality có kiểm soát

Điều này khiến việc phân biệt “degrade có chủ đích” và “bug đang bị nuốt” rất khó.

## 2.7. Test hiện xác nhận fallback tồn tại, chưa xác nhận fallback có quan sát được đúng chuẩn

Các test hiện tại chủ yếu kiểm tra:

- response không vỡ
- `used_fallback` là `True`
- `message` có chứa exception text

Chưa có test cho các điểm quan trọng hơn của observability contract:

- nguồn nào degraded
- degraded level là gì
- operation có nên fail-open hay fail-closed
- readiness có phản ánh đúng business capability hay không
- client có thể dựa vào flags thay vì parse message

---

## 3. Mô hình degraded-state observability đề xuất

## 3.1. Nguyên tắc thiết kế

1. **Không phá compatibility ngay lập tức**
   - giữ `source`, `used_fallback`, `message` trong giai đoạn đầu
   - thêm contract mới dạng optional fields

2. **Tách biệt data correctness với transport success**
   - HTTP 200 không đồng nghĩa dữ liệu khỏe
   - response body phải nói rõ mức suy giảm

3. **Phân biệt dependency failure với legitimate empty state**
   - “không có dữ liệu” khác “không đọc được dữ liệu”

4. **Machine-readable trước, human-readable sau**
   - dashboard/alert/client dùng code + flags
   - `message` chỉ là lớp giải thích cho con người

5. **Fail policy phải explicit**
   - fail-open, fail-soft, fail-closed không được ngầm hiểu trong code

---

## 4. Contract response chuẩn đề xuất

## 4.1. Envelope observability dùng chung

Đề xuất thêm một khối chuẩn vào các response aggregate/report/degraded-capable, tên gợi ý:

- `observability`
- hoặc `degraded_state`

Khuyến nghị dùng `observability` vì đủ rộng cho source health, warnings, fail policy.

### Shape đề xuất

```json
{
  "observability": {
    "status": "healthy|degraded|unavailable",
    "degraded": true,
    "fail_policy": "fail_open|fail_soft|fail_closed",
    "decision": "serve_partial|serve_empty|reject_request",
    "warnings": [
      {
        "code": "DEPENDENCY_QDRANT_UNAVAILABLE",
        "severity": "warning",
        "scope": "source",
        "source_name": "qdrant_dossier_metadata",
        "message": "Không thể tải dossier documents từ Qdrant."
      }
    ],
    "sources": [
      {
        "source_name": "postgres_dossier_reviews",
        "role": "primary",
        "status": "healthy|degraded|unavailable|disabled|empty",
        "availability": "up|down|disabled|unknown",
        "data_state": "fresh|partial|empty|stale|unknown",
        "used_fallback": false,
        "record_count": 12,
        "error_code": null,
        "error_category": null,
        "latched": false,
        "message": null
      }
    ]
  }
}
```

## 4.2. Response flags tối thiểu

Cho các response hiện có, nên thêm tối thiểu các cờ sau:

- `status`: `healthy | degraded | unavailable`
- `degraded`: bool
- `fail_policy`: `fail_open | fail_soft | fail_closed`
- `decision`: `serve_full | serve_partial | serve_empty | reject_request`
- `warnings`: list các warning/error machine-readable
- `source_health`: danh sách trạng thái nguồn chi tiết

Nếu muốn tối thiểu hơn cho phase đầu:

- giữ top-level `used_fallback`
- thêm `status`
- thêm `degradation_reason_codes`
- thêm `source_health`

## 4.3. Mapping với fields hiện có

### Giữ lại
- `source`
- `used_fallback`
- `message`

### Bổ sung
- `status`
- `degraded`
- `source_health`
- `warnings`

### Quy ước tương thích ngược
- `used_fallback=True` không còn là tín hiệu duy nhất
- `used_fallback=False` vẫn có thể là `status="degraded"` nếu data partial nhưng không dùng fallback object rỗng
- `message` không còn được coi là nguồn chuẩn để client quyết định logic

---

## 5. Source health model chuẩn cho DSCons

## 5.1. Phân loại role của source

Mỗi source nên có `role` để quyết định fail policy:

- `primary`: mất là ảnh hưởng bản chất endpoint
- `secondary`: mất làm giảm chất lượng nhưng endpoint vẫn hữu ích
- `derived`: dữ liệu suy diễn/presentation
- `optional`: enrich thêm cho dashboard hoặc summary

### Mapping theo code hiện tại

#### `CompanyOperationalStateService`
- `projects`: `primary`
- `employee_logs`: `primary`
- `dossier_reviews`: `primary`
- `operations_cases`: `secondary`
- `operations_actions`: `secondary`
- `operations_blockers`: `secondary`
- `remediation_plans`: `secondary`
- `policy_overview`: `derived`
- `remediation_summary`: `derived`

#### `RemediationPlanningService`
- `postgres_dossier_reviews`: `primary`
- `coverage_manifest`: `secondary` với list mode, nhưng có thể là `primary` cho một số automation build
- specialized playbook theo `document_type`: `derived capability`

#### `DossierReportService`
- `qdrant_dossier_metadata`: `primary`
- `readiness_rules`: `local_static_primary`
- `stage_requirements`: `local_static_primary`

## 5.2. Source status enum

Đề xuất enum dùng chung:

- `healthy`: đọc thành công, dữ liệu dùng được
- `empty`: đọc thành công nhưng không có dữ liệu
- `degraded`: đọc được một phần hoặc phải hạ chất lượng
- `unavailable`: dependency lỗi hoặc transform lỗi, source không dùng được
- `disabled`: bị tắt do config
- `unknown`: chưa xác định

## 5.3. Data state enum

Tách khỏi availability:

- `fresh`
- `partial`
- `empty`
- `stale`
- `unknown`

Ví dụ:

- Postgres bật nhưng query lỗi: `availability=down`, `status=unavailable`, `data_state=unknown`
- Qdrant query thành công nhưng không có document: `availability=up`, `status=empty`, `data_state=empty`
- Aggregate build được nhưng remediation summary bị lỗi: endpoint `status=degraded`, source derived kia `status=degraded`, core source vẫn healthy

---

## 6. Error taxonomy đề xuất

## 6.1. Error category cấp cao

Đề xuất taxonomy chung:

- `configuration`
- `dependency_unavailable`
- `timeout`
- `data_missing`
- `data_invalid`
- `schema_mismatch`
- `transform_error`
- `partial_result`
- `permission`
- `unexpected`

## 6.2. Error code cụ thể theo DSCons

Nên chuẩn hóa code ngắn, stable, upper snake case.

### Company operational state
- `PROJECTS_SOURCE_UNAVAILABLE`
- `EMPLOYEE_LOGS_SOURCE_UNAVAILABLE`
- `DOSSIER_REVIEWS_SOURCE_DISABLED`
- `DOSSIER_REVIEWS_SOURCE_UNAVAILABLE`
- `OPERATIONS_CASES_SOURCE_UNAVAILABLE`
- `OPERATIONS_ACTIONS_SOURCE_UNAVAILABLE`
- `OPERATIONS_BLOCKERS_SOURCE_UNAVAILABLE`
- `REMEDIATION_PLANS_SOURCE_UNAVAILABLE`
- `POLICY_OVERVIEW_BUILD_FAILED`
- `REMEDIATION_SUMMARY_BUILD_FAILED`
- `REMEDIATION_PLAN_SUMMARY_BUILD_FAILED`

### Remediation planning
- `DOSSIER_REVIEW_SOURCE_DISABLED`
- `DOSSIER_REVIEW_SOURCE_UNAVAILABLE`
- `COVERAGE_MANIFEST_MISSING`
- `COVERAGE_MANIFEST_INVALID`
- `COVERAGE_PROJECT_NOT_FOUND`
- `SPECIALIZED_PLAYBOOK_UNAVAILABLE`
- `GENERIC_REMEDIATION_FALLBACK_USED`

### Dossier report
- `QDRANT_DOSSIER_METADATA_UNAVAILABLE`
- `QDRANT_DOSSIER_METADATA_TIMEOUT`
- `DOSSIER_REPORT_TRANSFORM_FAILED`

### Readiness / health
- `POSTGRES_UNAVAILABLE`
- `QDRANT_UNAVAILABLE`
- `STATIC_ASSET_MISSING`
- `COVERAGE_MANIFEST_UNAVAILABLE`
- `BUSINESS_CAPABILITY_DEGRADED`

## 6.3. Warning object shape

```json
{
  "code": "COVERAGE_MANIFEST_MISSING",
  "category": "data_missing",
  "severity": "warning",
  "scope": "source",
  "source_name": "coverage_manifest",
  "retryable": true,
  "message": "Không tìm thấy coverage manifest tại tmp/hd2026_project_coverage_manifest.json."
}
```

Khuyến nghị `severity` chỉ dùng tập nhỏ:

- `info`
- `warning`
- `error`
- `critical`

---

## 7. Chính sách fail-open / fail-soft / fail-closed

## 7.1. Định nghĩa chuẩn

- **fail-open**: dependency lỗi nhưng vẫn cho phép luồng tiếp tục với dữ liệu cũ/partial vì rủi ro nghiệp vụ thấp
- **fail-soft**: trả response degrade có cờ rõ ràng, không giả vờ healthy
- **fail-closed**: từ chối request vì tiếp tục sẽ gây hiểu sai hoặc phát sinh hành động sai

## 7.2. Ma trận policy theo endpoint

### `GET /v1/company/operational-state`
- mất `operations_*`, `remediation_plans`, `policy_overview`, `remediation_summary`:
  - **fail-soft**
  - trả partial response + warning rõ
- mất `projects` hoặc `dossier_reviews` đồng thời:
  - vẫn có thể **fail-soft** ở phase đầu để giữ UI hoạt động
  - nhưng status endpoint phải là `degraded` mức cao
- mất hầu hết primary sources cùng lúc:
  - chuyển sang **fail-closed** ở phase sau cho các consumer automation, còn dashboard có thể dùng endpoint khác chuyên “snapshot degraded”

### `GET /v1/remediation/plans`
- Postgres disabled/unavailable:
  - với API đọc danh sách: **fail-soft** trong phase đầu để tương thích
  - nhưng phải gắn `status="unavailable"` hoặc `degraded` nặng, không chỉ `items=[]`
- coverage manifest missing/invalid:
  - cho phép **fail-soft** nếu vẫn build được plan tối thiểu
- build phục vụ automation downstream như generation/publish:
  - nên **fail-closed** nếu readiness không đủ hoặc source chính unavailable

### `POST /v1/remediation/plans/build`
- không tìm thấy finding:
  - không phải degraded, trả kết quả domain phù hợp hoặc 404/400 tùy contract
- Postgres unavailable:
  - **fail-closed**
- coverage manifest invalid:
  - nếu chỉ là enrichment và plan còn meaningful: fail-soft
  - nếu build plan sẽ kích hoạt workflow tài liệu thật: fail-closed

### `GET /v1/dossiers/{project_code}/readiness`
- Qdrant unavailable:
  - không nên giả vờ report “rỗng nhưng hợp lệ” mãi
  - đề xuất phase đầu: fail-soft + explicit degraded contract
  - phase sau: cho phép query param hoặc header để chọn strict mode, strict thì fail-closed 503

### `/health/readiness`
- không fail-soft kiểu mơ hồ
- phải phản ánh đúng capability degradation
- nếu business capability quan trọng degraded thì endpoint phải 503 hoặc ít nhất có `overall_status=degraded` với capability map rõ

## 7.3. Quy tắc ra quyết định

Nên chuẩn hóa decision tree:

1. dependency/source thuộc `primary` hay `secondary`?
2. kết quả trả về còn đủ để người dùng đưa ra quyết định an toàn không?
3. endpoint chỉ để quan sát hay còn bị automation khác phụ thuộc?
4. degraded có thể nhận biết bằng machine-readable flags chưa?
5. nếu client bỏ qua flags thì có gây hành động sai không?

Nếu câu 5 là “có”, endpoint đó không nên fail-soft vô thời hạn.

---

## 8. API contract chi tiết theo endpoint

## 8.1. `CompanyOperationalStateResponse`

Hiện đã có `source_status`, nhưng còn thiếu semantics rõ. Đề xuất:

### Nâng cấp `source_status` thành `source_health`
Mỗi item nên có:
- `source_name`
- `role`
- `status`
- `availability`
- `data_state`
- `used_fallback`
- `record_count`
- `error_code`
- `error_category`
- `message`

### Top-level bổ sung
- `status`
- `degraded`
- `fail_policy`
- `warnings`

### Quy tắc status top-level
- `healthy`: tất cả primary healthy/empty hợp lệ, secondary không có lỗi blocking
- `degraded`: ít nhất một secondary unavailable hoặc derived build lỗi, hoặc một primary partial
- `unavailable`: primary sources không đủ để digital twin đáng tin

## 8.2. `RemediationListResponse` / `RemediationBuildResponse` / `RemediationGapsResponse`

Bổ sung:
- `status`
- `degraded`
- `fail_policy`
- `warnings`
- `source_health`

Đặc biệt với từng `RemediationPlanItem`, nên có thêm:
- `confidence_level`: `full|partial|fallback`
- `plan_mode`: `specialized|generic_fallback`
- `blocking_source_issues`: list error code

Điều này giúp phân biệt:
- plan generic fallback vì chưa có playbook
- plan degraded vì thiếu manifest
- plan không nên dùng cho automation

## 8.3. `DossierReadinessResponse`

Hiện route build response từ dict report. Đề xuất thêm:
- `source`
- `used_fallback`
- `status`
- `degraded`
- `warnings`
- `source_health`

Nếu chưa muốn sửa schema lớn ngay, ít nhất nên thêm fields này vào report dict rồi phase sau mới đẩy vào pydantic schema.

---

## 9. Health/readiness redesign

## 9.1. Vấn đề hiện tại

`/health/readiness` mới phản ánh infra readiness tối thiểu, chưa phản ánh capability readiness.

## 9.2. Mô hình health 3 lớp

### Lớp 1: Liveness
`GET /health`
- app process sống
- không chạm dependency nặng
- luôn rất đơn giản

### Lớp 2: Dependency readiness
`GET /health/readiness`
- postgres
- qdrant
- static assets
- coverage manifest path
- optional: MLX model loadability nếu có cheap probe

### Lớp 3: Business capability readiness
Đề xuất mở rộng cùng endpoint hoặc endpoint mới `/health/capabilities`:
- `company_operational_state`
- `dossier_readiness_report`
- `remediation_planning`
- `dossier_review_workflow`
- `knowledge_search`

Mỗi capability nên có:
- `status`
- `depends_on`
- `degraded_by`
- `fail_policy`
- `ready_for_automation`
- `message`

## 9.3. Status code policy

### `/health`
- 200 nếu process sống

### `/health/readiness`
- 200 nếu mọi critical dependency sẵn sàng
- 503 nếu Postgres/Qdrant hoặc capability critical unavailable

### `/health/capabilities` hoặc nội dung capability trong readiness
- body phải có `overall_status`
- không chỉ boolean `ready`

Ví dụ:
```json
{
  "overall_status": "degraded",
  "dependencies": {...},
  "capabilities": {
    "dossier_readiness_report": {
      "status": "unavailable",
      "ready_for_automation": false,
      "degraded_by": ["QDRANT_UNAVAILABLE"]
    },
    "company_operational_state": {
      "status": "degraded",
      "ready_for_automation": false,
      "degraded_by": ["OPERATIONS_CASES_SOURCE_UNAVAILABLE"]
    }
  }
}
```

---

## 10. Logging và metrics

## 10.1. Logging cấu trúc bắt buộc

Mỗi degraded event nên log structured với cùng field set:

- `event`: `degraded_response_emitted`
- `service`
- `endpoint`
- `operation`
- `request_id` nếu có
- `project_code` / `review_id` / `finding_code` nếu có
- `status`
- `fail_policy`
- `decision`
- `error_code`
- `error_category`
- `source_name`
- `exception_type`
- `exception_message`
- `used_fallback`
- `record_count`

Khuyến nghị:
- log warning cho degraded có chủ đích
- log error cho fail-closed hoặc unexpected transform bug
- giữ exception text nhưng không nhét stack trace vào response body

## 10.2. Metrics nên có

### Counter
- `dscons_degraded_responses_total{endpoint,status}`
- `dscons_fallback_used_total{service,source_name,error_code}`
- `dscons_dependency_failures_total{dependency,error_code}`
- `dscons_fail_closed_total{endpoint,error_code}`

### Histogram
- `dscons_dependency_latency_seconds{dependency,operation}`
- `dscons_endpoint_latency_seconds{endpoint,status}`

### Gauge
- `dscons_capability_status{capability}` map healthy=2 degraded=1 unavailable=0
- `dscons_source_health{source_name}`

Nếu chưa có metrics stack đầy đủ, phase đầu có thể:
- structured logs trước
- log-based dashboard/grep
- phase sau mới formalize metrics exporter

## 10.3. Cardinality guardrail

Không đưa trực tiếp raw exception message vào metric labels. Chỉ dùng:
- `error_code`
- `error_category`
- `source_name`
- `endpoint`

Exception message giữ trong logs.

---

## 11. Dashboard surfacing

## 11.1. Vấn đề hiện tại

Dashboard hiện có thể đọc response thành công nhưng không biết đang xem dữ liệu degrade trừ khi parse text `message`.

## 11.2. Chuẩn hiển thị đề xuất

Mỗi dashboard/api consumer nên surfacing tối thiểu:

- badge top-level:
  - `Healthy`
  - `Degraded`
  - `Unavailable`
- danh sách nguồn degraded
- hành động gợi ý:
  - “thử lại sau”
  - “kiểm tra Postgres”
  - “chạy lại ingest coverage”
- nhãn “partial data” trên widget bị ảnh hưởng

## 11.3. Mapping UI theo endpoint

### Company operational state dashboard
- banner vàng nếu secondary source lỗi
- banner đỏ nếu primary source lỗi
- từng widget `operations`, `remediation`, `policy overview` có badge source riêng
- tránh hiển thị số 0 như thể business state thật khi source đang unavailable; hiển thị `N/A` hoặc `0 (fallback)`

### Dossier readiness dashboard
- nếu Qdrant unavailable, hiển thị “Không tải được metadata hồ sơ” thay vì bảng missing documents trông như dữ liệu thật
- distinguish `không có tài liệu` với `không truy xuất được tài liệu`

### Remediation views
- đánh dấu plan `generic fallback`
- đánh dấu `not automation-safe`
- readiness blocked do source issue phải phân biệt với blocked do business input thiếu

---

## 12. Rollout plan ít rủi ro

## Phase 0 - Chuẩn hóa semantics trên giấy và inventory hiện trạng
**Mục tiêu:** không đổi code behavior, chỉ chốt contract và policy.

Việc làm:
- lập bảng inventory tất cả path fail-soft trong 3 service và routes liên quan
- gán role cho từng source: primary/secondary/derived
- chốt error taxonomy và fail policy matrix
- chốt response envelope tối thiểu cho phase 1

Deliverable:
- tài liệu contract
- bảng mapping error code ↔ source ↔ endpoint
- bảng readiness capability

Acceptance criteria:
- mọi fail-soft path trong code hiện tại map được vào một error code và một fail policy
- không còn trường hợp “message-only semantics”

## Phase 1 - Thêm observability fields nhưng chưa đổi HTTP semantics
**Mục tiêu:** an toàn nhất, ít phá client nhất.

Việc làm:
- thêm `status`, `degraded`, `warnings`, `source_health` vào response model của:
  - company operational state
  - remediation list/build/gaps
  - dossier readiness report
- giữ nguyên `source`, `used_fallback`, `message`
- build adapter từ current `_safe_*` path sang warning/source_health entries
- bổ sung structured logging ở các điểm catch exception

Tradeoff:
- payload response lớn hơn
- cần cập nhật schema/tests nhưng ít rủi ro nhất

Acceptance criteria:
- client cũ vẫn chạy
- client mới có thể dựa hoàn toàn vào fields machine-readable
- mọi fallback hiện có đều tạo warning code và source_health entry tương ứng

Rollback:
- có thể tạm ngừng populate fields mới nhưng vẫn giữ schema optional

## Phase 2 - Redesign health/readiness theo capability
**Mục tiêu:** readiness phản ánh đúng business capability.

Việc làm:
- mở rộng `/health/readiness` với dependency + capability map
- thêm probes cho:
  - Postgres
  - Qdrant
  - coverage manifest
- thêm mapping capability:
  - company operational state
  - remediation planning
  - dossier readiness
- phân biệt infra ready và automation ready

Tradeoff:
- readiness logic phức tạp hơn
- cần tránh probe quá nặng

Acceptance criteria:
- khi Qdrant hỏng, readiness phải phản ánh dossier report unavailable
- khi manifest hỏng, readiness phải phản ánh remediation degraded
- khi chỉ static assets thiếu, capability data APIs vẫn được phản ánh riêng

Rollback:
- giữ endpoint cũ song song hoặc giữ fields cũ `checks` trong khi thêm `capabilities`

## Phase 3 - Dashboard và consumer surfacing
**Mục tiêu:** người dùng nhìn thấy degraded state thay vì số 0 giả.

Việc làm:
- cập nhật dashboard đọc `status/warnings/source_health`
- hiển thị badge và banner
- không render “zero” như valid when unavailable
- README cập nhật semantics degraded-state

Acceptance criteria:
- dashboard phân biệt empty thật với unavailable
- remediation fallback plan được gắn nhãn rõ
- không còn phụ thuộc vào parse `message`

Rollback:
- UI có thể ẩn badge mới nhưng API contract giữ nguyên

## Phase 4 - Siết dần fail policy cho automation-sensitive flows
**Mục tiêu:** chặn các trường hợp đang nguy hiểm nhưng trước đó bị fail-soft.

Việc làm:
- xác định endpoint/consumer nào dùng cho automation thật
- thêm strict mode:
  - header hoặc query `strict=true`
- với strict mode:
  - primary source unavailable => 503 / fail-closed
  - automation-unsafe fallback plan => reject

Tradeoff:
- có thể làm lộ lỗi trước đây bị che
- nhưng chỉ bật cho luồng đã sẵn sàng

Acceptance criteria:
- interactive dashboard vẫn degrade được
- automation path không dùng nhầm fallback data
- error rate tăng lên là có chủ đích và quan sát được

Rollback:
- tắt strict mode mặc định, chỉ giữ opt-in

---

## 13. Test và verification plan

## 13.1. Unit test mở rộng

### `tests/test_company_operational_state_service.py`
Bổ sung case xác nhận:
- từng lỗi source tạo đúng `error_code`
- top-level `status` chuyển `degraded`
- primary source fail khác secondary source fail
- lỗi ở `policy_overview`/`remediation_summary` chỉ mark derived degraded, không hạ thành unavailable toàn bộ
- empty data hợp lệ không sinh `dependency_unavailable`

### `tests/test_remediation_planning_service.py`
Bổ sung case:
- Postgres disabled => `status=unavailable`, warning code `DOSSIER_REVIEW_SOURCE_DISABLED`
- manifest missing => warning `COVERAGE_MANIFEST_MISSING`
- manifest invalid => `data_invalid`
- generic fallback plan => `plan_mode=generic_fallback`, không nhầm với dependency failure
- build strict mode => fail-closed khi source chính unavailable

### `DossierReportService`
Cần test mới cho:
- Qdrant unavailable => degraded/unavailable rõ ràng
- empty result hợp lệ => `status=healthy` hoặc `empty`, không phải fallback failure
- transform error riêng với dependency error

## 13.2. Route/API contract tests

Tạo contract test cho:
- `/v1/company/operational-state`
- `/v1/remediation/plans`
- `/v1/remediation/plans/build`
- `/v1/remediation/gaps`
- `/v1/dossiers/{project_code}/readiness`
- `/health/readiness`

Các test phải xác nhận:
- fields mới tồn tại
- enum values hợp lệ
- HTTP status và body status không mâu thuẫn
- message là optional explanation, không phải nguồn semantics chính

## 13.3. Error-path verification

Bắt buộc cover:
- dependency disabled
- dependency unavailable
- invalid local file/manifest
- transform bug trong summary builder
- empty legitimate result
- partial result do secondary source fail

## 13.4. SSE verification

Với `GET /v1/company/operational-state/stream`:
- event payload phải mang observability fields giống REST
- degraded state không làm stream chết ngay
- nếu repeated failures liên tục, event nên phản ánh `status=degraded/unavailable` thay vì emit dữ liệu giả khỏe

## 13.5. Definition of Done cho nhánh observability này

Hoàn thành khi:
1. mọi endpoint aggregate/report trong scope có machine-readable degraded contract
2. readiness phản ánh được capability, không chỉ infra
3. dashboard không còn coi fallback empty là dữ liệu thật
4. logs/metrics cho phép đếm degraded events theo source/error code
5. ít nhất một automation-sensitive flow có strict fail-closed path
6. README mô tả rõ semantics degraded thay vì ca ngợi fallback chung chung

---

## 14. Anti-pattern cần tránh

## 14.1. Dùng `items=[]` để che mọi lỗi
Sai vì:
- caller không phân biệt empty thật hay dependency chết
- dashboard hiển thị sai tình hình vận hành

## 14.2. Nhét raw exception text vào `message` rồi coi như đủ observability
Sai vì:
- machine không dùng ổn định được
- dễ lộ chi tiết nội bộ không cần thiết
- không phân loại được theo taxonomy

## 14.3. Một cờ `used_fallback` cho mọi loại suy giảm
Sai vì:
- generic fallback capability khác dependency outage
- derived summary fail khác source primary fail

## 14.4. Health endpoint chỉ check infra, bỏ qua capability
Sai vì:
- người vận hành thấy “ready” nhưng nghiệp vụ thật đang unusable

## 14.5. Dashboard render số 0 thay cho unavailable
Sai vì:
- tạo false confidence
- quyết định điều hành có thể sai

## 14.6. Fail-soft vô thời hạn cho luồng automation
Sai vì:
- càng lâu càng tạo data debt
- automation sẽ hành động trên dữ liệu không đáng tin

## 14.7. Metric labels chứa full exception message
Sai vì:
- nổ cardinality
- dashboard monitoring vô dụng

## 14.8. Dùng cùng một error code cho dependency failure và business empty state
Sai vì:
- không thể phân biệt incident vận hành với trạng thái dữ liệu hợp lệ

---

## 15. Thứ tự triển khai tối ưu và phụ thuộc với các hướng khác

## Phụ thuộc trực tiếp
- phụ thuộc nhẹ vào nhánh #4 shared projection vì source semantics sẽ rõ hơn khi review-derived data được gom chuẩn
- phụ thuộc nhẹ vào nhánh #5 integration verification để mở rộng contract/E2E tests
- liên quan đến #1 workflow hardening ở chỗ workflow persisted thành công nhưng enrichment degrade phải được surfacing đúng
- liên quan đến #2 schema sync vì một phần degrade có thể thực chất do drift/schema mismatch, cần taxonomy bắt được

## Thứ tự đề xuất
1. chốt degraded contract + error taxonomy
2. thêm response fields machine-readable ở service/report endpoints
3. redesign readiness/capability health
4. cập nhật dashboard surfacing
5. thêm strict fail-closed cho automation-sensitive flows
6. tích hợp vào integration/E2E suite

Thứ tự này ít rủi ro vì:
- bắt đầu bằng additive contract
- chưa phá client cũ
- tạo nền cho test và health trước khi siết fail policy

---

## 16. Tóm tắt quyết định chính

- Giữ fail-soft nhưng không cho phép fail-soft “im lặng”
- Chuẩn hóa observability contract bằng `status`, `degraded`, `warnings`, `source_health`
- Phân biệt rõ primary/secondary/derived source
- Chuẩn hóa error taxonomy và fail policy matrix
- Redesign health/readiness theo capability, không chỉ infra
- Dashboard phải surfacing degraded state, không render giả như healthy
- Về dài hạn, automation-sensitive flows phải chuyển dần từ fail-soft sang strict fail-closed