# Docs, Tests, and Verification Audit Plan

## Scope audited

This audit reviewed the following existing files for refactor/test planning only:

### Docs
- `docs/step5-command-center-readiness-planning.md`
- `docs/policy-operations-playbook.md`
- `docs/operations-case-management-upgrade-proposal.md`
- `docs/dossier-review-full-flow-audit.md`

### Tests
- `tests/test_dossier_review_routes.py`
- `tests/test_workflow_persistence_service.py`

### Runnable verification / diagnostics scripts
- `tmp/check_health_endpoints.py`
- `tmp/diagnose_dossier_review_full_lifecycle.py`
- `tmp/verify_workflow_summary_contract.py`
- `tmp/verify_company_operational_state_contract.py`

---

## Executive summary

The repository already has useful planning and verification artifacts, but they are spread across three different durability levels:

1. **Durable architectural docs** in `docs/`
2. **Real unit-style tests** in `tests/`
3. **Live environment verification scripts** in `tmp/`

The main problem is not lack of material. It is **lack of structure and boundaries**. Right now:

- docs overlap in purpose and repeat concepts
- tests mix narrow unit behavior with route contract checks
- `tmp/` scripts behave like contract and smoke tests, but are stored as ad hoc diagnostics
- there is no index telling future refactor work:
  - which doc is normative
  - which test category to add for a change
  - which checks require a running server and seeded data
  - which checks are safe to run offline with mocks only

This creates memory risk for a future refactor: the team may accidentally rewrite contracts already described in docs, duplicate checks, or rely on fragile live scripts as if they were unit tests.

---

## Current-state audit

## 1. Documentation sprawl audit

### `docs/operations-case-management-upgrade-proposal.md`
Role today:
- high-level strategy and target architecture for operations case management
- proposes first-class entities, command center direction, and AI structured planning direction

Strengths:
- strong framing document
- good source of architectural intent
- useful as a durable "why this refactor exists" record

Problems:
- overlaps heavily with Step 5 planning
- mixes architecture, rollout, UI ideas, and entity contracts in one file
- some proposed contracts are now partially superseded by newer docs

Recommendation:
- keep as the **foundational proposal**
- do not keep extending it with implementation-phase detail
- add links from newer docs instead of duplicating content

Best future role:
- architecture rationale / original proposal archive

### `docs/step5-command-center-readiness-planning.md`
Role today:
- most detailed planning document for command center, readiness explainability, and AI structured planning
- contains API proposals, object contracts, rollout phases, and acceptance criteria

Strengths:
- strongest current implementation-planning artifact
- provides additive contract guidance
- ties together operational state, readiness, policy, and planning

Problems:
- too broad for one durable planning file
- mixes:
  - command center projection design
  - readiness explainability contracts
  - AI planning object model
  - SSE guidance
  - rollout acceptance criteria
- likely to become stale if used as both spec and backlog

Recommendation:
- split into smaller durable docs later:
  - command center projections
  - readiness explainability contracts
  - AI planning contracts
- preserve this current file as the umbrella planning narrative until split is done

Best future role:
- umbrella roadmap doc, later reduced to an index/overview

### `docs/policy-operations-playbook.md`
Role today:
- operational governance playbook for workflow policy behavior
- explains session-level vs finding-level policy, escalation, risk thresholds, circuit breaker behavior, audit expectations, and rollout discipline

Strengths:
- highly actionable
- much closer to operational runbook quality than most other docs
- directly useful for future refactor safety and policy regression prevention

Problems:
- mixes:
  - policy semantics
  - operational procedures
  - incident handling
  - future governance improvements
- could become too long as policy grows

Recommendation:
- keep as a durable runbook
- later split into:
  - policy semantics reference
  - policy operations runbook
  - policy change management checklist

Best future role:
- operations runbook and governance source of truth

### `docs/dossier-review-full-flow-audit.md`
Role today:
- audit of current end-to-end workflow behavior and lifecycle

Strengths:
- important snapshot of current system behavior
- useful baseline for regression-aware refactor planning

Problems:
- likely a point-in-time audit rather than a stable long-term spec
- risks becoming stale as route and workflow internals evolve
- should not be treated as the only source of truth once more targeted specs exist

Recommendation:
- keep, but classify as **audit snapshot**
- do not continue expanding it as a living spec
- later archive under an audit/history section or prefix

Best future role:
- behavioral baseline reference for refactor comparison

---

## 2. Test inventory audit

### `tests/test_workflow_persistence_service.py`
What it covers:
- deterministic review code format
- work-log status normalization
- work-log action type normalization
- mocked end-to-end `WorkflowPersistenceService.start_review_run()` behavior:
  - review session persisted
  - findings created
  - assignments created
  - work logs normalized/persisted
  - summary fields returned

Strengths:
- true unit/service-style tests
- uses fake postgres client, avoiding live dependencies
- validates internal normalization logic and orchestration outcomes
- durable starting point for refactor safety

Gaps:
- no direct policy-branch coverage
- no negative-path coverage for blocked/manual-review behavior
- no coverage for summary fields introduced in newer workflow contract docs, such as:
  - `policy_checks`
  - `policy_summary`
  - `blocked_reasons`
  - `lifecycle_outcomes`
  - `trigger_summary`
  - `post_close_actions`
  - `operational_state_snapshot`
- no explicit tests for resumability / partial lifecycle outcomes
- currently one file mixes utility-style tests and orchestration-flow tests

Recommendation:
- later split into focused test modules by responsibility

### `tests/test_dossier_review_routes.py`
What it covers:
- route for workflow start returns structured summary
- route-level dossier review CRUD-ish lifecycle:
  - create review
  - add finding
  - assign
  - submit
  - verify
  - close
- list reviews route returns schema-compatible payload

Strengths:
- route contract coverage exists
- uses patched services/clients to avoid real infra
- catches important API response-shape expectations

Gaps:
- many endpoints exercised in one long route lifecycle scenario
- limited assertions on response contract shape breadth
- no route-level error-path tests
- no policy/manual-review/block/circuit-breaker scenarios
- no SSE route checks
- no dedicated tests for operations endpoints or operational-state contract routes in `tests/`
- route tests are doing both contract validation and lifecycle scenario validation

Recommendation:
- later separate route contract tests from workflow lifecycle scenario tests

---

## 3. `tmp/` verification audit

### `tmp/check_health_endpoints.py`
Role today:
- smoke check for basic JSON endpoints and SSE preview

Strengths:
- easy manual operator smoke test
- useful for checking server wiring and endpoint availability

Problems:
- depends on a running local server
- no formal pass/fail assertions except implicit failure on exceptions
- mixes:
  - health checks
  - operational-state summary preview
  - SSE preview
- belongs to a smoke/live verification category, not `tmp/` forever

Recommended category:
- live smoke verification

### `tmp/diagnose_dossier_review_full_lifecycle.py`
Role today:
- posts a realistic workflow-start payload against live API
- expects automation to complete assign-submit-verify-close
- prints summarized outcome and returns non-zero on failure

Strengths:
- realistic integration diagnostic
- useful after behavior-changing refactors
- validates actual system wiring beyond mocks

Problems:
- assumes seeded data and active local server
- hard-codes environment and project/persona values
- behaves like a live integration scenario, not a temporary scratch script
- success criteria are scenario-specific and brittle if policy behavior changes intentionally

Recommended category:
- live integration diagnostic / scenario verification

### `tmp/verify_workflow_summary_contract.py`
Role today:
- live contract verifier for `WorkflowReviewRunSummary`
- checks presence of top-level and nested required fields

Strengths:
- directly protects response contract
- especially valuable while refactoring large workflow/persistence code
- returns structured failure report

Problems:
- requires live server and specific workflow behavior
- field-presence only; does not validate semantic invariants deeply
- stored in `tmp/` despite being a durable contract check

Recommended category:
- live contract verification

### `tmp/verify_company_operational_state_contract.py`
Role today:
- live contract verifier for `CompanyOperationalStateResponse`
- checks top-level fields, `policy_overview`, and required `source_status` source names

Strengths:
- good additive contract protection
- guards one of the most central aggregate payloads

Problems:
- requires live server and source availability assumptions
- mixes contract presence with expected source inventory
- belongs in a durable verification directory, not temporary scratch space

Recommended category:
- live contract verification

---

## Core structural problem

The repo currently lacks a formal distinction between these layers:

### A. Unit tests
Characteristics:
- no running server
- no real database/network
- fast, deterministic
- validate pure logic, transformations, orchestration boundaries

Current examples:
- most of `tests/test_workflow_persistence_service.py`

### B. Route/integration tests with mocks
Characteristics:
- instantiate FastAPI app or service layer
- patch infrastructure
- validate HTTP response shapes and orchestration wiring
- still deterministic and CI-friendly once pytest/unittest is available

Current examples:
- most of `tests/test_dossier_review_routes.py`

### C. Live contract verification
Characteristics:
- require running app
- often require seeded or real-ish data
- validate production-like response contracts remain intact
- should fail loudly but are not replacements for unit tests

Current examples:
- `tmp/verify_workflow_summary_contract.py`
- `tmp/verify_company_operational_state_contract.py`

### D. Live smoke / diagnostics
Characteristics:
- operator-focused checks
- endpoint reachability, SSE preview, scenario probes
- useful pre/post deploy or local debug
- may be less stable than contract verification

Current examples:
- `tmp/check_health_endpoints.py`
- `tmp/diagnose_dossier_review_full_lifecycle.py`

Until these are separated, future refactor work will continue to confuse:
- "test coverage"
- "contract verification"
- "manual diagnostics"
- "operator smoke checks"

---

## Proposed durable documentation structure

## Recommended doc categories

### 1. Architecture proposals
Purpose:
- why a subsystem exists
- intended direction
- target operating model

Recommended docs:
- keep `docs/operations-case-management-upgrade-proposal.md`
- later add or derive:
  - `docs/architecture/operations-case-management.md`
  - `docs/architecture/command-center.md`

### 2. Contracts/specs
Purpose:
- authoritative payload/object contracts
- additive response expectations
- safe reference during refactor

Recommended future docs:
- `docs/contracts/readiness-explainability.md`
- `docs/contracts/workflow-review-summary.md`
- `docs/contracts/company-operational-state.md`
- `docs/contracts/ai-planning.md`
- `docs/contracts/command-center-projections.md`

### 3. Operations playbooks/runbooks
Purpose:
- operator guidance
- guardrail interpretation
- rollout / incident handling

Recommended future docs:
- keep `docs/policy-operations-playbook.md`
- later split into:
  - `docs/runbooks/policy-operations.md`
  - `docs/runbooks/policy-change-checklist.md`

### 4. Audit snapshots
Purpose:
- point-in-time assessments
- regression baselines
- historical references, not living source of truth

Recommended future placement:
- `docs/audits/dossier-review-full-flow-audit.md`
- `docs/audits/refactor-baseline-YYYY-MM.md` for future snapshots

### 5. Refactor plans
Purpose:
- phased split plans
- risks, order, and validation expectations

Recommended future placement:
- `docs/refactors/backend-refactor-phases.md`
- `docs/refactors/test-strategy.md`
- this audit file can later move under that area if desired

---

## Recommended immediate doc actions

## Keep as primary docs
- `docs/operations-case-management-upgrade-proposal.md`
- `docs/policy-operations-playbook.md`
- `docs/step5-command-center-readiness-planning.md`

## Reclassify as audit snapshot
- `docs/dossier-review-full-flow-audit.md`

## Add an index layer
A future lightweight index doc should list:
- architecture docs
- contract docs
- runbooks
- audit snapshots
- verification scripts and when to run them

Recommended filename:
- `docs/README.md` or `docs/index.md`

## Highest-value future split
If only one large doc is split first, split:
- `docs/step5-command-center-readiness-planning.md`

Suggested split targets:
- `docs/contracts/command-center-projections.md`
- `docs/contracts/readiness-explainability.md`
- `docs/contracts/ai-planning.md`
- keep original Step 5 file as overview/roadmap

---

## Proposed durable test and verification structure

## 1. Unit tests under `tests/unit/`
Purpose:
- fast deterministic logic checks

Suggested future files:
- `tests/unit/services/test_workflow_persistence_review_code.py`
- `tests/unit/services/test_workflow_persistence_work_logs.py`
- `tests/unit/services/test_workflow_persistence_start_run_happy_path.py`
- `tests/unit/services/test_workflow_persistence_policy_outcomes.py`
- `tests/unit/services/test_workflow_persistence_lifecycle_outcomes.py`

Potential content migration:
- split current `tests/test_workflow_persistence_service.py` across these files

## 2. API/route tests under `tests/api/`
Purpose:
- route contracts with patched services/clients

Suggested future files:
- `tests/api/test_dossier_review_workflow_routes.py`
- `tests/api/test_dossier_review_session_routes.py`
- `tests/api/test_dossier_review_assignment_routes.py`
- `tests/api/test_dossier_review_close_routes.py`
- `tests/api/test_company_operational_state_routes.py`
- `tests/api/test_operations_routes.py`

Potential content migration:
- split current `tests/test_dossier_review_routes.py` by endpoint group

## 3. Integration tests under `tests/integration/`
Purpose:
- higher-level app/service integration with fakes, seed fixtures, or temporary DB later
- still more CI-oriented than live scripts

Suggested future files:
- `tests/integration/test_dossier_review_full_lifecycle.py`
- `tests/integration/test_workflow_to_operations_projection.py`
- `tests/integration/test_company_operational_state_aggregation.py`

Note:
- these should differ from live scripts by controlling dependencies as much as possible

## 4. Live verification scripts outside `tmp/`
Purpose:
- explicit, opt-in environment-backed verification
- not confused with CI unit coverage

Suggested future directory:
- `tools/verification/`

Suggested future files:
- `tools/verification/check_health_endpoints.py`
- `tools/verification/verify_workflow_summary_contract.py`
- `tools/verification/verify_company_operational_state_contract.py`
- `tools/verification/diagnose_dossier_review_full_lifecycle.py`

Optional subfolders if desired:
- `tools/verification/contracts/`
- `tools/verification/smoke/`
- `tools/verification/scenarios/`

---

## Recommended separation rules

## Unit tests
Use when validating:
- helper normalization
- mapping logic
- policy summary aggregation
- blocked reason collection
- lifecycle outcome construction
- transformations from persistence records to response models

Should not require:
- running FastAPI server
- network calls
- real Postgres

## Route/API tests
Use when validating:
- endpoint status codes
- response schema shape
- route-to-service wiring
- error mapping
- backward-compatible field presence

May use:
- `TestClient`
- patched services
- fake client classes

Should not require:
- live local server
- manually seeded environment

## Integration tests
Use when validating:
- multi-step workflow behavior across service boundaries
- operational-state aggregation from multiple collaborators
- policy + persistence + route interactions together

Prefer:
- controlled test fixtures
- fake or temp infrastructure
- deterministic data setup

## Live contract verification
Use when validating:
- deployed/local-running API still exposes required fields
- critical response contracts survive refactor
- source inventory and aggregate payloads remain operable

Should be clearly labeled:
- requires running server
- requires seed data
- not a replacement for unit tests

## Live smoke/diagnostic checks
Use when validating:
- endpoint reachability
- SSE stream emits
- one realistic scenario still executes end to end

Should be treated as:
- operational tooling
- local debugging or pre-release checks

---

## Specific test gaps to fill

## Highest-priority gaps

### Workflow persistence service
Missing future unit tests:
- policy blocked session result produces expected `automation_status`
- finding-level block creates manual follow-up instead of auto progression
- `blocked_reasons` composition from multiple scopes
- `policy_summary` counts and `highest_risk_level`
- `trigger_summary` echo and metadata handling
- `post_close_actions` generation
- `operational_state_snapshot` attachment shape
- mixed outcome runs where some findings auto-progress and others do not

Suggested future files:
- `tests/unit/services/test_workflow_persistence_policy_outcomes.py`
- `tests/unit/services/test_workflow_persistence_contract_fields.py`

### Dossier review routes
Missing future route tests:
- invalid request returns validation error shape
- blocked workflow still returns structured summary contract
- start route includes expanded workflow summary fields
- list/detail routes preserve backward-compatible payload shape
- error path when persistence client raises exceptions
- SSE operational-state route content type and event framing

Suggested future files:
- `tests/api/test_workflow_start_route_contract.py`
- `tests/api/test_workflow_start_route_policy_paths.py`
- `tests/api/test_dossier_review_error_paths.py`
- `tests/api/test_operational_state_stream_route.py`

### Operational-state contract
Missing future tests:
- mocked route test for additive top-level fields:
  - `operations_cases`
  - `operations_actions`
  - `operations_blockers`
  - `dossier_reviews`
  - `policy_overview`
- summary count consistency checks
- source fallback behavior
- policy overview aggregation logic

Suggested future files:
- `tests/api/test_company_operational_state_contract.py`
- `tests/unit/services/test_company_operational_state_policy_projection.py`

### Operations endpoints
Current gap:
- no audited `tests/` coverage for:
  - `/v1/operations/cases`
  - `/v1/operations/actions`
  - `/v1/operations/blockers`

Suggested future files:
- `tests/api/test_operations_cases_routes.py`
- `tests/api/test_operations_actions_routes.py`
- `tests/api/test_operations_blockers_routes.py`

---

## Recommended handling for current `tmp/` scripts

## Reclassify, do not delete
These scripts are useful. The issue is their location and implied status.

### Move later to durable verification area
- `tmp/check_health_endpoints.py` -> `tools/verification/check_health_endpoints.py`
- `tmp/diagnose_dossier_review_full_lifecycle.py` -> `tools/verification/diagnose_dossier_review_full_lifecycle.py`
- `tmp/verify_workflow_summary_contract.py` -> `tools/verification/verify_workflow_summary_contract.py`
- `tmp/verify_company_operational_state_contract.py` -> `tools/verification/verify_company_operational_state_contract.py`

### Add a simple convention header later
Each live script should state:
- purpose
- requires running server: yes/no
- requires seeded data: yes/no
- expected exit codes
- whether it is smoke, contract, or scenario verification

---

## Phase-based refactor/test planning for docs and verification

## Phase A - Classification and indexing
Goal:
- create one durable map of docs/tests/verification assets

Actions:
- add docs index
- classify docs as proposal, contract, playbook, or audit
- classify scripts as smoke, contract, or scenario verification

Risk reduced:
- avoids accidental reliance on stale docs
- gives future refactor one navigation layer

## Phase B - Separate unit vs API vs live verification
Goal:
- stop mixing deterministic test coverage with manual live checks

Actions:
- split `tests/test_workflow_persistence_service.py`
- split `tests/test_dossier_review_routes.py`
- define `tests/unit`, `tests/api`, `tests/integration`
- move `tmp/` verification scripts into `tools/verification`

Risk reduced:
- improves CI readiness
- makes failures easier to interpret

## Phase C - Contract-first protection for large refactors
Goal:
- protect high-risk response surfaces before file splitting starts

Actions:
- add route-level tests for:
  - workflow summary contract
  - company operational state contract
  - operations route contracts
- preserve live verifiers as additional safety net

Risk reduced:
- large-file refactor can proceed with less contract drift

## Phase D - Historical doc cleanup
Goal:
- reduce duplication and stale narrative drift

Actions:
- split Step 5 doc into contract-focused docs
- archive audit snapshots under `docs/audits/`
- leave proposal/playbook as stable top-level references or move under structured folders

Risk reduced:
- future implementers read one authoritative contract per concern

---

## Risks and cautions

## 1. False confidence from live scripts
Current `tmp/` scripts can pass while unit coverage remains weak. They should supplement, not replace, deterministic tests.

## 2. Contract drift hidden in narrative docs
Several payload and behavior expectations live only inside long prose docs. Those expectations should eventually be mirrored by focused contract docs and route tests.

## 3. Over-centralized mega-docs
`step5-command-center-readiness-planning.md` is valuable, but if it keeps growing it becomes harder to tell:
- what is current contract
- what is exploratory
- what is deferred roadmap

## 4. Stale audit snapshots treated as live truth
`dossier-review-full-flow-audit.md` should be retained as baseline evidence, but not used as the sole living contract during implementation.

## 5. Mixed test intent
A single test file that tries to validate utility logic, orchestration, and full route lifecycle becomes harder to maintain during refactor.

---

## Recommended future inventory summary

## Docs to keep and classify
- `docs/operations-case-management-upgrade-proposal.md` — foundational proposal
- `docs/step5-command-center-readiness-planning.md` — umbrella roadmap/spec to later split
- `docs/policy-operations-playbook.md` — operational playbook/runbook
- `docs/dossier-review-full-flow-audit.md` — audit snapshot to later archive

## Tests to split later
- `tests/test_workflow_persistence_service.py`
- `tests/test_dossier_review_routes.py`

## Live verification scripts to relocate later
- `tmp/check_health_endpoints.py`
- `tmp/diagnose_dossier_review_full_lifecycle.py`
- `tmp/verify_workflow_summary_contract.py`
- `tmp/verify_company_operational_state_contract.py`

---

## Parent-agent-ready recommendations

1. **First doc organization move**: introduce a docs index and classify current docs by type before adding more planning prose.
2. **First test organization move**: split route tests and workflow persistence tests by responsibility, not by subsystem name only.
3. **First verification move**: treat current `tmp/` scripts as durable live verification assets and move them under `tools/verification` in a later implementation pass.
4. **First contract hardening move**: add deterministic route tests for workflow summary and company operational state contracts before refactoring large service files.
5. **First archival move**: preserve `docs/dossier-review-full-flow-audit.md` as a historical audit snapshot, not an evolving spec.

---

## Suggested future command matrix

### Fast deterministic tests
- unit/service logic
- route contract tests with mocks

### Controlled integration tests
- multi-step workflow lifecycle with fakes/fixtures

### Manual or pre-release live verification
- health and SSE smoke check
- workflow summary contract check
- company operational state contract check
- end-to-end lifecycle scenario diagnostic

This separation is the key improvement needed to make later refactor work safer and more durable.