diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 2d2eb16..ae23694 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -27,8 +27,20 @@ jobs: - name: Test (unit — no live tenant/model/network) run: pytest -m "not live and not bench" --cov=vpcopilot --cov-report=term-missing --cov-fail-under=65 - # Nightly + manual only: the live smoke and the discovery bench need real XC/model creds and are - # never gated on a PR. The bench fails the job if discovery quality drops below BASELINE.md. + # Nightly + manual only. What this job checks is the BENCH-GATE MECHANISM: that a scored run still + # clears the floors in `bench/BASELINE.md`. It scores a deterministic fixture, so it needs no + # credentials and spends no model quota. + # + # It used to describe itself as a "live smoke + bench-gate" and was handed ANTHROPIC_API_KEY, + # XC_API_URL, XC_API_TOKEN and XC_NAMESPACE. It never used any of them: `pytest -m "live or bench"` + # collects exactly ONE test of 739, because **no test in the suite carries `@pytest.mark.live`** — + # the marker is declared in pyproject.toml and unused. So a green tick every morning attested to + # tenant and model coverage that did not exist, which is the same "we did not check this" / + # "this is clean" confusion the pipeline itself refuses to make. The selector, the name and the + # secrets now match what actually runs. + # + # Real live tests behind the `live` marker are tracked in BACKLOG.md; restore the credentials and + # widen the selector back to "live or bench" when they exist. nightly: if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch' runs-on: ubuntu-latest @@ -38,10 +50,5 @@ jobs: with: python-version: "3.12" - run: python -m pip install -e ".[deploy,console,dev]" - - name: Live smoke + bench-gate - env: - ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} - XC_API_URL: ${{ secrets.XC_API_URL }} - XC_API_TOKEN: ${{ secrets.XC_API_TOKEN }} - XC_NAMESPACE: ${{ secrets.XC_NAMESPACE }} - run: pytest -m "live or bench" + - name: Bench-gate (scores a fixture against BASELINE.md — no tenant, no model) + run: pytest -m bench diff --git a/BACKLOG.md b/BACKLOG.md index 20c4c69..6a611d6 100644 --- a/BACKLOG.md +++ b/BACKLOG.md @@ -26,5 +26,29 @@ Loose ideas, not yet scheduled. Anything with a shape clear enough to plan again `correlate.py` `coverage_key` collapses LB-wide controls and keys `service_policy` per endpoint; the pipeline skips a band-aid an earlier finding already covers and writes `correlations.json`. +- **Real live tests behind the `live` marker.** The marker is declared in `pyproject.toml` + ("hits a real XC tenant / model / network — excluded from CI (run manually / nightly)") and + **nothing uses it**: `grep -rn "mark.live" tests/` returns zero hits, so `pytest -m "live or bench"` + collects exactly one test, and that one scores a synthetic fixture. Every roadmap item so far has + been proven by hand against the tenant — the G2 canary, I1's origin probe, I2's `banknimbus-dev` + diff, H1/H2's OSV behaviour, K1's MCP client handshake, K2's 76-second PR review — and **not one of + those proofs is repeatable by anyone but the person who ran it**. That is the gap: the demo's + central claim is "we ran it for real", and nothing in the repository re-runs it. + - Shape: a `tests/test_live_*.py` set marked `@pytest.mark.live`, each **restoring what it + mutates** and skipping cleanly (`pytest.skip`) when its credential is absent, so a contributor + without a tenant sees "skipped", never "failed" and never a false pass. + - Candidates, in the order they would pay off: the safety spine end to end on `vpcopilot-lab` + (create → attach → validate → rollback, asserting the LB returns to its snapshot); `drift.check` + read-only against a real LB; the OSV client against `api.osv.dev` (alias-following, CVSS v4, + `upgrade_target_for`) — needs no XC and no model, so it is the cheapest first step; and the MCP + server's handshake against a real client. + - **Deliberately NOT simply widening the nightly selector.** Doing that without writing the tests + reinstates exactly the problem this entry exists to record: a job whose name promises live + coverage it does not have. When these land, restore the four secrets to the nightly job and widen + `pytest -m bench` back to `pytest -m "live or bench"`. + - Note the cost: a nightly that mutates a real tenant needs the lab LBs to be free, and + `nimbus-www` stays protected. Budget for flakiness that is the tenant's, not the code's. + _(Requested 2026-07-30, after the nightly job was found to be attesting to coverage it never had.)_ + _The four evidence items — sign the bundle, `export --verify`, audit event sink, attribution backfill — moved to `ROADMAP.md` **J1–J4**._