CI Compatibility Audit
Use this skill before pushing workflow, dependency, test, build, lint, Docker, or tooling changes.
Current CI Map
Primary workflow: .github/workflows/ci.yml
Frontend job:
cd frontendnpm ci --engine-strict- frontend toolchain graph check with
node --version,npm --version,npm ls vite vitest @vitejs/plugin-react jsdom,npm exec vite -- --version, andnpm exec vitest -- --version npm run typechecknpm run lint- API smoke tests for shared API, CSRF/SSE streaming, wiki SSE UL, and auth API-base behavior
npm testnpm run build- subpath build with
VITE_APP_BASENAME=/knowledgevaultandVITE_API_URL=/knowledgevault/api
Backend job:
cd backendpip install -r requirements-ci.txt(a reduced set — it deliberately excludeslancedb,pyarrow,unstructured[all-docs], andsentence-transformers; those are stubbed at test time, see caveats below)pip install -r requirements-dev.txtruff check .pytest --tb=short -v --timeout=300 tests/— full backend suite (3918 tests since PR #215 / FR-4). Job timeout is 60m (raised from 30m in PR #215 to accommodate the 3918-test full suite + coverage step on slow CI Linux runners; local baseline is ~3-5m, CI is ~18m, plus coverage runs another ~18m).--timeout=300caps per-test hangs at 5 min. The conftest.py has 4 fixtures (CSRF bypass, rate-limiter reset, SQLite pool reset, bcrypt cache for 'pass123' test password) that the test suite relies on.- informational coverage (
continue-on-error: true) over the same suite, also with--timeout=300
Repository contract job:
python scripts/check_config_contract.pypython scripts/check_pr_scope_drift.pypython scripts/check_sast_baseline.py— fails if the SAST baseline grew on the PR (suppresses new findings withoutSAST_ALLOW_BASELINE_EXPANSION=1), if thesastCI job no longer runsrun_bandit.py, or if the scan scope (TARGET) drifted frombackend/app.python scripts/check_skill_sync.py— fails if repo-specific skills drifted across the three runner trees (.claude/skills/,.agents/skills/,.opencode/skills/). Seedocs/engineering/skill-conventions.mdfor the mirror rule andscripts/sync_skills.pyfor propagation.python scripts/check_secretscan.py— fails if.secretscanignorehas unparseable globs, adversarial positive samples are not ignored, or adversarial negative samples are over-matched. Stale and overly-broad globs emit advisory (non-fatal) warnings.python scripts/check_test_collection_scope.py— fails if a pytest test file (test_*.py/*_test.py) exists outsidebackend/tests/, wherepytest tests/would never collect it (issue #563 / C11).
SAST job:
pip install -r backend/requirements-dev.txt(getsbandit)python scripts/run_bandit.py— runs bandit against the committed baseline atbackend/security/bandit-baseline.jsonand fails only on NEW findings (pre-existing findings are suppressed by the baseline). Seebackend/security/README.mdfor the suppressed-finding counts and the baseline-regeneration workflow (python scripts/run_bandit.py --update-baseline).
Note (post-PR #215): Earlier versions of this skill documented a "narrow pytest subset" (8 test files). That subset was the pre-FR-4 state. PR #215 (issue #209) expanded CI to the full
pytest tests/suite as part of the defense-in-depth hardening. Adding files to the suite is now automatic — just add the test file. The legacy "narrow subset" concept is no longer applicable. The local mirror command below reflects this.
Checks
- Lockfiles exist and match the package manager used by CI.
- CI commands exist in package manifests or requirements files.
- Cache paths point at real lockfiles.
- Scripts do not depend on local-only absolute paths.
- Workflow shell syntax is valid on the configured runner.
- Pull request diff checks have enough fetch depth.
- Local validation commands mirror CI when possible.
- Truncated CI output does not hide the command exit status.
- SAST: if
backend/appchanged, runpython scripts/run_bandit.pylocally. If it reports NEW findings, either fix them or (if acceptable pre-existing debt) regenerate the baseline with--update-baselineand justify the newly-suppressed finding IDs in the PR. - Regression falsifiability: if the change adds a regression test, was it verified falsifiable (revert the fix, confirm the test fails, restore)?
- For the 60m job timeout: tests with
pytest-timeout=300per-test are bounded, but the cumulative suite (~36m with coverage) MUST fit. If you add tests that take cumulatively >20m, the job will fail. Profile slow tests withpytest --durations=20.
Local Mirror Commands
cd frontend && npm ci --engine-strict && npm run typecheck && npm run lint
cd frontend && npm test -- src/lib/api.test.ts src/lib/api.csrf.test.ts src/lib/api.sse.test.ts src/pages/WikiPage.sse.test.tsx src/stores/useAuthStore.api-base.test.ts
cd frontend && npm test && npm run build
cd frontend && VITE_APP_BASENAME=/knowledgevault VITE_API_URL=/knowledgevault/api npm run build
cd backend && ruff check . && pytest --tb=short -v --timeout=300 tests/
python scripts/check_config_contract.py
python scripts/check_pr_scope_drift.py
python scripts/check_sast_baseline.py
python scripts/check_skill_sync.py
python scripts/check_secretscan.py
python scripts/check_test_collection_scope.py
python scripts/run_bandit.py
Run these before pushing so a CI-only lint/type failure doesn't cost a
push → fail → fixup-commit round trip. If frontend/node_modules is absent,
run npm ci --engine-strict first.
For the backend test step, the full pytest tests/ run takes ~3-5m locally and ~18m on CI Linux. Run your changed-area tests first for fast feedback:
cd backend && pytest -q tests/<file>::<Class>::<test>
Then run the full suite before pushing.
Environment caveats (so local results aren't misread)
- CI's dependency set is reduced — "locally green" ≠ "CI green". CI installs
only
requirements-ci.txt+requirements-dev.txt, which omitlancedb,pyarrow,unstructured, andsentence-transformers. A dev machine usually has the fullrequirements.txtinstalled, so a backend test can pass locally yet fail in CI at import (ModuleNotFoundError) or behave differently. To validate a backend test-scope change (e.g. adding a new test file) faithfully, reproduce the CI env instead of trusting the local run:
This is also faster than the local suite (no multi-GB model/db loads). Corollary: a test only passes under the reduced set because something stubs the missing packages — those per-filepython -m venv /tmp/civenv /tmp/civenv/bin/pip install -r backend/requirements-ci.txt -r backend/requirements-dev.txt cd backend && /tmp/civenv/bin/python -m pytest -q tests/<candidate_file>.pylancedb/pyarrow/unstructuredstubs are load-bearing for CI, not dead boilerplate. Do not "clean them up" without confirming the file still collects under the CI venv. assert_url_safe(SSRF guard) does real DNS + blocks loopback/private. It callssocket.getaddrinfoand rejects loopback/private/link-local hosts unlessALLOW_LOCAL_SERVICES=1. Putting it on a hot path or in a Pydantic validator makes tests that use fake hostnames (*.example) orlocalhostURLs fail or stall. Validate URL changes at change-time, not on every read. (.examplefails fast withgaierror, so a hang is heavy-dep loading, not DNS.)- Python: CI pins 3.11. On a newer local interpreter (e.g. 3.14) some
backend tests fail with
RuntimeError: There is no current event loop— the test harness uses the removed implicit-event-loop pattern. These are false failures from the local interpreter, not regressions. Prefer a 3.11 venv; theruff check .lint gate and CI-targeted tests are what matter. - Backend conftest.py fixtures (post-PR #215): 4 autouse fixtures now run for
every test — CSRF bypass (CSRF-naive modules), rate-limiter reset, SQLite
pool reset (clears the singleton pool between tests), and bcrypt cache
(pre-computes the bcrypt hash for 'pass123' once per session). If a new test
hangs in CI on what looks like a pool or bcrypt issue, check whether the test
relies on the pool or auth_service in a way that the fixtures don't handle.
Pattern reference:
tests/conftest.py. - Frontend jsdom gotchas (router context for
<Link>, driving RadixSelect, virtualized lists): seereferences/frontend-testing-gotchas.mdfor the repo's established mock patterns before improvising.
Cross-platform evidence-write fallback
Some automated QA gates write evidence to .swarm/evidence/; on Windows these
writes can fail with "parent directory already contains a .swarm/ folder" or
similar path errors. When this happens, the local run is not a valid CI signal;
do not treat the gate failure as a code defect.
Fallback protocol:
- Run
ruff check .(backend) ornpm run lint(frontend) manually. - Run the targeted pytest / vitest commands manually.
- If those pass, the code is likely CI-compatible; note the evidence-write failure in the PR description or a comment.
Keep using path-safe operations (e.g., pathlib.Path) in any code you write;
this guidance is about handling pre-existing tooling path issues, not excusing
sloppy paths.
Output
Classify each risk as:
BLOCKER: likely CI failure or invalid workflow.RISK: plausible CI instability requiring targeted validation.NOTE: useful context, not blocking.
Include the exact workflow step or command for every item.