# Forge Testing

> Evaluate whether tests provide reliable risk-based evidence across units, boundaries, workflows, and failure modes.

- Skill: `is-bo/forge-testing-3` (Agent Skill)
- Install (CLI): `npx skillmds@latest add is-bo/forge-testing-3`
- Raw SKILL.md: https://api.skillmd.com/api/skills/is-bo/forge-testing-3/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: is-bo (https://skillmd.com/u/is-bo)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/is-bo/forge-testing-3

---


# forge-testing: Testing strategy

Engine: Upstream-powered — Addy Osmani Agent Skills

## Purpose

Evaluate whether tests provide reliable risk-based evidence across units, boundaries, workflows, and failure modes.


## Deterministic runtime composition

Before loading any provider procedure, run:

Resolve `../../runtime/cli/src/composition-entry.js` relative to this `SKILL.md`, then run:

`node "<resolved-absolute-runner-path>" testing compose --workflow audit --root "<repository-root>" --dry-run --json`

Add one repeatable `--request <provider-or-source>` flag for each explicit user request. Add
`--condition <task-condition>` or `--risk-surface <surface>` only for a task fact you directly
proved; never infer one from generic wording. The command above is the default for this
audit-oriented module; for implementation use `--workflow build`, and for a fix, retest, or
release gate use `--workflow fix`, `verify`, or `ship` respectively. Read the JSON response,
keep the Forge contract at index zero, and resolve paths against the absolute `runtime_root`
reported in that response. Read `eager[].runtimePath` when entering the module. The full
`selected[]` list is availability/provenance; load only `deferred[].runtimePath` when the task
reaches that concern, in tier order. Refuse any path that escapes the root. Respect every reported
suppression and context budget. If `missing` is non-empty, stop and report the installation as
damaged; do not improvise a prose fallback. The runner and specialist content may live in a plugin
cache or global installation; never assume they are inside the audited repository.


Resolve and read `../fullstack-forge/references/shared/module-contract.md` (applicability,
execution, mutation, verification, completion) and
`../fullstack-forge/references/shared/evidence-rules.md` (statuses, standards, tools, findings via
`../fullstack-forge/references/PROTOCOL.md`) relative to this module `SKILL.md` before reporting.

Never hide failed checks or claim that an operation ran when it did not.

## Automatic activation signals

Activate when a request or direct repository evidence involves testing strategy, when
the user explicitly names `forge-testing`, or when discovery proves an applicable boundary.

- Executable software
- Release readiness

## When not to activate

- Non-executable documentation-only packages

## Automated support

Relevant discovery inputs are:

- test manifests and commands
- coverage configuration
- critical workflow inventory

Deterministic support, bounded evidence only:

- `detect-project-commands`
- `run-project-command`

## Agent inspection procedure

1. Map the test pyramid: what exists at unit, integration, API, end-to-end, and evaluation level, and what each layer actually asserts.
2. Trace the riskiest workflows from discovery to their covering tests; record critical paths with no failure-path or authorization test.
3. Inspect test quality: assertions that prove behavior versus existence, mock realism, isolation, and determinism (hunt flaky patterns).
4. Verify negative coverage: unauthorized access, invalid input, concurrency, retries, and idempotency for the paths that claim them.
5. Run the suite and record counts, duration, skips, and any tests that cannot fail.

Manual inspection requirements:

- Review omitted high-impact scenarios and test maintainability
- Inspect CI artifacts and quarantine policy

Stack-specific guidance:

- Use native test isolation and real boundary substitutes such as ephemeral databases where practical

## Evidence to collect

Standards used as criteria:

- NIST SSDF
- testing-pyramid and contract-testing concepts

## Common production failures

- Map critical risks to unit, integration, contract, end-to-end, migration, security, and accessibility tests
- Inspect determinism, isolation, data factories, assertions, cleanup, time control, concurrency, and flaky retries
- Verify tests can fail for the defect they claim to detect and do not overmock the boundary under test

## Missing-control checks

Each item needs direct evidence or one reasoned status.

- Unit tests
- Integration tests
- API tests
- Database tests
- Authorization tests
- Tenant-isolation tests
- Upload tests
- Malware-pipeline tests
- End-to-end tests
- Accessibility tests
- Visual-regression tests
- Migration tests
- Failure-path tests
- Retry tests
- Idempotency tests
- Concurrency tests
- Offline tests
- Payment tests
- AI evaluation tests
- Test isolation
- Flaky tests
- Mock quality
- Critical workflow coverage
- Production-like test configuration
- Risk-based adequacy rather than line coverage alone

## Commands and tools

- Run `forge testing audit --json` or `fullstack-forge testing audit --json` when
  an explicit audit is requested and the CLI is installed. Normal feature work does not require it.

## Safe fixes

- Add missing assertions and deterministic setup
- Remove an unnecessary retry only after proving the test is stable

## Approval-required changes

- Deleting coverage, weakening assertions, or changing product behavior to satisfy tests

## Verification

- Run targeted tests before and after inducing a representative failure
- Run the relevant full suite after final edits

## Completion contract

Follow `fullstack-forge/references/shared/completion.md` and the limitations below.

## Known limitations

- Coverage percentage alone does not establish behavior coverage

