# Test Design And Coverage

**Boundary against the error-handling discipline.** That skill writes the
failure path. This skill proves it. Decide the failure behaviour there;
assert it here. Never let a test define the behaviour it is checking.

**Pick one layer per test, by what the test owns.** Produce tests at four
layers and state the layer in the test's file or name. Unit: one module's
logic, with no I/O, no wall clock and no network. Contract: the shape and
the semantics of a boundary you publish or consume — HTTP schema, database
schema, event payload, CLI exit codes — run against the real counterpart or
a recorded one, never against a mock you wrote from memory. Integration:
your code plus one real dependency, wired the way production wires it.
End-to-end: one journey through the deployed system, asserted on what the
user sees. A reviewer checks that each new test sits at exactly one layer
and asserts at that layer only.

**Push each test down to the cheapest layer that can still fail.** Produce
an end-to-end test only for a journey whose breakage is a business failure.
Default: end-to-end tests stay under 10% of the suite [ASSUMPTION]. A test
that needs two real dependencies is an integration test at best; move it up
one layer or cut a dependency. A reviewer checks that no logic rule is
covered only by an end-to-end test.

**Trace every acceptance criterion to a named test.** Produce a mapping
table: criterion identifier on the left, test identifiers on the right.
Carry the criterion identifier inside the test name or its docstring so the
trace survives a file move. A reviewer checks that every criterion in the
story appears in the table with at least one test, and rejects any
criterion whose only evidence is a manual walk-through.

**Test the failure paths and the boundaries, not only the happy path.**
Produce, for each behaviour: the success case, one case per rejected input
class, one case per dependency failure the code handles (timeout, refusal,
malformed reply), and the edges of every range — zero, one, many, first,
last, one past the limit, empty, and the largest input the code claims to
accept. Default: a change that adds a failure path adds at least one test
that drives it [ASSUMPTION]. A reviewer counts failure-path tests against
success tests and returns any behaviour tested only on success.

**Assert on observable behaviour.** Produce assertions on returned values,
persisted state, emitted messages and rendered output. Do not assert on how
many times an internal collaborator was called. A reviewer checks that a
refactor with no behaviour change leaves the test bodies untouched.

**Name the test after the rule it defends.** Produce names that state the
condition and the expected outcome, so a reader of the CI log knows what
broke without opening the file. A reviewer checks that no test name is a
number, a bug id alone, or the name of the function under test.

**Write the test list before the code.** Produce the list of cases from the
acceptance criteria first, then implement. A reviewer checks that the test
list covers every criterion before any implementation is reviewed.
