Skip to content

5 Run one controlled regression

Freeze identities, establish a comparable baseline, change one surface and inspect Trial-level evidence.

Snapshot the Candidate, validate the Dataset, doctor the Stack, preview Context and run diagnostic mode before promotion-eligible work when the path is new.

You do

State one hypothesis and one allowed change. Keep Dataset, criteria, provider and runtime fixed unless the change deliberately creates a new baseline.

Harbor records

Manifest digests, Context v3, model/Judge identities, Trial output, criterion Evidence, validity, coverage and governance impact.

This does not prove

A higher raw reward is not improvement if evidence is invalid, coverage falls or the baseline is not comparable.