5 Run one controlled regression
Freeze identities, establish a comparable baseline, change one surface and inspect Trial-level evidence.
Snapshot the Candidate, validate the Dataset, doctor the Stack, preview Context and run diagnostic mode before promotion-eligible work when the path is new.
You do
State one hypothesis and one allowed change. Keep Dataset, criteria, provider and runtime fixed unless the change deliberately creates a new baseline.
Harbor records
Manifest digests, Context v3, model/Judge identities, Trial output, criterion Evidence, validity, coverage and governance impact.
This does not prove
A higher raw reward is not improvement if evidence is invalid, coverage falls or the baseline is not comparable.