Skip to content

From real Sessions to controlled promotion

A six-chapter guide from real failure evidence to one comparable Candidate regression and an external deployment handoff.

From real Sessions to controlled promotion

A six-chapter guide from real failure evidence to one comparable Candidate regression and an external deployment handoff.

This guide follows one honest evolution loop. It separates diagnosis from promotion and keeps three questions visible in every chapter: what you do, what Harbor records, and what the result does not prove.

Contents

1 Why self-editing is not progress

Improvement requires a fixed task, valid evidence, one controlled change and a policy outside the Optimizer.

An Agent can rewrite its prompt, tools or Evaluator and still become worse. “It changed” is not “it improved.”

A trustworthy loop fixes the business cases, execution identity and scoring contract before the change. It separates the Optimizer that proposes a change from the Gate that judges promotion.

You do

Name the business failure and the owner of the final deployment decision.

Harbor records

The identities and evidence chain needed to revisit the claim.

This does not prove

A quick diagnostic or rewritten Candidate does not prove promotion quality.

2 Define the four concepts

Understand Dataset, Generator, Evaluator, Optimizer, train/validation/test splits and meta-evaluation in business language.

Self-evolution is not “letting the model edit itself.” It is an inspectable learning chain: Dataset defines the problem, Generator produces results, Evaluator judges evidence, and Optimizer proposes one controlled change. The Evaluator is itself tested through Meta-Evaluation.

Meet the four roles

  • Dataset — business tasks and failure cases that answer what to evaluate.
  • Generator — runs the Candidate and produces answers or Artifacts: who answers and how.
  • Evaluator — applies criteria and Evidence requirements: what counts as good and whether evidence is sufficient.
  • Optimizer — proposes one new-version change from failure evidence: what to change next.

In an Agent system, each role is broader than a model. Generator includes prompts, Skills, tools and runtime; Evaluator includes rubric, Judge, parsing and validity rules; Optimizer may be a Coding Agent constrained by the project Skill.

Why separate train, validation and test data

Data splits prevent taking an exam after seeing its answers.

LayerPurposeMay it guide changes?
Training/fix setExpose known badcases and modify prompts, Skills, code or tool policyYes; that is its purpose
Validation/regression setCompare Candidates, tune thresholds and select an approachResults are visible, so it can be overfit
Test/holdout setEstimate unseen-case performance after identities are frozenAnswers should remain hidden from the Optimizer before the final run

Once a case changes the Candidate, it is no longer an unseen test case. Small datasets need not be mechanically split into equal thirds, but every case still needs an explicit purpose and leakage history.

Historical Sessions are useful for discovering real failures and constructing training/fix cases; they are not promotion evidence. After review, redaction and versioning, badcases can enter a Dataset. Final Candidate regression should use a fixed Dataset, Stack, Context and baseline.

How the Plugin runs the loop

  1. Curate Dataset — validate Task uniqueness, paths, instructions and immutable source digest.
  2. Freeze Generator — record Candidate manifest, model binding and Host/Docker environment.
  3. Run Evaluator — produce criterion observations, Evidence, validity, abstention and coverage.
  4. Constrain Optimizer — allow one reviewed change on an explicit surface and create a new identity.
  5. Regress again — preserve old Jobs so baseline and Candidate evidence remain comparable.
  6. Run Gate — return PROMOTE or REJECT for fixed inputs without deploying.

Why the Evaluator must be evaluated

An Evaluator is not an oracle. A Judge can be affected by wording, position, model version or parsing failures; a deterministic script can read the wrong field.

The Plugin supports independent Ground Truth and compares repeated Evaluator observations against it to compute ESF, SCE and RCR. Data used to change a rubric forms an evaluator tuning set; evidence that the Evaluator is reliable should come from an independent holdout. A Candidate Evaluator cannot create its own labels and use them to prove itself correct.

You do

Choose representative cases, label each as training, validation or holdout, define Agent outputs, write evidence requirements for criteria, constrain the allowed change surface and review every proposal.

Harbor records

Dataset and Stack manifests, criterion and Evaluator identities, Candidate/model binding, execution environment, Trial Evidence, coverage, meta-evaluation provenance and Gate receipt.

This does not prove

A valid manifest proves structure and identity. Training improvement does not prove generalization; a high test score does not prove the Evaluator correct; and a passing Gate does not mean deployment occurred. Trustworthy improvement requires all of these boundaries.

See Concepts and trustworthy scores for the detailed glossary and Plugin mapping.

3 Diagnose recent Sessions

Use a disclosed Historical Job to find recurring failures before building a promotion Dataset.

Historical evaluation lowers the cold-start cost. Preview bounded Session metadata, inspect what may reach the Judge, then confirm a non-promotion Job.

You do

Select recent completed Sessions that reflect the failure, review exclusions and consent to the disclosed Judge/data boundary.

Harbor records

A private redacted Batch, frozen Judge identity, one Trial per Session, criterion applicability, coverage and reason codes.

This does not prove

Historical scores are not comparable Candidate regression evidence and never go through Promotion Gate.

4 Turn badcases into a Dataset

Convert recurring failure patterns into reviewable Tasks without leaking raw private history.

A Dataset is curated business intent, not a dump of private transcripts. Abstract the failure, preserve the relevant constraint and create an independently reviewable expected behavior.

You do

Deduplicate failure patterns, remove private identifiers, separate tuning from holdout and review every instruction file.

Harbor records

Task ids, paths, instructions, sensitive-metadata checks and immutable source digest.

This does not prove

A clean Dataset does not prove the population is complete or that one metric captures all business risk.

5 Run one controlled regression

Freeze identities, establish a comparable baseline, change one surface and inspect Trial-level evidence.

Snapshot the Candidate, validate the Dataset, doctor the Stack, preview Context and run diagnostic mode before promotion-eligible work when the path is new.

You do

State one hypothesis and one allowed change. Keep Dataset, criteria, provider and runtime fixed unless the change deliberately creates a new baseline.

Harbor records

Manifest digests, Context v3, model/Judge identities, Trial output, criterion Evidence, validity, coverage and governance impact.

This does not prove

A higher raw reward is not improvement if evidence is invalid, coverage falls or the baseline is not comparable.

6 Read the Gate and hand off

Interpret PROMOTE or REJECT as a policy recommendation and hand deployment authority to external CI/CD.

Compare the Candidate Job with a comparable baseline under an explicit Promotion Policy. Read the overall decision together with regressions, coverage and invalid criteria.

You do

Review representative evidence, confirm the policy matches business risk and decide whether external CI/CD should consume the recommendation.

Harbor records

Baseline/Candidate Job identities, Policy identity, comparison details and deterministic Gate artifact.

This does not prove

PROMOTE is not a production deployment, Champion mutation or universal quality guarantee. It is a recommendation for the frozen evidence and policy.

The loop ends with an explicit handoff—and begins again only when new evidence justifies another controlled change.