This is the multi-page printable view of this section. .
Workflows
1 - Evaluate recent Sessions
Historical evaluation is the cold start when you have real Agent interactions but no curated Dataset. It is diagnostic, not promotion evidence.
Web path
- Open Historical Sessions in the DSH Harbor page.
- Preview up to three recently completed top-level Sessions visible to the current DSH workspace.
- Review the frozen Judge identity, projected evidence fields, redaction policy and cost/data disclosure.
- Select Sessions and confirm.
- The Plugin writes a private redacted Batch, materializes one Harbor Trial per Session and starts a non-promotion Historical Job.
The Agent-tool path can preview up to ten Sessions, but only from the exact current working directory. Preview returns safe metadata and a short-lived owner-bound selection token—not raw Session ids or transcripts.

What reaches the Judge
Bounded, projected and credential-shaped-redacted evidence can be sent to the selected Judge model. Raw Session ids, full tool payloads, reasoning and attachments are not directly used as Judge input. Private Batch and Job artifacts remain local and may contain business evidence or ordinary absolute paths.
Do not describe this as “data never leaves the machine.”
How to read the result
- one selected Session becomes one Trial;
- criteria can be valid, invalid or abstained;
- coverage and reason codes matter alongside scores;
completed-unscoredis a meaningful outcome, not success;- Candidate rerun, comparable baseline and Gate are not applicable to this path.
Use repeated failures to curate a Dataset, then move to the Candidate pipeline.
Recovery boundary
Historical Web operation state and locks are process-local today. A Host restart can lose reattachment even when artifacts remain on disk. This differs from durable @harbor snapshots and from other operation journals; recovery guarantees must be stated per operation.
2 - Candidate evaluation and promotion
Candidate evaluation answers a narrower question than Historical diagnosis: did one controlled change improve a fixed business task under comparable evaluation conditions?
Strict sequence
- Snapshot Candidate into an immutable manifest.
- Validate Dataset identity, task uniqueness, paths, sensitive metadata and source digest.
- Doctor Candidate, Dataset, Evaluation Stack and optional Promotion Policy.
- Preview Context v3 and discover comparable baselines before spending on a Job.
- Run diagnostic first when the stack or provider path is new.
- Change one controlled surface—Agent or Evaluator, not both silently.
- Run promotion-eligible regression with fixed identities.
- Inspect progress, Trial output, criterion evidence and governance impact.
- Compare and Gate against a comparable baseline.
Comparability
A baseline is not comparable merely because it used the same repository. Candidate manifest, Dataset manifest, Evaluation Stack, Context, execution environment and relevant model/Judge identities must satisfy the contract. A changed Dataset digest, stack version, provider identity or runtime boundary can require a fresh baseline.
Gate
The Gate is deterministic for fixed baseline Job, Candidate Job and policy inputs. It can return PROMOTE or REJECT based on valid score movement, minimum improvement, regressions, coverage and other policy conditions.
It does not deploy, mutate the Champion or bypass external CI/CD approval.
Version 0.9.7 accepts Candidate Context v3 in both Dashboard overview and Job detail. Compare/Gate still remain capability-gated by artifact validity, comparable identities, mode and policy.
What a release test does not prove
0.9.7 did not run a real provider model, real Candidate/Historical Session data, or a paid Harbor evaluation; automated tests do not establish a business-quality baseline. Your Dataset, Evaluator and production evidence must establish that baseline.
3 - Evaluator governance and meta-evaluation
Candidate quality and Evaluator quality are separate governance problems.
Interface and inspection
Formal Candidate Evaluators implement harbor-dsh-evaluator/v2. A descriptor identifies implementation kind (script or llm-as-judge), ternary Criteria and a bounded allowlist of editable source files. Inspection omits secret-shaped and local-path-shaped values.
Controlled update
An Evaluator update replaces one descriptor-authorized source file under optimistic concurrency. The caller supplies the expected digest and new Evaluator and Stack versions. The update never runs evaluation or Gate automatically.
Independent Ground Truth
Ground Truth may be human, programmatic, consensus, model or external, but provenance must be explicit and independent of the Candidate Evaluator. A draft is non-overwriting and identifies the criteria it covers.
Meta-evaluation
Repeated Evaluator observations are compared with independent Ground Truth to produce:
- ESF — evaluator score fidelity;
- SCE — score calibration error;
- RCR — ranking consistency/reliability.
The Evaluator needs train, validation and test boundaries too
Rubrics, Judge prompts, parsers and thresholds can all be “trained”:
- tuning set — find false positives/negatives and modify rubric, prompt or script;
- validation set — compare Evaluator versions and select thresholds or implementations;
- meta-evaluation holdout — reveal independent cases only after identities are frozen to estimate reliability on unseen cases.
If the same human labels guide Evaluator changes and then serve as final proof of accuracy, information has leaked. Raw human review needs independent provenance and must not be rewritten as if it came from the Candidate Evaluator.
Governance order in the Plugin
harbor_evaluator_inspectreads the interface, Criteria and editable source boundary.harbor_ground_truth_initcreates a non-overwriting independent Ground Truth draft with provenance.harbor_evaluator_meta_evaluatecompares repeated observations with Ground Truth and writes ESF, SCE and RCR.- Only when evidence supports a change,
harbor_evaluator_updatereplaces one authorized file under an expected digest and requires new Evaluator/Stack versions. - Rerun tuning and holdout meta-evaluation before using the new Evaluator for Candidate evaluation.
Meta-evaluation does not automatically update the Evaluator or run Candidate Gate. Historical evaluation is a distinct scenario: applicability, coverage and abstention matter because real Sessions may not exercise every criterion.