This is the multi-page printable view of this section. .
Documentation
- 1: Install and start
- 2: Workflows
- 3: Workbench, context, and reviewed actions
- 4: Concepts and trustworthy scores
- 5: Architecture and trust boundaries
- 6: Reference
- 7: Security, privacy, and execution boundaries
- 8: Troubleshooting
Choose the path that matches your question:
- Install and start — get the Plugin, Adapter and Skill into the selected DSH profile.
- Historical diagnosis — learn from recent completed Sessions without calling them promotion evidence.
- Candidate evaluation — freeze identities, run a comparable regression and apply policy.
- Plugin Workbench — understand Context, Evidence, Artifacts and reviewed actions.
- Concepts — learn why valid score, raw reward and coverage are not interchangeable.
- Architecture — inspect protocols, execution environments and trust boundaries.
- 19-tool reference — look up every Agent tool and approval boundary.
- Security — read the limits before running untrusted Tasks.
Version baseline: Harbor Self-Evolving 0.9.7 (Beta), compatible with Harbor >=0.21,<0.22. Development-preview behavior from an untagged checkout is labeled separately.
1 - Install and start
Requirements
- A working DeepSeek Harness installation and the business Agent workspace you want to evaluate.
- Node.js/npm for the DSH Plugin setup command.
- Python environment support used by the installed Adapter.
- Harbor
>=0.21,<0.22. - Docker only if you explicitly choose container execution; 0.10.3 defaults to Host.
Install
Run from the business Agent workspace, not from this source repository:
For a version-pinned installation, replace latest with 0.10.3. Setup writes the selected DSH profile dependency, configures the Harbor project integration, installs the compatible Python Adapter and exposes the bundled evolve-agent-with-harbor Skill. Follow the exact restart command printed by setup.
Do not use dsh plugin add ./packages/dsh-plugin for a normal installation. That creates a machine-local link: dependency and omits the Adapter setup.
Verify
After restart, confirm all three surfaces:
- the selected profile depends on exact registry version
"dsh-harbor-evolution": "0.10.3", notlink:...; harbor plugins listcontainsdsh-evolutionanddsh-historical-evaluation;- the bundled
evolve-agent-with-harborSkill is present.
Then open the DSH Harbor navigation entry. A healthy installation exposes Workbench, Historical Sessions, Context and Settings rather than a standalone web server.
Choose your first path
No Dataset yet? Start with Historical diagnosis. Preview up to three recently completed Sessions in Web (or up to ten exact-cwd Sessions through the Agent tool), review the redaction/Judge disclosure, then confirm a non-promotion Job.
Candidate and Dataset ready? Follow Candidate evaluation. Snapshot, validate, doctor, preview Context, choose a comparable baseline, run one controlled regression, then apply Gate.
Source development
Only contributors modifying this repository should clone it and run:
Source builds can contain unreleased behavior and must not be presented as the formal 0.10.3 package. The earlier untagged one-click updater preview was withdrawn before 0.9.7; the browser only checks versions and copies a complete, reviewable terminal command.
2 - Workflows
2.1 - Evaluate recent Sessions
Historical evaluation is the cold start when you have real Agent interactions but no curated Dataset. It is diagnostic, not promotion evidence.
Web path
- Open Historical Sessions in the DSH Harbor page.
- Preview up to three recently completed top-level Sessions visible to the current DSH workspace.
- Review the frozen Judge identity, projected evidence fields, redaction policy and cost/data disclosure.
- Select Sessions and confirm.
- The Plugin writes a private redacted Batch, materializes one Harbor Trial per Session and starts a non-promotion Historical Job.
The Agent-tool path can preview up to ten Sessions, but only from the exact current working directory. Preview returns safe metadata and a short-lived owner-bound selection token—not raw Session ids or transcripts.

What reaches the Judge
Bounded, projected and credential-shaped-redacted evidence can be sent to the selected Judge model. Raw Session ids, full tool payloads, reasoning and attachments are not directly used as Judge input. Private Batch and Job artifacts remain local and may contain business evidence or ordinary absolute paths.
Do not describe this as “data never leaves the machine.”
How to read the result
- one selected Session becomes one Trial;
- criteria can be valid, invalid or abstained;
- coverage and reason codes matter alongside scores;
completed-unscoredis a meaningful outcome, not success;- Candidate rerun, comparable baseline and Gate are not applicable to this path.
Use repeated failures to curate a Dataset, then move to the Candidate pipeline.
Recovery boundary
Historical Web operation state and locks are process-local today. A Host restart can lose reattachment even when artifacts remain on disk. This differs from durable @harbor snapshots and from other operation journals; recovery guarantees must be stated per operation.
2.2 - Candidate evaluation and promotion
Candidate evaluation answers a narrower question than Historical diagnosis: did one controlled change improve a fixed business task under comparable evaluation conditions?
Strict sequence
- Snapshot Candidate into an immutable manifest.
- Validate Dataset identity, task uniqueness, paths, sensitive metadata and source digest.
- Doctor Candidate, Dataset, Evaluation Stack and optional Promotion Policy.
- Preview Context v3 and discover comparable baselines before spending on a Job.
- Run diagnostic first when the stack or provider path is new.
- Change one controlled surface—Agent or Evaluator, not both silently.
- Run promotion-eligible regression with fixed identities.
- Inspect progress, Trial output, criterion evidence and governance impact.
- Compare and Gate against a comparable baseline.
Comparability
A baseline is not comparable merely because it used the same repository. Candidate manifest, Dataset manifest, Evaluation Stack, Context, execution environment and relevant model/Judge identities must satisfy the contract. A changed Dataset digest, stack version, provider identity or runtime boundary can require a fresh baseline.
Gate
The Gate is deterministic for fixed baseline Job, Candidate Job and policy inputs. It can return PROMOTE or REJECT based on valid score movement, minimum improvement, regressions, coverage and other policy conditions.
It does not deploy, mutate the Champion or bypass external CI/CD approval.
Version 0.9.7 accepts Candidate Context v3 in both Dashboard overview and Job detail. Compare/Gate still remain capability-gated by artifact validity, comparable identities, mode and policy.
What a release test does not prove
0.9.7 did not run a real provider model, real Candidate/Historical Session data, or a paid Harbor evaluation; automated tests do not establish a business-quality baseline. Your Dataset, Evaluator and production evidence must establish that baseline.
2.3 - Evaluator governance and meta-evaluation
Candidate quality and Evaluator quality are separate governance problems.
Interface and inspection
Formal Candidate Evaluators implement harbor-dsh-evaluator/v2. A descriptor identifies implementation kind (script or llm-as-judge), ternary Criteria and a bounded allowlist of editable source files. Inspection omits secret-shaped and local-path-shaped values.
Controlled update
An Evaluator update replaces one descriptor-authorized source file under optimistic concurrency. The caller supplies the expected digest and new Evaluator and Stack versions. The update never runs evaluation or Gate automatically.
Independent Ground Truth
Ground Truth may be human, programmatic, consensus, model or external, but provenance must be explicit and independent of the Candidate Evaluator. A draft is non-overwriting and identifies the criteria it covers.
Meta-evaluation
Repeated Evaluator observations are compared with independent Ground Truth to produce:
- ESF — evaluator score fidelity;
- SCE — score calibration error;
- RCR — ranking consistency/reliability.
The Evaluator needs train, validation and test boundaries too
Rubrics, Judge prompts, parsers and thresholds can all be “trained”:
- tuning set — find false positives/negatives and modify rubric, prompt or script;
- validation set — compare Evaluator versions and select thresholds or implementations;
- meta-evaluation holdout — reveal independent cases only after identities are frozen to estimate reliability on unseen cases.
If the same human labels guide Evaluator changes and then serve as final proof of accuracy, information has leaked. Raw human review needs independent provenance and must not be rewritten as if it came from the Candidate Evaluator.
Governance order in the Plugin
harbor_evaluator_inspectreads the interface, Criteria and editable source boundary.harbor_ground_truth_initcreates a non-overwriting independent Ground Truth draft with provenance.harbor_evaluator_meta_evaluatecompares repeated observations with Ground Truth and writes ESF, SCE and RCR.- Only when evidence supports a change,
harbor_evaluator_updatereplaces one authorized file under an expected digest and requires new Evaluator/Stack versions. - Rerun tuning and holdout meta-evaluation before using the new Evaluator for Candidate evaluation.
Meta-evaluation does not automatically update the Evaluator or run Candidate Gate. Historical evaluation is a distinct scenario: applicability, coverage and abstention matter because real Sessions may not exercise every criterion.
3 - Workbench, context, and reviewed actions
The Harbor page is injected into the existing DSH Web GUI. It reuses DSH locale, Session, conversation, composer, Settings and Tool View surfaces; it is not a second chat product or login system.
Object-first Workbench
A Job opens into eight connected sections: Summary, Trials, Pipeline, Optimization, Compare/Gate, Evaluator/Rubric, Artifacts and Audit. The Pipeline tracks Candidate, Dataset, Integration, Renderer, Judge, Meta, Reporter, Optimizer and Gate separately so a wiring failure is not mislabeled as a low score.
The Trial explorer supports server-side pagination, status and score-validity filters, sorting, criterion evidence focus and frozen Trial sets. Business artifacts stay attached to the Trial that produced them.

Page context
When supported, ordinary send freezes the current Harbor page context. Ask AI or an explicit @harbor reference takes precedence and is one-shot. The Host resolves the opaque token for the exact Session, project and revision, then returns typed refs and a navigation action. Evidence reading validates the complete ancestry and treats content as untrusted.

Reviewed actions
Seven draft kinds cover Candidate change, Evaluator change, Compare, Diagnostic Evaluation, Infrastructure Retry, Gate Request and Deployment Handoff. Proposal creation is not authorization.
- A user explicitly requests an action.
- AI creates an expiring draft from fresh page context.
- Deterministic preflight shows target, diff, identities and limits.
- The user confirms the exact action.
- The Host journals execution and exposes progress/recovery state.
Candidate/Gate/handoff drafts do not silently execute. Compare is read-only. Unsupported production actions are denied.
Settings and operations
Settings exposes project root source, Stack/jobs/CLI checks, credential policy, execution environment and version status. Agent tools continue to use the calling Session cwd as authority; the process-local Settings root mainly serves Web and fallback behavior.
Background operations can expose progress, partial evidence, cancellation and navigation, but durability varies. Unknown state must not be rendered as success, and failure must not trigger silent retry.
0.9.7 accepts Candidate Context v3 and explicitly recognizes Historical Context v2. Compare/Gate remain gated by artifact validity, comparable identities, mode and policy; see Plugin limits.
4 - Concepts and trustworthy scores
Translate the four concepts into business language
| Concept | Practical meaning | Question answered in Harbor |
|---|---|---|
| Dataset | Business cases with task instructions, inputs and expected evidence. It defines what to evaluate; it is not a raw dump of conversations. | Which capabilities must work, and which failures must be detected? |
| Generator | The execution process that receives a Task and produces an answer or Artifact under a fixed Candidate, model and runtime. It may be an LLM Agent, script, workflow or other program. | Who answers, how is it run, and what does it produce? |
| Evaluator | The criteria and evidence logic used to judge output quality. It may be a deterministic script or llm-as-judge. | What counts as good, and is the evidence sufficient to decide? |
| Optimizer | The role that reads Dataset-level results and Trial evidence, then proposes one constrained, reviewable change. It cannot declare success by itself. | What should change, and how do we avoid changing too much at once? |
Harbor compiles these user-facing concepts into a stricter Evaluation Stack with Integration, Renderer, Judge, Contract, Reporter, Policy, Context and Gate roles. Users describe the business problem first instead of filling in an internal architecture questionnaire.
A Harbor Dataset is primarily an evaluation dataset. Whether it serves training, validation or final testing is a governance decision—not a directory name. A case cannot secretly guide optimization and still be presented as never-seen final evidence.
Train, validation and test splits
Machine learning separates data to control information leakage: what the optimization process has already seen, and whether the final result still says anything about unseen cases.
Training set
A training set directly guides learning or improvement. Traditional ML updates parameters from it; Agent engineering may use it to diagnose badcases and modify prompts, Skills, tool policies, code or retrieval configuration.
- Inputs, outputs and failure reasons may be inspected repeatedly.
- Optimizers may derive change hypotheses from it.
- Its score shows whether known failures were fixed, not generalization by itself.
- Reviewed badcases derived from Historical Sessions usually enter this layer before they become any promotion dataset.
Validation or development set
A validation set compares alternatives, selects thresholds and determines when to stop. Even without gradient updates, repeated feedback lets the Optimizer overfit to it.
- Use it to select among Candidates or configurations.
- Record any use for tuning Evaluator rubrics, Policy thresholds or runtime parameters.
- It can be an iterative regression set, but not an indefinitely reused independent final proof.
- Comparable baseline and Candidate Jobs may run on a fixed validation set, while their purpose must remain labeled as development selection or promotion evidence.
Test or holdout set
A test set estimates final generalization. Before the final run, the Optimizer, Candidate author and Evaluator-tuning process should not see its answers, Ground Truth or per-case feedback.
- Run only after Candidate, Evaluator, Policy and runtime identities are frozen.
- Minimize repeated inspection and reruns.
- Once its results guide the next change, treat it as development data and create a new holdout.
- A Promotion Gate is comparable only when baseline and Candidate use the same immutable Dataset, Stack and Context with valid evidence.
A useful data layering for Agents
| Layer | Primary purpose | Who sees feedback | Typical Plugin path |
|---|---|---|---|
| Historical diagnostic samples | Find real failure patterns and hypotheses | Humans and the diagnostic Judge | Preview completed Sessions → disclose boundary → non-promotion Historical Job |
| Training/fix set | Repair known badcases | Optimizer and developers | Curate reviewed failures into Tasks and run targeted regressions |
| Validation/regression set | Compare Candidates and select changes | Optimizer may inspect aggregates and Trial evidence | Freeze Candidate/Dataset/Stack/Context and run comparable Jobs |
| Test/holdout set | Final generalization and promotion evidence | Answers remain hidden before the final run | Run a promotion-eligible Job, then deterministic Gate |
| Evaluator meta-evaluation set | Test whether the measuring instrument is reliable | Evaluator governance process | Independent Ground Truth + repeated observations → ESF/SCE/RCR |
Small projects should preserve these logical boundaries even when they lack enough cases for statistically strong three-way splits. Record which cases changed the Candidate, which selected among alternatives, and which remain held out.
Generator: the generation process under test
A Generator is more than a model name. It includes:
- prompts, Skills, code, tools and dependencies in the Candidate;
- Candidate model binding and provider/model identity;
- Task instructions, Context and input Artifacts;
- Host or Docker execution environment;
- final output, structured results and Artifacts.
The Python Adapter maps DSH Candidates and Tasks into the Harbor Generator interface. The Plugin and Skill freeze relevant identities before expensive execution. The same model with a different prompt, tool, dependency or runtime is a different generation condition.
Evaluator: make “good” inspectable
An Evaluator decomposes the business objective into identified Criteria and evidence requirements. harbor-dsh-evaluator/v1 supports deterministic script and llm-as-judge implementations, but both must preserve distinct states:
- valid score — contract and evidence requirements were met;
- invalid score — a number exists but must not enter quality aggregation;
- abstention — the Evaluator explicitly could not decide;
- coverage — how much of the target population produced usable evidence.
A Judge is one Evaluator implementation, not an inherently correct oracle. Model version, rubric, parsing logic and source implementation belong to Evaluator identity. Any change requires new Evaluator and Evaluation Stack versions.
Optimizer: propose one controlled change from evidence
The Optimizer consumes failure patterns and evidence—not an average score detached from context. The controlled path is:
- identify repeatable failures in Trial and criterion Evidence;
- constrain the surface that may change, such as a prompt, Skill, Evaluator source or code;
- propose one reviewable change and create a new Candidate or Evaluator identity;
- rerun on a fixed Dataset and Stack;
- let Policy and Gate decide whether promotion conditions were met.
Ask AI, an Action Draft or an Optimizer proposal remains a recommendation. It does not automatically write files, start Jobs, run Gate or deploy; high-impact actions still require exact preflight and user confirmation.
Meta-evaluation: prove the measuring instrument
Ordinary evaluation asks, “Is the Generator output good?” Meta-evaluation asks, “Is the Evaluator judgment reliable?”
The Plugin compares repeated Evaluator observations with independent Ground Truth. Ground Truth may be human, programmatic, consensus, model or external, but it needs explicit provenance and independence from the Evaluator under test. Reports include:
- ESF (Evaluator Score Fidelity) — agreement with Ground Truth;
- SCE (Score Calibration Error) — calibration between confidence and observed correctness;
- RCR (Ranking Consistency/Reliability) — stability and Ground-Truth consistency of rankings.
Evaluators can overfit too. Samples used to modify a rubric, prompt or threshold form an evaluator tuning set; final evidence should come from an independent holdout. Labels produced by the Candidate Evaluator cannot be reused to prove that Evaluator correct.
How the Plugin connects the loop
- Dataset — validate manifest, Task uniqueness, paths and immutable source digest.
- Generator — freeze Candidate manifest, model binding, Context and runtime, then produce Trial output.
- Evaluator — produce criterion observations, Evidence, validity and coverage.
- Optimizer — propose one new-version change from reviewed evidence without owning the verdict.
- Meta-evaluation — test the Evaluator itself against independent Ground Truth.
- Gate — apply deterministic Policy only to fixed, comparable inputs and return a
PROMOTEorREJECTrecommendation.
The Historical path creates diagnostic hypotheses and possible future Dataset cases, but is non-promotion evaluation. Only the Candidate path can become promotion evidence after comparability, validity and Policy requirements are met.
Identity chain
A comparable Job binds immutable or versioned identities: Candidate Manifest, Dataset Manifest, Evaluation Stack, Context, execution environment, Candidate model binding and Judge identity. A Trial belongs to one Job and carries output, criterion observations, Evidence and Artifacts.
Raw reward is not a valid score
A verifier may emit a numeric raw reward even when required evidence is missing, parsing failed or the criterion abstained. Averages without validity and coverage can reward a broken pipeline.
PROMOTE and REJECT
A deterministic Gate evaluates fixed Job and Policy inputs. PROMOTE means the Candidate satisfies that Policy relative to a comparable baseline. REJECT means it does not. Neither means deployed, and neither replaces human or CI/CD authority.
5 - Architecture and trust boundaries
Three product roles
- Plugin: the user-facing DSH integration and permission boundary.
- Skill: the workflow policy that chooses the smallest safe evaluation path.
- Adapter: the Harbor runtime bridge that materializes Jobs, Trials, Evidence and Context.
The Gate remains a separate deterministic policy function; it is not part of the Optimizer and has no deployment capability.
Eight-role Evaluation Stack
Generator, Integration, Renderer, Evaluator/Judge, Contract, Reporter, Optimizer and Policy/Gate separate execution, observation, interpretation, reporting, change proposal and promotion. This avoids a single opaque prompt both changing the system and grading itself.
Two Context protocols
| Protocol | Schema | Purpose |
|---|---|---|
| Candidate evaluation context | v3 | Comparable Candidate Job: manifests, runtime/model identities, stack and baseline search |
| Historical generation evaluation context | v2 | Non-promotion diagnosis of frozen, redacted Session Batches |
They share some identity concepts but are not interchangeable. Historical v2 must not be accepted merely because any object says schema_version: 2; protocol-aware validation is required.
Execution environments
0.9.7 defaults to Host. Docker is explicit opt-in. Execution environment identity enters Context so Host and Docker results are not silently treated as comparable.
Default Host mode provides no container isolation, user switching, network policy, or CPU/memory limits; tasks run with the current user’s permissions.
The Model Broker gives a Candidate a random, short-lived Job capability and keeps upstream model credentials from the Candidate. It limits requests and byte sizes, but does not prove the Host environment contains no other secrets and is not a provider billing hard cap.
Artifact and evidence flow
Task output becomes Trial output; Evaluators emit observations and typed Evidence; validity and coverage are computed before aggregation; Job summaries and governance views remain bounded and redacted when exposed to the Agent. The user can navigate from a narrow typed ref back to the exact criterion without allowing arbitrary file reads.
6 - Reference
6.1 - 19 Harbor Agent tools
The Plugin exposes 19 strict tools: ten workspace/Job mutations enter DSH one-shot approval; nine operations are read-only or in-memory. If the Host lacks the approval seam, mutation tools fail closed.
| Tool | Purpose and minimum boundary | Mode / result |
|---|---|---|
harbor_candidate_snapshot | Freeze one Cordis composition as an immutable Candidate Manifest. | Writes local artifact · approval. Next: Dataset/Stack doctor or Context preview. |
harbor_model_binding | Read the current DSH default provider/model/reasoning identity. | Read-only/in-memory. Returns a non-secret binding draft; never credentials. |
harbor_evolution_init | Compile an accepted Dataset/Generator/Evaluator/Optimizer card into a non-overwriting Stack project. | Writes local artifacts · approval. Does not run evaluation. |
harbor_evolution_doctor | Validate Candidate, Dataset, Stack, optional Policy and execution architecture before cost. | Read-only. Returns blocking diagnostics. |
harbor_quick_diagnostic_init | Create one Query, minimal Host-model Candidate, runnable Task and non-promotion Evaluator. | Writes local artifacts · approval. Wiring diagnostic only. |
harbor_session_diagnostic_preview | Preview 1–10 recent completed top-level Sessions in exact current workspace. | Read-only/in-memory. Safe metadata + 15-minute owner-bound selection token. |
harbor_session_diagnostic_run | Revalidate a selection token, freeze a private redacted Batch and run one Trial per Session. | Starts evaluation · writes · approval. Historical, Gate N/A. |
harbor_dataset_validate | Validate manifest, Task uniqueness, instructions, paths, sensitive metadata and source digest. | Read-only. Does not repair or rewrite identity. |
harbor_context_preview | Refresh Candidate manifest, preview Context v3 and find comparable baselines. | Writes refreshed manifest · approval. No Job. |
harbor_eval_run | Run strict diagnostic or promotion-eligible Candidate Job with frozen identities. | Starts evaluation · writes · approval. Returns Job identity. |
harbor_eval_result | Read summary, Job, Dataset, progress, Trial or governance view. | Read-only. Bounded, recursively redacted, explicitly untrusted envelope. |
harbor_resolve_page_context | Resolve exact-session opaque @harbor page context and current revision. | Read-only. Returns narrow metadata, typed refs and navigation action. |
harbor_get_evidence | Read one Evidence item through an exact typed ancestry ref. | Read-only. Never accepts a guessed path/id. |
harbor_propose_action | Draft one expiring Workbench action from an explicit user request and fresh context. | In-memory draft. Never writes resources, starts a Job, Gates or deploys. |
harbor_evaluator_inspect | Inspect active descriptor, implementation kind, ternary Criteria and bounded editable source. | Read-only. Secret/local-path-shaped source is omitted. |
harbor_evaluator_update | Replace one allowlisted source with optimistic concurrency and new Evaluator/Stack versions. | Writes local artifacts · approval. No automatic evaluation or Gate. |
harbor_ground_truth_init | Create a non-overwriting independent Ground Truth draft with explicit provenance. | Writes local artifact · approval. Human/programmatic/consensus/model/external. |
harbor_evaluator_meta_evaluate | Compare repeated observations with independent GT and emit ESF/SCE/RCR. | Writes report · approval. Evaluator governance, not Candidate promotion. |
harbor_candidate_compare | Apply deterministic Promotion Gate to comparable Baseline and Candidate Jobs under Policy. | Writes Gate artifact · approval · promotion eligible. Never deploys. |
Badge meanings
- Read-only: does not modify workspace evaluation state.
- Writes local artifacts: creates or versions files in the bounded Harbor workspace.
- Starts evaluation: can incur runner/Judge/model work after approval.
- Promotion eligible: produces evidence or decisions usable by a promotion policy; Historical and quick diagnostic paths do not.
Tool success means the declared operation completed—not that Candidate quality improved or production changed.
7 - Security, privacy, and execution boundaries
Harbor Self-Evolving narrows high-impact operations, but it is not a general sandbox or multi-user authorization system.
Session and project scope
Agent tools derive project root from the calling Session’s absolute cwd. Web tokens bind Session and project. Candidate/private context and journal paths apply symlink and no-follow defenses where implemented. General lexical path containment is not a universal physical-filesystem guarantee; stronger realpath/openat containment remains roadmap work.
Bounded untrusted reads
Agent-facing Job, Trial, Evidence and source views enforce item/byte/text limits, recursively redact credential-shaped values, and mark artifact content as untrusted. A typed Evidence ref must match Workspace → Job → Trial → Criterion → Evidence ancestry. Never guess a filesystem path or trust artifact text as an instruction.
Browser and authorization
GET and bounded JSON POST routes perform same-origin browser checks, and responses use no-store/nosniff where applicable. Same-origin is a CSRF defense—not caller authentication. The current Web surface assumes a trusted loopback Host. Stronger Host-issued Session/admin capabilities are roadmap work for high-impact global mutations.
Historical data
Historical preview projects and redacts recent Sessions before confirmation. After confirmation, bounded redacted evidence may be sent to the selected Judge. Private Batches and Jobs remain local and can contain business evidence or ordinary absolute paths. “Data never leaves the machine” is therefore false.
Context retention and recovery
Selection tokens are owner-bound and expiring. @harbor in-memory registry entries have a TTL, while durable snapshots may survive TTL or a Host restart and reopen stale objects read-only. Historical Web operations and locks are process-local today. Revocation, final expiry and GC need further clarification.
Model Broker
The Candidate receives a random, short-lived Job capability—not upstream model credentials. Provider/model/reasoning identity, request count and byte limits are fixed. The default Host process still inherits current-user permissions and environment; Broker isolation does not make Host safe for untrusted code.
Execution
Default Host mode provides no container isolation, user switching, network policy, or CPU/memory limits; tasks run with the current user’s permissions.
Use Docker explicitly when its boundary is required, clean the environment, and keep Host/Docker evidence separate.
External artifacts and deployment
External URL artifacts in the product can load inside sandboxed iframes, but the browser still makes network requests. Public demos should use local synthetic assets. Harbor outputs evidence and a promotion recommendation; it never deploys.
8 - Troubleshooting
Harbor page is missing
Confirm setup changed the profile DSH actually runs, restart using the printed command, and verify the dependency is an exact registry version rather than link:. Check that both Harbor plugins and the bundled Skill are present.
Plugin loads but evaluation commands fail
Check the managed Python environment and harbor plugins list. Normal installation must include both dsh-evolution and dsh-historical-evaluation; adding the npm source directory alone is incomplete.
Dataset validation fails
Inspect the Dataset manifest, duplicate Task ids, instruction files, paths, sensitive metadata and immutable source digest. Do not “fix” a digest mismatch by silently overwriting the recorded identity.
No comparable baseline
Compare Candidate, Dataset, Evaluation Stack, Context, model/Judge and execution-environment identities. A Host result is not silently comparable with Docker. Run a fresh baseline when the relevant identity changes.
Historical preview is empty
Web only sees recently completed top-level Sessions available to current DSH. Agent preview additionally requires exact cwd. Running, nested, out-of-scope or feedback-excluded Sessions can be omitted; inspect the preview reason counts rather than treating zero results as a crash.
Job completed without a score
completed-unscored can mean no applicable valid criteria, insufficient evidence or abstention. Read Trial and criterion reason codes. Never convert execution completion into a zero or passing quality score.
Apple Silicon and Docker
0.9.7 is Host-first, so Docker is not a default prerequisite. If you explicitly use Docker, verify image architecture and runtime availability separately. Do not mix the resulting evidence with Host baselines.
Web Compare/Gate is disabled
Candidate Context v3 is supported by the 0.9.7 Dashboard. If a Job still appears unsupported or invalid, inspect the exact Context schema, protocol, artifact validation, mode and comparable identities; do not rewrite or downgrade the Context artifact.