Skip to content

This is the multi-page printable view of this section. .

Return to the regular view of this page.

Documentation

Install Harbor Self-Evolving, understand the evidence model, run both evaluation paths and inspect exact contracts.

Choose the path that matches your question:

  • Install and start — get the Plugin, Adapter and Skill into the selected DSH profile.
  • Historical diagnosis — learn from recent completed Sessions without calling them promotion evidence.
  • Candidate evaluation — freeze identities, run a comparable regression and apply policy.
  • Plugin Workbench — understand Context, Evidence, Artifacts and reviewed actions.
  • Concepts — learn why valid score, raw reward and coverage are not interchangeable.
  • Architecture — inspect protocols, execution environments and trust boundaries.
  • 19-tool reference — look up every Agent tool and approval boundary.
  • Security — read the limits before running untrusted Tasks.
Important

Version baseline: Harbor Self-Evolving 0.9.7 (Beta), compatible with Harbor >=0.21,<0.22. Development-preview behavior from an untagged checkout is labeled separately.

1 - Install and start

Install the registry release into the selected DSH profile, restart, verify the Plugin and choose an evaluation path.

Requirements

  • A working DeepSeek Harness installation and the business Agent workspace you want to evaluate.
  • Node.js/npm for the DSH Plugin setup command.
  • Python environment support used by the installed Adapter.
  • Harbor >=0.21,<0.22.
  • Docker only if you explicitly choose container execution; 0.10.3 defaults to Host.

Install

Run from the business Agent workspace, not from this source repository:

npx --yes dsh-harbor-evolution@latest setup --project-root "$PWD"

For a version-pinned installation, replace latest with 0.10.3. Setup writes the selected DSH profile dependency, configures the Harbor project integration, installs the compatible Python Adapter and exposes the bundled evolve-agent-with-harbor Skill. Follow the exact restart command printed by setup.

Warning

Do not use dsh plugin add ./packages/dsh-plugin for a normal installation. That creates a machine-local link: dependency and omits the Adapter setup.

Verify

After restart, confirm all three surfaces:

  1. the selected profile depends on exact registry version "dsh-harbor-evolution": "0.10.3", not link:...;
  2. harbor plugins list contains dsh-evolution and dsh-historical-evaluation;
  3. the bundled evolve-agent-with-harbor Skill is present.

Then open the DSH Harbor navigation entry. A healthy installation exposes Workbench, Historical Sessions, Context and Settings rather than a standalone web server.

Choose your first path

No Dataset yet? Start with Historical diagnosis. Preview up to three recently completed Sessions in Web (or up to ten exact-cwd Sessions through the Agent tool), review the redaction/Judge disclosure, then confirm a non-promotion Job.

Candidate and Dataset ready? Follow Candidate evaluation. Snapshot, validate, doctor, preview Context, choose a comparable baseline, run one controlled regression, then apply Gate.

Source development

Only contributors modifying this repository should clone it and run:

./hse dsh-install-source web

Source builds can contain unreleased behavior and must not be presented as the formal 0.10.3 package. The earlier untagged one-click updater preview was withdrawn before 0.9.7; the browser only checks versions and copies a complete, reviewable terminal command.

2 - Workflows

Diagnose real Sessions, run Candidate regressions, and govern Evaluators.

2.1 - Evaluate recent Sessions

Turn bounded, redacted recent DSH Sessions into a non-promotion Historical evaluation Job.

Historical evaluation is the cold start when you have real Agent interactions but no curated Dataset. It is diagnostic, not promotion evidence.

Historical evaluation flow

Web path

  1. Open Historical Sessions in the DSH Harbor page.
  2. Preview up to three recently completed top-level Sessions visible to the current DSH workspace.
  3. Review the frozen Judge identity, projected evidence fields, redaction policy and cost/data disclosure.
  4. Select Sessions and confirm.
  5. The Plugin writes a private redacted Batch, materializes one Harbor Trial per Session and starts a non-promotion Historical Job.

The Agent-tool path can preview up to ten Sessions, but only from the exact current working directory. Preview returns safe metadata and a short-lived owner-bound selection token—not raw Session ids or transcripts.

Synthetic Historical preview

What reaches the Judge

Bounded, projected and credential-shaped-redacted evidence can be sent to the selected Judge model. Raw Session ids, full tool payloads, reasoning and attachments are not directly used as Judge input. Private Batch and Job artifacts remain local and may contain business evidence or ordinary absolute paths.

Do not describe this as “data never leaves the machine.”

How to read the result

  • one selected Session becomes one Trial;
  • criteria can be valid, invalid or abstained;
  • coverage and reason codes matter alongside scores;
  • completed-unscored is a meaningful outcome, not success;
  • Candidate rerun, comparable baseline and Gate are not applicable to this path.

Use repeated failures to curate a Dataset, then move to the Candidate pipeline.

Recovery boundary

Historical Web operation state and locks are process-local today. A Host restart can lose reattachment even when artifacts remain on disk. This differs from durable @harbor snapshots and from other operation journals; recovery guarantees must be stated per operation.

2.2 - Candidate evaluation and promotion

Freeze identities, run a comparable regression, and obtain a deterministic promotion recommendation.

Candidate evaluation answers a narrower question than Historical diagnosis: did one controlled change improve a fixed business task under comparable evaluation conditions?

Candidate evaluation pipeline

Strict sequence

  1. Snapshot Candidate into an immutable manifest.
  2. Validate Dataset identity, task uniqueness, paths, sensitive metadata and source digest.
  3. Doctor Candidate, Dataset, Evaluation Stack and optional Promotion Policy.
  4. Preview Context v3 and discover comparable baselines before spending on a Job.
  5. Run diagnostic first when the stack or provider path is new.
  6. Change one controlled surface—Agent or Evaluator, not both silently.
  7. Run promotion-eligible regression with fixed identities.
  8. Inspect progress, Trial output, criterion evidence and governance impact.
  9. Compare and Gate against a comparable baseline.

Comparability

A baseline is not comparable merely because it used the same repository. Candidate manifest, Dataset manifest, Evaluation Stack, Context, execution environment and relevant model/Judge identities must satisfy the contract. A changed Dataset digest, stack version, provider identity or runtime boundary can require a fresh baseline.

Gate

The Gate is deterministic for fixed baseline Job, Candidate Job and policy inputs. It can return PROMOTE or REJECT based on valid score movement, minimum improvement, regressions, coverage and other policy conditions.

It does not deploy, mutate the Champion or bypass external CI/CD approval.

Version 0.9.7 accepts Candidate Context v3 in both Dashboard overview and Job detail. Compare/Gate still remain capability-gated by artifact validity, comparable identities, mode and policy.

What a release test does not prove

0.9.7 did not run a real provider model, real Candidate/Historical Session data, or a paid Harbor evaluation; automated tests do not establish a business-quality baseline. Your Dataset, Evaluator and production evidence must establish that baseline.

2.3 - Evaluator governance and meta-evaluation

Inspect, version and independently test the Evaluator instead of treating the scoring rule as an oracle.

Candidate quality and Evaluator quality are separate governance problems.

Interface and inspection

Formal Candidate Evaluators implement harbor-dsh-evaluator/v2. A descriptor identifies implementation kind (script or llm-as-judge), ternary Criteria and a bounded allowlist of editable source files. Inspection omits secret-shaped and local-path-shaped values.

Controlled update

An Evaluator update replaces one descriptor-authorized source file under optimistic concurrency. The caller supplies the expected digest and new Evaluator and Stack versions. The update never runs evaluation or Gate automatically.

Independent Ground Truth

Ground Truth may be human, programmatic, consensus, model or external, but provenance must be explicit and independent of the Candidate Evaluator. A draft is non-overwriting and identifies the criteria it covers.

Meta-evaluation

Repeated Evaluator observations are compared with independent Ground Truth to produce:

  • ESF — evaluator score fidelity;
  • SCE — score calibration error;
  • RCR — ranking consistency/reliability.

The Evaluator needs train, validation and test boundaries too

Rubrics, Judge prompts, parsers and thresholds can all be “trained”:

  • tuning set — find false positives/negatives and modify rubric, prompt or script;
  • validation set — compare Evaluator versions and select thresholds or implementations;
  • meta-evaluation holdout — reveal independent cases only after identities are frozen to estimate reliability on unseen cases.

If the same human labels guide Evaluator changes and then serve as final proof of accuracy, information has leaked. Raw human review needs independent provenance and must not be rewritten as if it came from the Candidate Evaluator.

Governance order in the Plugin

  1. harbor_evaluator_inspect reads the interface, Criteria and editable source boundary.
  2. harbor_ground_truth_init creates a non-overwriting independent Ground Truth draft with provenance.
  3. harbor_evaluator_meta_evaluate compares repeated observations with Ground Truth and writes ESF, SCE and RCR.
  4. Only when evidence supports a change, harbor_evaluator_update replaces one authorized file under an expected digest and requires new Evaluator/Stack versions.
  5. Rerun tuning and holdout meta-evaluation before using the new Evaluator for Candidate evaluation.

Meta-evaluation does not automatically update the Evaluator or run Candidate Gate. Historical evaluation is a distinct scenario: applicability, coverage and abstention matter because real Sessions may not exercise every criterion.

3 - Workbench, context, and reviewed actions

Navigate the native DSH Workbench, bind page context, inspect evidence and review actions before execution.

The Harbor page is injected into the existing DSH Web GUI. It reuses DSH locale, Session, conversation, composer, Settings and Tool View surfaces; it is not a second chat product or login system.

Object-first Workbench

A Job opens into eight connected sections: Summary, Trials, Pipeline, Optimization, Compare/Gate, Evaluator/Rubric, Artifacts and Audit. The Pipeline tracks Candidate, Dataset, Integration, Renderer, Judge, Meta, Reporter, Optimizer and Gate separately so a wiring failure is not mislabeled as a low score.

The Trial explorer supports server-side pagination, status and score-validity filters, sorting, criterion evidence focus and frozen Trial sets. Business artifacts stay attached to the Trial that produced them.

Synthetic responsive Workbench

Page context

When supported, ordinary send freezes the current Harbor page context. Ask AI or an explicit @harbor reference takes precedence and is one-shot. The Host resolves the opaque token for the exact Session, project and revision, then returns typed refs and a navigation action. Evidence reading validates the complete ancestry and treats content as untrusted.

Synthetic native conversation context

Reviewed actions

Seven draft kinds cover Candidate change, Evaluator change, Compare, Diagnostic Evaluation, Infrastructure Retry, Gate Request and Deployment Handoff. Proposal creation is not authorization.

  1. A user explicitly requests an action.
  2. AI creates an expiring draft from fresh page context.
  3. Deterministic preflight shows target, diff, identities and limits.
  4. The user confirms the exact action.
  5. The Host journals execution and exposes progress/recovery state.

Candidate/Gate/handoff drafts do not silently execute. Compare is read-only. Unsupported production actions are denied.

Settings and operations

Settings exposes project root source, Stack/jobs/CLI checks, credential policy, execution environment and version status. Agent tools continue to use the calling Session cwd as authority; the process-local Settings root mainly serves Web and fallback behavior.

Background operations can expose progress, partial evidence, cancellation and navigation, but durability varies. Unknown state must not be rendered as success, and failure must not trigger silent retry.

Warning

0.9.7 accepts Candidate Context v3 and explicitly recognizes Historical Context v2. Compare/Gate remain gated by artifact validity, comparable identities, mode and policy; see Plugin limits.

4 - Concepts and trustworthy scores

Dataset, Generator, Evaluator, Optimizer, train/validation/test splits, meta-evaluation, trustworthy scores and promotion semantics.

Translate the four concepts into business language

ConceptPractical meaningQuestion answered in Harbor
DatasetBusiness cases with task instructions, inputs and expected evidence. It defines what to evaluate; it is not a raw dump of conversations.Which capabilities must work, and which failures must be detected?
GeneratorThe execution process that receives a Task and produces an answer or Artifact under a fixed Candidate, model and runtime. It may be an LLM Agent, script, workflow or other program.Who answers, how is it run, and what does it produce?
EvaluatorThe criteria and evidence logic used to judge output quality. It may be a deterministic script or llm-as-judge.What counts as good, and is the evidence sufficient to decide?
OptimizerThe role that reads Dataset-level results and Trial evidence, then proposes one constrained, reviewable change. It cannot declare success by itself.What should change, and how do we avoid changing too much at once?

Harbor compiles these user-facing concepts into a stricter Evaluation Stack with Integration, Renderer, Judge, Contract, Reporter, Policy, Context and Gate roles. Users describe the business problem first instead of filling in an internal architecture questionnaire.

Important

A Harbor Dataset is primarily an evaluation dataset. Whether it serves training, validation or final testing is a governance decision—not a directory name. A case cannot secretly guide optimization and still be presented as never-seen final evidence.

Train, validation and test splits

Machine learning separates data to control information leakage: what the optimization process has already seen, and whether the final result still says anything about unseen cases.

Training set

A training set directly guides learning or improvement. Traditional ML updates parameters from it; Agent engineering may use it to diagnose badcases and modify prompts, Skills, tool policies, code or retrieval configuration.

  • Inputs, outputs and failure reasons may be inspected repeatedly.
  • Optimizers may derive change hypotheses from it.
  • Its score shows whether known failures were fixed, not generalization by itself.
  • Reviewed badcases derived from Historical Sessions usually enter this layer before they become any promotion dataset.

Validation or development set

A validation set compares alternatives, selects thresholds and determines when to stop. Even without gradient updates, repeated feedback lets the Optimizer overfit to it.

  • Use it to select among Candidates or configurations.
  • Record any use for tuning Evaluator rubrics, Policy thresholds or runtime parameters.
  • It can be an iterative regression set, but not an indefinitely reused independent final proof.
  • Comparable baseline and Candidate Jobs may run on a fixed validation set, while their purpose must remain labeled as development selection or promotion evidence.

Test or holdout set

A test set estimates final generalization. Before the final run, the Optimizer, Candidate author and Evaluator-tuning process should not see its answers, Ground Truth or per-case feedback.

  • Run only after Candidate, Evaluator, Policy and runtime identities are frozen.
  • Minimize repeated inspection and reruns.
  • Once its results guide the next change, treat it as development data and create a new holdout.
  • A Promotion Gate is comparable only when baseline and Candidate use the same immutable Dataset, Stack and Context with valid evidence.

A useful data layering for Agents

LayerPrimary purposeWho sees feedbackTypical Plugin path
Historical diagnostic samplesFind real failure patterns and hypothesesHumans and the diagnostic JudgePreview completed Sessions → disclose boundary → non-promotion Historical Job
Training/fix setRepair known badcasesOptimizer and developersCurate reviewed failures into Tasks and run targeted regressions
Validation/regression setCompare Candidates and select changesOptimizer may inspect aggregates and Trial evidenceFreeze Candidate/Dataset/Stack/Context and run comparable Jobs
Test/holdout setFinal generalization and promotion evidenceAnswers remain hidden before the final runRun a promotion-eligible Job, then deterministic Gate
Evaluator meta-evaluation setTest whether the measuring instrument is reliableEvaluator governance processIndependent Ground Truth + repeated observations → ESF/SCE/RCR

Small projects should preserve these logical boundaries even when they lack enough cases for statistically strong three-way splits. Record which cases changed the Candidate, which selected among alternatives, and which remain held out.

Generator: the generation process under test

A Generator is more than a model name. It includes:

  • prompts, Skills, code, tools and dependencies in the Candidate;
  • Candidate model binding and provider/model identity;
  • Task instructions, Context and input Artifacts;
  • Host or Docker execution environment;
  • final output, structured results and Artifacts.

The Python Adapter maps DSH Candidates and Tasks into the Harbor Generator interface. The Plugin and Skill freeze relevant identities before expensive execution. The same model with a different prompt, tool, dependency or runtime is a different generation condition.

Evaluator: make “good” inspectable

An Evaluator decomposes the business objective into identified Criteria and evidence requirements. harbor-dsh-evaluator/v1 supports deterministic script and llm-as-judge implementations, but both must preserve distinct states:

  • valid score — contract and evidence requirements were met;
  • invalid score — a number exists but must not enter quality aggregation;
  • abstention — the Evaluator explicitly could not decide;
  • coverage — how much of the target population produced usable evidence.

A Judge is one Evaluator implementation, not an inherently correct oracle. Model version, rubric, parsing logic and source implementation belong to Evaluator identity. Any change requires new Evaluator and Evaluation Stack versions.

Optimizer: propose one controlled change from evidence

The Optimizer consumes failure patterns and evidence—not an average score detached from context. The controlled path is:

  1. identify repeatable failures in Trial and criterion Evidence;
  2. constrain the surface that may change, such as a prompt, Skill, Evaluator source or code;
  3. propose one reviewable change and create a new Candidate or Evaluator identity;
  4. rerun on a fixed Dataset and Stack;
  5. let Policy and Gate decide whether promotion conditions were met.

Ask AI, an Action Draft or an Optimizer proposal remains a recommendation. It does not automatically write files, start Jobs, run Gate or deploy; high-impact actions still require exact preflight and user confirmation.

Meta-evaluation: prove the measuring instrument

Ordinary evaluation asks, “Is the Generator output good?” Meta-evaluation asks, “Is the Evaluator judgment reliable?”

The Plugin compares repeated Evaluator observations with independent Ground Truth. Ground Truth may be human, programmatic, consensus, model or external, but it needs explicit provenance and independence from the Evaluator under test. Reports include:

  • ESF (Evaluator Score Fidelity) — agreement with Ground Truth;
  • SCE (Score Calibration Error) — calibration between confidence and observed correctness;
  • RCR (Ranking Consistency/Reliability) — stability and Ground-Truth consistency of rankings.

Evaluators can overfit too. Samples used to modify a rubric, prompt or threshold form an evaluator tuning set; final evidence should come from an independent holdout. Labels produced by the Candidate Evaluator cannot be reused to prove that Evaluator correct.

How the Plugin connects the loop

  1. Dataset — validate manifest, Task uniqueness, paths and immutable source digest.
  2. Generator — freeze Candidate manifest, model binding, Context and runtime, then produce Trial output.
  3. Evaluator — produce criterion observations, Evidence, validity and coverage.
  4. Optimizer — propose one new-version change from reviewed evidence without owning the verdict.
  5. Meta-evaluation — test the Evaluator itself against independent Ground Truth.
  6. Gate — apply deterministic Policy only to fixed, comparable inputs and return a PROMOTE or REJECT recommendation.

The Historical path creates diagnostic hypotheses and possible future Dataset cases, but is non-promotion evaluation. Only the Candidate path can become promotion evidence after comparability, validity and Policy requirements are met.

Identity chain

A comparable Job binds immutable or versioned identities: Candidate Manifest, Dataset Manifest, Evaluation Stack, Context, execution environment, Candidate model binding and Judge identity. A Trial belongs to one Job and carries output, criterion observations, Evidence and Artifacts.

Evaluation identity chain

Raw reward is not a valid score

A verifier may emit a numeric raw reward even when required evidence is missing, parsing failed or the criterion abstained. Averages without validity and coverage can reward a broken pipeline.

PROMOTE and REJECT

Deterministic Gate decision

A deterministic Gate evaluates fixed Job and Policy inputs. PROMOTE means the Candidate satisfies that Policy relative to a comparable baseline. REJECT means it does not. Neither means deployed, and neither replaces human or CI/CD authority.

5 - Architecture and trust boundaries

DSH Plugin, Skill, Python Adapter, two Context protocols, execution environments and the deterministic Gate.
Harbor Self-Evolving architecture

Three product roles

  • Plugin: the user-facing DSH integration and permission boundary.
  • Skill: the workflow policy that chooses the smallest safe evaluation path.
  • Adapter: the Harbor runtime bridge that materializes Jobs, Trials, Evidence and Context.

The Gate remains a separate deterministic policy function; it is not part of the Optimizer and has no deployment capability.

Eight-role Evaluation Stack

Generator, Integration, Renderer, Evaluator/Judge, Contract, Reporter, Optimizer and Policy/Gate separate execution, observation, interpretation, reporting, change proposal and promotion. This avoids a single opaque prompt both changing the system and grading itself.

Two Context protocols

ProtocolSchemaPurpose
Candidate evaluation contextv3Comparable Candidate Job: manifests, runtime/model identities, stack and baseline search
Historical generation evaluation contextv2Non-promotion diagnosis of frozen, redacted Session Batches

They share some identity concepts but are not interchangeable. Historical v2 must not be accepted merely because any object says schema_version: 2; protocol-aware validation is required.

Execution environments

0.9.7 defaults to Host. Docker is explicit opt-in. Execution environment identity enters Context so Host and Docker results are not silently treated as comparable.

Warning

Default Host mode provides no container isolation, user switching, network policy, or CPU/memory limits; tasks run with the current user’s permissions.

The Model Broker gives a Candidate a random, short-lived Job capability and keeps upstream model credentials from the Candidate. It limits requests and byte sizes, but does not prove the Host environment contains no other secrets and is not a provider billing hard cap.

Artifact and evidence flow

Task output becomes Trial output; Evaluators emit observations and typed Evidence; validity and coverage are computed before aggregation; Job summaries and governance views remain bounded and redacted when exposed to the Agent. The user can navigate from a narrow typed ref back to the exact criterion without allowing arbitrary file reads.

6.1 - 19 Harbor Agent tools

Every Plugin tool, its input boundary, mutation/approval semantics, output evidence and intended next step.

The Plugin exposes 19 strict tools: ten workspace/Job mutations enter DSH one-shot approval; nine operations are read-only or in-memory. If the Host lacks the approval seam, mutation tools fail closed.

ToolPurpose and minimum boundaryMode / result
harbor_candidate_snapshotFreeze one Cordis composition as an immutable Candidate Manifest.Writes local artifact · approval. Next: Dataset/Stack doctor or Context preview.
harbor_model_bindingRead the current DSH default provider/model/reasoning identity.Read-only/in-memory. Returns a non-secret binding draft; never credentials.
harbor_evolution_initCompile an accepted Dataset/Generator/Evaluator/Optimizer card into a non-overwriting Stack project.Writes local artifacts · approval. Does not run evaluation.
harbor_evolution_doctorValidate Candidate, Dataset, Stack, optional Policy and execution architecture before cost.Read-only. Returns blocking diagnostics.
harbor_quick_diagnostic_initCreate one Query, minimal Host-model Candidate, runnable Task and non-promotion Evaluator.Writes local artifacts · approval. Wiring diagnostic only.
harbor_session_diagnostic_previewPreview 1–10 recent completed top-level Sessions in exact current workspace.Read-only/in-memory. Safe metadata + 15-minute owner-bound selection token.
harbor_session_diagnostic_runRevalidate a selection token, freeze a private redacted Batch and run one Trial per Session.Starts evaluation · writes · approval. Historical, Gate N/A.
harbor_dataset_validateValidate manifest, Task uniqueness, instructions, paths, sensitive metadata and source digest.Read-only. Does not repair or rewrite identity.
harbor_context_previewRefresh Candidate manifest, preview Context v3 and find comparable baselines.Writes refreshed manifest · approval. No Job.
harbor_eval_runRun strict diagnostic or promotion-eligible Candidate Job with frozen identities.Starts evaluation · writes · approval. Returns Job identity.
harbor_eval_resultRead summary, Job, Dataset, progress, Trial or governance view.Read-only. Bounded, recursively redacted, explicitly untrusted envelope.
harbor_resolve_page_contextResolve exact-session opaque @harbor page context and current revision.Read-only. Returns narrow metadata, typed refs and navigation action.
harbor_get_evidenceRead one Evidence item through an exact typed ancestry ref.Read-only. Never accepts a guessed path/id.
harbor_propose_actionDraft one expiring Workbench action from an explicit user request and fresh context.In-memory draft. Never writes resources, starts a Job, Gates or deploys.
harbor_evaluator_inspectInspect active descriptor, implementation kind, ternary Criteria and bounded editable source.Read-only. Secret/local-path-shaped source is omitted.
harbor_evaluator_updateReplace one allowlisted source with optimistic concurrency and new Evaluator/Stack versions.Writes local artifacts · approval. No automatic evaluation or Gate.
harbor_ground_truth_initCreate a non-overwriting independent Ground Truth draft with explicit provenance.Writes local artifact · approval. Human/programmatic/consensus/model/external.
harbor_evaluator_meta_evaluateCompare repeated observations with independent GT and emit ESF/SCE/RCR.Writes report · approval. Evaluator governance, not Candidate promotion.
harbor_candidate_compareApply deterministic Promotion Gate to comparable Baseline and Candidate Jobs under Policy.Writes Gate artifact · approval · promotion eligible. Never deploys.

Badge meanings

  • Read-only: does not modify workspace evaluation state.
  • Writes local artifacts: creates or versions files in the bounded Harbor workspace.
  • Starts evaluation: can incur runner/Judge/model work after approval.
  • Promotion eligible: produces evidence or decisions usable by a promotion policy; Historical and quick diagnostic paths do not.

Tool success means the declared operation completed—not that Candidate quality improved or production changed.

7 - Security, privacy, and execution boundaries

What the Plugin protects, what remains trusted, what can leave the machine, and why Harbor never deploys.

Harbor Self-Evolving narrows high-impact operations, but it is not a general sandbox or multi-user authorization system.

Harbor trust boundaries

Session and project scope

Agent tools derive project root from the calling Session’s absolute cwd. Web tokens bind Session and project. Candidate/private context and journal paths apply symlink and no-follow defenses where implemented. General lexical path containment is not a universal physical-filesystem guarantee; stronger realpath/openat containment remains roadmap work.

Bounded untrusted reads

Agent-facing Job, Trial, Evidence and source views enforce item/byte/text limits, recursively redact credential-shaped values, and mark artifact content as untrusted. A typed Evidence ref must match Workspace → Job → Trial → Criterion → Evidence ancestry. Never guess a filesystem path or trust artifact text as an instruction.

Browser and authorization

GET and bounded JSON POST routes perform same-origin browser checks, and responses use no-store/nosniff where applicable. Same-origin is a CSRF defense—not caller authentication. The current Web surface assumes a trusted loopback Host. Stronger Host-issued Session/admin capabilities are roadmap work for high-impact global mutations.

Historical data

Historical preview projects and redacts recent Sessions before confirmation. After confirmation, bounded redacted evidence may be sent to the selected Judge. Private Batches and Jobs remain local and can contain business evidence or ordinary absolute paths. “Data never leaves the machine” is therefore false.

Context retention and recovery

Selection tokens are owner-bound and expiring. @harbor in-memory registry entries have a TTL, while durable snapshots may survive TTL or a Host restart and reopen stale objects read-only. Historical Web operations and locks are process-local today. Revocation, final expiry and GC need further clarification.

Model Broker

The Candidate receives a random, short-lived Job capability—not upstream model credentials. Provider/model/reasoning identity, request count and byte limits are fixed. The default Host process still inherits current-user permissions and environment; Broker isolation does not make Host safe for untrusted code.

Execution

Warning

Default Host mode provides no container isolation, user switching, network policy, or CPU/memory limits; tasks run with the current user’s permissions.

Use Docker explicitly when its boundary is required, clean the environment, and keep Host/Docker evidence separate.

External artifacts and deployment

External URL artifacts in the product can load inside sandboxed iframes, but the browser still makes network requests. Public demos should use local synthetic assets. Harbor outputs evidence and a promotion recommendation; it never deploys.

8 - Troubleshooting

Diagnose installation, profile, Dataset, Context, Historical and execution-environment failures without overstating success.

Harbor page is missing

Confirm setup changed the profile DSH actually runs, restart using the printed command, and verify the dependency is an exact registry version rather than link:. Check that both Harbor plugins and the bundled Skill are present.

Plugin loads but evaluation commands fail

Check the managed Python environment and harbor plugins list. Normal installation must include both dsh-evolution and dsh-historical-evaluation; adding the npm source directory alone is incomplete.

Dataset validation fails

Inspect the Dataset manifest, duplicate Task ids, instruction files, paths, sensitive metadata and immutable source digest. Do not “fix” a digest mismatch by silently overwriting the recorded identity.

No comparable baseline

Compare Candidate, Dataset, Evaluation Stack, Context, model/Judge and execution-environment identities. A Host result is not silently comparable with Docker. Run a fresh baseline when the relevant identity changes.

Historical preview is empty

Web only sees recently completed top-level Sessions available to current DSH. Agent preview additionally requires exact cwd. Running, nested, out-of-scope or feedback-excluded Sessions can be omitted; inspect the preview reason counts rather than treating zero results as a crash.

Job completed without a score

completed-unscored can mean no applicable valid criteria, insufficient evidence or abstention. Read Trial and criterion reason codes. Never convert execution completion into a zero or passing quality score.

Apple Silicon and Docker

0.9.7 is Host-first, so Docker is not a default prerequisite. If you explicitly use Docker, verify image architecture and runtime availability separately. Do not mix the resulting evidence with Host baselines.

Web Compare/Gate is disabled

Candidate Context v3 is supported by the 0.9.7 Dashboard. If a Job still appears unsupported or invalid, inspect the exact Context schema, protocol, artifact validation, mode and comparable identities; do not rewrite or downgrade the Context artifact.