This is the multi-page printable view of this section. .
Product
Harbor Self-Evolving adds continuous evaluation and controlled self-evolution to DeepSeek Harness. It is delivered as three collaborating parts:
- DSH Plugin — native tools, Workbench, Host services and permission boundaries.
evolve-agent-with-harborSkill — the official orchestration policy maintained by this project.- Python Adapter — Harbor Generator, Evaluator and Optimizer wiring plus deterministic Gate integration.
From Dataset to meta-evaluation
The Plugin does more than call a Judge. It separates the evaluation loop into governable roles:
| Role | How the Plugin supports it |
|---|---|
| Dataset | Validate Tasks, instructions, paths and source digest; distinguish training/fix data, validation data and final holdout. |
| Generator | Run the Agent through the Adapter under frozen Candidate, model binding, Context and Host/Docker identities, preserving output and Artifacts. |
| Evaluator | Unify script and llm-as-judge, recording Evidence, validity, abstention and coverage for each criterion. |
| Optimizer | Let the Skill and Agent propose one reviewed change from Trial evidence without owning scoring, Gate or deployment authority. |
| Meta-Evaluation | Compare the Evaluator with independent Ground Truth for fidelity, calibration and ranking reliability, so the Agent is not optimized against an untested measuring instrument. |
Training, validation and test sets are information boundaries rather than file formats. A case that changes the Candidate is training information; repeatedly selecting with a set turns it into validation information; only a holdout whose answers remain hidden from the Optimizer can estimate final generalization. See Concepts and trustworthy scores and Evaluator governance.
Two evaluation paths
| Path | Starts from | Designed to answer | Promotion evidence? |
|---|---|---|---|
| Historical diagnosis | Recent completed DSH Sessions | Where does the current Agent fail? | No—diagnostic only |
| Candidate regression | Frozen Candidate + Dataset + Stack + Context | Is one controlled change better on comparable evidence? | Eligible when policy and Gate requirements hold |
Skill
The bundled Skill chooses the lowest-friction safe path: use recent Sessions when no Dataset exists, or enter the strict Candidate pipeline when the user supplies one. It fixes identities before expensive execution, reads typed evidence rather than guessing from files, and never treats Gate as deployment authority.
Adapter
The Python package harbor-dsh-evolution maps DSH Candidates and Tasks into Harbor’s evaluation interfaces. Version 0.9.7 uses Candidate Context v3 and Historical Context v2, and the Web Workbench recognizes both current contracts. Host execution is default; Docker is opt-in.
Principles
- Identity before score.
- Validity and coverage before averages.
- One controlled change per iteration.
- Evidence and policy remain separate from optimization.
- Gate recommendation remains separate from deployment.
- Permissions and provider boundaries are disclosed at the point of action.
Status and roadmap
Shipped in 0.9.7: package setup, 19 Agent tools, Workbench, Historical diagnostics, Evaluator governance, Host-first execution, Candidate Context v3/Historical Context v2 Web support, and a review-only version command bound to complete installation identity.
Withdrawn preview: the untagged browser one-click updater was removed before release. The browser does not execute registry packages; users review and run the exact setup command in a terminal.
Roadmap: durable operation recovery, stronger mutation authorization, retention/GC, safer external artifact previews and broader browser/accessibility coverage.
Same-origin is a browser CSRF defense, not caller authentication. Host execution is not a sandbox, and Gate remains a deterministic recommendation for fixed inputs—not deployment authority.
From prototype to 0.9.7
- 0.1–0.8: establish the Candidate/Dataset/Stack model, evaluation loop and DSH integration.
- 0.9.0–0.9.4: native Workbench, Historical Session cold start, reviewed actions, context and evidence navigation.
- 0.9.5: consolidated Workbench and stronger release evidence.
- 0.9.6: Host-first execution with Docker opt-in, while preserving explicit safety boundaries.
- 0.9.7: repair current Context Web contracts, preserve update identity and publish the bilingual product/documentation site.
See Releases for evidence and limitations attached to each formal version.
1 - Native DSH evaluation workbench
The npm package dsh-harbor-evolution is more than a tool registry. It owns five layers:
- Setup lifecycle — configure the selected DSH profile, install the Python Adapter, expose the bundled Skill and request a restart.
- Agent orchestration — register 19 strict tools with typed boundaries and approval semantics.
- Native Web experience — Workbench, Historical launcher, Context, Evidence, Artifacts, operations and Settings.
- Host services — bounded project/job reads, Session selection, model brokerage and background operation coordination.
- Trust boundary — redaction, same-origin browser checks, explicit confirmation and untrusted evidence envelopes.
Tool surface
The Plugin registers 19 tools. Ten mutation-capable tools require one-shot DSH approval; nine are read-only or in-memory. The complete per-tool contract is in Tool reference.
Workbench
The Web experience brings the strict pipeline into DSH:
- Dashboard and Job/Trial status;
- Pipeline stages and preflight;
- Context, comparable baselines and Gate readiness;
- Trial output, criterion evidence and artifacts;
- Historical Session selection and disclosure;
- reviewed action drafts, deterministic preflight and explicit confirmation;
- background operations and recovery states;
- Settings, execution environment and version UI.

Context and reviewed actions
Ordinary chat remains ordinary chat. @harbor page context resolves a short-lived, owner-bound reference; typed Evidence refs then enforce Workspace → Job → Trial → Criterion → Evidence ancestry. Artifact text is always treated as untrusted data.
Ask AI can prepare a proposal draft, but it cannot write files, start Jobs, change Evaluators, run Gate or deploy. The user separately reviews deterministic preflight and confirms the exact action.
Current limits
Version 0.9.7 accepts Candidate Context v3 and explicitly recognizes Historical Context v2 by protocol. Candidate Compare/Gate still depends on a comparable baseline, valid artifacts, promotion-eligible mode and policy; support does not imply a PROMOTE result.
Same-origin checks are a browser CSRF defense, not caller authentication. Historical Web operation locks are process-local, while some durable @harbor snapshots can outlive memory TTL; retention and recovery therefore vary by operation.
The earlier untagged one-click updater preview was withdrawn before release. Version 0.9.7 performs a version check and offers an exact copyable setup command only when the complete installation identity is available; the browser never executes a registry package.