0.10.1 — trustworthy scientific evaluation loop
Released 2026-09-26. Formal source tag: v0.10.1.
Shipped
- Formal Candidate Experiments execute only the exact
harbor-dsh-evaluator/v2bundle through a strict private Dataset adapter; Dataset verifiers and v1 fallback cannot become score authority. - Trusted Evaluation Reports, content-addressed Job bundles, cross-artifact Evaluator identity, score suppression and execution-failure handling keep invalid evidence out of quality conclusions.
- Repeat-aware Experiments, paired uncertainty, immutable Promotion reports, exact Policy snapshots and stronger Docker/environment invariants make governed comparison auditable.
- Historical diagnosis, Candidate-bound business observations, independent Ground Truth meta-evaluation and confirmed badcase drafts preserve their distinct evidence boundaries.
- Diagnostic, experiment and governed product profiles reveal only the controls justified by their evidence level.
- Plugin, Skill and Python Adapter identities align at 0.10.1, with 34 public schemas in both packages.
Release repair
0.10.0 published to npm, but clean Linux tag/PyPI workflows exposed two release-test portability issues before any Python artifact was built. Public tags and package versions were not moved or overwritten. 0.10.1 contains the same runtime behavior with stable test-helper imports and a bounded shared-runner response budget, and supersedes 0.10.0.
Verification
The 0.10.1 archive records local automated evidence, the 0.10.0 failed-release evidence, and the coordinated 0.10.1 publication status.
Limits
Automated source verification does not establish real-provider business quality. Authenticated manual GUI acceptance and a paid/real-provider Candidate Job were not performed.
Default Host mode is unrestricted and not sandboxed. Historical results never become Promotion Gate evidence, and a Job seal is a content-integrity receipt rather than an external signature.