<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Workflows on Harbor Self-Evolving</title><link>https://istarwyh.github.io/harbor-self-evolving/docs/workflows/</link><description>Recent content in Workflows on Harbor Self-Evolving</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sat, 26 Sep 2026 11:28:06 +0800</lastBuildDate><atom:link href="https://istarwyh.github.io/harbor-self-evolving/docs/workflows/index.xml" rel="self" type="application/rss+xml"/><item><title>Evaluate recent Sessions</title><link>https://istarwyh.github.io/harbor-self-evolving/docs/workflows/historical/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/docs/workflows/historical/</guid><description>&lt;p&gt;Historical evaluation is the cold start when you have real Agent interactions but no curated Dataset. It is &lt;strong&gt;diagnostic&lt;/strong&gt;, not promotion evidence.&lt;/p&gt;&#10;&lt;img class="td-image" src="https://istarwyh.github.io/harbor-self-evolving/images/diagrams/historical-flow.svg" alt="Historical evaluation flow" title="Current Historical Context v2 flow" loading="lazy" decoding="async"&gt;&lt;h2 id="web"&gt;Web path&#10;&lt;/h2&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Open &lt;strong&gt;Historical Sessions&lt;/strong&gt; in the DSH Harbor page.&lt;/li&gt;&#10;&lt;li&gt;Preview up to &lt;strong&gt;three&lt;/strong&gt; recently completed top-level Sessions visible to the current DSH workspace.&lt;/li&gt;&#10;&lt;li&gt;Review the frozen Judge identity, projected evidence fields, redaction policy and cost/data disclosure.&lt;/li&gt;&#10;&lt;li&gt;Select Sessions and confirm.&lt;/li&gt;&#10;&lt;li&gt;The Plugin writes a private redacted Batch, materializes one Harbor Trial per Session and starts a non-promotion Historical Job.&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;The Agent-tool path can preview up to &lt;strong&gt;ten&lt;/strong&gt; Sessions, but only from the exact current working directory. Preview returns safe metadata and a short-lived owner-bound selection token—not raw Session ids or transcripts.&lt;/p&gt;</description></item><item><title>Candidate evaluation and promotion</title><link>https://istarwyh.github.io/harbor-self-evolving/docs/workflows/candidate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/docs/workflows/candidate/</guid><description>&lt;p&gt;Candidate evaluation answers a narrower question than Historical diagnosis: &lt;strong&gt;did one controlled change improve a fixed business task under comparable evaluation conditions?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;img class="td-image" src="https://istarwyh.github.io/harbor-self-evolving/images/diagrams/candidate-pipeline.svg" alt="Candidate evaluation pipeline" title="Current Candidate Context v3 pipeline" loading="lazy" decoding="async"&gt;&lt;h2 id="sequence"&gt;Strict sequence&#10;&lt;/h2&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;&lt;strong&gt;Snapshot Candidate&lt;/strong&gt; into an immutable manifest.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Validate Dataset&lt;/strong&gt; identity, task uniqueness, paths, sensitive metadata and source digest.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Doctor&lt;/strong&gt; Candidate, Dataset, Evaluation Stack and optional Promotion Policy.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Preview Context v3&lt;/strong&gt; and discover comparable baselines before spending on a Job.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Run diagnostic first&lt;/strong&gt; when the stack or provider path is new.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Change one controlled surface&lt;/strong&gt;—Agent or Evaluator, not both silently.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Run promotion-eligible regression&lt;/strong&gt; with fixed identities.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Inspect progress, Trial output, criterion evidence and governance impact.&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Compare and Gate&lt;/strong&gt; against a comparable baseline.&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;h2 id="comparability"&gt;Comparability&#10;&lt;/h2&gt;&#10;&lt;p&gt;A baseline is not comparable merely because it used the same repository. Candidate manifest, Dataset manifest, Evaluation Stack, Context, execution environment and relevant model/Judge identities must satisfy the contract. A changed Dataset digest, stack version, provider identity or runtime boundary can require a fresh baseline.&lt;/p&gt;</description></item><item><title>Evaluator governance and meta-evaluation</title><link>https://istarwyh.github.io/harbor-self-evolving/docs/workflows/evaluator/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/docs/workflows/evaluator/</guid><description>&lt;p&gt;Candidate quality and Evaluator quality are separate governance problems.&lt;/p&gt;&#10;&lt;h2 id="interface"&gt;Interface and inspection&#10;&lt;/h2&gt;&#10;&lt;p&gt;Formal Candidate Evaluators implement &lt;code&gt;harbor-dsh-evaluator/v2&lt;/code&gt;. A descriptor identifies implementation kind (&lt;code&gt;script&lt;/code&gt; or &lt;code&gt;llm-as-judge&lt;/code&gt;), ternary Criteria and a bounded allowlist of editable source files. Inspection omits secret-shaped and local-path-shaped values.&lt;/p&gt;&#10;&lt;h2 id="update"&gt;Controlled update&#10;&lt;/h2&gt;&#10;&lt;p&gt;An Evaluator update replaces one descriptor-authorized source file under optimistic concurrency. The caller supplies the expected digest and &lt;strong&gt;new Evaluator and Stack versions&lt;/strong&gt;. The update never runs evaluation or Gate automatically.&lt;/p&gt;</description></item></channel></rss>