<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>From real Sessions to controlled promotion on Harbor Self-Evolving</title><link>https://istarwyh.github.io/harbor-self-evolving/book/</link><description>Recent content in From real Sessions to controlled promotion on Harbor Self-Evolving</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sun, 20 Sep 2026 09:42:09 +0800</lastBuildDate><atom:link href="https://istarwyh.github.io/harbor-self-evolving/book/index.xml" rel="self" type="application/rss+xml"/><item><title>Why self-editing is not progress</title><link>https://istarwyh.github.io/harbor-self-evolving/book/01-why/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/book/01-why/</guid><description>&lt;p&gt;An Agent can rewrite its prompt, tools or Evaluator and still become worse. “It changed” is not “it improved.”&lt;/p&gt;&#10;&lt;p&gt;A trustworthy loop fixes the business cases, execution identity and scoring contract before the change. It separates the Optimizer that proposes a change from the Gate that judges promotion.&lt;/p&gt;&#10;&lt;h2 id="you-do"&gt;You do&#10;&lt;/h2&gt;&#10;&lt;p&gt;Name the business failure and the owner of the final deployment decision.&lt;/p&gt;&#10;&lt;h2 id="harbor-records"&gt;Harbor records&#10;&lt;/h2&gt;&#10;&lt;p&gt;The identities and evidence chain needed to revisit the claim.&lt;/p&gt;</description></item><item><title>Define the four concepts</title><link>https://istarwyh.github.io/harbor-self-evolving/book/02-four-concepts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/book/02-four-concepts/</guid><description>&lt;p&gt;Self-evolution is not “letting the model edit itself.” It is an inspectable learning chain: &lt;strong&gt;Dataset defines the problem, Generator produces results, Evaluator judges evidence, and Optimizer proposes one controlled change.&lt;/strong&gt; The Evaluator is itself tested through Meta-Evaluation.&lt;/p&gt;&#10;&lt;h2 id="four-roles"&gt;Meet the four roles&#10;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Dataset&lt;/strong&gt; — business tasks and failure cases that answer what to evaluate.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Generator&lt;/strong&gt; — runs the Candidate and produces answers or Artifacts: who answers and how.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Evaluator&lt;/strong&gt; — applies criteria and Evidence requirements: what counts as good and whether evidence is sufficient.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Optimizer&lt;/strong&gt; — proposes one new-version change from failure evidence: what to change next.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;In an Agent system, each role is broader than a model. Generator includes prompts, Skills, tools and runtime; Evaluator includes rubric, Judge, parsing and validity rules; Optimizer may be a Coding Agent constrained by the project Skill.&lt;/p&gt;</description></item><item><title>Diagnose recent Sessions</title><link>https://istarwyh.github.io/harbor-self-evolving/book/03-diagnose/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/book/03-diagnose/</guid><description>&lt;p&gt;Historical evaluation lowers the cold-start cost. Preview bounded Session metadata, inspect what may reach the Judge, then confirm a non-promotion Job.&lt;/p&gt;&#10;&lt;h2 id="you-do"&gt;You do&#10;&lt;/h2&gt;&#10;&lt;p&gt;Select recent completed Sessions that reflect the failure, review exclusions and consent to the disclosed Judge/data boundary.&lt;/p&gt;&#10;&lt;h2 id="harbor-records"&gt;Harbor records&#10;&lt;/h2&gt;&#10;&lt;p&gt;A private redacted Batch, frozen Judge identity, one Trial per Session, criterion applicability, coverage and reason codes.&lt;/p&gt;&#10;&lt;h2 id="does-not-prove"&gt;This does not prove&#10;&lt;/h2&gt;&#10;&lt;p&gt;Historical scores are not comparable Candidate regression evidence and never go through Promotion Gate.&lt;/p&gt;</description></item><item><title>Turn badcases into a Dataset</title><link>https://istarwyh.github.io/harbor-self-evolving/book/04-curate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/book/04-curate/</guid><description>&lt;p&gt;A Dataset is curated business intent, not a dump of private transcripts. Abstract the failure, preserve the relevant constraint and create an independently reviewable expected behavior.&lt;/p&gt;&#10;&lt;h2 id="you-do"&gt;You do&#10;&lt;/h2&gt;&#10;&lt;p&gt;Deduplicate failure patterns, remove private identifiers, separate tuning from holdout and review every instruction file.&lt;/p&gt;&#10;&lt;h2 id="harbor-records"&gt;Harbor records&#10;&lt;/h2&gt;&#10;&lt;p&gt;Task ids, paths, instructions, sensitive-metadata checks and immutable source digest.&lt;/p&gt;&#10;&lt;h2 id="does-not-prove"&gt;This does not prove&#10;&lt;/h2&gt;&#10;&lt;p&gt;A clean Dataset does not prove the population is complete or that one metric captures all business risk.&lt;/p&gt;</description></item><item><title>Run one controlled regression</title><link>https://istarwyh.github.io/harbor-self-evolving/book/05-regress/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/book/05-regress/</guid><description>&lt;p&gt;Snapshot the Candidate, validate the Dataset, doctor the Stack, preview Context and run diagnostic mode before promotion-eligible work when the path is new.&lt;/p&gt;&#10;&lt;h2 id="you-do"&gt;You do&#10;&lt;/h2&gt;&#10;&lt;p&gt;State one hypothesis and one allowed change. Keep Dataset, criteria, provider and runtime fixed unless the change deliberately creates a new baseline.&lt;/p&gt;&#10;&lt;h2 id="harbor-records"&gt;Harbor records&#10;&lt;/h2&gt;&#10;&lt;p&gt;Manifest digests, Context v3, model/Judge identities, Trial output, criterion Evidence, validity, coverage and governance impact.&lt;/p&gt;</description></item><item><title>Read the Gate and hand off</title><link>https://istarwyh.github.io/harbor-self-evolving/book/06-gate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://istarwyh.github.io/harbor-self-evolving/book/06-gate/</guid><description>&lt;p&gt;Compare the Candidate Job with a comparable baseline under an explicit Promotion Policy. Read the overall decision together with regressions, coverage and invalid criteria.&lt;/p&gt;&#10;&lt;h2 id="you-do"&gt;You do&#10;&lt;/h2&gt;&#10;&lt;p&gt;Review representative evidence, confirm the policy matches business risk and decide whether external CI/CD should consume the recommendation.&lt;/p&gt;&#10;&lt;h2 id="harbor-records"&gt;Harbor records&#10;&lt;/h2&gt;&#10;&lt;p&gt;Baseline/Candidate Job identities, Policy identity, comparison details and deterministic Gate artifact.&lt;/p&gt;&#10;&lt;h2 id="does-not-prove"&gt;This does not prove&#10;&lt;/h2&gt;&#10;&lt;p&gt;&lt;code&gt;PROMOTE&lt;/code&gt; is not a production deployment, Champion mutation or universal quality guarantee. It is a recommendation for the frozen evidence and policy.&lt;/p&gt;</description></item></channel></rss>