Skip to content

This is the multi-page printable view of this section. .

Return to the regular view of this page.

Engineering notes

Design notes about evidence, evaluation architecture and the boundaries behind Harbor Self-Evolving.

Engineering notes explain why the product is designed this way. Formal shipped claims remain anchored to Releases and source evidence.

1 - Evidence before optimization

Why an Agent needs validity, coverage and comparable identities before it needs another rewrite.

The fastest way to make self-evolution untrustworthy is to let the same opaque loop choose cases, rewrite itself and declare victory.

Harbor reverses that order. First freeze what is being tested and how. Then inspect whether evidence is valid and how much of the population it covers. Only then propose one change and compare it against a baseline whose identities still match.

This is why the Optimizer and Gate are separate, why Historical Sessions are diagnosis rather than promotion, and why a raw reward never silently becomes a valid score.

The result may feel slower than “rewrite until the demo looks good.” It is much faster than shipping a regression whose score cannot be explained.

2 - Host first, with honest boundaries

Why 0.9.6 made Host execution the default without calling it a sandbox.

Requiring Docker before a user can diagnose an Agent creates friction. Pretending direct execution is isolated creates risk. Version 0.9.6 chooses Host as the default and names the boundary precisely.

Host mode runs with the current user’s permissions and environment. It supplies no container isolation, user switching, network policy or CPU/memory limits. Docker remains explicit opt-in when that boundary is required.

Execution environment becomes part of Context identity. A result produced on Host is not silently compared with Docker. Lower friction does not require weaker evidence—as long as the runtime difference remains visible.