OURO

Blog · Evolution log

Method2026-10-10

Evolution Log #0: What we are going to publish

From the start of the private beta, every evolution round is published here de-identified; four fixed items per entry, failures included.

From the start of the engine’s private beta we will publish a de-identified report of every evolution round here. The reason is simple: OURO sells “verifiable improvement,” and if our own evolution process stays private, that claim carries no weight.

Every entry covers four fixed things. First, what kind of agent was evolved and what eval set was used (type and size only, never customer names or content). Second, the before-and-after scores, confidence intervals and holdout results. Third, what the engine changed: which layer, a summary of the change, and the one or two most interesting reasons rejected candidates failed. Fourth, how much compute and money it took, which nodes it ran on, and a link to the PoE record on-chain.

Failures get published too. A round with no gain, or one that failed the holdout, goes out all the same. We keep thinking about the case in the DGM paper where the system “deleted the hallucination-detection log”: the most likely failure of an evolving system is not stagnation but improving in the wrong way. Publishing failed rounds lets the community help us keep watch over the evaluation itself.

De-identification rules. Customer identity, raw eval samples and full prompts are never published; changes are shown as summaries and diff statistics; the on-chain record itself contains no personal information. Enterprise customers can opt out of the log entirely.

The first real entry will follow the start of the private beta in Q1 2027. Until then this space will carry method notes: how to write a good eval set, the difference between scaffolding evolution and fine-tuning, and how we read the incentive lessons from Bittensor.

← All posts