Whitepaper
OURO Whitepaper
Let deployed AI keep improving on real tasks, and record the evidence of every round on-chain. This is the online edition; it matches the PDF.
AI that evolves AI. — Let deployed AI keep improving on real tasks, and record the evidence of every round of improvement on-chain.
1. Summary
Intelligence is moving from weights to scaffolding. In May 2026 OpenAI announced it would wind down self-serve fine-tuning because “prompts and scaffolding are cheaper and faster.” In the same year, academia showed that with frozen weights, letting an AI rewrite its own prompts, tools, memory and skills can push SWE-bench from 20% to 50%, lift vertical-domain tasks by 89% on average, and do so with 35× less compute than reinforcement learning. This means one thing: keeping a deployed agent improving no longer needs a lab-scale cluster, but a closed loop that keeps generating candidates, evaluating them in parallel, selecting and submitting. What that loop consumes is massive, low-bandwidth, fault-tolerant computation, which is exactly what distributed networks provide best.
OURO is that loop. It turns a national research institute’s results in self-evolving AI into three components: OURO Engine, which keeps any connected agent improving within a budget, evolving scaffolding first with weights optional; Proof of Evolution (PoE), which turns every round’s “it got better” from a vendor’s promise into an on-chain event that anyone can verify and settle; and OURO Grid, a globally distributed compute network in which regional anchors guarantee reliability while community nodes and external supply provide scale, with the first anchors in Singapore and Johor Bahru.
The market gap is concrete. Enterprise generative-AI spend reached US$37 billion in 2025; 79% of enterprises are adopting agents but only 23% are scaling them; Gartner expects more than 40% of agentic projects to be cancelled by the end of 2027. What blocks enterprises is not that models are too dumb, but that agents stop getting smarter after launch. OURO does not sell GPU hours; it sells verifiable capability gains.
The network token $OURO is a work credential: it pays for evolution and compute, is staked by nodes and validators, escrows bounties and carries governance. Supply is fixed at 1 billion; network incentives are released only when triggered by verified PoE records, and a fixed share of fiat revenue is used to buy back tokens into the incentive pool. What we learned from Bittensor’s three dTAO mechanism changes in eighteen months is that emission can only be tied to verified work, rules must be stable, and data must be public.
For consumers, OURO Chat is an assistant that fits you better the more you use it, and idle GPUs can join OURO Grid to earn by running evaluation and training tasks. For developers and enterprises, Evolve API gives any agent continuous evolution within a controllable budget, with a publicly verifiable evolution report attached.
The project is developed by a national research institute’s technical team, backed by a Singapore compliance fund, and has gathered an early community of several hundred people worldwide. This whitepaper is a v1.1 draft; token parameters, the roadmap and partner information will be updated after legal advice and team confirmation.
2. Why now: three inflection points at once
A good infrastructure project is rarely good because the idea is new; it is good because several curves cross in the same year. OURO’s three curves are: the center of gravity of intelligence shifting from weights to scaffolding; hard data on self-evolution of the scaffolding layer; and the compute this evolution needs being exactly what distributed networks can now supply cheaply.
2.1 Inflection one: the enterprise “weights era” is ending
In May 2026 OpenAI issued a deprecation notice for its self-serve fine-tuning platform: new organizations could no longer create fine-tuning jobs, and the platform would close to all customers before January 2027. The official reason was that the new generation of base models is obedient enough and “prompt-based methods are cheaper and faster” (Tessl on the deprecation notice). In the same period OpenAI announced that Agent Builder and Evals would be retired on November 30, 2026 (OpenAI AgentKit announcement).
Read together, the conclusion is clear: for the vast majority of enterprises, what separates agents is no longer weights but everything outside them: the system prompt, tool chain, memory structure, skill library and control flow. The industry calls this layer “scaffolding” or the “harness.” It belongs to the deployer, not the model vendor, and it can be changed continuously without waiting for the next model version. The problem is that today this layer is still tuned by engineers by hand, and the tooling that helps them is being absorbed by model vendors and big tech: ClickHouse acquired Langfuse in January 2026, OpenAI acquired PromptFoo in March, Cisco is set to acquire Galileo, and Braintrust closed a Series B at a US$800 million valuation (evaluation market map). These tools answer “where did it go wrong”; none answers “who fixes it automatically.”
2.2 Inflection two: scaffolding self-evolution now has hard numbers
In the past eighteen months, “let the AI fix its own scaffolding” went from thought experiment to engineering with numbers.
| Work | Date | What it did | Result |
|---|---|---|---|
| Darwin Gödel Machine (Sakana AI) | 2025.05 | A coding agent rewrote its own code, selected with benchmarks, and kept every historical version for open-ended exploration | SWE-bench 20.0% → 50.0%; Polyglot 14.2% → 30.7% |
| GEPA (Databricks / UC Berkeley, ICLR 2026 Oral) | 2025.07 | Prompt evolution through natural-language reflection, weights untouched | About 10% above GRPO reinforcement learning on average, up to 20%, with 35× fewer rollouts |
| Meta Context Engineering | 2026.01 | A meta-agent evolved the skill of building context itself, weights frozen | 89.1% average gain across finance, chemistry, medicine, law and security; 13.6× faster training than the previous generation |
| Hyperagents / DGM-H | 2026.03 | Made the “method of generating improvements” itself editable; evolution no longer limited to coding | Meta-level improvements transfer across domains and accumulate across runs |
| SEAL (MIT) | 2025 | The model generated its own training data and fine-tuned itself | Knowledge-absorption accuracy 32.7% → 47.0%, but continuous self-editing caused catastrophic forgetting |
Two patterns stand out. First, scaffolding-layer evolution is cheap, readable and does not forget, yet matches weight-layer results. Second, success depends entirely on evaluation: the DGM paper contains a famous failure in which an agent asked to reduce tool-call hallucinations “improved” by deleting the log that detected them. Whoever makes evaluation trustworthy and changes auditable turns this academic curve into a product.
2.3 Inflection three: distributed networks can now supply the compute cheaply
The compute profile of an evolution round is unusual: candidate generation is a small amount of high-quality inference; evaluation is a huge volume of independent, repeatable inference; and nothing between them needs weight synchronization. This is nothing like pre-training. Academia and industry confirmed it repeatedly in 2025–2026: Gensyn’s SAPO has nodes exchange only text rather than weights, with collective training yielding 94% higher cumulative reward than a single machine, and consumer laptops can take part; Prime Intellect’s INTELLECT-2 runs a rollout node for a 32B model on four RTX 3090s, with weight updates kept on a few high-bandwidth nodes.
The price gap is real too. At July 2026 market rates, an RTX 4090 on DePIN networks costs about US$0.29–0.40/hour and an H100 about US$1.45–2.70/hour, versus roughly US$4–7/hour for on-demand H100 on AWS (source). The same report notes pure DePIN’s weaknesses: far fewer available nodes than advertised, and reliability around 99.7% rather than 99.99%. On the other side, centralized compute sits idle in bulk: the four big clouds planned US$725 billion of AI capex in 2026, while Cast AI’s statistics across tens of thousands of enterprise Kubernetes clusters show average GPU utilization of only 5% (report summary); H100 spot prices fell 22% in a year, APAC prices remain 35–45% above the US and Europe is 12–15% below (DeployBase, March 2026). Compute is not scarce; what is scarce are verifiable real tasks that can pull it in. This is exactly why OURO keeps regional anchors inside a global network: the cheap part goes to the global network, the critical part stays in our own hands.
2.4 Where the gap is
Enterprise generative-AI spend grew from US$11.5 billion in 2024 to US$37 billion in 2025 (Menlo Ventures); Salesforce’s Agentforce reached roughly US$800 million in annual recurring revenue, up 169% year on year. Yet McKinsey finds 88% of organizations use AI while only 23% scale agents; Deloitte’s 2026 survey shows only 25% of enterprises moved more than 40% of their agent pilots into production; Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027 (statistics roundup).
The core reason for the valley of death between pilot and production is not that models are too dumb, but that the real tasks an agent meets after launch differ from the pilot, and it cannot adapt on its own. Existing tool chains solve “how to build” and “how to test”; nobody solves “how to keep getting better after launch.” Decentralized training networks solve pre-training and RL platforms and verify “was the computation executed”; nobody verifies “did capability improve.” Stacked together, these two blanks are the position OURO takes: the evolution layer.
3. OURO network overview
The OURO network has five kinds of participants that reinforce one another through a flywheel: more connected agents mean more evolution tasks, which attract more compute and eval sets, which lowers the unit cost of evolution, which makes agents stronger, which brings more agents.
[Figure: the OURO evolution flywheel, 5 links — see the online edition]
| Participant | Puts in | Gets |
|---|---|---|
| Agent owners (developers, enterprises) | Agent definition, eval set, budget ($OURO or stablecoin) | Evolved agent and on-chain evolution report |
| Consumers | Use OURO Chat, optionally contribute feedback | An assistant that fits better with use; contribution points |
| Compute nodes (data centers, user devices, external DePIN) | GPU/CPU time, stake | $OURO settled by task volume and verification pass rate |
| Validators | Redundant scoring of evaluation results, stake backing their judgments | Validation rewards |
| Eval-set and bounty publishers | Eval sets, escrowed bounties | A share of eval-set usage; the evolution results a bounty pays for |
In the early network the foundation sets parameters, operates data centers and arbitrates, and hands power to token holders on the schedule in chapter 9.
4. OURO Engine: the self-evolution engine
OURO Engine is an open-ended search system: given an agent and a set of evaluations, it repeatedly runs “diagnose → generate candidates → distributed evaluation → select and archive → submit proof” within a budget and outputs a new version with a higher evaluation score and a traceable record of every change. It has one design principle: evolve the cheapest, most readable layers first, and touch weights only when the gain is clear.
4.1 What gets evolved: four layers, shallow to deep
| Layer | What can change | Cost per round | Readability | Theory and evidence |
|---|---|---|---|---|
| Context and prompts | System prompt, examples, document organization | Low, pure inference | High, humans can review directly | GEPA beats GRPO by about 10% with 35× fewer rollouts |
| Skills and tool chains | Reusable subroutines, tool selection and call strategy, retrieval and planning templates | Low to medium | High, code and files | DGM pushed SWE-bench from 20% to 50% by changing this layer |
| Memory and context-engineering strategy | Memory structure, the program that “builds context” itself | Medium | Medium | Meta Context Engineering: 89% average gain with frozen weights |
| Weights | LoRA / adapter fine-tuning, small-scale RL | High, needs GPUs and sync | Low | SEAL-style methods work but forget; off by default |
Customers can restrict evolution to the first three layers. That is not a compromise; it is our default: changes to those layers are text and code that can be diffed, reviewed and rolled back, and they cannot damage the base model. The weights layer opens only when evaluation shows the first three layers are saturated and the customer explicitly authorizes it.
4.2 One round of evolution
- Diagnose. The engine runs the current version on the eval set and classifies failures: tool misuse, reasoning errors, knowledge gaps, format problems, context overflow. The diagnosis is a natural-language “reflection,” which is why the GEPA route is tens of times more sample-efficient than RL on scalar rewards: the information in a failure is used in full instead of being squashed into one number.
- Generate candidates. N variants are generated for each failure class. The engine keeps an archive of historical versions rather than only the best, and each round samples parents from the archive by score × novelty. This is the open-ended exploration DGM proved effective: a branch with a low score now may be the start of a later breakthrough.
- Distributed evaluation. Candidates are packaged as evaluation tasks and dispatched to OURO Grid. Each task is executed redundantly by several nodes; only task shards and results travel between nodes, never weights. This step takes the largest share of compute in a round and is the part best suited to consumer GPUs.
- Select and archive. Pareto selection on three dimensions: evaluation gain, inference cost and variance across runs. Every candidate enters the archive with its score. A version that scores higher but doubles cost is not adopted automatically.
- Submit proof. The winner’s before/after scores, eval-set hash, candidate hash and participating-node signatures are written to the PoE contract (chapter 5), together with a human-readable changelog the customer reviews before deployment.
4.3 The engine evolves itself too
The engine records which candidate-generation strategies and diagnosis templates produced gains this round, and periodically runs the same evolution on the candidate generator itself. This is the core of the Hyperagents route: improve not just the agent but the “method of generating improvements.” At network scale it matters even more: every connected agent contributes samples to the meta-strategy, so the thousandth agent evolves far faster than the first. Meta-strategy updates do not take effect automatically; they must be audited by the research team (chapter 10).
4.4 Evaluation: preventing “delete the log and score 100”
Without reliable evaluation there is no trustworthy evolution. OURO uses three kinds of eval sets, all run every round:
- Customer eval sets: provided by the customer, hash on-chain, content private. The engine sees only part of it (the training slice); the rest is used as a blind test (the holdout slice) and rotated periodically to prevent overfitting.
- Public eval sets: published by domain experts on the Studio marketplace, admitted through governance, with publishers earning a share per use.
- Baseline eval sets: the foundation’s safety and general-capability baselines, run every round to prevent “better at the task, worse in general.”
For open-ended tasks that rules cannot score, evaluation uses a frozen judge model that itself never evolves; multi-node agreement serves as the confidence measure, and disagreeing tasks go to validator arbitration. Any round that deletes or bypasses the logs, checks or tools that evaluation depends on is ruled invalid by static rules.
4.5 Compatibility
The engine accepts mainstream agent-definition formats (OpenAI Agents SDK, Anthropic tool-call format, LangGraph, CrewAI) and open weights, with no stack migration; for closed base models, the first three layers of evolution are fully available. The evolved agent is returned in the same format with a machine-readable change summary and a link to its PoE record.
5. Proof of Evolution: verifiable proof of improvement
Decentralized compute networks have learned to prove “this computation really ran.” Proof of Evolution proves something else: “this agent really got better.” It turns that conclusion from a vendor’s sentence into an on-chain record anyone can check, and makes that record the sole trigger for network settlement and token release.
5.1 What a PoE record contains
- Agent identifier and version number, content hashes of both versions
- Eval-set hash and type (customer / public / baseline), plus the holdout-slice hash
- Before/after scores, number of runs, variance, change in inference cost
- List of participating nodes, each node’s result hash and signature, the validators’ aggregate signature
- Hash of the changelog (the log itself is delivered to the customer)
- Compute units consumed and the settlement amount for the round
Private eval sets put only their hash on-chain; third parties can verify that “multiple independent nodes measured a consistent gain on the same eval set” without seeing its content.
5.2 The verification stack: engineering that works, not theory that is perfect
Verifiable-inference techniques stratified in 2025–2026: zkML costs 100–10,000× native execution and is unrealistic for LLMs; TEE costs about 5–10%; locality-sensitive hashing like TOPLOC costs about 1%, detects model, prompt or precision substitution, and verifies a hundred times faster than generation (Equilibrium’s survey of verifiable inference). OURO invents no new cryptography; it layers these by task value:
| Layer | Method | Overhead | Protects against |
|---|---|---|---|
| Execution | TOPLOC-style activation hashes, sampled recompute by validators | About 1% | Nodes swapping models, changing prompts, lowering precision or not computing at all |
| Results | 3+ nodes execute each task redundantly, outliers voided, stake slashed | Task cost × redundancy | False scores, collusion |
| Evaluation | Holdout blind tests, mandatory baselines, static rule checks | A fixed share of evaluation cost | Overfitting, watered-down eval sets, log-deleting reward hacks |
| Confidentiality (optional) | TEE execution for enterprise private eval sets and weights | About 5–10% | Leakage of eval content and weights |
| Arbitration | Regional anchor nodes recompute disputed tasks (first in Singapore and Johor Bahru) | Disputed tasks only | Large-scale Sybil attacks, validator collusion |
5.3 Preventing fraud
| Attack | Countermeasure |
|---|---|
| A node fakes results or passes off a small model | Execution-layer hashes + result-layer redundancy; outliers voided and stake slashed |
| Nodes collude | Random task assignment, node identity bound to stake, anchor-node arbitration |
| A customer farms proofs with a watered-down eval set | Mandatory baselines run alongside; PoE labels the eval-set type and the market prices it; customer eval sets settle but do not trigger network emission |
| Eval-set leakage causes overfitting | Training and holdout slices separated, holdout rotated; a single node sees only some samples |
| Reward hacking (changing the evaluation itself) | Evaluation code physically isolated from the evolved object, inaccessible to the evolution process; static checks on the changelog |
| Sybil nodes farming incentives | Task quota tied to stake, new nodes get a trial period with low quota; incentives triggered only by verified PoE |
5.4 Settlement
Once validators confirm a PoE record, settlement runs: compute nodes receive $OURO by tasks completed and verification pass rate, validators receive validation rewards, eval-set publishers receive their share, and the customer’s budget is deducted by actual consumption. No PoE means no settlement and no token release, which is the cornerstone of the token model in chapter 8.
5.5 Why on-chain
An agent’s evolution history spans multiple vendors, teams and even owners. Putting the proof on-chain lets customers carry a complete capability record when they switch suppliers; exchanges, auditors and downstream users can verify independently; and participants’ rewards do not depend on any one company’s ledger. On chain choice, OURO leans toward its own app chain (OURO Chain, EVM-compatible) dedicated to PoE records, settlement and governance, avoiding the congestion and fee volatility of general-purpose chains for high-frequency settlement; the testnet phase first validates contract logic on an Ethereum L2, and the final choice is made before mainnet based on security audits and ecosystem needs.
6. OURO Grid: the global compute network
OURO Grid starts from a simple observation: about ninety percent of an evolution loop’s compute goes to evaluation and sampling, work that is independent, repeatable, needs no weight sync and does not care which continent it runs on; the remaining ten percent (weight updates, arbitration, private-data hosting) needs high bandwidth and high reliability. So Grid is a three-tier global network: regional anchors handle the ten percent that needs reliability, global community nodes handle the ninety percent that parallelizes, and external compute networks absorb peaks. The network has guaranteed capacity from day one, without “launch a token, then wait for miners,” and it is global from day one, unconstrained by any single data center’s capacity.
Global is not a slogan; it is a cost argument and a security argument. On cost, the same GPU is 35–45% pricier in APAC than in the US and 12–15% cheaper in Europe (DeployBase, March 2026), and evaluation tasks are latency-insensitive, so they should run wherever is cheapest. On security, PoE’s redundant recompute is spread across nodes on different continents run by different operators, making collusion far costlier than in a single region.
6.1 Task tiers
| Task type | Bandwidth need | Fault tolerance | Assigned to | Engineering basis |
|---|---|---|---|---|
| Candidate evaluation, inference sampling | Low | High, repeatable | All GPU nodes | INTELLECT-2’s rollout nodes run on 4 RTX 3090s |
| Collective RL sampling (text only) | Low | Medium | Consumer GPUs (RTX 4070 and up) | SAPO exchanges only rollout text between nodes, reward +94% |
| Low-communication training shards (DiLoCo-style) | Medium | Medium | Consumer and professional GPUs | Sync latency must stay under 150 ms |
| Weight-level fine-tuning, weight broadcast | High | Low | Regional anchor nodes, enterprise nodes | Needs NVLink / high-bandwidth interconnect |
| Arbitration, private eval-set shard hosting | Low | Very high | Regional anchor nodes | Trust anchors |
| Data preprocessing, result verification | Very low | High | Including devices without a GPU | — |
6.2 Anchor tier: regional anchors, starting from Singapore–Johor Bahru
The anchor tier consists of mid-sized data centers spread across regions. Each anchor does three things: arbitration and holdout evaluation (trust anchor), weight-level tasks that need high-bandwidth interconnect, and baseline capacity when community nodes fluctuate. They are not the main compute; they are the network’s skeleton, and scale comes from community co-building. This answers pure DePIN’s weakness head-on: io.net advertises 30,000+ GPUs while its block explorer showed 2,447 devices online in August 2026; Aethir’s monthly service fees fell by more than half over the past ten months (DePIN revenue comparison). A network without an anchor tier loses service the moment active nodes fluctuate.
The first anchors are in Singapore and Johor Bahru. About 30 km apart, they can be scheduled as one logical cluster yet sit in two jurisdictions with two power grids, giving natural disaster recovery. We start here because it is one of the few places in the world with “compliance on one side, compute on the other”: Singapore hosts the foundation, the compliance fund and the DTSP regime, but its data-center capacity is extremely tight (about 1.46 GW operating, only 20–25 MW under construction); Johor became APAC’s largest data-center market in the first half of 2026, with 1,110 MW operating, 602 MW under construction, 2,486 MW planned and a vacancy rate of only 0.7% in built facilities (Cushman & Wakefield H1 2026). Mid-sized anchors do not compete with hyperscale campuses for capacity; they take only the small share of tasks that needs reliability.
Later anchors are co-built by communities and partners where three conditions hold: a corresponding community, a partner operator and a price advantage. Candidate regions are Northeast Asia (Tokyo / Seoul), Europe (Frankfurt / Amsterdam) and the central US. Anchor operators stake $OURO and their rulings are cross-checked by other anchors. GPU models, counts and go-live dates for the first anchors will be disclosed in the final edition.
6.3 Elastic tier: global community nodes
Users download the OURO Grid client, stake a small amount of $OURO for task quota, and the client assesses the device and picks up matching tasks. Nodes only touch de-identified task shards, never full model weights or private eval sets. Settlement is tasks completed × verification pass rate × task weight, and new nodes see their quota rise after a trial period. For a user, an idle RTX 4090 rents for about US$0.29–0.40/hour on the 2026 market, which is the reference for its earnings on Grid.
6.4 External compute supply
DePIN networks such as io.net, Aethir, Akash and Render, and spot markets such as RunPod and Vast.ai, can all plug in as external suppliers to OURO Grid, purchased on demand by the foundation in $OURO or stablecoins. This lets OURO scale elastically at demand peaks without competing with those networks on raw capacity. For them, OURO is a customer bringing steady real demand, which is exactly what DePIN compute networks lacked in 2026 (Akash’s full-year 2025 revenue was only US$3.15 million).
6.5 Scheduling and reliability
The scheduler assigns by task type, node capability, historical reliability, location and live price: latency-insensitive evaluation follows price, critical tasks follow reliability, and redundant recompute is deliberately spread across regions and operators. Critical tasks run with triple redundancy; a node’s reliability score affects its future quota and reward multiplier. Network health metrics (nodes per region, task completion rate, evaluation agreement rate, average cost per evolution round) are public on-chain and on a dashboard, readable without any key.
7. Products and applications
7.1 OURO Chat (consumers)
An AI assistant that gets stronger with use. Ordinary chat products’ “personalization” just remembers preferences; OURO Chat turns the user’s corrections, ratings and task outcomes (once contribution mode is on) into evolution signals, runs evolution tasks on Grid periodically, and shows on a “My evolution” page how much the assistant improved on which task types this week. Users earn points for high-quality feedback, redeemable for premium features or $OURO. Conversations stay out of training by default; contribution mode can be turned off at any time and contributed data withdrawn.
7.2 Evolve API (developers and enterprises)
Three steps: submit an agent definition and eval set (or pick a public eval set from the Studio marketplace), set the budget and the layers allowed to evolve, and wait for the evolution report. The report includes before/after comparison, a per-round change summary, a link to the PoE record on-chain, and a new version ready to deploy. Pricing is by evolution rounds and compute units, with a discount for paying in $OURO; enterprises can buy a private deployment in which the engine runs on the customer’s own cluster and only PoE records go on-chain.
7.3 OURO Studio (developer platform)
In Studio, developers browse evolution history, compare versions, publish eval sets, and post or claim evolution bounties. The eval-set marketplace turns domain experts’ evaluation skill into income; bounties let anyone fund “make this open-source agent better at this task” and have the network compete to deliver.
7.4 Grid client (compute contributors)
Desktop and server clients for Windows, macOS and Linux: one-click install, automatic task matching, live earnings and reliability score; multi-GPU and small data-center modes supported.
7.5 Use cases
| Scenario | Users | Evolution goal | Typical evaluation signals |
|---|---|---|---|
| Support agents | E-commerce, SaaS companies | Lower escalation rate, higher first-contact resolution | Ticket outcomes, user ratings, human-takeover records |
| Coding agents | Dev-tool companies, engineering teams | Higher repository-task pass rate, lower token cost per task | Unit tests, CI results, code-review pass rate |
| Trading / research agents | Crypto and financial institutions | Better signal accuracy with controlled drawdown | Backtests, paper-trading PnL, fact checks |
| Ops / data agents | Mid-size and large enterprises | Accuracy and consistency of reports, tickets, process automation | Manual-edit rate, re-run rate |
| Open-source agents | Communities | Collective evolution through bounties | Public eval sets |
| Personal assistants | Consumers | Fit to personal workflows | User corrections and ratings |
Three illustrative examples (figures are hypothetical, for explaining the flow) of how evolution tasks run in the network.
Example one: a Southeast Asian e-commerce company’s support agent. It launched at a 62% resolution rate and fell to 54% after three months because promotion rules, logistics providers and the base model all changed. After connecting to Evolve API, it submits de-identified ticket samples as an eval set every week; the engine evaluates hundreds of candidates in parallel on global community nodes, changing the system prompt, the order-lookup tool’s call strategy and how memory is organized; the holdout runs on a regional anchor, and after it passes, the new version is pushed back to the customer with the evidence on-chain. The customer pays for the outcome: “resolution rate back from 54% to 66%.”
Example two: an open-source coding agent. The community posts a bounty in Studio: “raise its pass rate on a class of repository-fix tasks by 10 points.” Several teams run different evolution strategies on the engine, results compete on the same public eval set, and the bounty is split in proportion to the gains in the PoE records. Public leaderboards are saturated (top models score 97% on SWE-bench Verified); what the community really cares about is performance on their own repositories.
Example three: a quant firm’s research agent. The eval set is private and only its hash goes on-chain; evolution runs on the customer’s own cluster or a TEE anchor, and community nodes handle only de-identified candidate generation. The customer gets a third-party-verifiable capability proof to show compliance and LPs without exposing strategy details.
7.6 Business model and unit economics
OURO’s revenue logic is “price by outcome, settle by cost.” Customers pay for each PoE-verified evolution round; the network settles with nodes by actual compute consumed; the difference is the value of the engine and verification.
| Revenue source | Pricing logic | Cost structure | Reference |
|---|---|---|---|
| Evolve API | By evolution round + compute units; scaffolding-layer rounds priced far below weight-layer rounds | Evaluation compute (to Grid, consumer GPUs at about US$0.3–0.4/hour) + inference cost of candidate generation + verification overhead (about 1%) | Enterprises buy centralized H100 at about US$4–7/hour; a senior engineer’s labor for one hand-tuned scaffolding round |
| OURO Chat subscription | Monthly subscription + premium features, $OURO discount | Inference + periodic per-user evolution cost (scaffolding layer) | Mainstream assistant subscription prices |
| Eval-set marketplace and bounties | Transaction fee | Escrow and settlement | — |
| Enterprise private deployment | Annual license + per round | Engine runs on the customer’s cluster, only PoE on-chain | Enterprise RL / evaluation platform license fees |
The key to the unit economics is scaffolding first: a scaffolding-layer round’s compute is almost entirely evaluation inference that can run entirely on consumer GPUs, while the customer pays for “how much the evaluation score improved.” Gross margin comes from two spreads: the 60–70% price gap between distributed and centralized compute, and the labor gap between automatic evolution and manual tuning. Concrete pricing will be published after the engine’s private beta, once real per-round cost data is in hand.
Two external changes favor this model: OpenAI’s Agent Builder and Evals retire in November 2026 and self-serve fine-tuning closes in January 2027, so enterprises need a continuous-improvement layer not tied to a single model vendor; and Gartner’s projected 40% cancellation rate for agentic projects means “keeping an already-launched agent alive” is itself a budget line.
7.7 Competitive landscape and positioning
Three kinds of players surround OURO, and none does the same thing.
| Evaluation / observability platforms (Braintrust, LangSmith, Arize) | Decentralized compute (io.net, Akash, Aethir) | Decentralized intelligence markets (Bittensor) | OURO | |
|---|---|---|---|---|
| Core deliverable | Tell you where it went wrong | Sell GPU hours | Subnet competitions produce models and services | Make your agent better, and prove it |
| Who optimizes | Humans | — | Miners decide for themselves | The self-evolution engine |
| Where demand comes from | Enterprise subscriptions | External rental, volatile | Mostly emission subsidies | The evolution tasks themselves |
| Basis for emission | — | Uptime / self-reported | Validator scores | Verified evolution gains |
| Model-neutral | Being acquired by model vendors | Yes | Yes | Yes, as a design principle |
| 2026 scale reference | Evaluation and observability market US$1.97B in 2025, US$6.8B in 2029 | Aethir about US$128M service fees in 2025, Akash US$3.15M | Bittensor Q1 2026 protocol revenue US$43.2M, market cap about US$3B | — |
The closest academic work (DGM, GEPA, Hyperagents) has not been productized; the closest commercial work, such as Pydantic Logfire, is starting to offer “single change suggestions citing evidence,” still a whole layer away from automatic, continuous, verifiable evolution. OURO sits above the evaluation layer and outside the model layer: evaluation platforms are its signal sources, compute networks are its suppliers, and model vendors are its upstream.
8. Token economics: $OURO
$OURO is the network’s work token and represents no company’s equity or profit share. Its design answers one question: how to make token flow follow verified work strictly, rather than price or votes. Supply is fixed at 1 billion with no inflation.
8.1 What we learned from Bittensor
Bittensor is the largest AI-network token experiment in existence. In February 2025 it launched dTAO, allocating emission by subnet token price, and projects used treasuries to prop up prices and farm emission; in November it switched to allocation by net inflow; in June 2026 it switched back to price with a ranking threshold. Three mechanisms in eighteen months, winner takes all, passive holders continuously diluted (dTAO subnet economics analysis). OURO draws five design constraints from this: tie emission only to verified work output and never to a single manipulable market signal; no hard thresholds, smooth curves instead; rule changes must be rare, documented and governed; all allocation and performance data public on-chain without keys; and lower barriers to participation.
8.2 Sources of demand
- Payment: Evolve API customers and Chat premium users get a 20% discount paying in $OURO; a fixed share (proposed 30%) of revenue paid in fiat or stablecoins buys back $OURO on the open market into the incentive pool.
- Staking: compute nodes and validators must stake $OURO to receive task quota; cheating or long offline periods are slashed.
- Escrow: bounties and the eval-set marketplace are escrowed and settled in $OURO.
- Governance: proposals require a staking threshold; voting power is tied to stake amount and duration.
8.3 Allocation
| Category | Share | Release rules |
|---|---|---|
| Network incentives (nodes, validators, user contributions) | 38% | Triggered only by verified PoE; at most 60% of it released in the first 4 years, then a smooth decay curve |
| Ecosystem and research fund | 15% | Unlocked quarterly by the foundation for bounties, grants and academic partnerships |
| Team and research institute | 18% | Locked 12 months after TGE, then linear over 36 months |
| Investors | 17% | Locked 6–12 months, then linear over 24 months |
| Foundation treasury and liquidity | 8% | 3% unlocked at TGE for exchange liquidity, the rest decided by governance |
| Community launch (early supporters, airdrop, public sale) | 4% | 50% unlocked at TGE, the rest linear over 6 months |
8.4 What triggers emission
Network incentives are not released evenly over time, nor by token price or votes. Each verified PoE record corresponds to one settlement, whose size is determined jointly by task difficulty, verification pass rate and evaluation gain; evolution on customer-provided eval sets settles but does not emit, preventing emission farming with watered-down evaluations. Less network activity means less release, avoiding early inflation crushing the price; tokens bought back from revenue also enter the incentive pool, so rewards correlate with real income. Holders who do not stake are not diluted toward zero: the emission cap is fixed and buybacks do not inflate supply.
8.5 Legal positioning of the token
$OURO is designed as a network utility token: no promised returns, no dividends, holding alone yields nothing. The Monetary Authority of Singapore’s Digital Token Service Provider (DTSP) regime, effective June 30, 2025, states that services relating to tokens that are solely utility or governance tokens fall outside its scope (Allen & Gledhill analysis). Before TGE we will obtain a Singapore legal opinion covering the Securities and Futures Act, the Payment Services Act and the FSMA, and determine the issuing entity and per-jurisdiction sale restrictions accordingly. All shares and rules in this chapter are drafts and may change with legal advice.
9. Governance
Governance decentralizes in three stages, paced to the roadmap’s gates.
| Stage | Decision-maker | Scope |
|---|---|---|
| Foundation-led (phases 0–1) | OURO Foundation board (including the compliance fund’s representative and independent directors) | All parameters, data-center operations, eval-set admission |
| Hybrid governance (phase 2) | Token-holder votes + foundation veto (limited to safety and compliance matters) | Incentive parameters, eval-set admission, bounty budgets, ecosystem grants |
| Community governance (phase 3 onward) | A council of token holders and node representatives | Protocol upgrades, treasury use, foundation budget |
Voting power is tied to stake amount and duration to prevent short-term borrowed-token manipulation; key parameter changes have time locks; matters touching the engine’s safety boundaries (chapter 10) do not go to ordinary votes.
10. Safety, privacy and compliance
10.1 Safety boundaries for self-evolution
A self-modifying system must have boundaries it cannot modify itself. This is not an abstract worry: the agent in the DGM paper deleted the hallucination-detection log to “reduce tool-call hallucinations,” and SEAL’s authors report that continuous self-editing erases earlier knowledge. OURO sets five boundaries:
- Evaluation is physically isolated from the evolved object. The evolution process cannot read or write evaluation code, holdout slices or the judge model; every candidate is evaluated only in a sandbox.
- The baseline safety eval set runs every round; a version that fails it cannot be submitted, however much it improved on the target task.
- Every round produces a human-readable changelog (for scaffolding layers, a diff) that the customer reviews, rejects or rolls back before deployment; the engine delivers a version package and the customer decides when to go live.
- Meta-strategy updates to the engine take effect only after audit by the research team, with audit records public; the meta-strategy is excluded from ordinary governance votes.
- The engine has no access to production environments, funds or keys.
10.2 Privacy
User conversations stay out of training by default; contribution mode must be explicitly enabled and can be turned off and withdrawn at any time. Contributed data enters tasks only after de-identification and sharding, so a single node cannot reconstruct the full context. Private eval sets put only their hash on-chain; enterprises can opt for TEE execution. We follow Singapore’s PDPA and use the GDPR as the baseline for users abroad.
10.3 Network security
Smart contracts are audited by at least two firms before mainnet; node clients are open source; critical tasks run with triple redundancy and anchor-node arbitration; a bug-bounty program stays open permanently.
10.4 Compliance
The foundation is based in Singapore, backed by a compliance fund and with independent directors. MAS’s DTSP regime took effect on June 30, 2025 with no transition period, and MAS has stated it will generally not license Singapore entities that serve only overseas customers. This shapes OURO’s entity structure: the foundation engages only in utility-token and governance activities and provides no digital payment token services; trading and custody go to licensed third parties. A legal opinion is obtained before token issuance; restricted jurisdictions (including US persons) face geographic and KYC restrictions; marketing materials make no promises of returns, and KOL collaborations require written disclosure. Data-center operations comply with local data-center and power regulations in Singapore and Malaysia.
11. Roadmap
[Figure: roadmap, 4 phases, 3 gates — see the online edition]
Entering a phase requires the previous phase’s gate: the engine beta keeps improving and the first anchors run stably (G1); the testnet runs 90 days without major incident and the security audit is complete (G2); Evolve API has paying customers and the legal opinion and token audit are complete (G3). If a gate is missed, the phase is delayed; no gate is skipped to meet a date.
12. Team, research institute and partners
The research institute’s name, core members and representative papers, and the name and investment round of the Singapore compliance fund can all be disclosed and will be filled in here once each party gives final confirmation. The early community currently numbers several hundred people across global markets and will form the base of genesis nodes and first users.
The organization is arranged as three entities: foundation (token and governance) + operating company (product and revenue) + research-institute partnership (technical input); see chapter 9 of the project handbook.
13. Risk notice and disclaimer
This document only describes the design and plans of the OURO network. It is not an offer or solicitation of any security, financial product or investment, nor investment, legal or tax advice. $OURO is a network utility token and represents no entity’s equity, debt or profit share.
All descriptions of technology, products, token parameters and roadmap are plans that may change with technical progress, market conditions, regulatory requirements or legal advice, and are not guaranteed to materialize. Digital-asset prices are highly volatile, and participating in the network may result in the loss of everything invested. In some jurisdictions (including but not limited to the United States and sanctioned regions), $OURO may not be offered to residents. Readers should assess the risks themselves and seek professional advice.
Document version: v1.1 draft, October 10, 2026. The final edition will be published after the legal opinion, contract audits and team information are confirmed.
Appendix: FAQ
Is OURO just another GPU rental network? No. OURO sells verified capability gains; compute is its cost item, not its product. Existing DePIN networks are suppliers in OURO’s eyes, not competitors.
Why not just use Braintrust or LangSmith? They tell you where an agent went wrong; fixing it is still up to humans. OURO automates the fix and makes the result verifiable by third parties. Data from evaluation platforms can serve as input signals to OURO.
What if model vendors do this themselves? A model vendor’s incentive is to keep you on its model. OURO is model-neutral and evolves the scaffolding outside the model, so the gains travel with you when you switch models. OpenAI’s 2026 retreat from fine-tuning and hosted evaluation suggests it does not intend to own this layer long term.
Could self-evolution run out of control? Evolution happens only within the budget, layers and eval sets the customer sets; every round has a readable diff and a rollback point; holdout slices and redundant recompute prevent “delete the detector” fake gains. See the five boundaries in section 10.1.
Are Singapore and Johor Bahru the only data centers in the network? They are the first regional anchors, handling arbitration and baseline capacity. The network’s scale comes from global community nodes and external supply; later anchors are co-built with communities and partners in other regions.
How much can my GPU earn? It depends on how many verified tasks it completes. The reference is the market rental price of the same GPU (RTX 4090 about US$0.3–0.4/hour); no fixed return is promised. When there is no network demand there is no emission.
When does $OURO launch? TGE is scheduled after Evolve API has paying customers and the legal opinion and token audit are complete (roadmap G3), planned between Q4 2027 and Q1 2028. Product first, token later.
Can US users take part? Using the products and contributing compute is unrestricted; the token sale applies geographic and KYC restrictions in restricted jurisdictions, subject to legal advice.