Eval set format
JSONL, one sample per line. Four expectation types: exact match, rule, rubric score and code test. The engine automatically holds out 20% as a blind slice that never takes part in evolution and is used only for final verification.
# evals.jsonl — one sample per line; expected can be an exact answer, a rule or a scoring rubric{"id": "t-0001", "input": {"messages": [...]}, "expected": {"type": "exact", "value": "ORDER-8821 refunded"}}{"id": "t-0002", "input": {"messages": [...]}, "expected": {"type": "rule", "must_call": "lookup_order", "must_not": ["escalate_to_human"]}}{"id": "t-0003", "input": {"messages": [...]}, "expected": {"type": "rubric", "criteria": ["cites the correct return policy", "polite tone"], "judge": "ouro-judge-v1"}}{"id": "t-0004", "input": {"messages": [...]}, "expected": {"type": "code", "tests": "tests/test_refund.py"}}