Skip to content

Eval set format

JSONL, one sample per line. Four expectation types: exact match, rule, rubric score and code test. The engine automatically holds out 20% as a blind slice that never takes part in evolution and is used only for final verification.

# evals.jsonl — one sample per line; expected can be an exact answer, a rule or a scoring rubric
{"id": "t-0001", "input": {"messages": [...]}, "expected": {"type": "exact", "value": "ORDER-8821 refunded"}}
{"id": "t-0002", "input": {"messages": [...]}, "expected": {"type": "rule", "must_call": "lookup_order", "must_not": ["escalate_to_human"]}}
{"id": "t-0003", "input": {"messages": [...]}, "expected": {"type": "rubric", "criteria": ["cites the correct return policy", "polite tone"], "judge": "ouro-judge-v1"}}
{"id": "t-0004", "input": {"messages": [...]}, "expected": {"type": "code", "tests": "tests/test_refund.py"}}