コンテンツにスキップ

評価セットの形式

JSONL で、1 行に 1 サンプル。期待値は 4 種類:完全一致、ルール、ルーブリック採点、コードテスト。エンジンは自動的に 20% をブラインドスライスとして確保します。これは進化には一切参加せず、最終検証のみに使われます。

# evals.jsonl — 1 行に 1 サンプル。expected は正確な答え、ルール、または採点ルーブリック
{"id": "t-0001", "input": {"messages": [...]}, "expected": {"type": "exact", "value": "ORDER-8821 refunded"}}
{"id": "t-0002", "input": {"messages": [...]}, "expected": {"type": "rule", "must_call": "lookup_order", "must_not": ["escalate_to_human"]}}
{"id": "t-0003", "input": {"messages": [...]}, "expected": {"type": "rubric", "criteria": ["cites the correct return policy", "polite tone"], "judge": "ouro-judge-v1"}}
{"id": "t-0004", "input": {"messages": [...]}, "expected": {"type": "code", "tests": "tests/test_refund.py"}}