跳到內容

評測集格式

JSONL,每行一個樣本。四種期望類型:精確比對、規則、rubric 評分、程式碼測試。引擎會自動留出 20% 作為盲測切片,不參與進化,只用於最終驗證。

# evals.jsonl — 每行一個樣本;expected 可以是精確答案、規則或評分 rubric
{"id": "t-0001", "input": {"messages": [...]}, "expected": {"type": "exact", "value": "ORDER-8821 已退款"}}
{"id": "t-0002", "input": {"messages": [...]}, "expected": {"type": "rule", "must_call": "lookup_order", "must_not": ["轉真人"]}}
{"id": "t-0003", "input": {"messages": [...]}, "expected": {"type": "rubric", "criteria": ["引用了正確的退貨政策", "語氣禮貌"], "judge": "ouro-judge-v1"}}
{"id": "t-0004", "input": {"messages": [...]}, "expected": {"type": "code", "tests": "tests/test_refund.py"}}