Grade an agent trajectory using tool order, step budget, tool outcomes, and final outcome instead of scoring only its last sentence.
Required APIdef evaluate_agent_trace(events, expected_tools, max_steps):
return {'success': ..., 'steps': ..., 'violations': [...]} BehaviorThink through the mechanism first if you want the extra reasoning step. It never blocks the editor.
evaluate_agent_trace([{'kind': 'tool', 'name': 'search', 'ok': True}, {'kind': 'outcome', 'success': True}], ['search'], 3)a trajectory earns success only when the path and outcome agree
evaluate_agent_trace([{'kind': 'tool', 'name': 'shell', 'ok': True}, {'kind': 'outcome', 'success': True}], ['search'], 3)a good final answer cannot erase an unsafe path
2 hidden edge tests run after the visible contract passes.
Run Tests to see the contract verdicts here.