Agent Evaluation & Observability
📌 Planned placeholder page: this page is reserved for "Agent evaluation and observability — trajectory-level evaluation methodology, tracing and failure analysis", and the content will be filled in progressively along with learning.
Planned Content
- [ ] Layered metrics: outcome / process / safety / efficiency
- [ ] Evaluation sets: templated tasks, verifiers, sandbox resets, golden trajectories
- [ ] LLM-as-judge for trajectories with evidence extraction and rubrics
- [ ] Task / plan / LLM / tool span design and failure analysis loop
- [ ] Evaluation-observability-gate closed loop
Next Steps
- [ ] Complete this page item by item against the planning checklist in the Domain Overview
- [ ] Add runnable examples and pitfall records for each entry
- [ ] Change the status from "Planned" to "Collected" when done
For writing guidelines, please refer to the Domain Overview.