Skip to content

Agent Evaluation & Observability

📌 Planned placeholder page: this page is reserved for "Agent evaluation and observability — trajectory-level evaluation methodology, tracing and failure analysis", and the content will be filled in progressively along with learning.

Planned Content

  • [ ] Layered metrics: outcome / process / safety / efficiency
  • [ ] Evaluation sets: templated tasks, verifiers, sandbox resets, golden trajectories
  • [ ] LLM-as-judge for trajectories with evidence extraction and rubrics
  • [ ] Task / plan / LLM / tool span design and failure analysis loop
  • [ ] Evaluation-observability-gate closed loop

Next Steps

  • [ ] Complete this page item by item against the planning checklist in the Domain Overview
  • [ ] Add runnable examples and pitfall records for each entry
  • [ ] Change the status from "Planned" to "Collected" when done

For writing guidelines, please refer to the Domain Overview.

Built with VitePress · Knowledge shared openly