Skip to content

GUI Agent (Computer Use)

📌 Planned placeholder page: this page is reserved for "GUI Agent (Computer Use) — Perception, action space, verification loop, sandboxing and evaluation for agents that operate interfaces", and the content will be filled in progressively along with learning.

Planned Content

  • [ ] Three kinds of automation: scripts, GUI agents and RPA, and when to use each
  • [ ] Perception: screenshots, accessibility trees, Set-of-Mark, hybrid approaches
  • [ ] Action space and the verification loop (terminal state must be checked programmatically)
  • [ ] Sandboxing and security: prompt injection, credential handling, action tiers, audit trails
  • [ ] Benchmarks, self-built evaluation sets and honest capability limits

Next Steps

  • [ ] Complete this page item by item against the planning checklist in the Domain Overview
  • [ ] Add runnable examples and pitfall records for each entry
  • [ ] Change the status from "Planned" to "Collected" when done

For writing guidelines, please refer to the Domain Overview.

Built with VitePress · Knowledge shared openly