GUI Agent (Computer Use)
📌 Planned placeholder page: this page is reserved for "GUI Agent (Computer Use) — Perception, action space, verification loop, sandboxing and evaluation for agents that operate interfaces", and the content will be filled in progressively along with learning.
Planned Content
- [ ] Three kinds of automation: scripts, GUI agents and RPA, and when to use each
- [ ] Perception: screenshots, accessibility trees, Set-of-Mark, hybrid approaches
- [ ] Action space and the verification loop (terminal state must be checked programmatically)
- [ ] Sandboxing and security: prompt injection, credential handling, action tiers, audit trails
- [ ] Benchmarks, self-built evaluation sets and honest capability limits
Next Steps
- [ ] Complete this page item by item against the planning checklist in the Domain Overview
- [ ] Add runnable examples and pitfall records for each entry
- [ ] Change the status from "Planned" to "Collected" when done
For writing guidelines, please refer to the Domain Overview.