图形用户界面 (GUI) 自动化在现实环境中仍然具有挑战性,,其中动态布局, 意外对话框, 和不断变化的界面状态可能会导致自主代理偏离用户意图。最近基于视觉的多模式代理通过直接在屏幕截图和自然语言指令,上操作来提高灵活性,但规划和适应通常仍然是内部,,限制了用户'检查,监督,或纠正系统行为的能力。我们提出了 Plover, 一个以计划为中心、基于视觉的 GUI 自动化系统,它将任务计划和重新计划具体化为持久的,、可检查的, 和可修改的工件。通过计划器 - 执行器架构, Plover 支持通过可编辑计划, 自然语言指导, 和基于屏幕截图的干预, 来明确监督不断演变的执行, 本地化修正,,同时保留修复期间的先前进度。一项由六名参与者参与的形成性研究为交互设计提供了信息。然后,我们通过基准故障案例修复和基于场景的工作流程分析来评估 Plover。我们的结果表明,当计划保持可见且干预措施局部化, 时,许多自主 GUI 代理故障在结构上是可修复的,并且明确的重新规划有助于使 GUI 自动化更加透明, 可控, 和适应性强。
Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users ability to inspect, supervise, or correct system behavior. We present Plover, a plan-centric vision-based GUI automation system that externalizes task plans and replanning as persistent, inspectable, and revisable artifacts. Through a planner--executor architecture, Plover supports explicit supervision of evolving execution, localized correction through editable plans, natural-language guidance, and screenshot-grounded interventions, while preserving prior progress during repair. A formative study with six participants informed the interaction design. We then evaluate Plover through benchmark failure-case repair and scenario-based workflow analyses. Our results show that many autonomous GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning helps make GUI automation more transparent, controllable, and adaptable.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)