当应用组相对策略优化 (GRPO) 进行 GUI 基础, 部署时,从单个屏幕截图视图中采样; 组通常要么在困难实例上全部失败,要么在简单实例上全部成功,,不会产生有用的相对优势。我们提出了 VISTA (View 一致的自我验证训练), 一个基于 GRPO 的训练框架,该框架从同一 GUI 的多个目标保留视图构建每个比较组。该 http URL 视图由保持目标元素可见并精确重新映射其框的裁剪生成。, 因此模型的推出可以在语义上等效但几何上不同的输入之间进行比较。为了稳定短坐标生成而不将强化学习变成无条件模仿, VISTA 进一步添加了一个自我验证的交叉视图锚:,这是一个通过优势加权损失优化的预言机答案,,该答案被排除在组基线之外,并且仅在模型产生最大奖励推出时才激活。在五个 GUI 接地基准和多个 Qwen 主干, VISTA 持续改进此 http URL ScreenSpot-Pro, 的接地时,它将 Qwen3-VL 4B/8B/30B-A3B 从 55.5/52.7/53.7 提高到 63.4/65.8/67.0。稳健性分析进一步显示出更高的最差视图准确度和更低的预测翻转率。

When applying Group Relative Policy Optimization (GRPO) for GUI Grounding, rollouts are sampled from a single screenshot view; groups often become either all failures on difficult instances or all successes on easy ones, yielding no useful relative advantage. We propose VISTA (View-Consistent Self-Verified Training), a GRPO-based training framework that constructs each comparison group from multiple target-preserving views of the same GUI this http URL view is generated by a crop that keeps the target element visible and remaps its box exactly, so model rollouts are compared across semantically equivalent but geometrically different inputs. To stabilize short coordinate generation without turning reinforcement learning into unconditional imitation, VISTA further adds a self-verified cross-view anchor: an oracle answer optimized with an advantage-weighted loss, excluded from the group baseline and activated only when the model has produced a maximum-reward rollout. Across five GUI-grounding benchmarks and multiple Qwen backbones, VISTA consistently improves grounding this http URL ScreenSpot-Pro, it raises Qwen3-VL 4B/8B/30B-A3B from 55.5/52.7/53.7 to 63.4/65.8/67.0. Robustness analyses further show higher worst-view accuracy and lower prediction flip rates.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)