我们提供 MobileGym, 一个浏览器托管的, 轻量级, 完全可控的环境,适合日常移动使用, 目标是交互保真度,无需复制专有后端。它通过对结构化 JSON 状态进行基于确定性状态的判断来实现日常应用程序之前无法实现的两种功能: 可验证结果信号, 以及通过低成本并行推出实现可扩展的在线强化学习。捕获完整的环境状态, 配置, 分叉, 并与结构化 JSON, 进行比较,单个服务器可以托管数百个并行实例,,每个实例大约 400 MB 内存,冷启动时间大约为 3 秒。分层状态模型和声明性任务定义框架使状态可编程性和任务创建在规模,上保持实用,并且单一编程判断机制提供确定性评估结论和密集的强化学习奖励。随附的 MobileGym-Bench 提供 416 个参数化任务模板,,包括 256 个测试模板和 160 个训练模板,,超过 28 个应用程序,,具有确定性判断和结构化 AnswerSheet 协议,可避免自由文本匹配失败。在 Sim-to-Real 案例研究中,Qwen3-VL-4B-Instruct 上的, GRPO 在 256 任务测试集, 和 59 任务真实设备信号子集上, GRPO 获得了 +12.8 个百分点,, 真实设备执行保留了 95.1% 的模拟端训练增益。项目页: 此 https URL。
We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without replicating proprietary backends. It enables two capabilities previously out of reach for everyday apps: verifiable outcome signals through deterministic state-based judging over structured JSON state, and scalable online RL through low-cost parallel rollouts. The full environment state is captured, configured, forked, and compared as structured JSON, and a single server can host hundreds of parallel instances, with about 400 MB memory per instance and about 3 s cold start. A layered state model and a declarative task-definition framework keep state programmability and task creation practical at scale, and a single programmatic judging mechanism delivers both deterministic evaluation verdicts and dense RL rewards. The accompanying MobileGym-Bench provides 416 parameterized task templates, including 256 test and 160 train templates, over 28 apps, with deterministic judges and a structured AnswerSheet protocol that avoids free-text matching failures. In a Sim-to-Real case study, GRPO on Qwen3-VL-4B-Instruct gains +12.8 percentage points on the 256-task test set, and on a 59-task real-device signal subset, real-device execution retains 95.1% of the simulation-side training gain. Project page: this https URL.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)