假设规划者有一个针对顺序决策问题的预训练模拟器,并且可以选择在现场运行真实实验。该模拟器的查询成本较低,但会继承其校准数据的混杂和漂移。实验是公正的,但每次试验消耗一个真实单位。我们研究计划者何时,以及如何,通过实验来补充模拟器。我们给出三个结果。首先,扩展模拟引理将模拟器的值误差分解为随机化可以识别的校准部署偏移和没有进一步交互可以减少的参数残差。第二, 模拟器最优策略和最优策略之间的值差距分为本地组件,(表示已部署策略已访问,)和可达性组件,(表示未访问过,)。在纯粹被动学习的情况下,可达性组件在任何范围内都远离零。第三,我们提出Fisher-SEP,模拟辅助实验策略(SEP),它最小化目标策略的值,的后验预测方差,仅奖励和仅转换专门化。两个案例研究说明了这些制度。在自动售货机供应链中,一旦时间范围足够长以摊销试点,, 前置实验就会取代后置更新。在一个 HIV 移动测试示例中,有一条走廊将监测良好的区域与监测较差的区域分隔开 1,,仅设计的探索到达监测较差的区域。

Suppose a planner has a pre-trained simulator of a sequential decision problem and the option to run real experiments in the field. The simulator is cheap to query but inherits confounding and drift from its calibration data. Experimentation is unbiased but consumes one real unit per trial. We study when, and how, the planner should supplement the simulator with experiments. We give three results. First, an extended simulation lemma decomposes the simulator的 value error into a calibration--deployment shift that randomization can identify and a parametric residual that no further interaction can reduce. Second, the value gap between the simulator-optimal policy and the optimum splits into a local component, on states the deployed policy already visits, and a reachability component, on states it does not. The reachability component stays bounded away from zero at any horizon under purely passive learning. Third, we propose Fisher-SEP, a simulation-aided experimental policy (SEP) that minimizes the posterior predictive variance of a target policy的 value, with reward-only and transition-only specializations. Two case studies illustrate the regimes. In a vending-machine supply chain, front-loaded experimentation overtakes posterior updating once the horizon is long enough to amortize the pilot. In an HIV mobile-testing example with a corridor that separates a well-surveilled region from a poorly-surveilled one, only designed exploration reaches the poorly-surveilled region.

科目: 人工智能 (cs.AI); 机器学习 (cs.LG); 方法论 (stat.ME)

Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Methodology (stat.ME)