数据主权法规越来越要求公共机构部署开源, 本地 LLM 代理,这些代理跨实时政府 API 链接多个工具调用。然而,, 开源模型在这种多步骤设置, 中始终表现不佳,并且没有现有基准衡量差距。我们引入了韩国开放公共 API 基准(KOPA-Bench),,其中包含 145 个实际任务。为了缩小这一差距,,我们提出了 EDGE, 一个基于执行的动态图,用于由实时执行驱动的工具调用数据合成。 EDGE 构建一个图表,显示每个工具的 输出如何提供另一个的 输入, 仅保留实际调用实时 API, 时成功的链接,并遍历这些经过验证的链接以合成可执行的多步轨迹。通过 GRPO 对结果数据集, 进行微调,我们的 9B 模型几乎与同一系列中未调整的 27B 模型相匹配, 不仅在 KOPA-Bench 上而且在 BFCL 基准上都有显着改进。

Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-Bench), comprising 145 real-world tasks. To close this gap, we present EDGE, an Execution-grounded Dynamic Graph for tool-calling data synthEsis driven by live execution. EDGE builds a graph of how each tool的 output can feed another的 input, keeps only the links that succeed when actually called against the live APIs, and traverses these verified links to synthesize executable multi-step trajectories. Fine-tuned via GRPO on the resulting dataset, our 9B model nearly matches the untuned 27B model from the same family, improving substantially not only on KOPA-Bench but also on the BFCL benchmark.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)