工具的使用将大型语言模型扩展到参数知识之外,,但可靠的执行需要平衡适当的推理深度和严格的结构有效性。我们从基于案例的角度解决这个问题,提出 CAST, 一个案例驱动的框架,将历史执行轨迹视为结构化案例。 , CAST 不是重复使用原始样本输出,而是提取源自案例的信号来识别复杂性概况,以估计最佳推理策略, 与故障概况一起映射可能的结构故障。该框架将这些知识转化为细粒度的奖励设计和自适应推理,,使模型能够在强化学习期间自主内化基于案例的策略。 BFCLv2 和 ToolBench 上的实验表明,CAST 提高了模式忠实执行和任务级工具使用的成功率,同时减少了不必要的考虑。该方法使整体执行准确度提高了 5.85 个百分点,并将平均推理长度减少了 26%,,从而显着减少了高影响力的结构错误。最终, 这展示了历史执行案例如何为校准工具的使用提供可重用的适应知识。

Tool use extends large language models beyond parametric knowledge, but reliable execution requires balancing appropriate reasoning depth with strict structural validity. We approach this problem from a case-based perspective to present CAST, a case-driven framework that treats historical execution trajectories as structured cases. Instead of reusing raw exemplar outputs, CAST extracts case-derived signals to identify complexity profiles for estimating optimal reasoning strategies, alongside failure profiles to map likely structural breakdowns. The framework translates this knowledge into a fine-grained reward design and adaptive reasoning, enabling the model to autonomously internalize case-based strategies during reinforcement learning. Experiments on BFCLv2 and ToolBench demonstrate that CAST improves both schema-faithful execution and task-level tool-use success while reducing unnecessary deliberation. The approach achieves up to 5.85 percentage points gain in overall execution accuracy and reduces average reasoning length by 26%, significantly mitigating high-impact structural errors. Ultimately, this demonstrates how historical execution cases can provide reusable adaptation knowledge for calibrated tool use.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)