代理人工智能架构通过外部工具, 增强了法学硕士,释放了强大的功能,但可能会产生大量成本。此外,工具的使用并不总是有益的:冗余或低效调用甚至会损害任务性能。有效的工具使用, 因此, 取决于LLM 的核心决策: 在执行任务时是否调用工具。我们引入了一个受决策理论启发的原则框架,以根据三个关键因素:必要性,效用,和可承受性来理解工具使用决策。我们的分析结合了两个互补的视角:,一个是规范性视角,推断最佳工具调用的真实需求和效用,,另一个是描述性视角,根据观察到的行为推断模型'自我感知的需求和效用。我们跨原生和定制工具,、两个工具, 和六个任务评估了六个开放模型和一个专有的 OpenAI 模型。模型 感知的需求和效用仍然与其真实价值, 不一致,特别是在预算限制下。这种不一致会导致代价高昂的过度使用和性能下降。为了改进工具决策,,我们从模型隐藏状态训练需求(LNE) 的轻量级潜在估计器。 LNE 通常比模型自我报告更准确地预测真实需求,并改进跨模型规模和工具类型的预算工具分配。代码和数据集可在此 https URL 获取。
Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities but potentially incurring substantial costs. Moreover, tool use is not always beneficial: redundant or low-utility calls can even harm task performance. Effective tool use, therefore, hinges on a core LLM decision: whether to call or not call a tool when performing a task. We introduce a principled framework inspired by decision-making theory to understand tool-use decisions along three key factors: necessity, utility, and affordability. Our analysis combines two complementary lenses: a normative perspective that infers true need and utility for optimal tool calls, and a descriptive perspective that infers the model的 self-perceived need and utility from their observed behaviors. We evaluate six open models and a proprietary OpenAI model across native and customized harnesses, two tools, and six tasks. Models perceived need and utility remain misaligned with their true values, particularly under budget constraints. This misalignment produces both costly overuse and performance-degrading calls. To improve the tool decisions, we train lightweight latent estimators of need (LNEs) from model hidden states. LNEs generally predict true need more accurately than model self-reports and improve budgeted tool allocation across model scales and tool types. Code and dataset available at this https URL.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)