自主人工智能代理可以保持完全授权,但随着行为漂移,对手适应,并且决策模式发生变化而变得不安全,而无需任何代码更改。我们提出\textbf{信息可行性原则}:,管理代理减少到估计未观察到的风险的界限$\hat{B}(x) = U(x) + SB(x) + RG(x)$,并且仅在其容量时允许操作$S(x)$ 超过 $\hat{B}(x)$ 安全裕度。 \textbf{Agent 生存能力框架}, 以 Aubin的 生存能力理论, 为基础,建立了三个属性 - 监控 (P1), 预期 (P2), 和单调限制 (P3) - 对于记录的故障模式来说,对于单独必要的和集体足够的属性。 \textbf{RiskGate} 使用专用统计估计器实例化框架 (KL 散度, 分段与休息 $z$-tests, 顺序模式匹配), 故障安全单调管道, 和闭环自动驾驶仪,形式化为 Aubin的 调节图的实例,以终止开关为最后手段; a标量生存指数 $VI(t) \in [-1,+1]$ 与一阶 $t^*$ 预测将治理从被动转变为预测。贡献包括理论框架,、参考实现, 和针对已发布的代理故障分类法的分析覆盖范围; 定量实证评估属于后续工作。
Autonomous AI agents can remain fully authorized and still become unsafe as behavior drifts, adversaries adapt, and decision patterns shift without any code change. We propose the \textbf{Informational Viability Principle}: governing an agent reduces to estimating a bound on unobserved risk $\hat{B}(x) = U(x) + SB(x) + RG(x)$ and allowing an action only when its capacity $S(x)$ exceeds $\hat{B}(x)$ by a safety margin. The \textbf{Agent Viability Framework}, grounded in Aubin的 viability theory, establishes three properties -- monitoring (P1), anticipation (P2), and monotonic restriction (P3) -- as individually necessary and collectively sufficient for documented failure modes. \textbf{RiskGate} instantiates the framework with dedicated statistical estimators (KL divergence, segment-vs-rest $z$-tests, sequential pattern matching), a fail-secure monotonic pipeline, and a closed-loop Autopilot formalised as an instance of Aubin的 regulation map with kill-switch-as-last-resort; a scalar Viability Index $VI(t) \in [-1,+1]$ with first-order $t^*$ prediction transforms governance from reactive to predictive. Contributions are the theoretical framework, the reference implementation, and analytical coverage against published agent-failure taxonomies; quantitative empirical evaluation is scoped as follow-up work.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)