零射击对象目标导航 (ZS-OGN) 需要具体代理在没有任何事先训练的情况下探索和定位目标对象。为此,, 最近的方法利用了基础模型。但它们通常依赖于静态先验,缺乏适应性,,这会导致重复错误和代价高昂的试错。在本文,中,我们提出了一个自我进化的 ZS-OGN 框架,可以持续改进测试时间。具体来说,, 我们通过从过去的轨迹中提取可操作的知识来构建代理规则记忆。然后,我们提出了一种基于置信上限,的检索策略,通过平衡语义相关性和历史成功来选择有效规则。此外,我们引入了一个记忆引导的预想模块,可以在行动,之前预测潜在的结果,减少低效的探索。大量实验表明,我们的方法优于现有的零样本基线,,以更少的不必要步骤将成功率提高了 10.1\%。
Zero-Shot Object-Goal Navigation (ZS-OGN) requires embodied agents to explore and locate target objects without any prior training. To this end, recent methods leverage foundation models. But they typically rely on static priors and lack adaptation, which leads to repeated errors and costly trial and error. In this paper, we propose a self-evolving ZS-OGN framework that enables continuous test-time improvement. Specifically, we build an agentic rule memory by extracting actionable knowledge from past trajectories. Then, we propose a retrieval strategy based on upper confidence bound, selecting effective rules by balancing semantic relevance and historical success. In addition, we introduce a memory-guided preflection module that forecasts potential outcomes before action, reducing inefficient exploration. Extensive experiments show that our method outperforms existing zero-shot baselines, achieving a 10.1\% improvement in success rate with fewer unnecessary steps.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)