基础模型的最新进展使人工智能科学家能够自动化日益完整的研究工作流程,,从假设生成和代码执行到手稿准备。然而,仅工作流程覆盖并不能提供科学发现所依赖的完整证据。现有系统通常对文本,代码,标签,或预先计算的摘要,进行推理,从而使代理无法获得科学决定性的空间,时间,跨渠道,和程序关系。我们介绍 OmniScientist, 是一位端到端, 全模式人工智能科学家,可直接根据异构原始证据进行多学科研究。感知层和 3 个用于构思, 实验, 和写作的自主代理在确定性管道, 内运行,允许观察在整个研究生命周期中形成研究问题, 实验决策, 和最终主张。通过在代码,中运行想法,严格,和声明检查,系统强制执行新颖性筛选,统计有效性,执行来源,和数字可追溯性。我们根据涵盖 5 个学科家族, 4 个科学证据家族, 的 36 个真实数据案例对 OmniScientist 进行评估,包括图像, 信号, 音频, 视频, 3-D 结构, 轨迹, 表格, 公式, 和图表。该系统完成了所有 36 个案例中从原始数据到编译稿件的完整路径,并通过参考推理主干获得了 6.3 的平均论文总分。在与仅接收预先计算的标量特征,的盲变体的配对比较中,直接感知改善了所有 7 个评估维度,并赢得了 85% 的面对面判断。这些结果表明,全生命周期的感知对于基于证据的科学发现至关重要,并为培养具有广泛能力的人工智能科学家提供了一条实用途径。

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)