科学思想很少是从白纸开始的。他们继承机制,修复已知的局限性,并重组早期工作,的片段,就像生物基因组一样。目前的基准测试仍然很少谈到人工智能系统是否可以遵循这种继承结构。我们推出 IdeaGene-Bench(IG-Bench), 作为科学谱系推理和基于谱系的想法生成的基准。 IG-Bench 围绕 IdeaGene 框架: 组织,每篇论文或提案都表示为一组最小, 类型, 循证的 Idea Genome 对象,,GenomeDiff 对齐这些对象以记录继承, 突变, 损失, 外部导入, 和六种操作进化动态下的新颖插入。该基准包含 1,961 黄金谱系痕迹, 1,085 策划的 Idea Genome 对象, 和 10 个科学领域的 920 条成对 GenomeDiff 记录。它支持两种评估。 IG-Exam (42 任务类型, 1,029 个实例) 测试跨思想基因组抽象, 继承追踪, 进化推理, 和谱系验证的封闭式谱系推理。 IG-Arena 使用谱系条件群体进化分数(PES), 评估一代,询问提案是否可以作为给定谱系群体的连贯后代插入: 它应该继承正确的想法基因组对象, 与附近的工作有有意义的变化, 并为未来的研究提供选择价值。对 14 名法学硕士科学家进行的实验暴露了成分瓶颈。最强的系统在谱系推理上仅达到 27.3% 准确率,,结构化谱系上下文重新洗牌了系统排名,而不是统一帮助每个参与者。

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)