在通过云, 推动科学数据存储和协作之前,寻求健康研究突破的医学研究人员必须克服协作和关键临床数据收集的重大障碍。
Before the advancement of scientific data storage and collaboration via the cloud, medical investigators seeking health research breakthroughs had to overcome significant obstacles to collaboration and key clinical data gathering.
数据是孤立的并且难以分发,,因此那些想要进行研究的人别无选择,只能自己收集数据。这不仅使研究成本更高,,而且比较跨数据集的研究结果也具有挑战性。
Data were siloed and difficult to distribute, so those looking to undertake research had no option but to gather them themselves. This not only made research more expensive, but it was challenging to compare findings across datasets.
1975, 在麻省理工学院和波士顿研究心律失常的研究人员的 Beth Israel Hospital 设想了另一种方式: 该团队开始收集心电图记录并将其数字化,目的不仅是研究它们,,而且还使它们可供更广泛的研究界使用。
In 1975, researchers studying arrhythmias at MIT and Boston的 Beth Israel Hospital envisioned another way: the team began collecting and digitizing electrocardiogram recordings with the intention of not only studying them, but of also making them available to the wider research community.
该团队为这个过程建立了自己的计算机,,煞费苦心地一一复制磁带,,并为录音创建了超过100,000个注释。这个过程花了几年,,但到了 1980, 夏天,磁带终于准备好了。该团队最初认为他们的工具将覆盖不到十几个学术和行业团体。但兴趣不断涌入。在接下来的十年, 中,他们继续邮寄了大约 100 份。
The team built their own computers for the process, painstakingly duplicated tapes one by one, and created more than 100,000 annotations for the recordings. The process took years, but by summer 1980, the tapes were finally ready. The team initially thought their tool would reach fewer than a dozen academic and industry groups. But interest kept pouring in. Over the next decade, they went on to mail about 100 copies.
这些数据最终成为全球平台 PhysioNet — 的第一个数据库,该平台于 1999 年在哈佛-麻省理工学院健康科学与技术项目 — 中创建,作为复杂生理信号的临床数据存储库。
The data eventually became the first database of the global platform PhysioNet — founded in 1999 at the Harvard-MIT program in Health Sciences and Technology — as a clinical data repository for complex physiological signals.
当时, 这种类型的数据共享, 在今天看来是默认的,,这是一个近乎革命性的想法。 PhysioNet的 “ 创立非常有远见,” Thomas Heldt, Richard J. Cohen (1976) 医学和生物医学物理学教授, 麻省理工学院医学工程与科学研究所的副主任,的, 是最近在《自然健康》杂志上发表的一篇论文的高级作者,该论文研究了平台的影响。
At the time, that type of data-sharing, which may seem like the default today, was a near-revolutionary idea. PhysioNet的 “founding was incredibly visionary,” says Thomas Heldt, Richard J. Cohen (1976) Professor in Medicine and Biomedical Physics, associate director of MIT的 Institute for Medical Engineering and Science, and the senior author of a recent paper in Nature Health examining the platform的 impact.
最终,那些通过邮件发送的磁带变成了刻录的CD-ROM,,然后演变成托管在新创建的互联网上的FTP服务器。如今,, PhysioNet 回顾了超过 25 年的运营,,该平台托管数百个数据库,,并已成为现有最全面的生物医学和临床数据存储库之一。去年, 超过 15,000 份科学出版物引用了 PhysioNet,,来自 180 多个国家的用户已在该平台上注册。它被研究人员,制造商,和临床决策者广泛使用。
Eventually, those magnetic tapes sent through the mail became burned CD-ROMs, which then evolved into FTP servers hosted on the newly minted internet. Today, as PhysioNet looks back at over 25 years of operation, the platform hosts hundreds of databases, and has become one of the most comprehensive biomedical and clinical data repositories in existence. Last year, more than 15,000 scientific publications cited PhysioNet, and users from more than 180 countries have registered on the platform. It is widely used by researchers, manufacturers, and clinical decision-makers.
“看到这样的愿景被证明是正确的,并且为这么多人带来帮助,真是太棒了。”
“It is really beautiful to see that such a vision has proven right and so enabling for so many people.”
2009, 左右,一位名叫 Tom Pollard 的博士生正在伦敦的 领先的医院系统之一对重症患者进行研究。尽管医院产生了大量有价值的临床数据,,但策划和支持其更广泛的研究用途所需的基础设施和流程仍在开发中。
Around 2009, a PhD student named Tom Pollard was conducting research on critically ill patients at one of London的 leading hospital systems. Although the hospital generated large volumes of valuable clinical data, the infrastructure and processes needed to curate and support their wider research use were still developing.
问题不仅仅是隐私问题。医院信息系统的建立主要是为了支持患者护理和管理,而不是研究。数据在各个系统中分散,并且很少考虑到未来的重用,,这使得将它们转变为连贯的研究资源变得困难且昂贵。
The problem was not simply privacy. Hospital information systems were built primarily to support patient care and administration, not research. Data were fragmented across systems and rarely curated with future reuse in mind, making it difficult and expensive to turn them into coherent research resources.
但波拉德需要数据来完成他的论文。在互联网, 上闲逛后,他最终发现了重症监护医疗信息集市(MIMIC),,这是一个由 PhysioNet 托管的去识别化电子健康记录数据库。认识到其潜力,,他的临床主管, Kevin Fong, 组织了一次对波士顿的访问。不久之后, Fong 和 Pollard 坐在 Roger Mark, 对面,讨论他们的团队如何协作。
But Pollard needed data to complete his dissertation. After poking around on the internet, he eventually discovered the Medical Information Mart for Intensive Care (MIMIC), a database of de-identified electronic health records hosted by PhysioNet. Recognizing its potential, his clinical supervisor, Kevin Fong, organized a visit to Boston. Soon afterward, Fong and Pollard were sitting across the table from Roger Mark, discussing how their teams might collaborate.
学术激励长期以来一直青睐出版物和独家分析,而不是为他人准备数据以供使用的不太明显的工作。这种紧张局势今天依然存在。 PhysioNet的创始人接受了不同的模型,,他说,他们相信共享研究资源可以加速发现并最终改善人类健康,。 MIMIC 成为 Pollard 论文, 的核心,在完成博士学位, 后,他来到麻省理工学院帮助构建下一代数据库。
Academic incentives have long favored publications and exclusive analyses over the less-visible work involved in preparing data for others to use. That tension persists today. PhysioNet的 founders embraced a different model, believing that sharing research resources could accelerate discovery and ultimately improve human health, he says. MIMIC became central to Pollard的 dissertation, and after completing his PhD, he came to MIT to help build the next generation of the database.
自 PhysioNet 成立, 以来,共享研究数据的价值已获得更广泛的认可。已故罗杰·马克, MIT的 杰出健康科学与技术名誉教授、PhysioNet的 创始人之一, 描述其目的是围绕数据建立一个 “ 可访问的跨国社区” 到 “ 对全球健康产生积极影响。”
In the years since PhysioNet was established, the value of sharing research data has gained much wider recognition. The late Roger Mark, MIT的 distinguished professor of health sciences and technology emeritus and one of PhysioNet的 founders, described its purpose as building an “accessible multinational community around data” to “positively impact global health.”
今年早些时候, Mark 和已故的 George Moody, PhysioNet的 联合创始人, 因其对 PhysioNet 和生物医学信号处理的贡献而共同获得了享有盛誉的 IEEE 生物医学工程奖。 IEEE 引用了其 �%9 在心电图信号处理和全球传播精选生物医学和临床数据库方面的领先地位,,从而加速了全球生物医学研究。”
Earlier this year, Mark and the late George Moody, PhysioNet的 co-founder, jointly received the prestigious IEEE Biomedical Engineering Award for their contributions to PhysioNet and biomedical signal processing. IEEE cited their “leadership in ECG signal processing and global dissemination of curated biomedical and clinical databases, thereby accelerating biomedical research worldwide.”
平台, 的源代码及其大部分数据, 都是公开的。根据 Nature 文章: “A,随着平台的发展, PhysioNet的 社区的范围大大超出了信号处理和心血管健康的起源,涵盖了临床信息学, 重症监护和健康机器学习。” 人们已经使用它来构建自己的 PhysioNet 式基础设施, Heldt 说。 Pollard 指出 Health Data Nexus 等类似平台是 PhysioNet的 传统平台的示例。
The source code for the platform, like much of its data, is public. According to the Nature piece: “As the platform evolved, PhysioNet的 community broadened substantially beyond its origins in signal processing and cardiovascular health to encompass clinical informatics, critical care and machine learning for health.” People have used that to build their own PhysioNet-esque infrastructure, says Heldt. Pollard points to similar platforms like Health Data Nexus as examples of PhysioNet的 legacy.
尽管现在有更多资源托管类似的电子健康数据,,根据 Google DeepMind 研究员 Vivek Natarajan,,PhysioNet 和 MIMIC “ 设定了标准,”,他说, “ 并且它 现在仍然是标准。”
Although there are now more resources out there hosting similar electronic health data, according to Google DeepMind researcher Vivek Natarajan, both PhysioNet and MIMIC “set the standard,” he says, “and it的 still the standard right now.”
“PhysioNet 改变了我对研究瓶颈的看法。这通常不是想法或天赋。这是摩擦力。当数据访问速度缓慢, 昂贵, 且困难, 时,最先消亡的想法是高风险的, 那些可能行不通的事情, 但如果行得通,将会带来变革。他说,如果你想要真正的进步,”,这完全是错误的模型。 “PhysioNet 降低了尝试雄心勃勃的想法的固定成本,,这改变了科学的可能性。”
“PhysioNet changed how I think about the bottleneck in research. It is often not ideas or talent. It is friction. When access to data is slow, expensive, and hard, the ideas that die first are the high-risk ones, the things that probably will not work, but would be transformative if they did. That is exactly the wrong model if you want real progress,” he says. “PhysioNet lowers the fixed cost of trying ambitious ideas, and that changes what science becomes possible.”
PhysioNet, 曾经是一个主要面向生物医学信号处理和医疗保健领域工作人员的存储库, 已经发展了 25 年。最初, 所持有的数据仅包含心血管心电图数据。现在,PhysioNet 主要是电子健康记录, 成像数据, 以及软件和人工智能模型的来源。
PhysioNet, once a repository mainly for those working in biomedical signal processing and the health-care fields, has evolved in its 25 years. Originally, the holdings consisted solely of cardiovascular ECG data. Now PhysioNet is a largely a source for electronic health records, imaging data, and software and AI models.
特别是随着人工智能方法的起飞,“社区发生了转变,” Heldt 解释道。那些需要信号处理数据的人仍然使用 PhysioNet 数据库,,但用户群体已扩大到包括大型科技公司的员工,、教师, 和所有医学领域的从业者, 以及健康相关的机器学习和人工智能研究人员。 Heldt 表示,如今, 后一组“ 主导了用户社区,”。
Particularly as artificial intelligence approaches took off, “the community shifted,” Heldt explains. Those in need of signal processing data still use PhysioNet databases, but the pool of users has expanded to encompass staff at large tech companies, teachers, and practitioners in all areas of medicine, as well as researchers in health-related machine learning and AI. Today, that latter group “dominates the user community,” says Heldt.
Natarajan, 表示,该平台托管可用于医疗保健 AI 研究, 的最高质量数据集,Natarajan, 的研究涉及 AI, 科学, 和医学,并发表了几篇使用其数据集的论文。
The platform hosts the highest-quality datasets available for health-care AI research, according to Natarajan, whose research involves AI, science, and medicine and who has published several papers that used its datasets.
Natarajan 表示,“它是过去十年中促进医疗保健人工智能所有进步的重要基石,”。除了使用 PhysioNet 数据, 之外,他和他的同事还向平台, 贡献了数据,帮助创建代表 PhysioNet 的自我维持的生态系统。
“It has been an important cornerstone that has catalyzed all the progress in health-care AI over the last decade,” says Natarajan. In addition to using PhysioNet data, he and his colleagues have contributed data to the platform, helping create the self-sustaining ecosystem that typifies PhysioNet.
展望未来几十年,Heldt 和 Pollard 等, 平台管理者设想通过年度会议继续扩大其影响范围。该团队还准备试验一个新系统,该系统将允许用户注释数据并贡献自己的专业知识,,为平台的下一阶段丰富 PhysioNet的 资源。
Looking toward the coming decades, stewards of the platform like Heldt and Pollard envision continuing to expand its reach with an annual conference. The team is also preparing to pilot a new system that will allow users to annotate data and contribute their own expertise, enriching PhysioNet的 resources for the next phase of the platform.
“人们现在想做的研究需要跨学科的。 Pollard 表示,统计学家, 计算机科学家, 临床医生, 药剂师, 和护士必须齐心协力,贡献他们的知识来开发对人们有用的算法”。 “社区已经扩大了,,人工智能的进步扩大了研究人员可以解决的问题以及他们认为可能的问题。”
“The kind of research that people want to do now needs to be interdisciplinary. Statisticians, computer scientists, clinicians, pharmacists, and nurses must all come together and contribute their knowledge to develop algorithms that are useful for people” says Pollard. “The community has broadened, and advances in AI have expanded both the questions researchers can address and what they believe is possible.”