Before the advancement of scientific data storage and collaboration via the cloud, medical investigators searching for health research breakthroughs had to beat significant obstacles to collaboration and key clinical data gathering.
Data were siloed and difficult to distribute, so those trying to undertake research had no option but to assemble them themselves. This not only made research dearer, but it surely was difficult to check findings across datasets.
In 1975, researchers studying arrhythmias at MIT and Boston’s Beth Israel Hospital envisioned one other way: the team began collecting and digitizing electrocardiogram recordings with the intention of not only studying them, but of also making them available to the broader research community.
The team built their very own computers for the method, painstakingly duplicated tapes one after the other, and created greater than 100,000 annotations for the recordings. The method took years, but by summer 1980, the tapes were finally ready. The team initially thought their tool would reach fewer than a dozen academic and industry groups. But interest kept pouring in. Over the subsequent decade, they went on to mail about 100 copies.
The information eventually became the primary database of the worldwide platform PhysioNet — founded in 1999 on the Harvard-MIT program in Health Sciences and Technology — as a clinical data repository for complex physiological signals.
On the time, that form of data-sharing, which can look like the default today, was a near-revolutionary idea. PhysioNet’s “founding was incredibly visionary,” says Thomas Heldt, Richard J. Cohen (1976) Professor in Medicine and Biomedical Physics, associate director of MIT’s Institute for Medical Engineering and Science, and the senior writer of a recent paper in Nature Health examining the platform’s impact.
Eventually, those magnetic tapes sent through the mail became burned CD-ROMs, which then evolved into FTP servers hosted on the newly minted web. Today, as PhysioNet looks back at over 25 years of operation, the platform hosts a whole lot of databases, and has grow to be one of the vital comprehensive biomedical and clinical data repositories in existence. Last yr, greater than 15,000 scientific publications cited PhysioNet, and users from greater than 180 countries have registered on the platform. It’s widely utilized by researchers, manufacturers, and clinical decision-makers.
“The research impact is really significant,” says Heldt, who can also be a professor within the MIT Department of Electrical Engineering and Computer Science and a principal investigator on the Research Laboratory of Electronics, “and quite humbling.”
“It is actually beautiful to see that such a vision has proven right and so enabling for therefore many individuals.”
Setting a typical
Around 2009, a PhD student named Tom Pollard was conducting research on critically ailing patients at considered one of London’s leading hospital systems. Although the hospital generated large volumes of helpful clinical data, the infrastructure and processes needed to curate and support their wider research use were still developing.
“Hospital data were collected primarily to support immediate patient care, with less attention given to how they could be curated and reused for research,” says Pollard, now a research scientist at MIT’s Laboratory for Computational Physiology (LCP), technical director of PhysioNet, and the lead writer on the Nature Health paper.
The issue was not simply privacy. Hospital information systems were built primarily to support patient care and administration, not research. Data were fragmented across systems and barely curated with future reuse in mind, making it difficult and expensive to show them into coherent research resources.
But Pollard needed data to finish his dissertation. After poking around on the web, he eventually discovered the Medical Information Mart for Intensive Care (MIMIC), a database of de-identified electronic health records hosted by PhysioNet. Recognizing its potential, his clinical supervisor, Kevin Fong, organized a visit to Boston. Soon afterward, Fong and Pollard were sitting across the table from Roger Mark, discussing how their teams might collaborate.
Academic incentives have long favored publications and exclusive analyses over the less-visible work involved in preparing data for others to make use of. That tension persists today. PhysioNet’s founders embraced a special model, believing that sharing research resources could speed up discovery and ultimately improve human health, he says. MIMIC became central to Pollard’s dissertation, and after completing his PhD, he got here to MIT to assist construct the subsequent generation of the database.
Within the years since PhysioNet was established, the worth of sharing research data has gained much wider recognition. The late Roger Mark, MIT’s distinguished professor of health sciences and technology emeritus and considered one of PhysioNet’s founders, described its purpose as constructing an “accessible multinational community around data” to “positively impact global health.”
Earlier this yr, Mark and the late George Moody, PhysioNet’s co-founder, jointly received the distinguished IEEE Biomedical Engineering Award for his or her contributions to PhysioNet and biomedical signal processing. IEEE cited their “leadership in ECG signal processing and global dissemination of curated biomedical and clinical databases, thereby accelerating biomedical research worldwide.”
The source code for the platform, like much of its data, is public. In accordance with the Nature piece: “Because the platform evolved, PhysioNet’s community broadened substantially beyond its origins in signal processing and cardiovascular health to encompass clinical informatics, critical care and machine learning for health.” People have used that to construct their very own PhysioNet-esque infrastructure, says Heldt. Pollard points to similar platforms like Health Data Nexus as examples of PhysioNet’s legacy.
Although there at the moment are more resources on the market hosting similar electronic health data, in keeping with Google DeepMind researcher Vivek Natarajan, each PhysioNet and MIMIC “set the usual,” he says, “and it’s still the usual at once.”
That standard, in keeping with those that use the platform, modified how research is conducted. Access to data shouldn’t be the determinant for which ideas are possible, in keeping with Ziad Obermeyer, an associate professor on the University of California at Berkeley School of Public Health and the College of Computing, Data Science, and Society.
“PhysioNet modified how I feel in regards to the bottleneck in research. It is usually not ideas or talent. It’s friction. When access to data is slow, expensive, and hard, the ideas that die first are the high-risk ones, the things that probably won’t work, but could be transformative in the event that they did. That is strictly the incorrect model for those who want real progress,” he says. “PhysioNet lowers the fixed cost of trying ambitious ideas, and that changes what science becomes possible.”
The AI boom
PhysioNet, once a repository mainly for those working in biomedical signal processing and the health-care fields, has evolved in its 25 years. Originally, the holdings consisted solely of cardiovascular ECG data. Now PhysioNet is a largely a source for electronic health records, imaging data, and software and AI models.
Particularly as artificial intelligence approaches took off, “the community shifted,” Heldt explains. Those in need of signal processing data still use PhysioNet databases, however the pool of users has expanded to encompass staff at large tech firms, teachers, and practitioners in all areas of drugs, in addition to researchers in health-related machine learning and AI. Today, that latter group “dominates the user community,” says Heldt.
The platform hosts the highest-quality datasets available for health-care AI research, in keeping with Natarajan, whose research involves AI, science, and medicine and who has published several papers that used its datasets.
“It has been a crucial cornerstone that has catalyzed all of the progress in health-care AI during the last decade,” says Natarajan. Along with using PhysioNet data, he and his colleagues have contributed data to the platform, helping create the self-sustaining ecosystem that typifies PhysioNet.
Looking toward the approaching many years, stewards of the platform like Heldt and Pollard envision continuing to expand its reach with an annual conference. The team can also be preparing to pilot a brand new system that can allow users to annotate data and contribute their very own expertise, enriching PhysioNet’s resources for the subsequent phase of the platform.
“The form of research that folks wish to do now must be interdisciplinary. Statisticians, computer scientists, clinicians, pharmacists, and nurses must all come together and contribute their knowledge to develop algorithms which can be useful for people” says Pollard. “The community has broadened, and advances in AI have expanded each the questions researchers can address and what they imagine is feasible.”

