Researchers aim to prevent identity theft from medical records

Join Our Community of Science Lovers!

This article was published in Scientific American’s former blog network and reflects the views of the author, not necessarily those of Scientific American


Over time, patients end up providing a wealth of information to their health care providers, and when all our data are aggregated, they are also a boon to researchers studying trends in diseases and demographics for clues in how to better treat illness. And nowadays, as more patient health care records go digital, patient information becomes more widely shared among researchers—which can be a good thing or a bad thing, depending upon who has access to it.

Electronic medical record (EMR) systems contain detailed, yet anonymous patient-level data represented in codes that correspond to different health conditions, including disease, symptom or injury. Lately, EMRs are increasingly being used to provide data for genome-wide association studies (GWAS) used to identify relationships among specific genomic variants and health-related phenomena, a key to delivering on the promise of personalized medicine. However, patient privacy can be threatened when personal information is linked to genetic information using codes that are available through public databases and electronic medical records, a team of Vanderbilt University researchers in Nashville conclude in a study published Monday in the Proceedings of the National Academy of Sciences.


On supporting science journalism

If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.


The researchers claim to have illustrated this problem as part of their research, where they identified 96 percent of a group of 2,762 patients with the help of the diagnosis codes in the patients' records.

A possible solution, according to Vanderbilt researchers Grigorios Loukides, Aris Gkoulalas-Divanis and Bradley Malin, is to use a method for creating anonymous records that replaces the current system—known as the International Statistical Classification of Diseases and Related Health Problems (ICD)—with a series of related codes. The researchers created an algorithm that generalizes clinical information so that patients remain anonymous, while providing the medical and genetic connections needed by researchers.

Loukides and his colleagues tested the algorithm's data protection performance against simulated malicious computer hacker attacks using actual information from more than 2,600 patients, assuming a potential hacker knew a patient's identity, some or all of a patient's ICD codes, and whether the patient record was included in released data. The technique foiled attempts to uncover a patient's private information, the researchers wrote, and maintained the data integrity necessary to retain useful information for validating genome-wide studies.

Image ©iStockphoto.com/ DNY59

Larry Greenemeier is the associate editor of technology for Scientific American, covering a variety of tech-related topics, including biotech, computers, military tech, nanotech and robots.

More by Larry Greenemeier

Subscribe to Support Independent Journalism

Great science journalism requires human expertise, time, effort and creativity. And it costs money. That’s why I and the journalists here at Scientific American hope you’ll join our community.

When you subscribe, you are supporting staff and freelance journalists who are passionate about telling science stories that are true, important and compelling. Our editors and reporters are often experts in their fields, which means they understand the nuances of big discoveries and can untangle the breakthroughs from the hype. With a subscription, you are also supporting rigorous fact-checking to ensure the words we publish are precise and accurate. And you’re supporting original illustrations, graphics and photos that bring you closer to an advanced laboratory, an ice sheet in Antarctica or a space mission in orbit. You’re helping us craft other types of high-quality journalism as well: Our newsletters are carefully written, edited and curated by staffers you have or will come to know and love. Our Science Quickly podcast is based on original reporting, collaboration with editors and scientists and exacting production.

Subscriptions keep this engine running so we can continue to deliver thoughtful, rigorous and independent science journalism to you. In an era of viral misinformation, this work is crucial. If you value what we do, I hope you’ll consider joining us as a subscriber

Thank you,

Jeanna Bryner, Editor in Chief, Scientific American

Subscribe