{"id":3553,"date":"2023-01-03T21:55:45","date_gmt":"2023-01-03T21:55:45","guid":{"rendered":"https:\/\/medschool.prd.vanderbilt.edu\/vanderbilt-medicine\/?p=3553"},"modified":"2025-03-18T16:47:58","modified_gmt":"2025-03-18T16:47:58","slug":"the-expert-from-nowhere","status":"publish","type":"post","link":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/the-expert-from-nowhere\/","title":{"rendered":"The  Expert  from Nowhere"},"content":{"rendered":"<figure id=\"attachment_3554\" aria-describedby=\"caption-attachment-3554\" style=\"width: 600px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-3554\" src=\"https:\/\/cdn.vanderbilt.edu\/t2-main\/medschool-prd\/wp-content\/uploads\/sites\/82\/2023\/01\/feature5.jpg\" alt=\"\" width=\"600\" height=\"400\" srcset=\"https:\/\/cdn.vanderbilt.edu\/t2-main\/medschool-prd\/wp-content\/uploads\/sites\/82\/2023\/01\/feature5.jpg 600w, https:\/\/cdn.vanderbilt.edu\/t2-main\/medschool-prd\/wp-content\/uploads\/sites\/82\/2023\/01\/feature5-300x200.jpg 300w\" sizes=\"auto, (max-width: 600px) 100vw, 600px\" \/><figcaption id=\"caption-attachment-3554\" class=\"wp-caption-text\">Illustration by Matt Carlson<\/figcaption><\/figure>\n<p>To understand a protein\u2019s structure is to understand its function, says structural and chemical biologist Jens Meiler, PhD, distinguished research professor of Chemistry.<\/p>\n<p>It can take a PhD student up to five sleep-deprived years to determine the structure of a single protein, and of the 20,000 human proteins, only about 17% are considered to have had their structure determined experimentally with very high accuracy.<\/p>\n<p>Meanwhile, AlphaFold 2, a deep learning program owned by Google\u2019s parent company, can in minutes compute a protein\u2019s structure with an accuracy competitive with experiment.<\/p>\n<p>Meiler, who for decades has used machine learning (ML) to predict protein structure, notes that an excess of biomolecular data amassed over two decades by experimentalists was used to train AlphaFold 2, and he acknowledges that the resulting performance is indeed impressive.<\/p>\n<p>\u201cBut the real hard problems are problems of limited data,\u201d said Meiler, who in addition to his Vanderbilt faculty appointment holds an Alexander von Humboldt Professorship at Leipzig University in Germany. Meiler is collaborating with other Vanderbilt researchers to advance precision medicine, where treatment is to be tailored more than ever to individual differences among patients, including down to the molecular level.<\/p>\n<p>And he\u2019s using machine learning to do it.<\/p>\n<p><strong>The difference between machine learning and AI<\/strong><\/p>\n<p>Among scientists and engineers interviewed for this story, research interests range from basic science, where experiment and observation are comparatively unfettered, to drug discovery, which straddles basic science and clinical science, to clinical phenotyping and prediction, where health information policy is apt to impinge upon scientific observation. A phenotype, resulting from the interaction of genetics and the environment, is an observable characteristic of an organism. Clinical phenotyping and prediction characterize patients and research subjects in terms of their health \u2014 an area of science where stigma and privacy notably come into play.<\/p>\n<p>ML is a component of artificial intelligence (AI), says computer scientist Bradley Malin, PhD, the Accenture Professor and professor of Biomedical Informatics, whose research broadly explores how to make data accessible for biomedical research.<\/p>\n<p>\u201cIf you knew the way the world worked, you wouldn\u2019t need to perform machine learning, because all the rules of how the universe works would be defined,\u201d said Malin, whose team also publishes ML demonstration projects in the clinical data domain. \u201cMachine learning says, \u2018I don\u2019t know all of the rules, I don\u2019t know all of their relationships. So, let me try to learn a model for how the world works, and I will use data to drive that learning process.\u2019 Artificial intelligence is this bigger, all-encompassing perspective of how to create and represent intelligent environments.\u201d<\/p>\n<p>ML comprises a zoo of learning models, that is, different sorts of computer programs that can identify meaningful patterns in previously unseen data. Linear regression, logistic regression, support vector machines, decision trees and artificial neural networks figure among the major classes. (ML models are variously derived through supervised, semi-supervised or unsupervised learning \u2014 see the sidebar on page 25.)<\/p>\n<p>Among biomedical researchers, a surge of new interest in ML dates to around 2010, when breakthroughs brought a decades-old neural network model called deep learning into large-scale industrial application for speech recognition. And their interest continued to grow with the advent of transformers, a deep learning model introduced by a team at Google Brain in 2017.<\/p>\n<p>John McLean, PhD, Stevenson Professor of Chemistry and chair of the department, says early on he patterned his use of ML for analytical chemistry on neural network models developed for internet commerce. In McLean\u2019s research lab there are 10 mass spectrometry platforms, and as he says, mass spectrometry is incredibly fast. The spectrometers that he and his team help to conceptualize, design and build are primarily used to characterize biological samples \u2014 tissue, blood, urine. Given a sample having 10,000 to 80,000 molecules, spectrometers in this lab can sort the flurry of molecules and register their relative abundance in a few seconds.<\/p>\n<p>Imagine a Microsoft Word document 833 million pages long written in 60 minutes: that\u2019s how quickly raw biomolecular data can build up in McLean\u2019s lab.<\/p>\n<p>\u201cWithout machine learning, there\u2019s no hope of interrogating that data,\u201d McLean said. \u201cWe\u2019ve used these tools for the better part of a decade now, where we let the data inform us what is important about itself. Being able to reduce complex data \u2014 and I hate to put it this way, but \u2014 to infographics, where somebody can actually act on that information really quickly, is kind of our whole reason for using things like AI and machine learning.\u201d<\/p>\n<p>He mentions his work with so-called organ-on-a-chip technology, where cells harvested from donor cadavers are used to simulate organs in miniature. With ML standing by, exposing a simulated liver to a drug is like ringing a bell, he said. \u201cOn a molecular basis, we would analyze everything that the liver would secrete, like listening to how the bell rang \u2014 some molecules going up, some going down, some having other relationships.\u201d<\/p>\n<p>The lab\u2019s neural net programs wind up producing a type of graph called a heat map, showing clusters of molecules behaving similarly to one another. \u201cIn about 24 hours of experiments, we could recapitulate the known literature around how the liver would respond to a drug like acetaminophen. And not only that, but we can see all the molecules that behave just like the ones that were known in the literature but had never been discovered before.<\/p>\n<p>\u201cAnd we could not do that without machine learning.\u201d<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Wet-and-dry lab work<\/strong><\/p>\n<p>Monoclonal antibodies are biologic drugs used to treat diseases such as cancer, autoimmune disorders and infections. In his research lab, computer scientist and computational biologist Ivelin Georgiev, PhD, explores questions such as how antibodies recognize pathogens and cancers, and based on his findings he uses ML to design vaccines and antibody therapeutics.<\/p>\n<p>\u201cImmunology, virology, microbiology, single cell biology \u2014 all of these are areas that have been around for a while, and we still don\u2019t understand a lot of fundamental rules and interactions in these different processes,\u201d said Georgiev, associate professor of Pathology, Microbiology and Immunology. \u201cThe goal, at least in my mind, with what we do is to be able to really get into the personalized medicine area, where it\u2019s not just going to be sufficient to generate vaccines or antibodies for general use, but where you actually tailor those drugs to each particular person.\u201d<\/p>\n<p>Jens Meiler agrees that ML is poised to speed not only drug discovery in general, but also the precision drug discovery that Georgiev talks about. Meiler has already begun to help clinical teams exploit therapeutic opportunities newly discernable at the molecular level in individual patients. Working with scientists and clinicians at Vanderbilt- Ingram Cancer Center, he\u2019s helping to pioneer ML-assisted precision cancer therapy: Based on DNA from cancer tissue, Meiler uses ML to predict the structure of mutant proteins, which has helped clinical teams determine which drugs to prescribe.<\/p>\n<p>\u201cWe\u2019ve begun doing that on a regular basis for drugs that are already approved, and 10 years from now we\u2019re going to be using AI to engineer the best possible molecule for your specific mutation,\u201d Meiler said.<\/p>\n<p>He stressed that AI\u2019s role in precision medicine is to be only one component of a larger undertaking. \u201cIt will not work unless you put it in a pipeline with experts who will use the artificial intelligence outputs as one portion of judgment, subject to double checking and rigorous evaluation. And Vanderbilt is probably one of the three best places in the United States to do this kind of research, having some of the very best researchers on individual proteins targeted by drugs \u2014 transporters, ion channels, signaling proteins.\u201d<\/p>\n<p>Like Meiler, other basic scientists appear to value ML for its potential interplay with experimentation.<\/p>\n<p>McLean said, \u201cMachine learning\u2026is beautiful for showing you correlations, but it is not good for showing causation, or showing you why did something happen.\u201d<\/p>\n<p>When viewed by their advocates as stand-ins for theories about how the world works, learning models are often termed \u201clearned\u201d models. When derived from highly complex data, as is the case with neural networks and decision trees, the correlations that drive learning models tend to be obscured. Though they may work gangbusters and far outstrip human capacity, as models of the world they tend to be inscrutable.<\/p>\n<p>Speaking recently at Oxford University in England, Demise Hassabis, PhD, the founder and CEO of DeepMind, the company behind AlphaFold 2, suggested that learned models will increasingly come to define biology. \u201cA lot of these emergent and complex phenomena are just too complicated to be described with a few equations. I don\u2019t really see how you can, say, come up with Kepler\u2019s laws of motion \u2026 of a cell,\u201d Hassabis said.<\/p>\n<p>Meanwhile, computer scientists are said to be making headway in addressing the so-called explainability problem that attends complex ML.<\/p>\n<p>\u201cIt\u2019s completely conceivable that in the near future, those equations that for us are formidable and probably intractable will not be so anymore,\u201d Georgiev said.<\/p>\n<p>Andr\u00e9 Bastos, PhD, assistant professor of Psychology, uses ML to study cognition, down to the level of individual neurons and neuronal networks.<\/p>\n<p>\u201cCurrent application of machine learning is perhaps going to give us the ability to categorize different types of cognition, but it won\u2019t give us mechanistic insight,\u201d Bastos said in a recent Vanderbilt webinar on AI and biomedical research. \u201cI think promising work is going to combine the so-called wet lab approach and the dry lab approach.\u201d<\/p>\n<p>That seems to be the emerging paradigm in drug discovery: Wedded with experimental validation, ML has come to assist the hunt for molecular drug targets, provide target-based virtual screening of potentially therapeutic molecules (in their billions), and closely aid in drug design, safety prediction, and safety surveillance for drugs already at market.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>On a precipice<\/strong><\/p>\n<p>ML has proven devilishly good at well-defined vision tasks, and for that reason was thought by one prominent computer scientist to have put at least some physicians out of work by now. \u201cI think that if you work as a radiologist you are like Wile E. Coyote in the cartoon. You\u2019re already over the edge of the cliff, but you haven\u2019t yet looked down,\u201d deep learning pioneer Geoffrey Hinton, PhD, told The New Yorker in 2017. He gave radiologists five, perhaps 10 years to reach their obsolescence. He also recounted a talk he had delivered at a medical school where he suggested that they should stop training radiologists.<\/p>\n<p>Five years later, while some radiologists may have begun using commercial AI to assist their work, none has apparently lost a job to a computer.<\/p>\n<p>\u201cComputers replacing doctors, or learned models replacing textbooks and theory, I don\u2019t see that happening,\u201d said informaticist and internal medicine specialist Colin Walsh, MD, MA, associate professor of Biomedical Informatics. \u201cFor one thing, we simply never are going to have all the necessary data quantified. As much as we\u2019d like to fantasize devices sensorizing everything in our world, that\u2019s not going to be the reality. Probably ever.\u201d<\/p>\n<p>Meanwhile, recent medical literature has seen a welter of projects (including some contributed by Walsh) showing the power of complex ML for clinical phenotyping\/prediction. Walsh attributes this research output to new availability of data \u2014 which in large degree is thanks to widespread adoption of electronic health records, or EHRs \u2014 and to ML tools having become more democratized. \u201cAnyone reading this article can on their laptop or home computer download software to run machine learning on a data set they may have sitting around on their hard drive,\u201d Walsh said.<\/p>\n<p>The EHR, it should be said, will often be characterized by informaticists, statisticians and others as a billing document that\u2019s only rather poorly disguised as a health record, and researchers have apparently learned to appreciate its vagaries and limitations and to hold it suspect as a source of data for drawing general inferences about human health.<\/p>\n<p>Nevertheless, the EHR from an ML perspective first of all features a lot of easy pickings, so-called structured data, meaning information that lends itself to tabular representation and thus to computation \u2014 lab results, billing codes representing diagnoses and procedures, drug prescriptions.<\/p>\n<p>Other data that could be helpful for clinical phenotyping\/prediction is buried in EHR notes written by the clinical team and increasingly by patients themselves.<\/p>\n<p>The transformer, that aforementioned deep learning model, burst into biomedical research in 2018, and one area where it shines is natural language processing, or NLP.<\/p>\n<p>\u201cWith the use of machine learning methods, we can now pretty accurately capture all sorts of information that figures in EHR notes,\u201d said computer scientist Cosmin Bejan, PhD, assistant professor of Biomedical Informatics. Bejan has used NLP of EHR text for tasks such as finding firefighters, homelessness, suicidal behavior and nonprescription drug use. Some of the research reports he has co-authored have established that certain drug exposures correlate with reduced COVID severity; that, compared to other firefighters, those exposed to toxins in the World Trade Center attack developed high rates of mutations associated with blood cancer and cardiovascular disease; that homelessness is poorly documented in the EHR, and that suicide attempt and suicidal ideation are likewise undercoded.<\/p>\n<p>Clinical phenotyping\/predictive learning models are tested on historical data, with an overwhelming majority never having made it to clinical testing for decision support or health care outcomes improvement. Bejan\u2019s colleague, computer scientist You Chen, PhD, assistant professor of Biomedical Informatics, has used ML and longitudinal EHR data to predict, among other things, preterm birth and the timing of hospital discharge. \u201cI\u2019m confident that machine learning will help health care in the near future, but we have a lot of challenges that need to be addressed before that can happen,\u201d said Chen, whose research extends to finding solutions to support the transfer of EHR-based predictive models between different health care institutions.<\/p>\n<p>Whenever talk turns to the application of ML in health care, there tends to be no lack of attention to (or trepidation around) what has come to be known in some circles as the fairness problem.<\/p>\n<p>\u201cAt least in the clinical domain, you have structural inequities and misconceptions that will permeate society for years to come,\u201d said Malin. When ML is trained for phenotyping\/prediction on routine health care data, applying those learning models for clinical decision support or outcomes improvement will carry some risk of perpetuating misconceptions that haunt society and health care delivery. \u201cProblems with data in the health care domain are not so much data problems as problems with how society is currently structured,\u201d Malin said.<\/p>\n<p>Bringing AI into health care will in part mean engaging with the public over issues of AI fairness. Two large new National Institutes of Health research projects are grappling with fairness from various directions, and Malin is helping to lead parts of both.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>From ML to AI<\/strong><\/p>\n<p>Two researchers interviewed for this story, Walsh and Michael\u00a0Matheny, MD, MS, MPH, bear the hard-earned distinction of having led development of EHR-based ML predictive models that have reached the clinical testing phase for decision support and patient outcomes improvement.<\/p>\n<p>Matheny, an internal medicine specialist and observational data scientist, built a model to predict risk of acute kidney injury (AKI) following cardiac catheterization. He used logistic regression and some 20 patient features in EHR data, collected from all 76 health care centers of the Veterans Health Administration. Results are pending from a large multisite randomized controlled trial at the VA, where periodic ML-powered reports were issued to cardiac catheterization teams, showing their observed-to-expected AKI rates.<\/p>\n<p>\u201cIf you\u2019re actually going to impact clinical care, you have to get out of data science and get into the actual art and practice of health care and medicine,\u201d said Matheny, professor of Biomedical Informatics. \u201cIn recent years, there has been a real awareness and acknowledgment that the hardest part of deploying AI and ML into health care isn\u2019t the AI and ML, it\u2019s their integration into the workflow and the care and the profit of the institution.\u201d<\/p>\n<p>Starting with 1,537 EHR features, Walsh used a decision tree-based learning model to predict suicidal behavior among adults at VUMC. The model can provide immediate decision support to clinical teams at the start of any return patient encounter. In an observational study, among patients figuring in the top 10% for suicide risk in universal face-to-face screening in the adult emergency room, while one in 200 went on to attempt suicide within 30 days, adding the outputs from Walsh\u2019s model to the face-to-face risk evaluation tripled that rate to three in 200. Results are now pending from a pragmatic randomized controlled trial using the model to prompt face-to-face suicide risk screening in three outpatient clinics at VUMC.<\/p>\n<p>\u201cIf we\u2019re going to throw a bunch of new predictions at clinicians, we better prove to them that those predictions are actually helpful,\u201d Walsh said. \u201cThat\u2019s the most important step. And it\u2019s often not done, because it\u2019s easier to just turn it on and say, \u2018Go, I\u2019m helping you, and I hope that it\u2019s better.\u2019\u201d<\/p>\n<p>For truly massive observational studies of health and health care, for pragmatic clinical trials, for clinical decision support and systematic outcomes improvement, the EHR, with all its known flaws, would appear to be the only game in town. Computing power and the understanding of health continue to advance \u2014 and the capacity of the EHR to equitably support improvement of health care outcomes might likewise advance, particularly with the addition of genotypes, Twitter profiles and whatever other new sorts of patient data might be in the offing.<\/p>\n<p>Meanwhile, AI\u2019s apparent failure to launch in the health care domain might better be interpreted as warranted caution.<\/p>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>To understand a protein\u2019s structure is to understand its function, says structural and chemical biologist Jens Meiler, PhD, distinguished research professor of Chemistry. It can take a PhD student up to five sleep-deprived years to determine the structure of a single protein, and of the 20,000 human proteins, only about 17% are considered to have&#8230;<\/p>\n","protected":false},"author":232,"featured_media":3554,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"jetpack_post_was_ever_published":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"_links_to":"","_links_to_target":""},"categories":[45,14,15],"tags":[],"class_list":["post-3553","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-vm-fall-2022","category-vm-features","category-vm-homepage-highlights"],"acf":[],"jetpack_publicize_connections":[],"jetpack_featured_media_url":"https:\/\/cdn.vanderbilt.edu\/t2-main\/medschool-prd\/wp-content\/uploads\/sites\/82\/2023\/01\/feature5.jpg","jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/pcDnub-Vj","_links":{"self":[{"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/posts\/3553","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/users\/232"}],"replies":[{"embeddable":true,"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/comments?post=3553"}],"version-history":[{"count":3,"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/posts\/3553\/revisions"}],"predecessor-version":[{"id":4186,"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/posts\/3553\/revisions\/4186"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/media\/3554"}],"wp:attachment":[{"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/media?parent=3553"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/categories?post=3553"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/medschool.vanderbilt.edu\/vanderbilt-medicine\/wp-json\/wp\/v2\/tags?post=3553"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}