Proceedings of a Workshop—in Brief
Convened October 28, 2025
As the use of artificial intelligence (AI) rapidly expands, the potential uses for these tools in genomics research and its translation into clinical care continue to take shape. AI is emerging to support researchers in genomic discovery during data processing, sequence quality assessment, and variant identification and interpretation (Aradhya et al., 2023). AI also holds promise for helping clinicians make more precise treatment plans, identify complex genetic diagnoses, and prevent health risks when used in a transparent and reproducible way (Walton et al., 2024). Considering these rapid developments, on October 28, 2025, a planning committee under the auspices of the National Academies of Sciences, Engineering, and Medicine's Roundtable on Genomics and Precision Health convened a workshop to explore current and potential future applications for AI in genomics and precision health, from translational research to clinical applications.1 Describing how AI would be considered during the workshop, Grant Wood, a former chief executive officer of the Global Genomic Medicine Collaborative, said that AI "enables the integration and interpretation of genomic, clinical, and environmental datasets to identify disease risks, optimize treatments,andadvance the quality of health care far beyond levels to date. For academia and research, health care delivery, industry, and policy, AI represents a technological frontier and a collaborative challenge ensuring that its deployment is validated, trustworthy, [and] widely available."2
The goals of this workshop, as outlined by Wood, were to explore how AI has been implemented in genomics and precision health settings and to discuss ways in which AI may be applied in the future, including for multi-modal diagnostics and translational genomics research, while considering the benefits and challenges related to data harmonization and security, workforce, and usefulness. The workshop was also intended to consider how the accuracy of, and bias inherent to, AI technologies are evaluated and their potential impacts on AI applications in genomics-related research and clinical care. The goal was also to examine lessons learned from other fields that may be transferable to genomics across the translational research process. Wood encouraged participants to explore how they see themselves in the future of AI during this transformative time in genomics, where AI has the potential to address previously unexamined challenges.
1 See materials from the workshop here: https://www.nationalacademies.org/projects/HMD-HSP-24-21/past-events (accessed January 28, 2026).
2 Wood noted that the AI definition [excerpted above] was adapted after querying Microsoft Copilot.
"The future is working with biology, because biology sets the rules," said Jim Weinstein, a senior vice president at Microsoft Health. AI should be viewed as "actionable information," reflecting the principle that technology should serve people, not the other way around. To provide a broad overview of AI, Weinstein highlighted six key advances in general AI capability areas:
One challenge for AI in clinical practice is managing temporal relationships, specifically, adapting to evolving clinical science and changing patient conditions over time, Weinstein said. However, he continued, an iterative, multimodal, whole-patient AI approach is now "enabling reasoning across patients, populations, and biology." Emerging AI tools are also enabling basic and applied researchers to interact with molecules, cells, and proteins using natural language queries. These interactions support applications such as cell type annotation, protein discovery, and drug design. New tools such as the Microsoft AI Diagnostic Orchestrator (MAI-DxO) can help improve diagnostic accuracy. In a benchmarking study using 304 complex medical cases, the AI agent made the correct diagnosis up to 85 percent of the time, versus 20 percent accuracy by general physicians.3
AI is being applied to the diagnosis of rare disorders. Weinstein said that less than 40 percent of patients with a suspected rare disease receive a diagnosis and that the average time to diagnosis is 5 years. Microsoft is developing Evidence Aggregator, an AI tool that extracts data on genetic variants from the biomedical literature to accelerate the diagnosis of rare diseases.
Several ethical and equity considerations for applying AI in genomics and clinical care should be considered, Weinstein said, including protecting the privacy of sensitive genomic data; ensuring equitable access to AI-driven health care tools and technologies; building public trust in AI-enhanced health care through transparency, explainability, and community engagement; and developing ethical data practices for AI implementation. AI-based strategies and tools can help promote ethical and equitable health care and ensure that patients are not excluded based on geography. As examples of such tools, he mentioned digital twins4 and a nationwide hub-and-spoke model of digital interconnectivity incorporating federated learning and multimodal AI (Weinstein and Allen, 2025).
Weinstein offered his perspective on key questions about the status and potential of AI in the current health care ecosystem. He expressed optimism that AI can "improve outcomes and well-being across society" and be integrated into care processes and the workforce to enhance efficiency and effectiveness. However, he voiced less confidence about ensuring AI safety and preventing harm in health care settings. He also noted that the extent to which AI can be trusted to be equitable and fair depends on the specific population being served. Microsoft's perspective is that patients own their data. Weinstein emphasized that meaningful progress requires not simply substituting one technology for another but fundamentally transforming the health ecosystem to incorporate patients' health data, environmental factors, socioeconomic information, and other contextual elements into their care.
3 For the details of this study, see Nori, H., M. Daswani, C. Kelly, S. Lundberg, M. Tulio Ribeiro, M. Wilson, X. Liu, V. Sounderajah, J. Carlson, M. P. Lungren, B. Gross, P. Hames, M. Suleyman, D. King, and E. Horvitz. 2025. Sequential diagnosis with language models. arXiv [Preprint] https://doi.org/10.48550/arXiv.2506.22405 (accessed December 12, 2025).
4 For more on digital twins, see National Academies of Sciences, Engineering, and Medicine. 2024. Foundational Research Gaps and Future Directions for Digital Twins. Washington, DC: The National Academies Press, Chapter 2: https://www.nationalacademies.org/read/26894/chapter/4 (accessed February 17, 2026).
It is estimated that 400 million people worldwide have a rare disease and that many patients go undiagnosed, said Melissa Haendel, the director of precision health and translational informatics and the Sarah Graham Kenan Distinguished Professor at the University of North Carolina Chapel Hill. For over four decades the number of known rare diseases has been reported to be around 7,000. Haendel said that current efforts to understand the heterogeneity across and within rare diseases place that number at more than 10,000 (Haendel et al., 2020). One way to address the variation in how organizations define a particular rare disease and the lack of standard clinical terminology used in many of those definitions, Haendel said, is the Monarch Initiative's Mondo Disease Ontology, which is "a global initiative to reconcile the world's rare disease definitions," incorporating genetic and phenotypic characteristics and other information. Monarch is working with ClinGen, patient groups, electronic health record (EHR) companies, and others to incorporate these harmonized definitions at the point of care. She said that of the 10,000 rare diseases, around 1,600 rare diseases were found in only one of five commonly used knowledge databases and that only 333 were found in all five sources, emphasizing the importance of the data sources selected for use in diagnostic applications.
In early 2025, the rare disease community came together to discuss overcoming barriers to the diagnosis of rare disease, and Haendel highlighted some of the critical bottlenecks identified across the patient diagnostic odyssey, saying that AI might be useful in addressing them.5 The prevalence of a given rare disease varies across health care systems and databases, which she said is due in part to the rarity of the diseases, but is also due to "differences in coding and billing biases." Haendel added that many AI models are trained on data extracted from EHRs, which oftentimes have undergone lossy compression (i.e., data size was reduced in a way that some original information was lost, so the data quality can be compromised). For example, she showed how "mesocardia" entered by the patient's clinician is transformed into "congenital heart disease" when using a common indirect mapping approach, losing the original meaning of the patient–clinician interaction and opportunities for AI to operate on the right data.6
As AI-powered precision medicine takes hold in some clinical areas (e.g., cancer, newborn screening), Haendel said, there are other areas that fall into the "chasm of approximate medicine" (e.g., areas where AI and precision medicine are not yet fully applied such as in rare diseases, chronic complex conditions, mental health conditions, and autoimmune diseases). To realize the promise of precision medicine, she said, patients must be considered in the context of the wide range of variables that influence health (environmental, social, and behavioral factors). There are also "data modalities that are relevant to health that are increasingly high-throughput and must be analyzed in the context of their temporal flow," Haendel said. However, she noted that most current AI models are built largely on "data that happens to be available," which are primarily structured data and notes extracted from the EHR and not other information that could be valuable, such as data from patients' experience seeking or encountering care, other linkable data (e.g. dental, genomics, prescriptions, wearables, vision), and population data. Haendel suggested that AI could be used for precision medicine and genomics across the diagnostic journey, including patient identification, diagnosis and triage, management and care coordination, and research.
The implementation of AI in precision medicine "trails far behind the technology," Haendel said, and the challenges to bringing technological advances to the point of care "are more social and policy than they are technology." Data are needed at the point of care to provide equitable access to AI tools for health, she said. This will require modernization of "data governance, stewardship, and policies…AI alone cannot fix discordant or missing data or knowledge," she said, and a socio-technical engineering approach is needed to make sure AI-based tools are "grounded in reality." There are also privacy and regulatory considerations for the integration of different types of patient data in AI models.
Haendel and Weinstein discussed the importance of a multidisciplinary workforce for achieving the potential of AI in health care. Haendel suggested that allied health professionals need training on the essentials of using AI-based health care
5 Meeting white paper available at https://zenodo.org/records/14906644. M. Haendel, J. McMurry, and A. Sizer. 2025. Critical bottlenecks in rare disease research and care: A community perspective. Zenodo.
6 For more information about direct versus indirect mapping approaches for rare diseases, see https://www.ohdsi.org/2025showcase-141/ (accessed January 8, 2026).
tools, which could be accomplished by working with data scientists and AI experts as partners in providing AI-based health care. Weinstein said there are workforce issues across all disciplines, and he highlighted the projected shortages of both physicians and nonphysician health care providers over the coming decade. Some medical students are very interested in new technology and the applications of AI, while others, including some faculty, can feel threatened by technology that could diminish their role. No one person has all the foundational knowledge in each discipline that AI can bring, Weinstein said. To build trust in the use of AI models, Haendel said, there needs to be an understanding of everything that happens at the point of care and to know—and understand—the evidence that led to the conclusion.
Hematoxylin and eosin (H&E)–stained pathology slides can provide diagnostic information about a patient's cancer, such as its cell and tissue structure. Subsequent molecular testing can then be conducted to identify predictive and prognostic genomic biomarkers (e.g., driver mutations, gene expression signatures, pathway activation status), which can inform treatment decisions. Molecular oncology testing is more costly than H&E staining ($1,000–5,000 vs. approximately $100 for H&E), slower (2–4 weeks vs. 2–3 days), and often limited by tissue availability, all of which can delay treatment decisions, said Laura Acqualagna, the director of AI and machine learning (ML) engineering at research and development at GSK.
"What if H&E already contains the answers?" Acqualagna asked. Current research suggests it is possible to apply AI tools to predict molecular signatures directly from H&E slides, she said. Such methods could be used as a cost-effective and timely screening approach to identify patients for confirmatory molecular testing. One approach uses multiple-instance learning to predict genetic signatures from H&E whole-slide images (e.g., see Allen et al., 2025). As a use case, Acqualagna described the application of this approach to predict microsatellite stability/instability (MSS/MSI) signature in tumors from patients with metastatic colorectal cancer. This could be used as a screening tool to identify patients with these mutations who might be responsive to immunotherapy, she said (Li et al., 2025). Another approach that bridges genomics and histology is the prediction of spatial transcriptomics from H&E slides.7 When both histology and molecular profiling data are available for a patient, they can be used synergistically for multimodal learning, along with other patient data, to better characterize the patient and the disease, make better predictions, and enhance clinical decision making.
There are challenges in using AI and multimodal learning for translational medicine, and Acqualagna highlighted three main areas. Data integration requires "harmonizing diverse data sources (e.g., genomics, imaging, EHRs) and ensuring interoperability," she said, and this is hindered by a lack of standardization and by datasets that are incomplete or biased. High-quality datasets are also needed for validation of AI/ML models, Acqualagna said. Ensuring reproducibility and reducing the bias of models can be challenging, particularly when some populations are underrepresented in datasets. Challenges for the translation of AI models into clinical settings include regulatory considerations, clinician education and training, addressing privacy risks, ensuring transparency when AI is used for clinical decision making, and educating patients and earning their trust. Acqualagna said that meeting these challenges requires broad collaboration across regulatory agencies, industry, academia, researchers, and clinicians.
More than 99 percent of genetic variants are of unknown clinical significance based on the interpretation of clinically relevant variants in databases such as ClinVar, said Kyle Farh, a vice president and distinguished scientist in the AI laboratory at Illumina. He spoke about how Illumina is developing new deep learning methods "to decipher the effects of variants in the human genome" and to use this information to improve drug target discovery.
One challenge for building AI to study variants of unknown significance (VUS) in humans is developing a training database of genetic sequences. To solve this problem, a database of 1,000 genomes was created by sequencing 233 primate species from around the world to train the PrimateAI-3D deep neural network. From an evolutionary perspective, Farh explained, variants that are tolerated in
7 See DeepSpot: Leveraging Spatial Context for Enhanced Spatial Transcriptomics Prediction from H&E Images, https://www.medrxiv.org/content/10.1101/2025.02.09.25321567v3.full-text (accessed February 17, 2025).
these primate species are unlikely to be pathogenic. Because the protein coding sequences of these primate species are about 99.6 percent identical to humans, it can be inferred that these VUS are likely also benign in humans, he said. As a result, 4.5 million human VUS are now classified as not causing severe disease in humans. Farh showed examples of how the benign primate VUS were validated against known pathogenic variants in the ClinVar database of human variants. PrimateAI-3D is trained only on the primate genome database and does not require any human genetic training data, Farh said. The algorithm identifies potentially pathogenic human variants "based on the absence of benign common primate variants in that region" (Fiziev et al., 2023).
PrimateAI-3D variant interpretation supports therapeutic target discovery research, and Farh showed how its prediction scores correlate with clinical biomarkers. For example, using data from individuals in the UK Biobank, he showed how a predicted variant score of 0 (benign) for LDLR was associated with normal blood levels of LDL cholesterol, and as the variant score increased to 1 (pathogenic), patient LDL levels were correspondingly higher. Patients with a variant score of 0 for PCSK9 had normal levels of LDL, while those with a score of 1 had lower cholesterol levels, indicating that PCSK9 is protective when mutated. Farh said the identification by PrimateAI-3D of "genes which when naturally mutated create a protective effect" allows pharmaceutical researchers to reproduce the effect therapeutically (e.g., with small molecule drugs or monoclonal antibodies). One could "apply this to any common disease and then identify genes which increase or decrease your risk of that disease," he said. However, the ability to identify novel drug targets this way is dependent on both cohort size and the quality of variant interpretation. PrimateAI-3D is also being applied to improve polygenic risk score (PRS) prediction for all populations because "rare variant PRS generalizes well to non-European populations."
Real-world evidence includes "unstructured and semi-structured patient data from disparate real-world sources and settings" such as EHRs, claims data, patient registries, patient reported data, and published literature, said Mark Kiel, the chief science officer and cofounder of Genomenon. Challenges in working with real-world data include the completeness, quality, and complexity of the data; the cost to access and time to aggregate the data; and lack of certainty that the analysis will provide useful information. Despite these obstacles, real-world data offer promise for more inclusive findings (broader populations and wider scope of diseases, including rare diseases), better patient outcomes, and accelerated health care innovation, he said. Ideal sources for real-world data will offer both breadth (in number and types of patients) and depth (i.e., granularity), Kiel said. Such data can be used to determine disease prevalence, identify biomarkers, and develop trial inclusion criteria and clinical endpoints.
Kiel emphasized the value of clinical literature as a source of real-world data. The published literature, which spans many decades of biomedical research, includes case reports, disease pedigrees, case series, and patient cohorts which contain patient demographic and phenotypic data, laboratory results, and information about treatment and outcomes. These real-world data can be aggregated into real-world evidence (RWE) variant landscapes which Kiel described as a comprehensive database of genetic variants classified according to clinical standards and presented with disease tags, functional insights, and citations.
"Genomic intelligence is provided for clinical diagnostics and precision therapeutic development" through advanced computation and human curation of the clinical literature at Genomenon, Kiel said. The Genomenon Genomic Graph (G3) process involves indexing the corpus of clinical genomic literature using advances in natural language processing to recognize terms within a set of 15 entities (e.g., variants, diseases, phenotypes, therapies) and then creating a semantic knowledge graph of their relationships using large language models (LLMs). The variant data are then curated by a team of around 100 genetic experts. Computation can provide scale and sensitivity, while human curation is needed to mitigate bias and ensure "specificity, accuracy, and applicability to the question at hand" Kiel said. LLM tools can then be applied to generate custom databases of RWE that can be interrogated to produce actionable "real-world insights" for clinical and pharmaceutical purposes.
There is a wide range of applications of real-world evidence for the pharmaceutical development of precision therapeutics. Kiel mentioned examples of various products, such as the calculation of disease prevalence and development of inclusion
criteria for clinical trials; the expansion of trial inclusion criteria based on identification of additional variants from the literature, enabling broader patient enrollment; the identification of clinical and functional biomarkers for both clinical and regulatory applications; and patient stratification to identify subpopulations.
Kiel said that one potential application of AI would be improving the interpretation of VUS in newborn screening by sequencing. Pilot studies of newborn sequencing are underway, looking at diagnosing treatable genetic diseases at birth for long-term monitoring or early intervention. He noted that the intent is to complement current newborn screening, not to replace it, and that sequencing requires appropriate informed consent from the patient and family.
The field is moving rapidly, with new models and generating novel insights from large scale data for biomarker discovery and patient stratification, Acqualagna said, while the challenge now is how to translate those discoveries into clinical practice to inform decision making. AI models must be trained on data from the populations that they are intended to serve. Kiel said that with the expansion of more newborn sequencing pilots, it is likely that more VUS will be identified, which will be a big challenge for genetics. This presents technical, predictive, and infrastructure challenges that need to be addressed with coordination to aggregate data to predict outcomes related to the variants, he said. Some of the variants that unlock knowledge about disease are often rare in the population, Farh said, so enormous sample sizes are needed for statistical power. It may be necessary to sequence hundreds of millions of people to fully realize the therapeutic insights from genetics, he said.
The total volume of data is increasing exponentially, and one-third of that data is generated by health care, said James Chen, the senior vice president for medical informatics at Tempus AI and an associate professor of medical oncology and bioinformatics at The Ohio State University. These data are often siloed, not connected, unstructured, and unorganized. Clinical practice guidelines are updated with increasing frequency, often multiple times each year in oncology, for example, to keep pace with the approval of new precision therapeutics. "It's no longer about how much you know," Chen said. "It's how fast you can look up the information."
Health care is "drowning in data and starving for insights," Chen continued, and AI can be used to "quiet the noise." Tempus AI is both a provider of genomic sequencing services and a developer of AI-based solutions to make everyday tasks easier in precision medicine. Data harmonization and the development of an information architecture are essential to using these volumes of data for AI algorithms effectively, he said.
The core of AI is machine learning. Chen noted that this is not new technology and that neural networks have been in use for decades. The evolution of AI has led to generative AI algorithms, including LLMs, which can create new content from existing data. The next level of AI is interactive AI, which is facilitated by foundation models trained on large datasets. A challenge for developing foundation models is assembling data for the model. One approach to addressing this challenge that Chen highlighted is the GenomeX initiative to develop a standard "language" to facilitate the exchange of clinical genomic data. Unidimensional foundation models are insufficient, he said, and Tempus is taking an approach that incorporates AI, multimodal data, and unified tooling to integrate workflows. Tempus also has partnerships with 4,500 institutions for the sharing of multimodal patient data through secure data exchange information architecture.
Together, these elements support foundation models, and Chen described several of their AI-enabled solutions. Tempus Link facilitates clinical trial matching and accelerates enrollment. For example, Chen said that the TriHealth system in Ohio achieved a 64 percent annual increase in trial enrollment and "a ninefold increase in efficiency of conversion from screened patients to enrolled subjects" by using Link. Tempus Next helps to close care gaps through guidance-aligned clinical decision support. The use of Next resulted in a 42 percent reduction in the time to clinical follow-up for patients seen for cardiac concerns, and it reduced disparities in time to follow-up cardiac care between black and white patients, he said.
Chen said that "to trust AI you need IA [information architecture]." Information architecture includes "expertise
in cleaning and harmonizing multimodal data; deploying AI solutions in clinical care workflows; plug-and-play connection with minimal support needed from clients; [and] transferring organized data back to partner systems." He emphasized the importance of bi-directional exchange of data and providing value back to data partners.
Nephi Walton, an associate professor at Wake Forest University, spoke about three patient encounters to illustrate how clinicians face new challenges as more patients use AI to research their health concerns. After seeing several specialists, Patient 1, a software engineer in his 20s, used an LLM to evaluate his health data and concluded that he needed a genetic test for an endocrine disorder. His self-diagnosis was confirmed after his endocrinologist acceded to his requests for sequencing. Patient 1s use of an LLM is what navigated him to clinical genetics. Patient 2, a mother with no background in medicine or computer technology, used an LLM to evaluate her child's complex clinical findings. Based on the results, she asked the pediatrician to order a single-gene test, which confirmed the diagnosis she had arrived at using the LLM. Patient 3 was "a business professional on a 10-year diagnostic odyssey for her child," Walton said. Using an LLM, she concluded that her child needed to see a geneticist. When she was finally referred to Walton, she told him she would have been fine with the specialists using an LLM during her child's appointment and would have preferred it because she would have felt more confident that they were able to access the best information.
Clinicians need to be able to use the same LLM technologies that patients are using, Walton said. Otherwise, "patients are going to lose confidence in us, . . . we are going to miss diagnoses, and we will become obsolete." He added that the use of LLMs is especially applicable to the diagnosis of very rare genetic disorders as clinicians cannot know all 10,000 rare genetic disorders, but LLMs can.
The accuracy of the genomic information delivered by LLMs is continuously improving, Walton said, and as accuracy improves, LLMs become more structured and less conversational. However, LLMs are trained on existing knowledge, and a challenge is how to update a model's training and potentially "tell it this new information is better than that old information," Walton said. Furthermore, LLM responses are influenced by data frequency, and information that is more frequent due to it being controversial or simply longer standing can end up being prioritized over more accurate information (McGrath et al., 2024). He added that LLM responses can be very convincing, sometimes causing clinicians to second-guess their own knowledge. To be used clinically, LLMs must be evaluated to ensure they can be trusted. Walton suggested that LLMs need "experience in the field with oversight," not unlike the residency a physician does. Expectations will need to be established, for example, concerning whether the performance target for an LLM should be no errors versus outperforming humans, he said.
Walton highlighted several emerging applications of AI to clinical genomics. A generative AI trained to replicate an individual's personality showed 85 percent accuracy in answering survey questions compared with the actual person, which Walton said is comparable to the accuracy with which individuals themselves provided the same answers to the same survey 2 weeks later.8 This capability could be applied to calculating polygenic risk scores that could also consider the individual's predicted behaviors that might influence their risk for disease as well as their response to an intervention. This would allow for targeted interventions taking likely behavior into account. Another developing area is the application of the transformer architectures that underlie LLMs to genetic data. Such models, when incorporated with phenotype data, have the potential to elucidate the complex patterns associated with human disease and help researchers understand disease's genetic origins. This would shed light on why different patients receiving the same treatment can experience different results. Such models would require 100 million patients for training but would have the potential to revolutionize medicine, Walton said. Powerful models may be built with fewer patients by using creative approaches to reduce data dimensionality.
If the potential of AI for precision medicine at scale is to be realized, AI must be available to primary care providers and incorporated into the workflow at the point of care, Walton said. "[I]t needs to be streamlined and simple [and] has to fit into a 10- to 15-minute encounter." While medical schools are working to incorporate AI into student curricula, more
8 For the details of this study, see Joon Sung Park, Carolyn Q. Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, Michael S. Bernstein. 2024. Generative Agent Simulations of 1,000 People. arXiv [Preprint] https://doi.org/10.48550/arXiv.2411.10109 (accessed February 17, 2026).
attention is needed on upskilling the current workforce to use AI-based tools, he said.
As the volume of genomic data has increased, guideline-based variant interpretation has become increasingly more complex, resource intensive, and time consuming, said Mullai Murugan, the director of software engineering at Baylor College of Medicine. She described using large reasoning models (LRMs) to accelerate variant interpretation by geneticists. LRMs, she said, are "an advanced form of LLM specifically fine-tuned for complex reasoning and problem solving."
The study she discussed was designed using LRMs to automate evidence extraction in accordance with the American College of Medical Genetics and Genomics (ACMG) and the Association for Molecular Pathology (AMP) PS4 criterion for variant classification.9 As explained by Murugan, this ACMG criterion requires genetic reviewers to "extract and count PS4 eligible cases from the literature" and determine if the prevalence of the variant is significantly increased in patients expressing the disease phenotype versus controls. Establishing whether a variant meets PS4 criterion has thus far required expert human review. The process is complicated by differences and ambiguities in how the variant and associated phenotypes are referred to and described in the publications.
Murugan described a study to benchmark LRMs against expert human curation for the ability to detect a variant in publications (task 1) and "identify PS4-eligible cases using ACMG and ClinGen [Variant Curation Expert Panels] guidelines"10 (task 2). The study was benchmarked against "expert-curated ground truth" which included three large frontier-scale models (OpenAI GPT-5, OpenAI o3, and Google Gemini 2.5 Pro) and two smaller efficiency-oriented models (Anthropic Claude Sonnet 4, and OpenAI o4-mini).
The ground truth benchmark dataset (based on Murugan et al., 2024) "consisted of 281 publications . . . cover[ing] 58 genes and 128 variants," Murugan said. The LRM pipeline was designed as three components. Part 1 generated the publications knowledge base, part 2 summarized the ACMG and ClinGen PS4-relevant guidelines, and part 3 was a directed acyclic graph (DAG)–based "PS4 case counting and evidence summarization." The output of each model was then benchmarked against the ground truth dataset. For task 1, Murugan reported, all five models achieved greater than 90 percent accuracy in detecting variants in the literature. She added that the large models performed better than the small models, which were thwarted by the use of non-standard nomenclature to describe the variant, or when the variant was only presented in tables, for example. For task 2, the large models achieved 85 to 90 percent concordance with the ground truth dataset in counting PS4-eligible cases. The small models did not perform as well, although Murugan noted that they were faster and cheaper than the large models. Errors occurred across all models, Murugan said. Examples of such errors included not finding the variant in the publication and overcounting or undercounting PS4 cases. Profiling model error patterns, she said, can inform the design of "model-aware prompting, reproducibility safeguards, and human-in-the-loop guardrails."
Based on the high level of concordance with the expert-curated ground truth, "these large reasoning models offer plenty of potential," Murugan said. The large models produce verbose reasoning, explaining their decision-making processes in detail, which she noted appealed to the geneticists. Work is ongoing to refine model design and expand the scope to other ACMG criteria. Murugan added that "these reasoning models can be generalized and applied to reasoning over unstructured biomedical data," such as EHR data.
Panelists discussed the potential for AI to enhance clinical genomics practice. Chen said that AI will reshape how clinicians engage with patients, enabling the clinicians to assess the patient's data in real time in the clinic and allowing them to "focus more on the art of medicine." As more genomic data are included in the medical record, there will be a revolution in the understanding of disease physiology, Walton
9 The PS4 criterion refers to the ACMG/AMP standards and guidelines for classifying pathogenic genetic variants. PS4 is considered to have "strong" evidence of pathogenicity where the "prevalence of the variant in the affected individual is significantly increased compared with the prevalence in controls." See https://www.acmg.net/docs/standards_guidelines_for_the_interpretation_of_sequence_variants.pdf (accessed January 9, 2026).
10 To learn more about the Criteria Specification Registry used by ClinGen, see https://cspec.genome.network/cspec/ui/svi/ (accessed January 20, 2026).
said. He suggested that AI agents could be inserted across the clinical workflow, providing concise information to support the management of some genetic concerns in the primary care setting. Murugan said that "knowledge-bridging AI assistants" are simple tools that could be implemented now to help physicians understand and explain genetic testing results to their patients.
Chen added that AI agents could help clinicians prioritize which patient concerns to focus on in the brief office visit and could help identify social determinants of health relevant to the patient's condition and care which could then be addressed. Walton agreed and emphasized the need to deliver care to where the patients are. He suggested AI solutions will soon be embedded in EHRs and said that there is a market opportunity for EHR vendors to provide these solutions in the systems used at both large academic centers and rural clinics. Chen said that for some AI applications, different delivery models are being developed for use in an institutional EHR (e.g., Epic) or in a community practice EHR system (e.g., OncoEMR) and that there is a web-based model for sites with legacy medical records systems.
Chen reiterated the potential of AI-enabled, guidance-aligned clinical decision support as a tool for closing care gaps (e.g., quickly identifying patients who meet criteria for genomic testing). A challenge, Walton noted, is that guidelines can be implemented differently by different health systems and that there can be conflicting guidelines for the same disorder from different organizations. Another challenge is incorporating new information into LLM tools and ensuring the LLM is applying that information, he said. Chen said it is important to know the source of the data for an LLM, the limitations of those data, and what guardrails are in place. He noted that some companies, such as his, are building their own "cordoned off" datasets for foundation models. Walton and Murugan also emphasized the importance of and challenges of incorporating "the right information" or high quality, accurate, and unbiased datasets in the models and agents.
Will Greene, a rare disease researcher, patient advocate, and board member of the Foundation for Prader-Willi Research (FPWR), discussed AI in genomics from his perspective as the parent of a child living with Prader-Willi syndrome (PWS). When his son, Ari, was born in 2021 "he had many of the signs and symptoms of a genetic disease," including "low birthweight, inability to cry or drink, [and] hypotonia." A month later Ari was diagnosed with PWS. Symptoms of PWS include hyperphagia (often leading to morbid obesity), anxiety, intellectual disability, developmental delays, and other medical problems, Greene said.
After numerous visits to specialists, several surgeries, and getting Ari settled on a treatment plan including growth hormones, Greene and his wife became involved with the work of the FPWR. He said this was before the use of public-facing AI tools became widespread, and his research and outreach activities were laborious and "done by hand." However, he now uses AI nearly every day in both his personal and professional life.
Greene said he first tried AI tools to help him understand some of the highly technical content in grant proposals he was reviewing for FPWR and added that he found it incredibly helpful. He soon discovered, however, that FPWR policies limit the use of AI for grant reviews, including restrictions on uploading of grant applications' confidential information to AI systems. Greene said that this might be due in part to concerns about reviewers submitting reviews done primarily or entirely by AI and revealed he was aware of one such case. The privacy of the information in the proposals and the potential for models to leak intellectual property is another possible concern. Greene said he understands these concerns but feels the risk is low.
Today, Greene uses AI to manage his son's care, saving him significant time and effort. He uploads his notes from clinical interactions and all his son's data into the models in search of insights. For example, he is using AI to help manage growth hormone dosing, as there are multiple formulas that can be used, and also to analyze his son's hormone level history when his levels are out of range.
AI also plays a large supporting role in Greene's research and advocacy work. For example, he often starts drafting articles and fundraising campaign materials by having a conversation with a chatbot about ideas and possibly getting a first draft
that he then refines (as he did for a recent magazine article on this topic11).
Greene said, "patients and caregivers are using AI for everything and are unlikely to stop." And while data privacy is important, "quick answers matter more" when managing a rare disease. Still, AI tools are not perfect, and there is still a need for human intelligence in the loop.
Robert Freimuth, an associate professor of biomedical informatics in the department of AI and informatics at Mayo Clinic, said he envisions genomic medicine as a "virtuous cycle" that enables discoveries, accelerates translation, provides the latest knowledge to clinicians and patients, and facilitates continually updated personalized care across care settings. He added that patients should also be able to share relevant information with family members. Freimuth outlined three areas for achieving this vision. The first is consumers of genomic information (patients, clinicians, researchers, computers), each with different needs and preferences for viewing data. The next is the exchange of data across systems, each with different capabilities, management, and upgrade schedules. The third area is interpretation and impact over time as new data are collected, knowledge advances, and clinical guidelines evolve.
At the core of this vision is the seamless exchange of genomic data, from its generation to annotation, interpretation, reporting, storage, and use. This requires that the data be interoperable, Freimuth said. One challenge for achieving data interoperability is that ambiguous semantics persist in clinical genomic reports, in part because of the need to simplify the communication of complex genomic information. He showed examples of a wide range of ambiguous genomic nomenclature found in reports and asked, if records containing ambiguous notations were fed into an LLM, could the output be trusted?
Tesler's Law states that "for any system there is a certain amount of complexity which cannot be reduced," Freimuth said. Different data systems handle the complexity of genomic information in different ways, which creates incompatibilities when attempting to exchange data. Mapping data across systems can be lossy, he said, and can result in degraded data. Standards play an important role, but the specifications alone will not address challenges with interoperability. Freimuth cited Grahame Grieve who said, "The technical challenges . . . are miniscule compared to the problem of getting the various people to agree with each other."12
For interoperability, Freimuth suggested developing a knowledge graph that would be a "graph of graphs," incorporating intersecting biomedical knowledge, patient data, and clinical guidance knowledge graphs. This "ecosystem of standards that are harmonized on the touchpoints" could potentially facilitate more successful flow of data among systems, he said. For example, he mentioned work done to harmonize the touchpoints between two disparate genomic data standards (HL7 FHIR and GA4GH) to establish a level of interoperability. This exercise was labor and time intensive, he said, but still more work is needed before AI could potentially take over the task. "Robust data standards are necessary but not sufficient for interoperability," Freimuth said. It is also necessary for people to "agree on semantics, to design and implement systems, and then to use those systems" as intended. What AI can be used for, he said, is quality control of the data. AI can suggest corrections and help to "improve quality, computability, and semantic expressivity," features which he said are essential for translating data into knowledge and action.
Freimuth offered three actions needed for implementing AI in genomics and precision medicine at scale: at the local level, "invest in data as an asset and an enabling resource;" at the community level, "prioritize robust semantics;" and at a global level, "actively work towards harmonization."
The clinical toolbox is expanding to include such AI tools as patient-facing apps, clinical decision support tools, and ambient AI scribes, said Tina Hernandez-Boussard, the associate dean of research and a professor of biomedical informatics at Stanford University. The data ecosystem is also expanding, including clinical history and patient characteristics, as well as omics, social determinants of
11 See https://rarerevolutionmagazine.com/three-ways-ai-is-changing-paediatric-genomic-medicine/ (accessed January 9, 2026).
12 See https://www.healthintersections.com.au/2011/04/09/law-1-interoperability-its-all-about-the-people.html (accessed January 9, 2026).
health, environmental measures, mechanistic models, and novel data streams. Digital twins are on the horizon to help clinicians build models that represent individual patients. The promise of AI tools is in bridging siloed data and incorporating real-world evidence, to "turn lived experience into actionable insights," Hernandez-Boussard said. The peril, she cautioned, "is blurred boundaries between technical innovation and ethical responsibility."
One challenge is how to use the available volumes of data to benefit the individual patient, and Hernandez-Boussard discussed her work on a prototype individual digital twin model (Hernandez-Boussard et al., 2021; Sadée et al., 2025). This continuous learning model uses multimodal individual patient data to develop retrospective trajectories and then infer predictive trajectories to inform and accelerate care decisions (e.g., simulated patient response to different treatment choices).
AI tools are currently regulated by the U.S. Food and Drug Administration (FDA) as a device or a drug, Hernandez-Boussard said, and the output of a device or a drug does not change. However the output of an LLM is continually changing, and she said the current regulatory framework "affect[s] the tools we receive and how we can use them." Another issue to be addressed is the ownership and governance of personal health information and of technologies built using personal data. She also noted that AI chatbots are being developed and implemented as wellness tools and are therefore not subject to FDA regulation. For example, Hernandez-Boussard said chatbots for mental health are promising as they are accessible anywhere and at any time, offer privacy, and reduce the stigma of mental health care. However, the general lack of regulatory oversight is a particular concern. She emphasized that an AI chatbot is not "thinking" but instead is an algorithm running probabilistic statistics to provide the next best response, and she mentioned a recent case of a teen suicide related to chatbot interactions.
"Human-in-the-loop architectures" are critical if these AI tools are going to influence clinical decisions, Hernandez-Boussard said. AI tools must be "built, tested, and interpreted with domain experts," including "clinicians and patients who define utility [and] technologists who ensure safety and transparency," she said. Training data curation is also essential because models learn patterns and incentives from data, and that can weaken guardrails when real-world inputs resemble the behavior that is trying to be prevented. Transparency and explainability concerning how LLM tools for clinical use were trained and evaluated and what safeguards are embedded are essential for trust, she concluded.
"Preventive health begins with insights," said Shivani Nazareth, the vice president of digital health strategy at Myriad Genetics and a board-certified genetic counselor. While working as a genetic counselor, she said, she felt that by the time patients got to her "they were making decisions under duress." This led her to move to industry to become involved in projects that could expand access to genetic testing and potentially "push medicine into preventative care."
Nazareth has long promoted a preventative health approach of sequencing all individuals at birth and then "unlocking information at different phases of life." Newborns are already screened at birth for several treatable genetic conditions. Having complete sequencing data could provide valuable insights across one's lifespan, for example, during adolescence for sports screenings for genetic cardiac conditions, during reproductive years for carrier status, and throughout adulthood to implement personalized cancer risk management or medication optimization. Nazareth acknowledged that there are challenges to this life phase approach, including establishing coverage and reimbursement and developing a framework for accessing an individual's data at key points in their lifespan in such a way that it does not impact their access to or coverage of care (i.e., if the data indicate a predisposition for a condition). Another issue is defining when it is appropriate to share information from the sequencing of children with other family members. Nazareth said there are current examples of genetic information "flowing backwards," where a heritable mutation is identified in a child, which leads to sequencing and identification of the same mutation in a parent, who can then take potentially lifesaving preventative action. Taking advantage of the volume of personal genotypic and phenotypic information for health decision making will also be a challenge. Nazareth mentioned Hippocratic AI, which is developing patient-facing AI that helps navigate patients to care and performs non-diagnostic tasks. She suggested
that this type of tool could provide anticipatory guidance to individuals with an identified genetic predisposition for a health condition.
Nazareth said that she is conflicted about knowing her own genetic status. After her mother began to exhibit worrisome behavioral changes, Nazareth requested the neurologist test for a mutation of the C9orf72 gene, which is associated with frontotemporal dementia and amyotrophic lateral sclerosis (ALS). Her mother's neurologist was dismissive and did not want to order the test, so Nazareth acquired and paid for the test kit and brought it to the next appointment, along with the test requisition and signed consent. When the mutation was confirmed, Nazareth said it was good to have this information to inform her mother's care. Professionally, Nazareth said, she knows she should be tested for this mutation. Personally, she said, she feels "paralyzed" and undecided about whether she wants to know her own status, adding, "There is not too much I can do with this information." She is using AI and the OpenEvidence platform to learn about her options and said she and her sister have enrolled in two clinical trials. Nazareth said this experience has made her rethink her "assumptions about what people will do with information." "Human behavior is complicated and messy," and not always predictable, she said. She is also thinking more about what genetic information should be shared with patients. Discussions of preventive health often assume that having the information is better, she said, but perhaps informed choice should be considered, where patients decide which information they want and when.
AI tools are changing how patients interact with clinicians and their expectations. Greene highlighted the value of AI in ensuring that communications are clear and mistake-free. He said he uses AI to check draft emails and suggested that clinicians should do the same since he receives many unclear communications. He also said that when a clinician does not know the answer, it is reasonable to assume that the patient (or the patient's advocate) has done extensive research on his or her condition and that the clinician should defer to the patient's suggestions. Nazareth agreed that clinicians need to be able to say they don't know but will do the research. She emphasized that the dismissiveness she experienced with her mother's doctor is unnecessary, especially now that patients arrive in the office very well informed.
Freimuth asked whether patients would be more comfortable if the clinician left the room and then consulted with an LLM or if the clinician used an LLM in the room together with the patient. Nazareth felt it would be fine for the clinician to consult the LLM in her presence. Greene was less sure how he would feel but said it would be best if the LLM was designed for and embedded in the clinical workflow, lest patients think the clinician is simply searching the internet or asking ChatGPT or otherwise using tools that are widely available to patients and their caregivers. Hernandez-Boussard said there's a need for a phased approach to implementation of AI-based tools in the clinical workflow to ensure end-user buy-in and to avoid increasing clinician burnout. Early testing involves asking several clinicians whether the model has answered their questions correctly and "when, where, and how [they] want to see this information," Hernandez-Boussard said. The model is then tested more widely in the clinic, followed by broad dissemination and surveillance.
Panelists also discussed the role of AI in democratizing patient-centered care. Nazareth suggested that transcription of family history by ambient AI tools could provide cleaner histories for pattern recognition by genetic counselors. In addition, AI can translate patient content, such as sequencing results, into other languages. Hernandez-Boussard agreed that AI provides opportunities to connect with more communities, but realizing these benefits will require attention not only to technical capabilities, but also to data representativeness, transparency, and patient and clinician skills to use AI tools safely and effectively. Greene also raised the need to enhance digital literacy to ensure that users are aware of the limitations, challenges, and risks of AI systems.
Leo Anthony Celi, a senior research scientist at the Massachusetts Institute of Technology and a medical intensive care unit doctor at Beth Israel Deaconess Medical Center, said there is interest in developing foundation models for single-cell biology, such as the scGPT model for single-cell multi-omics which was released in 2024. A more recent example is Cell2Sentence-Scale 27B, a foundation model built on Gemma that understands scRNA-seq profiles as "cell sentences" and generated a novel hypothesis about cancer cell behavior which sheds insights on new therapeutic
pathways. This approach using LLMs does not require custom architectures, he explained, and "interprets cellular data in natural language, making it accessible to biologists and clinicians, not just computational experts." Celi also mentioned the scFLUENT-seq model which can "capture nascent RNA as it's being transcribed." Using this single-cell approach it was discovered that "each cell is using only a tiny fraction of its genomic potential at any moment," which provides a very different picture from "bulk RNA-seq where over 80 percent of the genome appears active." Other recent models he mentioned included Tahoe-x1, "a 3 billion parameter open-source, single-cell foundation model," and Boltz-Gen, "a generative model for designing proteins and peptides." An issue for these types of omics foundation models is that "the bias of omics data translates into the bias of omics GPT," he said.
"Bias pervades every level of omics research," Celi said. Furthermore, data collection biases, technical and instrumental biases, and computational and analytical biases "reinforce each other, compounding risks of misinterpretation across the entire omics pipeline." Celi emphasized the need for a multilevel, integrated bias mitigation strategy that includes "standardization of instrumentation protocols, data protection protocols, [and] analytical workflows across platforms." Transparency is also essential, and he stressed the importance of "open systems with documentation of all preprocessing decisions, . . . and ensuring reproducible computational pipelines." He said, "AI cannot deliver reliable precision medicine until it is supported by bias-aware omics data." This will require ensuring equitable access to sequencing and computational resources; rigorous validation; and "robust governance frameworks for data-sharing, consent, and international standardization."
Celi summarized the current, emerging, and potential future horizons of AI. Currently, AI is useful for retrieving existing digitized information and making it more widely accessible. However, disparities in access persist, and equitable access is a key metric of success. The next horizon will be the recombination of siloed and non-digitized knowledge (e.g., from religious texts, indigenous wisdom, multilingual sources) to develop new hypotheses. Success can be measured by the "number of testable cross-domain hypotheses generated." The future promise of AI will be the discovery of "new knowledge by detecting weak signals and patterns in rare data about underserved populations," he said, as well as identifying solutions to long-standing problems "by interfacing previously disconnected knowledge." Success will be defined by the impact of these discoveries on "underserved, multilingual and multimodal communities."
Generative AI has rapidly disrupted "the foundations of education, knowledge creation, and human expertise," Celi said, and he stressed that "we cannot navigate this transformation with the frameworks of the past century." What is required he concluded, is the "courage to fundamentally reimagine the systems through which we learn, we discover, and we innovate."
Cora Han, the chief health data officer for University of California (UC) Health, said UC Health has centralized its health records for more than 10 million patients across its six academic health centers into a single data warehouse in its Center for Data-Driven Insights and Innovation. She mentioned several examples of AI use cases across UC Health and its academic health centers in the areas of reducing clinician burden, improving the detection of rare diseases, care management, and quality reporting.
AI governance is essential for facilitating the adoption of AI in health care and should be considered as an enabler of adoption, Han said. Governance means developing and implementing AI in a safe, responsible, and ethical way to build trust with providers, patients, administrators, and the community. Besides enabling more rapid, transparent, and replicable vetting and authorization of AI tools, AI governance "reduces the risk of unexpected harm and reputational damage, . . . promotes compliance with existing and evolving laws and regulation, . . . [and] promotes a safe and ethical innovation ecosystem," she said.
Han listed eight key elements of responsible AI: validity and reliability, safety, accountability and transparency, security and resiliency, explainability and interpretability, privacy, fairness, and addressing workforce impacts. She said that transparency is needed regarding the training and testing of models (e.g., training data, testing outcomes) and in communications with patients about how AI is or might be used in their care. She noted that California law now requires
health care providers to include a disclaimer when generative AI is used for clinical-based patient communications. Privacy is a particular concern for genomics. Genomic data are unchangeable and convey personal health information about patients as well as their families. There also is a risk of re-identification of deidentified data. Han said there are privacy-enhancing technologies designed to facilitate the sharing and use of genomic data while preserving individual privacy (e.g., synthetic data, federated learning). There are limitations to these techniques, she said, including balancing the trade-off between utility and privacy, and managing the technical complexities and practical deployment barriers associated with implementing these technologies.
Beyond initial validation, models should be evaluated for clinical effectiveness in practice, Han said. Specifically, does the model improve "patient outcomes, treatment and workflow efficiencies, and cost-effectiveness?" In this regard she said there is a role for implementation science, and UC San Francisco recently received a major philanthropic gift to develop an autonomous AI monitoring platform for clinical care.
In the near term there is potential for responsible AI in health care to reduce clinician burden, increase access to care and improve the coordination of care, and enhance the translation of research into practice, Han said. Attention needs to be paid to "workforce readiness and change management for AI workflows" and to "monitoring of AI tools for safety, equity, and performance." There is "extraordinary potential" for AI to facilitate what the late Atul Butte called "scalable privilege," using data to enable better care and outcomes for all patients, she concluded.
Ben Busby, the global alliances manager for omics at NVIDIA, said the company's accelerated genomics ecosystem is one of wide-ranging partnership with cloud and service platforms; instrument, diagnostics, pharmaceutical, and sequencing companies; and research institutions to develop multimodal solutions for accelerating health care discoveries. He mentioned several NVIDIA services and solutions which he said make "it easy and tractable for people to do things that are relatively computationally sophisticated." For example, the tools in the AI and graphics processing unit (GPU)-Accelerated Software Suite for Omics Analysis are designed to increase speed, reduce cost, and improve the accuracy of bioinformatics. He added that all the software he discussed at the workshop is open use or open source.
Busby said his personal vision of AI in genomics is one where a patient whose routine blood tests suggest a concern can have the standard phenotypic workup done in parallel with whole-genome sequencing, the results of which will inform pharmacological intervention. A challenge, he said, is that "the vast majority of our variant annotation tools are focused on 8 percent of diseases" which stems from the old "one variant/one disease" concept, whereas most diseases are multigenic or multifactorial, or both. To begin to address this, he said, the OpenCRAVAT platform "has 300 variant annotators" available to use, but most of those variants are uncontextualized and not easy to understand.
Multifactorial diseases can involve "polygenicity, background genomic effects (cis- or trans-), environmental effects mediated by epigenetics, [and] direct environmental effects (immunological, receptor binding, etc.)," Busby said. He added that "humans do not like thinking about multifactorial causation" and suggested that this is an area of genomics where deep learning could be especially helpful. One way of addressing this problem would be through traditional graph genomes. Historically, graph genomes have been computationally tricky to work with, but the advent of methods like reference guided graph assembly and proposed ones, like comparison to disease specific references have made the situation somewhat more tractable. That said, even faster and more efficient methods are desirable for large populations. He described an example which looked at background genomic effects on the penetrance of variants of TNF-alpha and HLA-A in three different populations from the 1000 Genomes Project. In summarizing the findings, Busby said that "the background genome is discretizable." The fact that the genome is discretizable means that we can look at both cis and trans effects on variant penetrance. He also described an example of "maximizing performance with massively parallel hash maps on GPUs" and then looking at haplotypes – and their constituent variants – for a population of people with a particular disease, which he said is "how we start to understand the paths that are associated with multiple etiologies."
Knowledge graphs of phenotypic data are another computational approach for genomics that can provide "more advanced clustering of disease subtypes," Busby said, and that enable "prediction of adequate pharmacology for these subtypes." He discussed examples related to breast and colorectal cancers. Looking to the future, Busby said retrieval-augmented generation can use knowledge graphs when providing support to LLMs. A graph neural network can be used "to parse just a specific part of the knowledge graph" for use by the LLM. He added that the LLM can also output a knowledge graph response, which can then be used for validation testing. Busby noted that this capability is available on GitHub for anyone to use. NVIDIA is currently looking into combining hash maps with phenotypic knowledge graphs "to find multiple paths at the same time that are correlated with specific sub-phenotypes," he said.
"Changing genomic data structures will help analyze multifactorial genomic etiologies," Busby concluded. "Building knowledge graphs of phenotypic data . . . will help to contextualize phenotypic information," and "putting effort into the above community initiatives is likely to help us deliver more precise treatments faster."
Panelists discussed the potential for AI to facilitate accessible precision care for all. One issue is incorporating the viewpoints of diverse populations in the development of increasingly human-like AI tools. Han highlighted the need to include patients and communities in conversations about what problems they need AI tools to help solve. Celi reiterated the need to include "wisdom from different philosophies and ideologies, including those that are outside the western realm" in the development of LLMs and said that ethical insights are being missed. It is important to include philosophers to think about what people really want from AI (e.g., should it replace humans in jobs it can do better?) because technology can also scale exclusion. Busby highlighted the opportunity to include a variety of technical experts in discussions of model development and to motivate the sequencing of genetically diverse populations.
Haendel asked what she said might be the most pressing question of the day: While there is excitement about using huge amounts of data for AI models, some populations are hesitant to share their data, so what policies, regulations, and governance should be put into place to protect privacy? AI will not fundamentally fix the country's already broken systems, Celi said. AI is already disrupting norms, from university education to scientific publication, and it is important to think urgently about the balance of harm to society with the speed at which AI-based approaches can be automated, he said. Han agreed with these societal questions and said that people "do want to share their data with their clinician . . . for their own care, for the care of those they love, often for the ‘public good.' But how do you define the public good, and who is actually defining it?" She said privacy-preserving technologies such as federated learning can help earn the trust of potential data sharers. Walton said it is also important to demonstrate the value of data sharing for communities and to communicate information of value back to those who shared their data. Busby said communities should maintain autonomous regulatory control over their data. Ideally, all sequencing would also be done locally, keeping all samples and data within the community confines, but that is often not possible. He also said that end-to-end solutions are needed that track what data were used in a model, the source of that data, and the importance of that data to the model so that credit—and potentially remuneration—can be given to those whose data were used. Another challenge, Celi said, is educating clinicians at scale about the uses, risks, and limitations of AI in genomics and precision care, and keeping them up to date as this technology is continually and rapidly changing.
Wood and Kunal Sanghavi, the associate director for genetic counseling at The Jackson Laboratory, highlighted observations and suggestions made by individual speakers on the applications of AI in genomics and precision health. The challenges and opportunities that were shared during the workshop are summarized in Box1, and the emerging uses of AI and potential actions that could be taken next are summarized in Box2. In closing the workshop, Wood called upon workshop participants to mentor others in genomics and precision medicine to "help them find their future in AI."
NOTE: This list is the rapporteurs' summary of suggestions made by one or more individual speakers as identified. These statements have not been endorsed or verified by the National Academies of Sciences, Engineering, and Medicine. They are not intended to reflect a consensus among workshop participants.
NOTE: This list is the rapporteurs' summary of suggestions made by one or more individual speakers as identified. These statements have not been endorsed or verified by the National Academies of Sciences, Engineering, and Medicine. They are not intended to reflect a consensus among workshop participants.
Aradhya, S., F. M. Facio, H. Metz, T. Manders, A. Colavin, Y. Kobayashi, K. Nykamp, B. Johnson, and R. L. Nussbaum. Applications of artificial intelligence in clinical laboratory genomics. American Journal of Medical Genetics Part C: Seminars in Medical Genetics 193(3):e32057.
Fiziev, P. P., J. McRae, J. C. Ulirsch, J. S. Dron, T. Hamp, Y. Yang, P. Wainschtein, Z. Ni, J. G. Schraiber, H. Gao, D. Cable, Y. Field, F. Aguet, M. Fasnacht, A. Metwally, J. Rogers, T. Marques-Bonet, H. L. Rehm, A. O'Donnell-Luria, A. V. Khera, and K. K. Farh. 2023.
Haendel, M., N. Vasilevsky, D. Unni, C. Bologa, N. Harris, H. Rehm, A. Hamosh, G. Baynam, T. Groza, J. McMurry, H. Dawkins, A. Rath, C. Thaxton, G. Bocci, M. P. Joachimiak, S. Köhler, P. N. Robinson, C. Mungall, and T. I. Oprea. 2020. How many rare diseases are there? Nature Reviews Drug Discovery 19(2):77–8.
Hernandez-Boussard, T., P. Macklin, E. J. Greenspan, A. L. Gryshuk, E. Stahlberg, T. Syeda-Mahmood, and I. Shmulevich. 2021. Digital twins for predictive oncology will be a paradigm shift for precision cancer care. Nature Medicine 27(12):2065–6.
Li, H., J. Qin, A. Li, R. Ouyang, Z. Chen, S. Huang, S. Qin, and Q. Huang Q. 2025. Systematic review and meta-analysis of deep learning for MSI-H in colorectal cancer whole slide images. NPJ Digital Medicine 8(1):456.
McGrath, S. P., B. A. Kozel, S. Gracefo, N. Sutherland, C. J. Danford, and N. Walton. 2024. A comparative evaluation of ChatGPT 3.5 and ChatGPT 4 in responses to selected genetics questions. Journal of the American Medical Informatics Association 31(10):2271–83.
Murugan, M., B. Yuan, E. Venner, C. M. Ballantyne, K. M. Robinson, J. C. Coons, L. Wang, P. E. Empey, and R. A. Gibbs. 2024. Empowering personalized pharmacogenomics with generative AI solutions. Journal of the American Medical Informatics Association 31(6):1356–66.
Sadée, C., S. Testa, T. Barba, K. Hartmann, M. Schuessler, A. Thieme, G. M. Church, I. Okoye, T. Hernandez-Boussard, L. Hood, I. Shmulevich, E. Kuhl, and O. Gevaert. 2025. Medical digital twins: Enabling precision medicine and medical artificial intelligence. Lancet Digital Health 7(7):100864.
Walton, N. A., R. Nagarajan, C. Wang, M. Sincan, R. R. Freimuth, D. B. Everman, D. C. Walton, S. P. McGrath, D J. Lemas, P. V. Benos, A. V. Alekseyenko, Q. Song, E. G. Uzun, C. O. Taylor, A. Uzun, T. N. Person, N. Rappoport, Z. Zhao, and M. S. Williams. 2024. Enabling the clinical application of artificial intelligence in genomics: A perspective of the AMIA Genomics and Translational Bioinformatics Workgroup. Journal of the American Medical Information Association 31(2):536–41.
Weinstein, J. N., and R. W. Allen. 2025. The hub and spoke model: Patient empowerment. New York: Barnes and Noble.
Rare penetrant mutations confer severe risk of common diseases. Science 380(6648):eabo1131.
Disclaimer: This Proceedings of a Workshop—in Brief was prepared by Theresa M. Wizemann and Sarah H. Beachy as a factual summary of what occurred at the workshop. The statements made are those of the rapporteurs or individual workshop participants and do not necessarily represent the views of all workshop participants; the planning committee; or the National Academies of Sciences, Engineering, and Medicine.
Planning Committee:Kunal Sanghavi (Co-chair), The Jackson Laboratory; Grant Wood (Co-chair), Global Genomic Medicine Collaborative (until its recent closure); Bimal P. Chaudhari, Nationwide Children's Hospital; Robert R. Freimuth, Mayo Clinic; Maia Hightower, Veritas Healthcare Insights; Adriana Huertas-Vazquez, Illumina; Amanda Perl, American Society of Human Genetics; Nalini Raghavachari, National Institute on Aging; Lee Sanders, Stanford University; Angela Starkweather, Rutgers University School of Nursing, representing the American Academy of Nursing; and Zhongming Zhao, UTHealth. The National Academies' planning committees are solely responsible for organizing the workshop, identifying topics, and choosing speakers. Responsibility for the final content rests entirely with the rapporteurs and the National Academies.
Reviewers: To ensure that it meets institutional standards for quality and objectivity, this Proceedings of a Workshop—in Brief was reviewed by Aaron Goldenberg, Case Western University; Angela R. Starkweather, Rutgers School of Nursing; Ryan J. Taft, Genetic Alliance; and Joyce Y. Tung, 23andMe Research Institute. Kirsten Sampson-Snyder, National Academies of Sciences, Engineering, and Medicine, served as the review coordinator.
Sponsors: This workshop was supported by contracts or agreements between the National Academies of Sciences and American Academy of Nursing; American College of Medical Genetics and Genomics; American Medical Association; American Society of Clinical Oncology; American Society of Human Genetics; Association for Molecular Pathology; Association of Public Health Laboratories; Biogen; College of American Pathologists; Geisinger Health; Genome Medical, Inc.; Health Resources and Services Administration (Contract 75R60221D00002, Task Order No. 75R60225F34010); Illumina, Inc; Kaiser Foundation Health Plan, Inc.; Meharry Medical College; Myriad Genetics; National Institutes of Health (Contract No. HHSN263201800029I, Task Order No. 75N98023F00019 and 75N98023F00022; Purchase Order No. 75N92024P00357), including the All of Us Research Program, National Cancer Institute, National Human Genome Research Institute, National Institute of Mental Health, and National Institute on Aging; National Society of Genetic Counselors; Parkinson's Foundation; the University of California, San Francisco; University of Central Florida; and the University of Vermont Health Network Medical Group. Any opinions, findings, conclusions, or recommendations expressed in this publication do not necessarily reflect the views of any organization or agency that provided support for the project.
Staff:Sarah H. Beachy, Roundtable Director, Kathryn Asalone Shively, Associate Program Officer (through August 2025), Michelle Drewry, Associate Program Officer (through November 2025), and Ashley Pitt, Senior Program Assistant.
Suggested citation: National Academies of Sciences, Engineering, and Medicine. 2026. Exploring Applications of AI in Genomics and Precision Health: Proceedings of a Workshop—in Brief. Washington, DC: National Academies Press. https://doi.org/10.17226/29392.
Copyright 2026 by the National Academy of Sciences. All rights reserved.