Skip to main content

NLM - National Library of Medicine Grants

Browse 31 open grants from NLM - National Library of Medicine. Find eligibility requirements, award amounts, and deadlines for each opportunity.

Showing 24 of 31 grants from NLM - National Library of Medicine

24 grants worth up to $12.5M match your search

Enter your email to see grant names, funders, and application links

Unlocking Medical AI: A Scalable, Privacy-Preserving Annotation Platform for Clinical and Physiological Data

open

NLM - National Library of Medicine

PROJECT SUMMARY This project addresses a critical challenge in medical artificial intelligence (AI): the lack of high-quality, annotated physiological datasets necessary for developing robust and generalizable models. Current methods for annotating clinical time-series data, such as signals from wearables, bedside monitors, and electronic health records, are insufficient due to irregular sampling rates, frequent missing data, and multi-modal complexity. Our goal is to develop and rigorously evaluate an advanced annotation platform that enables efficient and secure labeling of complex clinical datasets while ensuring compliance with HIPAA and interoperability standards. By addressing key gaps in data quality, scalability, and reproducibility, this platform will accelerate the development of AI-driven healthcare solutions. In Phase 1, we will focus on demonstrating technical feasibility by solving key challenges related to multi-signal data visualization, privacy-preserving annotation workflows, and user-centric design. Specifically, we will enhance existing tools to enable sub-second latency for efficient review of irregularly sampled, multimodal data. We will also design and implement robust privacy frameworks to ensure secure handling of sensitive health information. Finally, we will engage clinicians, researchers, and industry experts to identify critical platform features and vali- date the usability of our solution. This phase will lay the foundation for a scalable, clinically validated annotation platform capable of supporting diverse healthcare applications. In Phase 2, we will extend these efforts to validate the platform's scalability and utility in real-world healthcare settings. This will involve building scalable data architectures to support concurrent users, implementing standard- ized EHR integration using FHIR protocols, and deploying production-grade annotation workflows with advanced compliance controls. Pilot deployments at clinical sites will be conducted to evaluate platform performance, us- ability, and its impact on AI model development. The outcomes will provide robust evidence for the platform's effectiveness in generating high-quality datasets critical for AI innovation. By creating an advanced framework for medical data annotation, this project will contribute to improving the reproducibility and quality of AI models used in healthcare. The platform's innovative ability to generate reliable datasets will support breakthroughs in predictive analytics, real-time monitoring, and personalized medicine, ultimately driving better patient outcomes and more efficient healthcare delivery.

Up to $307K
2027-01-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Bringing Up Baby: Race, Infant Mortality, and the Creation of Prenatal Care, 1900-1930

open

NLM - National Library of Medicine

Project Summary and Abstract I am applying for an NLM Grant for Scholarly Works in Biomedicine and Health to complete my history of medicine monograph: Bringing Up Baby: Race, Infant Mortality, and the Creation of Prenatal Care, 1900-1930. This book will trace the intertwining threads of public health, eugenics, racial science, Progressive Era philanthropy, and professionalizing obstetrics in a story of how Americans became aware of and sought to fix the problem of infant mortality in the early twentieth century. In the 1910s and 1920s myriad groups and organizations, both those interested in health and those interested in social reform, studied the extent and causes of infant mortality, lobbied state and federal governments for maternal and infant welfare funding, and attempted to convince the American public that pregnancy was a condition that required medical surveillance and intervention. This will be the first work of history to dive into these movements and determine how nationalism, race, and medical professionalism efforts shaped the development of prenatal health care in this country. There have been no historical studies devoted solely to prenatal care and my findings into the emergence of this medical specialty and public health concern show it to be rooted in particular racial politics and national health concerns of the early 1900s. Relying on a range of sources including federal infant mortality studies, public health journals, personal correspondence, medical reports, meeting transactions, sociological reports, and popular health pamphlets, I illustrate that prenatal health care originated in a time of eugenics, Jim Crow, and medical misogyny, and perhaps never fully left those values behind. Investigating the history of prenatal health care will expand the historical field of American reproduction and medicine as well as inform current racial disparities in maternal and infant health care and mortality. In addition, this study draws together and speaks across multiple fields in history including women’s history, medical history, political history, and social history. I have plans to publish with Rutgers University Press, the press that published my well-received and widely-read first book Lost: Miscarriage in Nineteenth-Century America.

Up to $114K
2027-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

11th Machine Learning for Healthcare Conference (MLHC) 2026

open

NLM - National Library of Medicine

Project Summary Advances in machine learning and artificial intelligence (AI) are reshaping the landscape of healthcare, offering new opportunities to enhance diagnosis, personalize treatment, optimize clinical workflows, and ultimately im- prove patient outcomes. The Machine Learning for Healthcare (MLHC) Conference is dedicated to accelerating scientific progress in this rapidly evolving field by providing a premier forum for the exchange of ideas, dissemi- nation of cutting-edge research, and development of interdisciplinary collaborations among computer scientists, engineers, clinicians, and healthcare innovators. The 11th Annual MLHC Conference, to be held at Johns Hopkins University in Baltimore, MD, on August 12th – 14th, marks a significant milestone for the community. In celebra- tion of this anniversary, the meeting will highlight state-of-the-art advances in machine learning for healthcare while also placing special emphasis on the challenges and opportunities involved in translating these technolo- gies from research settings to clinical practice. Through its emphasis on biomedical informatics and data science methodologies that make health data and machine learning models more findable, interoperable, reusable, and trustworthy in clinical use, MLHC is directly aligned with the NLM’s mission. Specifically, sessions will focus on methodological innovation, rigorous evaluation, deployment considerations, and pathways toward achieving meaningful, real-world impact at the bedside. Core objectives of the conference are to (1) update attendees on timely developments across the spectrum of AI-driven healthcare research, (2) foster cross-disciplinary commu- nication and collaboration, and (3) strengthen the pipeline of future leaders in the field. A central component of MLHC is its commitment to supporting and engaging students/trainees through opportunities such as free con- ference registration, travel awards, best student paper awards, poster and oral presentation sessions, mentoring events, and training-focused workshops. This application seeks funding to support such trainee participation, en- abling students and early-stage researchers to present their work, attend scientific sessions, and engage directly with experts who are shaping the future of AI in healthcare. Support from NLM through this R13 mechanism, will enhance trainee activities, broadening access, particularly for those with limited travel and professional devel- opment resources. While in the previous years, the conference has mainly sought funding from private entities, this year NIH’s support will be critical in expanding the conference’s capacity to reach graduate students, post- doctoral researchers, and early-career investigators across the country. As the MLHC Conference enters its second decade, this support will ensure that the meeting continues to advance rigorous, reproducible, and clin- ically grounded research while cultivating the next generation of innovators dedicated to improving patient care through machine learning.

Up to $25K
2027-07-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

A Scalable, Open-Source Generative LLM Tool for Automated Classification of Diagnostic Errors

open

NLM - National Library of Medicine

1 Medical errors are the third leading cause of death in the United States yet estimates of their total 2 burden and epidemiology remain largely unknown, with few comprehensive assessments 3 available. To address this gap, we propose leveraging the Retract-and-Reorder (RAR) method, 4 an existing health information technology (IT) tool that detects near-miss, self-caught order errors, 5 to better understand the underlying causes of medical errors. The RAR method has been reliably 6 used to detect wrong-patient and certain types of medication prescribing order errors. We 7 expanded its application to diagnostic imaging, identifying additional error types such as wrong- 8 site, wrong-contrast, wrong-side, and wrong-modality, using logic-based natural language 9 processing (NLP). However, over 42% of detected errors remained unclassified, requiring labor- 10 intensive manual review for further categorization. In this proposal, we aim to develop a scalable 11 pipeline that automatically classifies order errors and addresses unknown error types using 12 generative large language models (LLMs). To accomplish this, we will first (AIM 1) develop and 13 validate a generative LLM-based classification model for categorizing RAR events into predefined 14 error types, focusing on imaging order errors. We will compare its performance against the current 15 logic-based NLP approach, hypothesizing that the LLM will achieve equal or better accuracy by 16 correctly classifying known error and identifying previously missed error types, thereby improving 17 overall classification. Then, we will (AIM 2) demonstrate the scalability of the LLM pipeline by 18 applying it to medication order errors and developing a dissemination plan. We hypothesize that 19 LLMs can be readily adapted to diverse large sets of order types across various domains without 20 requiring fine-tuning. This study will establish the feasibility of developing an advanced, 21 automated, and scalable open-source tool for classifying and characterizing RAR events across 22 different medical orders. By identifying and understanding various order error types across 23 domains, this research will support the development of measures and targeted interventions to 24 improve patient safety. Furthermore, our privacy-preserving approach, achieved by deploying an 25 open-source LLM along with comprehensive documentation and structured dissemination, will 26 enable adoption across institutions and diverse healthcare settings. Beyond imaging and 27 medication orders, this framework could support cross-institutional implementation, facilitating its 28 expansion into other order domains.

Up to $82K
2028-04-30
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Computational and experimental framework for the integrated tissue analysis of spatial metabolomics and transcriptomics datasets

open

NLM - National Library of Medicine

SUMMARY The cell type composition and cellular metabolism jointly shape a tissue’s microenvironment, and its function as a consequence. Over the past few years, spatial transcriptomics (ST) has become an important and commonly used method for mapping cellular states in tissues. As powerful as this approach is, however, transcriptomes constitute only one of many crucial biological modalities, including DNA, protein, and small molecules, which, upon integration, would provide a more comprehensive view of tissue architecture. Recently, spatial metabolomics (SM) by mass spectrometry imaging has become available and promises to enable the elucidation of entire metabolomes at high spatial resolution. At present though a robust approach is missing for the integration of spatial metabolomics with spatial transcriptomics. In this proposal we aim to develop novel algorithms for such an integration in the context of an important problem in cancer biology. When studying drug-treated tumors, ST and SM can reveal cellular states and the precise concentration of the drug that they are experiencing, respectively. We previously found that as cancer cells adapt to therapy, they undergo a set of cell state transitions that we have referred to as the ‘resistance continuum’. We also found in vivo evidence for the states along this continuum, however new insight requires an integration of spatial and temporal analysis. In our preliminary results, we showed the power of joint ST and SM analysis, but we were not able to track the clonal and drug treatment history of the cells over time. Thus an open question, with immense clinical relevance, is what is the effective concentration of a drug experienced by an adapting cell lineage. We propose to address this question here by deploying two independent frameworks. In Aim 1 we describe a method using an optimal transport framework to integrate data from a time-course comprising tumors adapting to a drug from different animals. Our computational framework will be designed to identify cellular transitions and propose specific hypotheses for testing. A second approach described in Aim 2 exploits novel lineaging technology and our established serial passaging approach for studying the same tumor over time. Analyzing this data will allow us to reveal the longitudinal history of a clone and reveal whether cells with lower or higher dose concentrations in their early adaptations were selected for higher drug resistance. Overall, the approaches developed in this proposal specifically address the challenges of spatial metabolomics and spatial transcriptomics data and we expect them to be of high value for many in the large community of researchers using spatial analyses.

Up to $229K
2028-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Hopi Language Information Resources to Reduce Health Disparities

open

NLM - National Library of Medicine

Project Summary The Hopi Tribe is a sovereign nation of approximately 20,000 enrolled members, of whom approximately 15,000 reside in the Hopi tribal land in northeastern Arizona. In response to the high prevalence of cancers in the Hopi population, the Hopi Tribe established the Hopi Cancer Support Services programs with a focus on increasing screening uptake and providing supportive services for Hopi people navigating cancer care. Despite robust community engagement, there is persistently low uptake of screening services and increased rates of cancer mortality among tribal members, representing an ongoing health disparity for the Hopi Tribe. In the evaluation of an evidence-based practice implementation of digital library resources for Hopi Cancer Support Services, patient navigators reported information uptake barriers due to the lack of cancer terminology in the Hopi language (Hopilavayi). Early efforts to create a Hopilavayi medical lexicon yielded basic terminology that is helpful for communication. However, the extant cultural focus on health that avoids topics of anatomical sites for cancers, morbidity, and mortality remains a communication barrier for Tribal health workers. Infrastructure limitations including limited availability of printed or digital Hopilavayi dictionaries and limited internet connectivity on the Hopi Tribal land pose a challenge to health literacy and communication, contributing to suboptimal uptake of screening services and unnecessarily high cancer morbidity and mortality. Additionally, Hopi cultural preservation priorities limit the use of internet resources to disseminate resources in Hopilavayi including dictionaries and health information. We have previously used an offline digital health library (SolarSPELL) to curate and disseminate culturally relevant cancer education materials to Hopi Tribal members, meeting both the infrastructure conditions and cultural preservation priorities for the Hopi Tribe. Building on that pilot success, we propose to leverage the offline digital library platform to disseminate dictionary and health content to 1) Develop a Hopilavayi lexicon for cancer terminology, and 2) Develop linguistically and culturally congruent cancer education resources for the Hopi Tribe. The proposed project uses a culturally-modified Delphi method approach to create novel Hopi language terms to convey key cancer-related concepts. Evidence-based cancer education materials will be developed and tested with community members to evaluate understanding and cultural congruence for refinement and future dissemination throughout the Hopi community. In addition to creating culturally-relevant cancer education materials appropriate for future dissemination and evaluation using offline digital libraries, the successful completion of the proposed work can serve as a model for creating culturally- and linguistically-relevant health promotion materials for a variety of communities that experience health disparities.

Up to $289K
2028-06-30
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

The Promise of Innovation: Organ Chips and the Making of Novel Biomedical Models

open

NLM - National Library of Medicine

PROJECT SUMMARY The use of animal models in biomedical research, and specifically in pharmaceutical safety and efficacy testing, has been foundational to biomedicine. However, the translational failures of animal models in pharmaceutical testing have been well documented, and developing new technologies to augment the use of animal models and better predict human relevance, safety, and efficacy in pre-clinical testing has been an area of significant investment over the past several decades. Researchers have focused on constructing new tools—leveraging developments in human-cell based models, in vitro platforms, and computational and machine learning technologies—to revolutionize pre-clinical drug testing. The proposed book project traces the emergence and development of translational technologies called organ chips to unpack the social, organizational, and scientific factors that shape whether and how novel technologies become successful. The project has three aims: (1) situate the emergence of organ chip technologies in a sociohistorical perspective, (2) document how scientific, social, and technical factors have shaped the design and uptake of organ chip technologies as models in biomedical research, and (3) analyze the ethical, legal and social implications (ELSI) of organ chips. The book will draw on an empirical, sociological study of the emergence and construction of organ chips, drawing on a research that has taken a rigorous mixed methods approach through extensive observations in an organ chip lab, in-depth interviews with organ chip researchers and other stakeholders, and content analysis of scientific publications and popular media. There have been very few scholarly books that bring together accounts of the emergence and construction of novel biomedical models and none have focused on the rising prominence of biomedical engineering approaches in medicine. In addition to contributing to this understudied area and analyzing the emergence of organ chips through a social scientific lens, this project has several other innovative dimensions, including how it situates technologies in sociohistorical perspective, examines biomedical model construction, and interrogates the role of hype building and scientific storytelling in the development of novel biotechnologies. The book will engage disciplinarily diverse audiences in science and technology studies, sociology, history of technology, bioethics, and engineering ethics. The book will appeal to those interested in biomedical innovation, the organizational and institutional dynamics that enable and constrain novel biotechnologies, and the ethics of innovation.

Up to $106K
2028-07-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Scalable Literature Curation, Summarization, and Monitoring of Genomic Variants

open

NLM - National Library of Medicine

PROJECT SUMMARY/ABSTRACT Genomic variant reclassification plays a key role in accurate diagnoses and appropriate clinical decisions, as new evidence can change a variant’s classification, impacting patient treatment options. However, the process of genomic variant reclassification is hindered by the sheer volume of rapidly expanding literature, the labor-intensive and error-prone nature of manual curation, and the lack of efficient mechanisms for continuously tracking and integrating new findings. This project aims to develop scalable and automated informatics solutions for genomic variant curation by extracting key metadata, ranking evidence by reliability, and implementing literature monitoring to support timely revisit to the evidence. Specifically, the project is organized into 3 specific aims. Aim 1 focuses on developing a pipeline to extract evidence and metadata related to genetic variants from the literature, while also categorizing evidence by study type, such as experimental, computational, epidemiological, and case reports. Additionally, a ranking model will be developed to prioritize extracted evidence based on relevance and recency. Aim 2 focuses on analyzing trends in variant reclassifications and their correlations with evidence metadata to improve the understanding of factors influencing classification changes. Based on this analysis, an automated tool will be developed to monitor, summarize, and highlight significant new evidence from the literature. Aim 3 focuses on developing a user-friendly web application to present the extracted metadata, ranking results, and literature updates in a format tailored to the needs of genetic researchers. In collaboration with domain experts, the application will be co-designed to enhance usability and integration into existing variant curation workflows. A pilot study will be conducted to evaluate its effectiveness in improving researchers’ efficiency, reducing costs, and minimizing the manual effort required for variant reclassification. During the K99 phase, Dr. Zhang will develop an automated pipeline for curating metadata of genomic evidence under the supervision of Dr. Chunhua Weng. In the R00 phase, Dr. Zhang will develop real-time genomic literature monitoring and a user-friendly web application for genomic researchers, collaborating with domain experts to co-design and evaluate its impact on variant reclassification. To ensure a successful transition to independence, Dr. Zhang will receive training in genomic medicine through coursework and collaborations with clinical experts. Additionally, Dr. Zhang will strengthen mentorship, leadership, and grant-writing skills through activities including co-mentoring students, managing research projects.

Up to $120K
2028-07-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Bridging the Terminology Gap: Building User-Friendly Tools for FAIR Data Creation

open

NLM - National Library of Medicine

Project Summary The NIH Strategic Plan for Data Science emphasizes the need for scalable, standards-driven infrastructure to support effective biomedical data sharing. A central component of that infrastructure is high-quality metadata— structured, semantically consistent annotations that make data findable, accessible, interoperable, and reusable (FAIR). While tools like the CEDAR Workbench and BioPortal have made it easier to use ontologies in metadata, a critical gap remains: the lack of robust infrastructure to manage the lifecycle of value sets, the curated collections of terms that underpin standardized metadata. This project will address that gap by building a general-purpose, modular infrastructure for value-set lifecycle management. The system will support value- set discovery, collaborative editing, validation, version tracking, and integration with ontology repositories. Designed for use across diverse NIH initiatives, the infrastructure will help ensure that metadata are consistent with domain standards, support reuse, and evolve reliably over time. By enabling researchers, curators, and informaticians to jointly create and manage reusable vocabularies, the new software will improve metadata quality and alignment with NIH data-sharing mandates. For metadata authors, it will streamline the process of defining terms that are semantically valid and contextually relevant. For repository users, it will enhance dataset comparability and enable better semantic search. For researchers seeking to reuse public data, it will reduce the need for manual harmonization and facilitate meaningful cross-study analysis. This work directly supports NIH’s vision of a connected, standards-based data ecosystem and advances the infrastructure needed to make that vision operational across the biomedical research landscape.

Up to $231K
2028-07-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Fair Allocation of Scarce Medical Resources in an Unfair World

open

NLM - National Library of Medicine

Project Summary This proposal will fund the first comprehensive, book-length bioethical analysis of allocating scarce medical resources. Historically, allocation challenges arose at the advent of penicillin, dialysis, and organ transplantation. During the 2000s, allocation frameworks were developed for pandemic influenza countermeasures, including vaccines and ventilators, in the event of a pandemic. More recently, allocation frameworks have been proposed or deployed for the allocation of various COVID-19 countermeasures and treatments, including vaccines, therapies, and intensive care bed space. Such frameworks have also been deployed for the allocation of other resources in limited supply, such as mpox countermeasures, RSV countermeasures, and scarce medicines such as chemotherapeutics and novel GLP-1 inhibitors. The study defines four core ethical goals for allocation: benefiting people and preventing harm; mitigating disadvantage; equal concern; and reciprocity. It then analyzes the ethical, legal, and policy dimensions of metrics, such as lives saved or years of life lost, used by policymakers to assess resource allocation frameworks. Likewise, it analyzes ethical, legal, and policy dimensions of considering various characteristics, such as age, health status, race, and sex, as proxies for expected outcomes. Last, it considers how ethical objectives can be combined into an allocation framework and explains how such a framework can be operationalized and effectively implemented. This work will be of academic value to medical and public health professionals, policymakers, and biomedical researchers (including scholars in bioethics, health law and policy) because medical innovations and medical crises will continue to prompt the need to fairly distribute scarce medical resources. This project will help improve the ethical justification, design, legality, and implementation of such frameworks, enabling them to more effectively mitigate the harms of scarcity while addressing health disadvantage. It will also advance public health by helping public health officials to navigate tradeoffs and avoid ethical and legal pitfalls as they design, operationalize, and implement frameworks for allocating scarce medical resources.

Up to $100K
2028-07-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Deep Learning of Child Abuse Imaging: Improving Outcomes of Children Evaluated for Physical Abuse

open

NLM - National Library of Medicine

PROJECT SUMMARY Fractures are a common manifestation of physical abuse, with children <2 years at highest risk. The identification of healing fractures is crucial in the evaluation of physical abuse in a young child as these can suggest ongoing violence within the home and have serious implications for child protection. However, estimating time-since-injury of healing fractures based on imaging is often difficult and imprecise. Although deep learning (DL) models could vastly improve accurate dating of healing fractures in children presenting with suspicious injuries, a critical gap remains for accessible large digital pediatric imaging datasets and needed artificial intelligence (AI) infrastructure. Notably, this gap has recently been designated a critical pediatric health priority by the American College of Radiology. This project closes this gap by establishing the framework for deidentified image sharing and storage between three PEDSnet sites (Nationwide Children’s Hospital, Cincinnati Children’s Hospital Medical Center, Riley Hospital for Children) via a Secure File Transfer Protocol and providing the AI infrastructure needed for better image interpretation and diagnosis. We will train and validate DL models with state-of-the-art transformers such as DINOv3 and benchmark to the well-established convolutional neural network architecture ResNet-50 using skeletal imaging of accidental fractures of long bones in children <4 years to directly and accurately age healing fractures. In parallel, we will use meta- learning with a combination of labeled accidental fractures and unlabeled abuse fractures, followed by few-shot learning to regress the age of abuse fractures. Deliverables include establishing the framework for image sharing within pediatric health systems and the development of DL algorithms for aging of healing fractures that could be implemented widely as a virtual consultant for radiologists faced with the task of interpreting imaging completed in children presenting with high-risk injuries. This is the first study to propose the development of DL algorithms for aging healing fractures by 1) training on multicenter imaging data and 2) using real-world data of patients evaluated for abuse. This proposal is a key first step towards development of a national resource to stimulate and support high-quality, collaborative imaging research within pediatrics, dramatically improving patient outcomes within both pediatric and community settings. By providing a mechanism for cross-site image sharing, this project enables future scalable multi-institutional model development and validation for improved interpretation of imaging completed in child abuse evaluations.

Up to $243K
2028-07-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Scaling clinical wearable foundation models for the detection of in-hospital deterioration

open

NLM - National Library of Medicine

This project supports a Research Software Engineer (RSE) to significantly advance in-hospital patient care by developing and disseminating cutting-edge, AI-driven software tools for the early prediction of clinical deterioration. The broad, long-term objective is to transform patient monitoring by enabling timely interventions, thereby improving patient outcomes and reducing healthcare costs associated with acute deterioration events in non-critical care settings. This work directly supports Aim 2 of NIH grant R01NR020774. A primary specific aim is to develop next-generation, personalized deterioration prediction models leveraging an extensive clinical wearable dataset. The research design involves employing generative foundation models and multimodal learning. Key methods include self-supervised pre-training on unlabeled physiological time series data using frameworks such as SimCLR and BYOL, followed by fine-tuning on labeled deterioration events. Multimodal foundation models will be developed to integrate continuous vital signs from wearables with Electronic Health Record (EHR) data, utilizing novel fusion techniques to capture complex interactions. To address data scarcity for rare clinical events, cross-location data synthesis techniques, including Generative Adversarial Networks (GANs) and optimal transport-based methods, will be investigated to generate realistic synthetic physiological data. These models will be rigorously validated using 10-fold cross-validation for both short-term (4-hour) and mid-term (24-hour) prediction windows, aiming for superior accuracy, timeliness, and generalizability. The models' capabilities will also be assessed for predicting missing continuous vital values and demographic features based solely on recorded vitals. A second major aim, directly aligning with the RSE's short-term career goal, is the creation and public release of a robust, open-source Python package for comprehensive validation of wearable sensor data against multiple ground truth sources. This package will incorporate time alignment algorithms, visualization tools (e.g., scatterplots, Bland-Altman plots), and automated statistical tests. Its development will adhere to research software engineering best practices, including a modular architecture for interoperability, extensive unit and integration testing for robustness, comprehensive documentation for user adoption and tools for distributed computing to handle large datasets. The RSE's long-term career objective involves building a platform to operationalize the developed deterioration foundation model, specifically the Continuous Clinical Alert System (CCAS). This platform will provide secure, scalable infrastructure for real-time data streaming, EHR integration, and seamless deployment of CCAS outputs into clinical workflows, supporting the entire AI/ML Software as a Medical Device (SaMD) lifecycle and eventual FDA submission. These RSE activities are critical for translating advanced AI research into clinically impactful tools, enhancing patient safety, and establishing sustainable research software.

Up to $140K
2029-06-30
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Accelerating FAIR4RS Clinical Patient Accrual through Standardized Infrastructure for In-Silico Clinical Phenotyping

open

NLM - National Library of Medicine

PROJECT ABSTRACT/SUMMARY The main objective of this proposal is to obtain support for Andrew Wen, MS, a research software engineer and data scientist at the University of Texas Health Science Center at Houston. Mr. Wen supports the development and dissemination of research tooling and infrastructure at the Center for Translational AI Excellence and Applications in Medicine (TEAM-AI) under the direction of Hongfang Liu, PhD. Mr. Wen's career has a demonstrated focus in the design, implementation, and dissemination of open-source research infrastructure, particularly with a focus on management, analysis on, and information retrieval from large repositories of unstructured clinical research data. This award would allow for 3 years of funding for Mr. Wen to lead development and dissemination of research tooling to facilitate accrual and management of patient cohorts via in-silico clinical phenotyping, with an initial use case on rare disease cohort identification. The benefits of development and dissemination of this tooling are two-fold: 1) by distributing tooling suited for the purpose, we accelerate data cohort accrual, which is often a bottlenecking step in clinical artificial intelligence research, particularly with respect to query generation/development from natural language and human-in-the-loop evidence review and relevance judgement, and 2) improving research reproducibility and compliance with FAIR4RS principles by ensuring that appropriate metadata for dataset curation is routinely stored and retained as part of standard, sharable, tooling.

Up to $160K
2029-07-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Preventing Medication Dispensing Errors in Pharmacy Practice with Risk-sensitive Artificial Intelligence

open

NLM - National Library of Medicine

PROJECT SUMMARY Preventable medication errors are a global problem that can cause significant patient harm and annually incur costs of $42 billion worldwide. In the United States, 3 million outpatient medical appointments, 1 million emergency department visits, and 125,000 hospital admissions each year are the result of medication errors. Medication errors result in 3 million outpatient medical appointments, 1 million emergency department visits, and 125,000 hospital admissions each year. Astoundingly, over 4 billion prescriptions are dispensed every year in the United States alone. Although dispensing error rates are generally low at 0.06%, the sheer volume of dispensed medications translates to 2.4 million incorrectly dispensed medications each year. In the pharmacy, dispensing errors arise when pharmacists do not detect that the medication filled inside a prescription vial is different from the medication ordered on the prescription’s label. These dispensing errors can result in patient harm, added strain on the healthcare system, and costly legal action against the pharmacy. Artificial intelligence (AI) can be employed to assist in the verification process to help avoid dangerous and costly pharmacy dispensing errors. However, for the human-AI partnership to function optimally, the AI should be capable of determining the relative risks of medication errors (e.g., warfarin vs. vitamin C, kidney function, pregnancy status) while encouraging providers to make sound cognitive decisions such that optimal trust is maintained (i.e., catching errors), and temporal and cognitive demand is reduced (i.e., improving efficiency and avoiding alert fatigue). Risk-sensitive classification is critical when misclassification errors widely vary in frequency and severity. Imperative to this goal is to design AI from which risk-sensitive information can be extracted and conveyed to calibrate user’s trust in AI as either over-trust or under-trust can lead to near miss and incident errors. This proposed project will further our knowledge for designing risk- sensitive AI outputs and inform the development of AI models that encourage pharmacy staff to make sound clinical decisions that lead to better patient outcomes while improving work-life at lower costs of care. This study develops risk-sensitive AI methods in the context of medication images classification and designs effective AI advice and reasoning that lead to lower cognitive demand and increased trust in the AI. Our hypothesis is that risk-sensitive AI will lead to improved pharmacist work performance and more calibrated trust. The objectives of this proposal are to: 1) design risk-sensitive artificial intelligence to double-check dispensed medication images in real-time; 2) evaluate changes in pharmacy staff trust due to the use of risk- sensitive artificial intelligence; and 3) determine the effect of risk-sensitive artificial intelligence on pharmacy staff work performance.

Up to $1.4M
2030-05-09
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Advancing AI for WeaklySupervised Multimodal Alignment and Query-Based Interpretation in Biological Imaging

open

NLM - National Library of Medicine

PROJECT SUMMARY A major challenge in modern biological data analysis is integrating and reasoning over the vast volumes of unstructured, multimodal data now available, such as images and text. While each modality offers complementary biological insight, they are not encoded in a shared format or language, making it difficult to align them or reason across them computationally. A core unmet need is the development of AI systems that can bridge this divide by learning shared representations across data types. This project targets building AI systems for aligning and reasoning jointly over biological image and text data. We focus on microscopy and the challenge of interpreting microscopy images often requiring integration with broader biological context—relating observed phenotypes to those seen in other experiments, identifying plausible mechanisms, and connecting to relevant prior studies. This knowledge is frequently buried in unstructured images and text scattered across publications, databases, and annotations. Conventional AI systems rely on supervised learning, which demands large amounts of manually annotated data and cannot scale to the complexity or breadth of modern biology. Training AI models using weak supervision offers a promising alternative: by learning from loosely aligned image-text pairs, models can capture cross-modal associations from noisy but abundant sources. Vision-language models (VLMs) built on this principle embed images and text into a shared semantic space and support flexible reasoning tasks such as retrieval and question answering. However, current models often suffer from “blurry vision”—they can identify broad semantic matches between images and text but fail to resolve the fine-grained visual distinctions essential for biological interpretation. The goal of our project is to overcome this limitation by advancing weak supervision methods that enable fine-grained alignment between biological images and text, with a focus on microscopy. We will curate a large and diverse dataset of fine-grained image-text pairs and train a visual encoder using multi-scale contrastive learning to integrate both global and local alignment signals. This encoder will power an agentic AI system for query-based interpretation of microscopy images, that can retrieve relevant biological evidence and generate natural-language interpretations of microscopy images in response to researcher queries. We will validate the system in expert-driven use cases spanning single-cell perturbation and tissue-level pathology, and disseminate it through integration into widely used imaging workflows. By building AI tools that help researchers connect microscopy image content to pathways, phenotypes, and prior studies, we aim to support flexible, biologically grounded exploration and accelerate data-driven discovery. By open-sourcing our datasets, methods, and trained models for fine-grained image-text alignment, we also aim to advance the broader capabilities of multimodal AI for biological data analysis.

Up to $1.4M
2030-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Using Multi-Agent System Powered by Large Language Model with Expert-in-the loop to Optimize Clinical Decision Support (MAPLE-CDS)

open

NLM - National Library of Medicine

PROJECT SUMMARY The widespread adoption of electronic health records (EHRs), supported by over $34 billion in government in- vestment, has significantly increased the use of clinical decision support (CDS) systems. CDS provides up-to- date information and recommendations to healthcare professionals and patients to reduce errors and improve healthcare quality. However, CDS effectiveness is often hindered by a low acceptance rate, which is typically below 10%. High rates of low-relevance alerts lead to alert fatigue, desensitizing clinicians to alerts of higher importance. At Vanderbilt University Medical Center (VUMC), over 1,000 CDS alerts generate 70,000 user com- ments annually, and manual reviews of alerts based on medical literature are time-consuming and prone to delays, creating an urgent need for an automated system to generate suggestions to optimize CDS alerts. Large language models (LLMs) and multi-agent systems are promising tools to address this need. LLMs achieve high efficiency in processing large volumes of text, while multi-agent systems can collaborate to solve complex problems from multiple perspectives. In addition, we will incorporate a CDS-focused medical knowledge graph into the system to better retrieve relevant content, manage complex relationships in clinical data, and provide metadata (e.g., evidence strength). The overall objective of this proposal is to develop and evaluate an LLM- powered multi-agent system that integrates alert content, user feedback, and external knowledge sources to generate suggestions to improve CDS. Our central hypothesis is that system-generated suggestions outperform suggestions generated in the current manual processes. Our work includes three specific aims: Aim 1) Create a CDS-focused medical knowledge base and develop a graph for external knowledge support using LLMs, Aim 2) develop and validate a multi-agent system with LLM guardrails to generate CDS optimization suggestions, and Aim 3) evaluate and refine the multi-agent system via expert-in-the-loop feedback. The expected outcomes of this work include a scalable knowledge graph that integrates the latest medical knowledge, ensuring that CDS tools remain current and evidence-based. Additionally, the creation of an innovative LLM-powered multi-agent CDS audit system will improve the accuracy, relevance, and efficiency of CDS alerts, significantly reducing the manual effort required to maintain and update CDS content. The modularized architecture of the system could facilitate continuous improvement with new AI technology and expansion to other EHR-related tasks. While this project will focus on auditing current CDS alerts, the system’s potential applications include 1) helping CDS experts develop new CDS tools and 2) monitoring medical literature for updates that impact existing CDS tools and notifying stakeholders of relevant changes. Ultimately, this project aims to contribute to more intelligent and efficient CDS, improving patient safety and healthcare quality.

Up to $401K
2030-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Statistical Methods for Integrating Irregularly Collected Longitudinal Multi-Modal data into Prediction Models

open

NLM - National Library of Medicine

PROJECT SUMMARY For individuals with chronic illnesses such as diabetes, heart failure, cancer, or obesity, early intervention prevents symptom escalation, acute care use, and mortality. The growing availability of longitudinal electronic health record (EHR), patient-reported outcome (PRO), and mobile health (mHealth) data, including passively collected accelerometer, smartphone, and sensor data, offers new opportunities for proactive intervention. However, the rapid expansion of routinely collected mHealth data has outpaced the research community's ability to interpret it effectively. In particular, the high dimensionality and irregular collection of longitudinal EHR, PRO, and mHealth data introduces key challenges for predictive modelling due to planned sparsity or unplanned missingness. Current methods fall short in three key areas: 1. Informative missingness: Data gaps often carry predictive signal, but are typically treated as nuisance, obscuring meaningful patterns in their timing and duration. 2. Loss of intra-day detail: Fine-grained mHealth data are often reduced to pre-specified daily or weekly summaries, discarding rich intra-day information with potential predictive value. 3. Population heterogeneity: Models trained on populations often perform poorly for underrepresented groups and fail to generalize to individuals, especially when only limited data are available per person. To address these gaps, we propose developing a robust methodological framework for predictive modelling using irregularly collected EHR, PRO, and mHealth data that improves upon imputation-based standard of care methods. In Aim 1, we will develop univariate and multivariate longitudinal models that account for delays between predictors and outcomes and incorporate detailed intra-day mHealth patterns using distributional learning with low-dimensional, near-lossless embeddings. In Aim 2, we address population heterogeneity by personalizing prediction through an embedding-based approach using landmark multidimensional scaling (MDS) and transfer learning, with reweighting of MDS landmarks to improve performance for underrepresented subgroups. In Aim 3, we validate these methods across diverse mHealth and EHR datasets, including NIH All of Us, UK Biobank, and other disease-agnostic and -specific retrospective and prospective datasets, using mixed-methods studies among clinicians to operationalize model outputs for clinical decision support. Though broadly applicable to multi-modal longitudinal data of all types and a range of disease settings, we focus on four chronic conditions with high clinical impact: cancer, congestive heart failure, diabetes, and obesity. The success of this project and its open-source tools will help close a critical methodological gap and enable effective use of multi-modal longitudinal data to improve clinical decision-making for chronic disease management.

Up to $1.5M
2030-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

A multi-agent AI system for automated curation and functional annotation of enzymes in human gut microbiome

open

NLM - National Library of Medicine

PROJECT SUMMARY Human gut microbiomes influence health by producing metabolites and enzymes that modulate immunity, transform drugs, and digest nutrients. However, most of these enzymes remain functionally unknown. Current annotation tools rely mainly on sequence similarity searches, which can only assign meaningful functions to less than 30% of microbial proteins. Although recent approaches incorporate protein language models and structural comparison, they still rely on predefined pipelines, manual literature or database searches, and specialized expertise in microbial research. This makes the annotations time-consuming without intelligent automation for context-aware insights and limits their scalability across diverse microbial ecosystems. Large language models (LLMs) have emerged as powerful tools in scientific research by analyzing data, answering complex questions, and generating new hypotheses. Building on these strengths, Artificial Intelligence (AI) agents, which combine LLMs with external resources like databases, tools and APIs, can automate tasks and workflows, mimicking human expert decision-making. Although they are widely used in industry, their potential in bioinformatics has only recently been explored. The overall objective of our project is to develop GENZ-AI (Gut ENZyme AI), a multi-agent AI system for automated curation and functional annotation of gut microbial enzymes. GENZ-AI will leverage LLM and advanced AI agents to autonomously delegate tasks, integrate diverse data sources, and deliver enriched annotations with relevant references. We will use advanced techniques, such as prompt optimization and imitation learning, to continuously refine its performance based on real-world annotation sample workflows and user feedback. The significance of GENZ-AI lies in leveraging these cutting-edge technologies to automate and enhance the data curation and workflow organization for enhanced enzyme annotation. This achievement will also improve gut microbiome-based diagnostics and therapeutics (e.g., dietary interventions, drug enhancement, immune modulation) while substantially reducing the time and effort required. The outcome will be a set of novel computational approaches implemented as user- friendly, reusable, open-source tools, including specialized applications for CAZymes, a class of glycan- metabolism enzymes critical to gut microbiome functions. The CAZyme annotation results and software tools will be integrated into dbCAN-PUL and dbCAN-sub databases. The key innovations of this project include a structure-informed protein language model for generalized EC number prediction, the application of CrewAI framework to build a multi-agent system optimized for enzyme annotation in microbiome, and the in-depth investigation of CAZyme and its glycan substrate utilization through GENZ-AI. The broader impact extends beyond the human gut microbiome, as GENZ-AI can be applied to any microbes, providing a scalable solution for diverse microbial ecosystems and pioneering the adaptation of LLM-powered AI agents in bioinformatics.

Up to $1.3M
2030-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Data-driven discovery of drug response mechanisms through biologically informed deep learning and large language models

open

NLM - National Library of Medicine

Summary/Abstract Developing more effective therapies requires a comprehensive understanding of the molecular mechanisms that govern drug responses. In pursuit of this goal, large-scale pharmacogenomic resources have been generated, capturing baseline multi-omics, systematic viability screens of drug and gene perturbations, and transcriptomic signatures of treatment-induced changes. Although each dataset reflects an important aspect of treatment response, these resources remain siloed across modalities and underutilized due to the lack of integrative computational frameworks. Building on our prior work in deep learning and bioinformatics tool development, we propose to address two critical gaps: i) the need for a unified framework to systematically integrate multi-modal pharmacogenomic data for modeling the drug–gene–pathway–response axis; and ii) the need to make these tools and data resources more accessible to biomedical researchers without programming expertise. Our central hypothesis is that biology-guided, multi-modal integration using deep learning will enable accurate prediction and interpretation of drug effects, from molecular perturbations to phenotypic outcomes. We will develop a novel deep learning architecture that uses transfer learning to combine knowledge from transcriptomics, drug features, drug and CRISPR screens, and perturbation signatures. The model will bridge drug-induced molecular changes and phenotypic viability effects through pathway-level representations, enabling mechanistic insight and generalizable prediction of drug responses across a wide range of biological contexts (Aim 1). To complement this systems-level model and enhance its real-world applicability, we will develop a scalable computational framework for inferring and evaluating drug mechanisms directly from transcriptomic perturbation signatures, leveraging embeddings derived from large language models. This approach enables gene set-free discovery of both known and novel drug mechanisms, particularly in under-annotated or complex settings (Aim 2). All models and findings will be rigorously validated using independent datasets. To maximize impact and accessibility, we will develop a user-friendly web platform that integrates these tools and data resources, allowing users to submit their own data, explore predictive outputs, and visualize pathway- and mechanism-level interpretations, without requiring programming expertise (Aim 3). Proposed in response to PAR-25-238, this study aligns closely with the National Library of Medicine’s mission to advance data-driven discovery in biomedical science. The project is supported by a multidisciplinary team with expertise in bioinformatics, pharmacogenomics, artificial intelligence, and software development. Successful completion will yield: i) the first deep learning framework to comprehensively model the drug–gene–pathway–response axis through systematic integration of multi-modal pharmacogenomic data; ii) deeper biological knowledge of how drugs drive outcomes; iii) a novel methodology for high-resolution, gene set-free drug mechanism discovery; and iv) an open-access platform that democratizes advanced pharmacogenomic modeling across a broad range of diseases.

Up to $358K
2030-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Eligibility Criteria Linking for Integrated Patient Selection and Emulation (ECLIPSE)

open

NLM - National Library of Medicine

ABSTRACT Clinical trials are essential to evidence generation, yet strict and complex eligibility rules often exclude large segments of real-world patients, limiting generalizability, delaying enrollment, and increasing costs. Target-trial emulations (TTEs) using real-world data offer a complementary approach for evidence generation, allowing investigators to test generalizability of clinical trials, assess how many patients may be eligible for hypothetical trials, and inform the design of future trials. However, TTEs are constrained by a core informatics challenge: most eligibility criteria are not end-to-end computable in structured and unstructured electronic health records (EHRs). Critical eligibility criteria, such as recurrence/progression, performance status, biomarker qualifiers, and substance-use patterns, are often embedded in free-text or are inconsistently encoded, undermining feasibility assessments and reproducible emulations. This project will develop and validate ECLIPSE (Eligibility Criteria Linking for Integrated Patient Selection and Emulation), a standards-aligned framework that compiles full protocol criteria into executable, auditable artifacts applied to structured and unstructured EHR data. Aim 1 will translate protocol eligibility text into computable phenotypes and query packages anchored to standard vocabularies; and validate criterion-level correctness and cohort concordance against expert references and prescreening logs. Aim 2 will build modular LLM pipelines for extracting eligibility criteria typically found in clinical notes (e.g., progression, performance status, biomarker qualifiers, substance use disorder indicators) with schema-constrained decoding and quantified uncertainty to enable auditable integration. Aim 3 will implement ECLIPSE across clinical trials in 3 conditions – kidney cancer, heart failure, and opioid use disorders – at three health systems in order to evaluate overlap with traditionally curated and enrolled cohorts, stability of effect estimates in TTEs, and operational gains in trial feasibility assessments. We will apply post-prediction inference (methods that adjust extracted variables for remaining extraction errors prior to analysis) so that feasibility counts and causal estimates reflect corrected inputs. The study leverages retrospective cohorts (~160,000 patients) across an academic network and a safety-net hospital, with pre-specified quantitative targets (e.g., ≥0.80 criterion-level agreement; ≤10% relative difference in effect estimates; ≥20% reduction in feasibility assessment time). Deliverables include open, versioned phenotype definitions and value sets, an implementation guide for OMOP and FHIR artifacts, validated extraction modules, and reproducible packages to execute artifacts where data reside. By making eligibility criteria computable, auditable, and portable across OMOP and FHIR, ECLIPSE will share interoperable methods, accelerate study feasibility assessment and emulation, and improve representativeness of trial populations, resulting in reusable, high-quality biomedical data resources for the research community.

Up to $349K
2030-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Developing Bioinformatics and Computational Biology Methods for Integrative Structure-Function Analysis of CRISPR Base-Editor TilingMutagenesis Experiments

open

NLM - National Library of Medicine

PROJECT SUMMARY This proposal aims to develop computational methods for the integrative protein structure-function analysis of base-editor (BE) tiling library and screening results. BE tiling mutagenesis is a high-throughput gene-editing technology to introduce transition variants (A→G and C→T) in cellular models, enabling interrogation of splice-altering, nonsense, and missense mutations. In vitro and in vivo BE experiments are increasingly employed by molecular and chemical biologists to study the mechanisms of mutations and therapeutic targets. However, the scientists face challenges to interpret the experimental readouts, accounting for the molecular effects of these genetic edits, and answer questions such as: “Are the edits unfolding the protein and destabilizing its function?” or “Which edits can recapitulate known drug-resistant mutations or overlap with clinically actionable pathogenic variants?” Currently, answering these questions often requires ad hoc scripting to integrate heterogeneous data types— mRNA-level BE libraries, protein-level mutations, and atomic-resolution structures—and manual queries across disparate databases. These challenges are further amplified when researchers perform multiple screens, such as across different editors and cell lines. Efficient computational methods are needed to harmonize, normalize, and aggregate results from multiple screens, identify a consensus set of hits, and extract meaningful biological insights. To address these gaps, our project will develop methods to link BE library data and screening results to proteins and prioritize and interpret hits in the context of protein sequence-structure-function relationships. The project is organized around two Specific Aims, each with two sub-aims. Aim 1 focuses on building scalable bioinformatics pipelines to link BE tiling data to protein sequences/structures and prioritize “hits” (i.e., sites predicted to cause molecular and functional changes). We will develop tools to map and visualize editing libraries and readouts from mRNA onto protein-level annotations (Sub-aim 1.1), and methods to identify an expanded set of “hits” by leveraging structural biology insights and aggregating results across multiple screens (Sub-aim 1.2). Aim 2 centers on applying data science and AI methods to characterize and interpret BE screening hits using protein features. We will identify key protein features enriched for impactful edits (Sub-aim 2.1) and develop AI- powered models to automate interpretation and summarization of base-editing outcomes using biologically informed reasoning (Sub-aim 2.2). For each aim, we will benchmark our methods using published datasets and expert-provided new datasets. All data and methods will be made publicly available through open-access repositories and, importantly, will be integrated into the web-based Genomics 2 Proteins portal, enabling researchers to interactively explore the data and use the methods for hypothesis generation. This work will produce reusable, generalizable, and scalable tools that advance the application of informatics and data science to accelerate data-powered biomedical discovery.

Up to $387K
2030-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Machine learning and statistical tools for subcellular spatial biology

open

NLM - National Library of Medicine

PROJECT SUMMARY Spatial omics is the new frontier in biotechnology – a series of innovations over the last ten years that give us exquisitely detailed views into molecular events and interactions inside cells, across all cells in a tissue sample. Some of these technologies can reveal a complete map of gene transcripts inside each cell and such “subcellular spatial transcriptomics” (SST) technology has immense and widely recognized potential for biomedical applications. Yet, current uses of this technology typically aggregate the available information at the level of an entire cell, rarely exploring the richness of subcellular information available from the assay. This project's goal is to develop a comprehensive toolkit for analyzing subcellular spatial transcriptomics (SST) data, extracting interpretable biological patterns and testable mechanistic insights into tissue function and pathology. The proposed approach will employ innovative spatial analysis techniques, leveraging state-of-the-art machine learning methods and robust statistical procedures. A major thrust will be on identifying subcellular spatial patterns involving individual genes, gene pairs and modules of genes, while being aware of biological variations from cell to cell. A new functionality in the toolkit will be to quantify changes in genes' subcellular distribution patterns between conditions, paving the way to a novel class of biomarkers. Planned approaches will build on recent publications from the PI's laboratory, improving the statistical power and scalability of state-of-the-art tools and exploring complementary modeling techniques. Another major goal will be to describe the subcellular space in useful ways, such as partitioning a cell's landscape into functionally distinct components, annotating axons and dendrites in brain data, and representing each cell's spatial transcriptome in a format that lends itself to machine learning algorithms. Tools developed for this goal will facilitate more accurate discovery of interpretable spatial patterns, charting of intercellular communication in brain SST data, and machine learning-based characterization of cells, ultimately leading to new ways of describing disease and biological conditions. The third plank of the proposed project is to discover how functional patterns at the subcellular level are encoded in gene sequences. For this task, machine learning tools will be implemented that relate gene sequence patterns to gene transcript distribution inside cells, and the discovered sequence patterns will then point to key regulators of those genes, thus providing potential targets for intervention. All functionalities of the proposed toolkit will be subjected to rigorous testing for robustness and reproducibility, and then applied to SST data sets from diverse biological systems, demonstrating their real-world utility. Furthermore, special attention will be given to software and data sharing, through adherence to “FAIR” (findable, accessible, interoperable, reusable) principles popularized by the NIH. This project will not only establish SST analytics on a firm footing, it will also generalize to other “omics” assays of subcellular resolution, that are under development today.

Up to $327K
2030-05-31
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Sustainable Algorithms for Generalizable and Effective (SAGE) Learning Prediction Systems That Deliver on the Promise of AI for All Patients and Communities

open

NLM - National Library of Medicine

PROJECT SUMMARY / ABSTRACT Temporal changes in clinical practice, patient populations, and information systems degrade performance of artificial intelligence (AI) and machine learning (ML) models. Lack of model generalizability across patient contexts—be they clinical, geographic, or sociodemographic—creates performance gaps that can lead to differences in access to care and unevenly distribute algorithmic benefits. Failing to proactively address performance drift and differences in performance across patient subgroups risks patient safety, undermines user trust, and fails to deliver on the promise of predictive analytics based either traditional ML or novel large language models (LLMs) to improve patient and population health. Learning prediction systems (LPS), an extension of the learning health system paradigm of data-driven continuous improvement, would conduct post- deployment surveillance to collect evidence of model success or deterioration and recommend changes that sustain prospective model performance, both overall and within patient subgroups. Existing model maintenance methods focus on population-level performance and have yet to explore how model updates may generalize across variable patient contexts and impact subgroup performance. Our central objective is to design LPS that promote sustainable AI/ML decision support tools (AI-DST) to ensure all patients benefit from the AI-enabled transformation of healthcare. Using data from the Department of Veterans Affairs (VA) and Vanderbilt University Medical Center (VUMC), we will apply complex simulation studies and real-world evaluations across 3 clinical domains to advance novel LPS metrics and methods that sustain generalizable and effective delivery of clinical AI/ML. In Aim 1, we characterize the complex spectrum of temporal changes in performance gaps between patient contexts and whether/how model updating practices and learning algorithms (ML and LLMs) impact performance gap drift. In Aim 2, we develop and benchmark novel methods to characterize and detect exacerbation of performance gaps between patient subgroups, as well as updating strategies that foster restoration of overall and subgroup performance. In Aim 3, we extend LPS methods in support of small patient subgroups, such as those served by rural healthcare centers, where limited sample sizes present unique challenges to effectively monitoring performance, detecting deterioration, and training updates. In Aim 4, we disentangle changes in outcomes associated with effective AI-DST from dataset shift, establishing LPS methods to monitor and update models while accounting for feedback interference, including monitoring of decision consistency across patient contexts. We will evaluate these methods prospectively in a deployed AI-DST preventing postpartum hemorrhage at VUMC. With expertise in informatics, AI, data science, ethics, and clinical care, our team is well-positioned to have a major impact on the development and implementation of LPS that sustain AI-DST for all patients, propelling the adoption of responsible AI/ML.

Up to $1.6M
2030-06-30
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Identify functional modules in spatial omics atlas with cellular community motifs

open

NLM - National Library of Medicine

Emerging spatial omics atlases, including Human Cell Atlas, HubMAP, and Human Tumor Atlas, open new opportunities to investigate the relationships between cell spatial arrangement and tissue functions in various biological systems and diseased microenvironments. However, topological coordinating rules among different cell types, such as tissue spatial patterns in functional modules, are still under-investigated as the general unit across various tissues and diseases. Different from clustering cell type composition from classical top-down multicellular neighborhood analysis, bottom-up methods formulate the cell organization from Cellular Community (CC) motifs as discrete conservative patterns of recurring interconnections of various cell types. We hypothesize that CC motifs surrogating functional modules serve as precise modular markers, enhancing interpretability, generalizability, and sensitivity in comparative analyses and facilitating detailed mechanistic understanding in spatial omics studies, including spatial transcriptomics, proteomics, epigenomics, and histopathological images. In this project, we propose building a scalable and generalizable computational framework to comprehensively study conservative spatial organizations as functional modules, alongside an AI-innate web portal to quantitatively inquire, analyze, and curate the knowledge on functional modules in atlas-level spatial omics studies. We first analyze spatial multicellular neighborhoods by proposing a computationally efficient search algorithm CC-index identifying various conservative CC motifs as functional modules. In the triangulated tessellation space from spatial omics at the atlas level, the proposed approach is specially optimized to make precise multi-size CC motif identification computationally feasible (Aim 1). These functional modules facilitate further spatial omics research expanding across multiple markers, samples, modalities, longitudes, resolutions, and scales. They can be further generalized to compare different biological statuses within multiple markers in continuous transcriptomics and proteomics expression. These functional modules will be used to facilitate various methodological and biomedical applications in spatial omics research and different applications in spatial omics research (Aim 2). A large language model (LLM)-powered web portal, FMportal, will provide researchers with a user-friendly interface to query, analyze, and organize functional modules through a vector database across spatial omics atlases. This AI-innate system includes natural language interaction and empowered information summarization with RAG and reasoning (Aim 3). All the generated results and knowledge will be carefully evaluated at multiple levels and through validations. Enabled by fast accumulating data, this work is likely a game changer, transforming spatial omics data into functional modules that reveal the broad and fundamental roles of spatial organization in tissue differentiation, organ development, disease progression, and drug response across diverse biomedical and pathological contexts.

Up to $1.4M
2030-06-30
health research

Free to search & build · $99 one-time to unlock the application pack · No subscription

Find grants matched to your organization

Answer a short questionnaire and get a personalized ranked list of grants you qualify for, with fit scores and application guidance.

Get Your Matches

Free to search · No account required