August 6, 2024

Share on

Mendel and UMass Amherst Unveil Groundbreaking Research on AI-driven Hallucination Detection in Healthcare

Innovative framework developed to ensure the accuracy and reliability of AI-generated medical summaries.

San Jose, CA, and Amherst, MA - August 6, 2024 – Mendel, a leader in Clinical AI, and the University of Massachusetts Amherst (UMass Amherst) have jointly published pioneering research addressing the critical issue of faithfulness hallucinations in AI-generated medical summaries. This collaborative effort marks a significant advancement in ensuring the safety and reliability of AI applications in healthcare settings.

Research Overview

In recent years, large language models (LLMs) such as GPT-4o and Llama-3 have shown remarkable capabilities in generating medical summaries. However, the risk of hallucinations—where AI outputs include false or misleading information—remains a significant concern. This study aimed to systematically detect and categorize these hallucinations to improve the trustworthiness of AI in clinical contexts.

The research team developed a robust hallucination detection framework, categorizing hallucinations into five subtypes of medical event inconsistency, incorrect reasoning, and chronological inconsistency. A pilot study of 100 summaries from GPT-4o and Llama-3 models revealed that GPT-4o produced longer summaries (>500 words) and often made bold, two-step reasoning statements, leading to hallucinations. Llama-3 hallucinated less by avoiding extensive inferences, but its summaries were of lower quality. The table below reports the number of model summaries, out of 50 summaries per model, that contains incorrect information according the source medical records:

“Our findings highlight the critical risks posed by hallucinations in AI-generated medical summaries,” said Andrew McCallum, Distinguished Professor of Computer Science, University of Massachusetts Amherst. “Ensuring the accuracy of these models is paramount to preventing potential misdiagnoses and inappropriate treatments in healthcare.”

The study also explored automated detection methods to mitigate the high costs and time associated with human annotations. The Hypercube system, leveraging medical knowledge bases, symbolic reasoning and NLP, played a crucial role in detecting hallucinations. It provided a comprehensive representation of patient documents, aiding in the initial detection step before human expert review.

“We are committed to continually enhancing Hypercube’s capabilities. The future of healthcare AI depends on reliable, accurate tools, and Hypercube’s evolving features, including real-time data processing and adaptive learning algorithms, will keep it at the forefront of clinical innovation,” said Dr. Wael Salloum, Chief Scientific Officer of Mendel AI.

Future Prospects

As AI continues to integrate into healthcare, addressing hallucinations in LLM outputs will be vital. Future research will focus on refining detection frameworks and exploring more advanced automated systems like Hypercube to ensure the highest levels of accuracy and reliability in AI-generated medical content. Hypercube's real-time data processing and adaptive learning algorithms will be essential in maintaining its position at the forefront of clinical innovation.

Accepted Paper

Mendel's work on Hypercube in detecting hallucinations is recognized by the academic community. The research paper is accepted for oral presentation at the KDD AI conference, August 2024: “Faithfulness Hallucination Detection in Healthcare AI. Prathiksha Rumale V*, Simran Tiwari*, Tejas G Naik*, Sahil Gupta*, Dung N Thai*, Wenlong Zhao*, Sunjae Kwon, Victor Ardulov, Karim Tarabishy, Andrew McCallum, Wael Salloum”. It details the methodologies and technologies underpinning Hypercube’s success.

For more information about the Hypercube platform try the Hypercube demo.

About Mendel

Mendel AI supercharges clinical data workflows by coupling large language models with a proprietary clinical hypergraph, delivering scalable clinical reasoning without hallucinations and ensuring 100% explainability. Headquartered in San Jose, California, Mendel is backed by blue-chip investors, including Oak HC/FT and DCM. For more information, visit Mendel or contact marketing@mendel.ai.

About UMass Amherst

UMass Amherst, the flagship campus of the University of Massachusetts system, is a nationally ranked public research university known for its excellence in teaching, research, and community engagement. The university fosters innovation and collaboration across a wide range of disciplines. For more information, visit UMass Amherst.

Press Contact:

Jessica McNellis

Gale Strategies

Jessica@GaleStrategies.com

‍

Exploring the Future of Healthcare AI: A Conversation with Kristin Maloney

The recent podcast featuring Kristin Maloney, hosted on Oncology Data Advisor, delves into Mendel AI's transformative role in healthcare. Kristin highlights how Mendel’s clinical AI solutions—such as Retina, Resolve, and Hypercube—are revolutionizing data-driven decision-making, empowering clinicians to extract critical insights from complex datasets quickly and accurately. Mendel AI's mission is clear: turning unstructured and structured healthcare data into actionable intelligence, bridging gaps in clinical care, and providing physicians with tools to deliver optimal patient outcomes.

Introducing Mendel's New Brand Focus: Supercharging Clinical Data Workflows in Healthcare

Mendel has evolved its brand to “Supercharge Your Clinical Data Workflows,” a shift that reflects our commitment to delivering AI solutions that genuinely enhance clinical data management. In healthcare, where talent shortages demand efficient and reliable tech, our Hypercube solution and neuro-symbolic AI bring unmatched cost-efficiency, speed, and accuracy to workflows. This shift emphasizes our focus on alleviating healthcare’s talent strain with tech that builds trust—eliminating errors and reducing the risk of hallucinations. Discover how Mendel’s transformative approach can optimize your workflows with validated solutions trusted by leaders in the industry.

Revolutionizing Patient Cohort Identification with AI – Insights from Mendel’s ACR Benchmark

Introducing ACR: A New Benchmark for Patient Cohort Retrieval This study introduces Automatic Cohort Retrieval (ACR), a novel task for efficiently identifying patient groups from large-scale medical data. Comparing AI-powered approaches, including large language models and neuro-symbolic systems, the research reveals promising advancements in automating cohort selection for clinical trials and studies. The findings highlight the potential of AI to revolutionize healthcare data analysis, while emphasizing the need for continued improvements in accuracy, efficiency, and reliability.

Introduction to Hypercube’s Ontology and Reasoning Engine

Large Language Models (LLMs) hold the potential to transform healthcare by generating clinical insights and supporting decision-making. However, LLMs face challenges such as hallucinations, lack of explainability, and limited reasoning capabilities, which restrict their effectiveness in clinical settings. Mendel's Hypercube platform addresses these limitations by integrating LLMs with structured clinical ontologies, enhancing both inference and decision-making. Unlike standard ontologies focused mainly on documentation, Mendel’s generative ontology prioritizes scalable reasoning through reductionism and emergentism, enabling more accurate clinical reasoning and streamlined data integration.

Mendel Unveils Groundbreaking Neuro-Symbolic AI System Outperforming GPT-4 for Automatic Cohort Retreival in New Study

“Our latest research at Mendel marks a significant milestone in the field of AI in general, and healthcare in particular,” said Wael Salloum, Cofounder and Chief Science Officer at Mendel. “We are the leader in clinical reasoning by coupling LLMs with our hypergraph reasoning, enhancing both the effectiveness and efficiency of patient cohort retrieval.

Improving Clinical Trial Participant Prescreening With Artificial Intelligence (AI): A Comparison of the Results of AI Assisted vs Standard Methods in 3 Oncology Trials

Delays in clinical trial enrollment and difficulties enrolling representative samples continue to vex sponsors, sites, and patient populations. Here we investigated use of an artificial intelligence-powered technology, Mendel.ai, as a means of overcoming bottlenecks and potential biases associated with standard patient prescreening processes in an oncology setting.

Coupling Symbolic Reasoning with Language Modeling for Efficient Longitudinal Understanding of Unstructured Electronic Medical Records

The application of Artificial Intelligence (AI) in healthcare has been revolutionary, especially with the recent advancements in transformer-based Large Language Models (LLMs). However, the task of understanding unstructured electronic medical records remains a challenge given the nature of the records (e.g., disorganization, inconsistency, and redundancy) and the inability of LLMs to derive reasoning paradigms that allow for comprehensive understanding of medical variables. In this work, we examine the power of coupling symbolic reasoning with language modeling toward improved understanding of unstructured clinical texts. We show that such a combination improves the extraction of several medical variables from unstructured records. In addition, we show that the state-of-the-art commercially-free LLMs enjoy retrieval capabilities comparable to those provided by their commercial counterparts. Finally, we elaborate on the need for LLM steering through the application of symbolic reasoning as the exclusive use of LLMs results in the lowest performance.

How to Approach De-Identification

Organizations that use patient data for internal or external research need to take steps to prevent the exposure of PHI to those who are not authorized to view it. They do this by redacting specific categories of identifiers from every patient document. Once the identifiers are masked, the risk profile of these datasets is significantly reduced. But how do you ensure that redaction engines are working to the highest accuracy?

Clinical Data Abstraction

Clinical Record OCR

PHI De-identification

Clinical Search Engine

Clinical Trial Matching

Clinical Data Assets

Mendel and UMass Amherst Unveil Groundbreaking Research on AI-driven Hallucination Detection in Healthcare

The Feed

Enhancing Oncology Clinical Trial Prescreening at UPenn with Mendel AI

Enhancing Oncology Clinical Trial Prescreening at UPenn with Mendel AI

Exploring the Future of Healthcare AI: A Conversation with Kristin Maloney

Exploring the Future of Healthcare AI: A Conversation with Kristin Maloney

Introducing Mendel's New Brand Focus: Supercharging Clinical Data Workflows in Healthcare

Introducing Mendel's New Brand Focus: Supercharging Clinical Data Workflows in Healthcare

Faithfulness Hallucination Detection in Healthcare AI: Ensuring Reliable Medical Summaries

Faithfulness Hallucination Detection in Healthcare AI: Ensuring Reliable Medical Summaries

Revolutionizing Patient Cohort Identification with AI – Insights from Mendel’s ACR Benchmark

Revolutionizing Patient Cohort Identification with AI – Insights from Mendel’s ACR Benchmark

Introduction to Hypercube’s Ontology and Reasoning Engine

Introduction to Hypercube’s Ontology and Reasoning Engine

Mendel Unveils Groundbreaking Neuro-Symbolic AI System Outperforming GPT-4 for Automatic Cohort Retreival in New Study

Mendel Unveils Groundbreaking Neuro-Symbolic AI System Outperforming GPT-4 for Automatic Cohort Retreival in New Study

Improving Clinical Trial Participant Prescreening With Artificial Intelligence (AI): A Comparison of the Results of AI Assisted vs Standard Methods in 3 Oncology Trials

Improving Clinical Trial Participant Prescreening With Artificial Intelligence (AI): A Comparison of the Results of AI Assisted vs Standard Methods in 3 Oncology Trials

Coupling Symbolic Reasoning with Language Modeling for Efficient Longitudinal Understanding of Unstructured Electronic Medical Records

Coupling Symbolic Reasoning with Language Modeling for Efficient Longitudinal Understanding of Unstructured Electronic Medical Records

How a diagnostic company was able to build a clinico-genomic database in a week

How a diagnostic company was able to build a clinico-genomic database in a week

How One Organization Changed The Way Patients are Identified for Clinical Trials with AI

How One Organization Changed The Way Patients are Identified for Clinical Trials with AI

How to Approach De-Identification

How to Approach De-Identification

Back to Top

Headquarters

Hypercube Copilots

Industry

Privacy & Legal

Company