Back to AI Research

AI Research

A Living Benchmark for Information Retrieval from E... | AI Research

Key Takeaways

  • A Living Benchmark for Information Retrieval from Electronic Health Records As hospitals increasingly integrate large language models (LLMs) into electronic...
  • Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records.
  • Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements.
  • We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes.
  • Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset.
Paper AbstractExpand

Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.

A Living Benchmark for Information Retrieval from Electronic Health Records
As hospitals increasingly integrate large language models (LLMs) into electronic health record (EHR) systems to help clinicians summarize patient histories and retrieve medical information, the need for reliable evaluation has become critical. However, current benchmarks for testing these tools are manually created, expensive to maintain, and quickly become outdated. This paper introduces a scalable, automated framework that generates question-answer pairs from patient records, allowing for a "living" benchmark that can be continuously updated to keep pace with evolving medical documentation and new technology. The ai agents story also surfaces in Google launches Gemini 3.8 Flash and..., adding another angle.

Creating a Living Benchmark

The researchers developed a two-stage generation process to create the Benchmark for Retrieving Information in EHRs (BRIE). First, an LLM extracts and de-duplicates key facts from a patient’s longitudinal medical notes, filtering out redundant information often caused by "copy-forward" documentation. Second, these facts are used to generate specific clinical questions and answers. To ensure the benchmark is accurate and clinically relevant, nineteen physicians reviewed and validated the generator. Because the generator itself is validated, it can be applied to new patient data over time, preventing the benchmark from becoming a static, obsolete artifact.

Evaluating Clinical LLMs

Using BRIE, the researchers tested nine state-of-the-art LLMs and five different inference strategies, such as retrieval-augmented generation (RAG) and agentic approaches. The evaluation focused on how well these models could retrieve and synthesize information across complex, multi-encounter patient records. The study found that while the models rarely hallucinated (fabricated information), they frequently suffered from "omission," meaning they often failed to surface clinically important facts, especially when the answer required connecting information across multiple documents. The same ai evaluation question is explored in AutoViewMem, which adds a research perspective.

Performance and Efficiency

The study compared different ways of feeding patient data to these models. They found that "Dense" retrieval—a method that ranks document chunks by semantic similarity—consistently outperformed simple keyword-based search and often matched or exceeded the performance of feeding the model the most recent notes. Crucially, this approach was significantly more cost-effective, reducing the number of tokens processed by an average of 69%. By filtering the record to only the most relevant documents, these retrieval methods helped models avoid getting lost in the "noise" of long, complex patient histories, leading to better clinical accuracy in identifying specific test results or prior treatment courses.

Key Takeaways for Deployment

The research highlights that single-reference evaluation—testing a model against only one "correct" answer—systematically underestimates a model's capabilities, as clinical reasoning can vary. By using a validated, scalable generator, the team was able to create multiple reference answers to better reflect real-world clinical judgment. Ultimately, the study demonstrates that as clinical LLMs become more deeply embedded in hospital workflows, healthcare organizations need these automated, refreshable evaluation frameworks to ensure that the tools assisting in patient care remain reliable, accurate, and up-to-date. The ai agents story also surfaces in Librarians launch viral workshops to help..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!