A Living Benchmark for Information Retrieval from Electronic Health Records
As hospitals increasingly integrate large language models (LLMs) into electronic health record (EHR) systems to help clinicians summarize patient histories and retrieve medical information, the need for reliable evaluation has become critical. However, current benchmarks for testing these tools are manually created, expensive to maintain, and quickly become outdated. This paper introduces a scalable, automated framework that generates question-answer pairs from patient records, allowing for a "living" benchmark that can be continuously updated to keep pace with evolving medical documentation and new technology. The ai agents story also surfaces in Google launches Gemini 3.8 Flash and..., adding another angle.
Creating a Living Benchmark
The researchers developed a two-stage generation process to create the Benchmark for Retrieving Information in EHRs (BRIE). First, an LLM extracts and de-duplicates key facts from a patient’s longitudinal medical notes, filtering out redundant information often caused by "copy-forward" documentation. Second, these facts are used to generate specific clinical questions and answers. To ensure the benchmark is accurate and clinically relevant, nineteen physicians reviewed and validated the generator. Because the generator itself is validated, it can be applied to new patient data over time, preventing the benchmark from becoming a static, obsolete artifact.
Evaluating Clinical LLMs
Using BRIE, the researchers tested nine state-of-the-art LLMs and five different inference strategies, such as retrieval-augmented generation (RAG) and agentic approaches. The evaluation focused on how well these models could retrieve and synthesize information across complex, multi-encounter patient records. The study found that while the models rarely hallucinated (fabricated information), they frequently suffered from "omission," meaning they often failed to surface clinically important facts, especially when the answer required connecting information across multiple documents. The same ai evaluation question is explored in AutoViewMem, which adds a research perspective.
Performance and Efficiency
The study compared different ways of feeding patient data to these models. They found that "Dense" retrieval—a method that ranks document chunks by semantic similarity—consistently outperformed simple keyword-based search and often matched or exceeded the performance of feeding the model the most recent notes. Crucially, this approach was significantly more cost-effective, reducing the number of tokens processed by an average of 69%. By filtering the record to only the most relevant documents, these retrieval methods helped models avoid getting lost in the "noise" of long, complex patient histories, leading to better clinical accuracy in identifying specific test results or prior treatment courses.
Key Takeaways for Deployment
The research highlights that single-reference evaluation—testing a model against only one "correct" answer—systematically underestimates a model's capabilities, as clinical reasoning can vary. By using a validated, scalable generator, the team was able to create multiple reference answers to better reflect real-world clinical judgment. Ultimately, the study demonstrates that as clinical LLMs become more deeply embedded in hospital workflows, healthcare organizations need these automated, refreshable evaluation frameworks to ensure that the tools assisting in patient care remain reliable, accurate, and up-to-date. The ai agents story also surfaces in Librarians launch viral workshops to help..., adding another angle. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!