Back to AI Research

AI Research

Synthetic Hospital: An Open, Verifiable, Physician-... | AI Research

Key Takeaways

  • What the paper is about Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them...
  • Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce.
  • Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented.
  • We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers.
  • Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface).
Paper AbstractExpand

Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.

What the paper is about

Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance. The same ai evaluation question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective.

What it covers

Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark Christine Park, Valerie Chen, Tim Dettmers Carnegie Mellon University Abstract Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient’s longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance. † † footnotetext: Code and data: https://github.com/sparkcpark/synthetic_hospital 1 Introduction Clinical AI has advanced rapidly on medical-knowledge benchmarks ( Singhal et al., 2023 ; Nori et al., 2023 ; Nori et al., 2024 ) , yet those gains have not translated into routine use in clinical workflows ( Gong et al., 2025 ) . A central reason is that the benchmarks used to develop and compare models do not reflect the settings in which clinical AI systems must operate. Real clinical work requires reasoning over patient records that span many encounters over months or years ( Fleming et al., 2024 ; Adams et al., 2025 ) , where electronic health records (EHR) are sparse and incomplete, may contain errors or inconsistencies ( Hripcsak & Albers, 2013 ; Weiskopf & Weng, 2013 ; Bell et al., 2020 ) , and clinically relevant evidence is distributed across notes, laboratory results, imaging reports, and other data stored in variable formats ( Rajkomar et al., 2018 ; Fleming et al., 2024 ) . However, most widely used public medical benchmarks instead evaluate isolated multiple-choice questions, short-answer vignettes, or single-note tasks ( Nori et al., 2024 ; Wang et al., 2025 ; Singhal et al., 2023 ; Nori et al., 2023 ; Hendrycks et al., 2021 ) . While these benchmarks are useful for measuring medical knowledge, they are a poor match for the longitudinal systems clinicians require in practice. However, using real EHR data as a starting point for a benchmark fails in two ways: it cannot be shared openly and it cannot consistently supply verifiable ground truth. On the first, privacy regulation, institutional review board (IRB) review, data use agreements (DUAs), de-identification requirements, and institution-specific governance all stand between a clinical dataset and an open benchmark, and even the de-identified EHR corpora that do exist are typically gated, narrow in scope, and difficult to redistribute ( Johnson et al., 2023 ; Wornow et al., 2023 ) . On the second, even with access to a real EHR, the chart reflects what was documented rather than the patient’s complete clinical picture, including all relevant findings and their temporal relationships ( Hripcsak & Albers, 2013 ) . Documentation errors ( Bell et al., 2020 ) , missing notes ( Weiskopf & Weng, 2013 ) , and inconsistent coding ( O’Malley et al., 2005 ) make it difficult to tell whether a model reasoned incorrectly or whether the record itself was incomplete ( Alaa et al., 2025 ) , so the scoring that evidence retrieval, diagnosis, and longitudinal synthesis require has limited reliable referent. In this paper, we introduce the Synthetic Hospital , an open, fully synthetic, knowledge-grounded longitudinal EHR benchmark that resolves both barriers. Although the patients are synthetic, their diagnoses, findings and clinical relationships are derived from medical educational sources and mapped to standard clinical ontologies, providing traceable provenance for the constructed patient state. Synthetic Hospital is built entirely from publicly available medical-education material containing no protected health information. Figure summarizes the construction pipeline. Because every diagnosis, finding, and temporal relation is derived from source medical-education material, mapped to standard clinical ontologies (ICD-10-CM, SNOMED CT, and LOINC), and retained with its provenance (each encounter traces to a distinct source case, and no two patients share source material), the benchmark provides explicit ground truth for the constructed patient state and evaluation tasks. Of note, the record shown to a model still reads like a chart; it is the ground truth behind the record, not the record itself, that is complete. A key concern for synthetic clinical data is realism, where rule-based approaches to generating data may lack the narrative and longitudinal complexity needed to evaluate frontier models ( Walonoski et al., 2018 ) . For an evaluation benchmark, our primary concern is whether individual records constitute plausible, clinically coherent test cases rather than whether the synthetic population reproduces real-world epidemiology. We evaluate this record-level realism directly: in blinded review, licensed physicians distinguished synthetic from real patient records only at near-chance rates ( 53 % 53% ; Section 4.1 ). We separately characterize distributional properties in Appendix H , showing that Synthetic Hospital preserves the education-derived case mix of its source corpus and clinically expected comorbidity structure while, by design, differing from a population-calibrated cohort. Beyond realism, two findings characterize the benchmark. First, the knowledge graph recovers 111 of 119 (93%) physician-specified clinical relationships; the eight missed relationships are real associations that are neither encoded in the ontologies nor visible through shared clinical findings and are documented as accepted gaps. Second, across 10 models and four clinical tasks, no model approaches ceiling performance and leadership varies across tasks, with the best model reaching only 0.73 severity-weighted F1 on longitudinal patient diagnosis and roughly 0.5 finding-level F1 on whole-patient summarization and imaging indication. Together, these results show that Synthetic Hospital enables open, verifiable evaluation of clinical AI across diagnostic reasoning, longitudinal synthesis, and retrieval. 2 Related Work Benchmarks for clinical AI. The most widely used medical benchmarks are single-vignette multiple-choice or short-answer datasets ( Jin et al., 2021 ; Jin et al., 2019 ; Pal et al., 2022 ; Singhal et al., 2023 ; Hendrycks et al., 2021 ; Tsatsaronis et al., 2015 ; Vilares & Gómez-Rodríguez, 2019 ; Kweon et al., 2024 ; Bae et al., 2023 ; Zhou et al., 2025 ) . Frontier models saturate these formats yet falter on practice tasks ( Gong et al., 2025 ) , and the multiple-choice format itself inflates apparent competence ( Griot et al., 2025 ; Singh et al., 2025 ; Cocchieri et al., 2026 ; Ma et al., 2025 ; Alaa et al., 2025 ) . Practice-oriented benchmarks add rubric-scored conversations ( Arora et al., 2025 ) , clinical-NLP task suites ( Wu et al., 2025 ) , simulated or conversational diagnosis ( Schmidgall et al., 2024 ; Schmidgall et al., 2026 ; Nori et al., 2025 ; Tu et al., 2024 ; Li et al., 2024 ; Fan et al., 2025 ) , multi-stage reasoning ( Qiu et al., 2025 ) , and emergency-room workflows ( Mehandru et al., 2025 ) ; focused benchmarks score note generation ( Yim et al., 2023 ) , problem-list summarization ( Gao et al., 2023 ) , structured querying ( Lee et al., 2022 ) , and clinical summarization ( Van Veen et al., 2024 ) , and MedHELM organizes existing benchmarks into a clinician-validated taxonomy without adding longitudinal chart tasks ( Bedi et al., 2026a ) . At the other extreme, benchmarks on de-identified real records ( Johnson et al., 2016 ; Johnson et al., 2023 ; Wornow et al., 2023 ; Cui et al., 2025 ; Zhao et al., 2025 ; Huang et al., 2023 ; Rajkomar et al., 2018 ) , including the clinician-written instruction benchmark MedAlign ( Fleming et al., 2024 ) , are clinically realistic but gated by credentialing and DUAs, and their ground truth is only what was charted, so a model failure cannot be separated from an incomplete record. Table maps these benchmarks across five dimensions of chart-based clinical work ( Sinsky et al., 2016 ; Weed, 1968 ) : no prior benchmark combines open data with coverage of all five, and longitudinal problem-list construction as a scored task remains unoccupied. Agentic and long-horizon EHR environments. A growing line evaluates agents that operate an EHR rather than answer from a curated prompt. FHIR-AgentBench ( Lee et al., 2025 ) and EHRAgent ( Shi et al., 2024 ) run over gated MIMIC data, and MedAgentBench ( Jiang et al., 2025 ) shares our FHIR framing but evaluates API-level task execution over 100 patients. Longer-horizon environments extend this to physician-reviewed workflows in a real-record EHR ( Liu et al., 2026 ) , large-scale interactive SQL and code tasks ( Qiao et al., 2026 ) , staged inpatient decision-making ( Lu et al., 2026 ) , triage conversations ( Zhu et al., 2026 ) , and computer use over clinical and administrative interfaces ( Bedi et al., 2026b ; Yu et al., 2026 ) . All are built on access-restricted real records or evaluate interface operation rather than chart content. Synthetic Hospital is complementary: it offers the same FHIR affordances over 1,268 openly redistributable patients with constructed, verifiable ground truth. Synthetic clinical data. Synthetic record generation is well established but has primarily targeted privacy rather than benchmark construction, from GAN-based structured records ( Choi et al., 2017 ) through neural generation of shareable notes ( Melamud & Shivade, 2019 ) and synthetic corpora for training openly releasable clinical LLMs ( Kweon et al., 2023 ) . Rule-based simulators such as Synthea ( Walonoski et al., 2018 ) produce standards-compliant FHIR records but structured codes rather than narratives; statistical, autoregressive, and knowledge-grounded trajectory generators ( Yoon et al., 2023 ; Theodorou et al., 2023 ; Pang et al., 2024 ; Zhou et al., 2026 ) achieve high fidelity but are trained on protected records and reproduce only documented observations; and LLM-generated records such as SimSUM ( Rabaey et al., 2024 ) produce fluent narratives without a provenance chain linking statements to underlying clinical facts. LongHealth ( Adams et al., 2025 ) is closest in spirit in constructing fictional patients, but comprises 20 single-encounter multiple-choice cases. Synthetic Hospital instead derives every patient from public educational material and links each diagnosis, finding, result, and narrative statement through a typed knowledge graph to ontology-grounded concepts, so its ground truth does not depend on what happened to be documented (Appendix Table I ). Appendix I gives an extended discussion of each of these lines of work. 3 Benchmark construction We introduce Synthetic Hospital, a deterministic five-stage pipeline that transforms publicly available medical educational resources into a longitudinal EHR benchmark with ontology-grounded provenance. Rather than treating synthetic records as the primary artifact, the pipeline first constructs a structured medical knowledge graph from source material and then renders that graph into realistic clinical documentation. Consequently, every diagnosis, finding, laboratory observation, and benchmark label remains traceable to its originating educational source rather than being inferred from generated text. Throughout this section we follow one released patient (released patient identifier 1973, a 58-year-old man with three encounters) as a running example; Appendices E and F trace it end to end. (1) Knowledge ingestion. We ingest USMLE-style board questions and supplementary medical knowledge from flashcard decks and reference documents. Rule-based parsers convert these heterogeneous sources into a unified structured representation while preserving metadata and provenance. Board questions become candidate clinical encounters, while supplementary fact cards provide supporting knowledge for retrieval, summarization, and diagnosis-finding relationships. (2) Ontology grounding and knowledge graph construction. From each board question, Kimi 2.5 extracts candidate primary, differential, and secondary diagnoses together with typed clinical findings. We then ground these concepts deterministically to ICD-10-CM, SNOMED CT, and LOINC; the LLM-proposed codes are treated only as candidates. The resulting graph links questions to diagnoses and findings, uses a separate LLM pass to type diagnosis–finding relationships, links supplementary fact cards to graph concepts, and segments each vignette into EHR sections while preserving the source text. Every node retains its source identifier and grounding method. For one source question of the running example, an emergency presentation with polyuria, polydipsia, and confusion, this stage yields hyperosmolar hyperglycemic state (validated to E11.01) as the primary diagnosis, type 2 diabetes and acute kidney injury among the secondary diagnoses, and 26 typed findings such as polyuria (symptom, key) and metformin therapy (medication, background). Appendix D provides the full mapping cascades, thresholds, coverage, and graph statistics. (3) Longitudinal patient generation. We assemble board questions into longitudinal patients using deterministic graph clustering before any narrative generation. Encounters may be grouped only if they are demographically compatible, and each added encounter must share a correct or secondary diagnosis with at least one existing cluster member. A constrained greedy clustering algorithm expands these connected groups while placing each source question in at most one patient, yielding 1,268 patients from 5,602 source questions. Kimi 2.5 then realizes each fixed cluster as a longitudinal record by generating a patient profile, encounter timeline, and limited cross-encounter HPI continuity; other clinical content remains grounded in the source vignettes, with problem and medication lists propagated deterministically. In the running example, three emergency and critical-care questions about a 52- and two 58-year-old men (within the 7-year tolerance of the middle-adult bucket) are linked through shared type 2 diabetes and acute kidney injury nodes into one patient; the profile call adds chronic diabetes, hypertension, and stage 3b kidney disease with matching home medications, and the timeline call orders the encounters as perforated appendicitis with septic shock, a hyperosmolar crisis eight months later, and hypercapnic respiratory failure in the ICU at month sixteen, with each later note’s problem list carrying the earlier diagnoses forward. A physician reviews generated records for plausibility and consistency. Thus, narrative generation occurs only after the patient state is fixed and does not determine benchmark ground truth. Appendix E gives the full clustering rules, generation prompts, and running example. (4) Benchmark construction. With the exception of the imaging-indication free-text reference, benchmark labels are deterministic functions of the underlying graph. Patient diagnosis uses the patient’s correct-answer diagnosis nodes and their acuity; evidence retrieval grades chart sections from finding relevance and diagnosis–finding relationships; context summarization scores graph-defined key findings; and specialty-conditioned summarization derives specialty-specific finding relevance from diagnosis ownership and graph relationships. For imaging indication, an LLM generates a terse order and reference clinical question from graph-fixed diagnoses and encounter context, and predictions are scored by clinical-concept overlap rather than surface wording. For the running example, this produces a patient-diagnosis reference of exactly the three acute diagnoses (severity weights 2, 2, and 3), a retrieval query over the patient’s 34 chart sections graded 0–3 (13 highly relevant), a summary reference of 20 key findings, specialty items for endocrinology, general surgery, gastroenterology, and pulmonology plus two absent-specialty abstention items, and an imaging item that pairs the terse order “CT abdomen/pelvis, stat: RLQ pain, fever” with a graph-anchored reference question. Every instance retains its source patient, encounter, and question identifiers. Appendix F provides the complete labeling and scoring rules. (5) Domain-faithful simulation. The completed records are served through a production-style clinical environment implementing FHIR R4 resources, role-based access control, Epic-style workflows, and a function-calling agent interface. Every element exposed through the simulated EHR retains a provenance link back to the underlying ontology-grounded knowledge graph and ultimately to the original educational source material. From these records we construct four longitudinal chart tasks: patient diagnosis, context summarization (including a specialty-conditioned variant), evidence retrieval, and imaging indication (Table 1 ). Patient diagnosis uses chart-neutral scoring: diagnoses documented in the chart but absent from the graph-derived reference are neither credited nor penalized. 3.1 Evaluation splits Evaluation splits. We partition the 1,268 patients at the patient level into three disjoint splits, with every benchmark instance inheriting its patient’s split. The public split contains 200 patients (1,859 instances) and is used for the experiments in Section 5 ; it is stratified by patient difficulty and encounter count to preserve broad clinical coverage. Of the remaining patients, 800 form a training split (7,619 instances) released with full ground truth, and 268 form a held-out split (2,536 instances) whose labels are accessible only through our scorer. Training and held-out patients are stratified by dominant ICD-10 chapter and encounter count and match within one percentage point across strata. Because each source question belongs to only one patient, no chart or source material crosses split boundaries. Diagnoses may recur across splits: 54% of held-out diagnoses also occur in training, while 46% are unseen, enabling evaluation of both patient- and disease-level generalization. The training split additionally supports learning with verifiable rewards. Each task produces a deterministic score in 0 , 1 0,1 from the underlying graph, so rollouts require no additional human labeling. The 7,619 training instances provide distinct starting states across 800 patients; k k rollouts per instance yield 7,619 ​ k 7{,}619k trajectories (approximately 30,000 at k = 4 k{=}4 ). Because additional rollouts do not create new patient states, generalization is evaluated on the patient-level held-out split. Patient diagnosis uses the chart-neutral scoring rule throughout, including when used as a training reward. Task Input Output Primary Metric Instances (public / held-out / train) Patient diagnosis Longitudinal EHR (multi-encounter) Longitudinal problem list (ICD-10 + acuity) Severity-weighted F1 200 / 268 / 800 Summarization Clinical question + EHR sections Structured clinical summary Finding-level F1 200 / 268 / 800 Specialty summarization EHR + target specialty Specialty-focused summary Specialty-relevance F1 983 / 1,325 / 4,037 Evidence retrieval Diagnosis + patient record Ranked evidence passages Precision@5, NDCG@10 200 / 268 / 800 Imaging indication Imaging order with vague indication + EHR Inferred clinical question + pre-read Question concept F1 276 / 407 / 1,182 Table 1: Synthetic Hospital benchmark tasks. For each task, we report the standardized input, expected output, primary evaluation metric, and the number of instances in the public, held-out, and training splits (12,014 in total). Secondary evaluation metrics are provided in Appendix A . EHR = electronic health record; NDCG = Normalized Discounted Cumulative Gain. 4 Assessing benchmark realism Before evaluating clinical AI systems, we first establish that Synthetic Hospital satisfies the two properties motivating its construction: clinically realistic records and verifiable benchmark ground truth. Section 4.1 evaluates whether physicians perceive the generated records as authentic clinical documentation, while Section 4.2 evaluates whether the ontology-derived labels agree with independent physician judgment. 4.1 Realism: synthetic records are indistinguishable from real To assess the realism of Synthetic Hospital, licensed physicians reviewed synthetic and real patient records presented through our Epic-like interface and labeled each chart as synthetic or real, a blinded discrimination design analogous to Turing-test evaluations of LLM text ( Jones & Bergen, 2024 ) and to the expert-review validation used for early synthetic-record generators ( Choi et al., 2017 ) . Because MIMIC-IV uses a markedly different documentation style and schema than Synthetic Hospital, we converted five real MIMIC-IV patients into the Synthetic Hospital schema with a deterministic, rule-based pipeline that normalizes section structure, laboratory and medication formatting, vital-sign rendering, radiology impressions and de-identification artifacts while leaving the clinical content unaltered. Concretely, each admission’s discharge note (with its section labels), admission record, and radiology reports are parsed into the same 18 section types and order used by the synthetic notes; de-identification masks are repaired by shifting MIMIC’s offset calendar into the 2020–2024 window while preserving inter-visit intervals and by substituting consistent fictional names for masked providers, or dropping a clause whose subject was masked; and presentation is re-rendered by fixed tables that expand abbreviated lab panels into one-per-line entries with full names, units, and reference ranges, expand medication frequency codes and tall-man lettering, cast numeric vital-sign strings into sentence form, and reduce multi-paragraph radiology reports to their impression. A final harmonization step applied to both arms removes whole sections that only one arm can carry (social history, which MIMIC redacts, from the synthetic arm; demographics and assessment/plan from the real arm) and drops trailing admissions beyond a shared chart-length budget, always at section or encounter boundaries so that no sentence is cut. The clinical prose itself is not rewritten, so the telegraphic register of real discharge notes is left for the physicians to detect. Without this conversion, physicians would be able to detect real vs synthetic samples from note structure rather than clinical content which would invalidate this evaluation. More importantly, it demonstrates that the benchmark representation is not tied to a single note format: clinical records originating from one hospital system can be normalized into any other structure used in a different hospital system while preserving their clinical content. Representative examples of the original, converted, and fully synthetic records are provided in Appendix G . Each real patient was paired with a synthetic patient matched on clinical domain, sex, age and chart length, yielding a balanced set of five real and five synthetic records with a 50% chance baseline. 10 licensed physicians each reviewed all 10 records in an individually randomized order, yielding 100 judgments. Record-level and distributional realism are distinct: the physician study tests whether individual records are plausible charts, not whether the cohort reproduces population epidemiology, which Synthetic Hospital does not attempt, since its case mix is education-derived by design. Appendix H characterizes this separately, showing that the benchmark preserves the ICD-10 chapter distribution of its source corpus (JSD = 0.029 =0.029 , Spearman ρ = 0.83 \rho=0.83 ) and the expected comorbidity structure while differing, as intended, from a population-oriented Synthea cohort. Synthetic and real records are not reliably indistinguishable. Across 100 judgments, physicians identified whether a record was real or synthetic with 53% accuracy, which did not differ from chance (95% bootstrap CI: 43–63%; two-sided binomial test, p = 0.62 p{=}0.62 ). Performance was similar for both record types: sensitivity for real records was 52% and specificity for synthetic records was 54%, and physicians selected "real" in 49% of judgments, indicating no systematic bias toward either label. Accuracy also did not differ between synthetic (mean 0.540) and real (mean 0.520) records within physicians (paired t ⁡ ( 9 ) = 0.20 t(9){=}0.20 , p = 0.85 p{=}0.85 ). No individual physician performed above chance after correction for multiple comparisons (best: 9/10; Holm-adjusted p = 0.22 p{=}0.22 ), and inter-rater agreement was no better than chance (Fleiss’ κ = − 0.05 \kappa{=}{-}0.05 ), suggesting that physicians did not rely on a consistent shared cue to distinguish the two sources. Every synthetic record was classified as real by at least one physician, and one was classified as real by 7/10 physicians. Finally, greater confidence did not reliably correspond to greater accuracy: judgments made with 80% stated confidence were 52% accurate, while those made with 100% confidence were 70% accurate. Together, these findings suggest that realism arises from the underlying longitudinal clinical representation rather than from reproducing the documentation style of a particular institution. 4.2 Verifiable ground truth: ontology-derived labels agree with physician judgment Realistic records alone are insufficient for a useful benchmark: the reference labels used for evaluation must also correspond to clinically meaningful reasoning. Because Synthetic Hospital generates benchmark labels directly from an ontology-grounded knowledge graph rather than manual annotation, we validate that these automatically constructed relationships agree with independent physician judgment. Ontology-derived relationships recover physician-recognized clinical associations. The specialty-conditioned benchmark decides which findings are relevant to a specialty by following links between diagnoses in the knowledge graph (for example, a renal finding is relevant to cardiology when the patient’s kidney disease is linked to their heart failure). To validate these links, a licensed physician created a pre-registered reference set of 119 clinically required diagnosis relationships (for example, diabetes–chronic kidney disease and hypertension–hypertensive heart disease), assigning each pair an expected relationship type before graph construction. The graph recovered 111 of the 119 relationships (93%): 96 through links derived from SNOMED CT relationships, shared anatomical sites, and shared findings, and 15 through physician-curated edges added for relationships that the physician had classed as definitional or associative but that no SNOMED relationship or shared site encodes (for example, atrial fibrillation–cardioembolic stroke and hyperlipidemia–coronary artery disease). The eight remaining relationships are real but are neither encoded in the ontologies nor visible through overlapping findings, because the two conditions present through disjoint findings: hyperemesis gravidarum and Wernicke encephalopathy share no finding (intractable vomiting versus confusion and ophthalmoplegia), and likewise ankylosing spondylitis and anterior uveitis, or dermatomyositis and occult malignancy; one pair (Stevens–Johnson syndrome and culprit drug exposure) has no diagnosis node for the exposure at all. These were documented as accepted gaps rather than recovered by lowering the shared-finding threshold, which would have admitted many spurious links. This evaluation shows that the graph captures the clinically important relationships needed for benchmark construction. Quantifying false-positive relationships remains future work. 5 Results 5.1 Experimental Setup We evaluate 10 models spanning frontier proprietary systems and open models from 27B to 1T parameters (listed with their sizes in Table 2 ). We controlled for prompting strategy, which can benefit models unequally and have an outsized effect on the evaluation: four strategies (zero-shot, few-shot, chain-of-thought, and ontology-grounded structured prompting) were compared on three pilot models spanning the panel’s strength range, and the best strategy per task was then locked and applied to all 10 models. Appendix C.1 defines the strategies and reports the full ablation (Table C.1 ). 5.2 Single-turn Evaluation Table 2 reports results for all 10 models across the four benchmark tasks. Three findings stand out. First, no model approaches ceiling performance on any task, indicating that the benchmark meaningfully separates systems and that current frontier models re The same ai evaluation question is explored in An Open Pipeline and Dashboard for..., which adds a research perspective. as detailed in the full paper on Arxiv The same ai evaluation question is explored in How Good Are Frontier Models at..., which adds a research perspective.

Comments (0)

No comments yet

Be the first to share your thoughts!