Back to AI Research

AI Research

LLMs as Post-hoc Auditors of Physiological Plausibi... | AI Research

Key Takeaways

  • LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study This research explores whether Large Languag...
  • Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data.
  • In addition to predictive accuracy, models are appreciated for their potential to provide interpretability, offering explicit equations that relate input variables to outcomes.
  • However, achieving interpretability and plausibility remains challenging, as evolved models may be complex or scientifically inconsistent.
  • In this study, we explore whether Large Language Models, can assist in improving the explainability of Symbolic Regression models generated by evolutionary computation methods.
Paper AbstractExpand

Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data. In addition to predictive accuracy, models are appreciated for their potential to provide interpretability, offering explicit equations that relate input variables to outcomes. However, achieving interpretability and plausibility remains challenging, as evolved models may be complex or scientifically inconsistent. In this study, we explore whether Large Language Models, can assist in improving the explainability of Symbolic Regression models generated by evolutionary computation methods. Building upon our previous work on estimating body fat percentage using grammar-based Genetic Programming , we investigate the use of LLMs as post-processing tools to analyze and rank evolved expressions according to their interpretability and medical plausibility. Four symbolic expressions are analysed by three LLMs over three repeated runs, and the resulting interpretations and rankings are assessed by a panel of three clinicians. Across the three LLMs, comparative model-ranking outputs received more favorable clinician assessments than isolated term-level interpretations. However, the LLMs also produced physiologically and mathematically questionable explanations, indicating that they are better suited to comparative auditing under expert oversight than to autonomous validation.\blfootnote{The present work is an extended version of a paper submitted into a journal.

LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study
This research explores whether Large Language Models (LLMs) can act as a bridge between complex mathematical models and medical experts. Symbolic Regression (SR) is a technique that generates explicit mathematical equations from data, which are often valued for their transparency. However, these equations can sometimes be overly complex or scientifically inconsistent. This study investigates if LLMs can serve as a post-processing tool to analyze these equations, ranking them based on how well they align with medical knowledge and how easy they are for clinicians to interpret. The same large language models question is explored in Molecular Déjà Vu, which adds a research perspective.

The Role of LLMs in Model Auditing

The researchers propose using LLMs as an "expert common-sense filter" to evaluate models generated by evolutionary computation methods. While traditional methods focus on mathematical accuracy, they often fail to account for biological or physiological reality. By providing the LLMs with specific prompts—defining their role as medical interpreters and asking them to explain equations in plain language—the study aims to see if these models can provide meaningful insights that a human doctor can verify.

Methodology and Evaluation

The study utilized a dataset of body fat percentage measurements to generate several candidate equations. These equations were then analyzed by three different LLMs over multiple runs. To validate the findings, a panel of three clinicians reviewed both the mathematical expressions and the explanations generated by the AI. The goal was to determine if the LLM’s reasoning matched the professional judgment of the medical experts. The same large language models question is explored in Rethinking On-Policy Distillation of Large Language..., which adds a research perspective.

Key Findings

The results suggest that LLMs are more effective at comparing different models against each other than at interpreting a single equation in isolation. When asked to rank multiple models, the LLMs provided outputs that were more favorably received by the clinicians. However, the study also identified significant risks: the LLMs occasionally produced explanations that were mathematically or physiologically questionable.

Important Considerations

The authors conclude that while LLMs show promise in assisting with the evaluation of symbolic models, they are not yet reliable enough for autonomous validation. Because LLMs can "hallucinate" or generate false information, they should be used strictly as a tool for comparative auditing under the direct oversight of human experts. This approach is intended to support, rather than replace, the critical role of medical professionals in verifying the scientific validity of AI-generated models. The same large language models question is explored in Making Alternative Data Work, which adds a research perspective. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!