Back to AI Research

AI Research

A report-grounded vision-language foundation model... | AI Research

Key Takeaways

  • EndoCLIP is a vision-language foundation model designed to bridge the gap between routine colonoscopy reports and the visual data captured during procedures....
  • Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports.
  • These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images.
  • Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records.
  • On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists.
Paper AbstractExpand

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.

EndoCLIP is a vision-language foundation model designed to bridge the gap between routine colonoscopy reports and the visual data captured during procedures. While colonoscopy reports contain detailed clinical information about lesion size, location, and appearance, they typically summarize an entire procedure rather than linking specific descriptions to individual video frames. EndoCLIP addresses this by using a progressive training approach to recover the correspondence between clinical text and specific images, allowing the model to perform tasks like lesion retrieval and classification using natural language.

Recovering Clinical Correspondence

The researchers developed a three-stage training process to align text with images from 280,476 routine colonoscopy records. First, the model localizes visual evidence of polyps within a case. Second, it selects "anchors" from single-lesion cases to establish a baseline for descriptive semantics. Finally, it uses these learned associations to disambiguate findings in complex, multi-lesion procedures. This process successfully generated 125,756 lesion-level image-text pairs, which were then used for contrastive pretraining.

Clinical Performance and Benchmarking

EndoCLIP was evaluated against general-purpose and biomedical vision-language encoders across several clinical tasks, including lesion-level retrieval, structured report generation, and six multi-center classification tasks. In zero-shot and linear-probe settings, EndoCLIP consistently outperformed existing models. Notably, in a blinded study involving 12 endoscopists, the model’s linear probe achieved performance comparable to expert readers in distinguishing between benign and malignant lesions. Furthermore, the model demonstrated the ability to generate structured reports—such as lesion diameter and Paris type—directly from frozen visual features.

Why This Matters

This research suggests that routine clinical documentation can be transformed into scalable supervision for AI models. By enabling clinical targets to be specified in language, EndoCLIP reduces the need for manual, task-specific annotations for every new clinical objective. This approach allows for a more flexible, language-driven interface for endoscopy tasks, potentially improving the efficiency of AI-assisted systems in clinical settings.

Limitations and Considerations

The researchers note that routine reports do not provide a direct, frame-by-frame link to images, which necessitated the development of their multi-stage recovery process. Additionally, while the model showed strong performance, the authors acknowledge that the effectiveness of this approach relies on the quality and consistency of the underlying routine documentation. The study also highlights that while EndoCLIP leads in many metrics, certain comparators performed better on specific, well-defined tasks, such as specific morphology classifications.

Comments (0)

No comments yet

Be the first to share your thoughts!