Cleveland Clinic / Case Western (Nakashima et al.)
Vision-language model that jointly embeds a cardiac MRI study, treated as video, with the impression section of its clinical report. Combines a video encoder over cine/LGE frame sequences with a Bio+ClinicalBERT text encoder using CLIP-style contrastive training. Supports zero-shot and few-shot classification of cardiomyopathies, amyloidosis, and LV dysfunction, plus image/report retrieval and structured report drafting. Trained on a private, single-institution corpus of roughly 11,000-14,000 CMR study-report pairs from Cleveland Clinic and Case Western.
Architecture
Hybrid
Video encoder over CMR cine/LGE frame sequences + Bio+ClinicalBERT text encoder (impression section), CLIP-style contrastive pretraining
Framework
PyTorch
Added to catalog
2026-07-10
MIT
License for model weights only. Associated code may be licensed seperately, check code source for specific terms.
Single-Institution CMR Study-Report Corpus (CMR-CLIP)
12,500 patients (14,214 paired CMR study-report pairs used for training; 41,936 studies pre-exclusion); Cleveland Clinic, Ohio
Disease/function screening (avg AUC)
Non-ischemic cardiomyopathy (zero-shot)
Ischemic cardiomyopathy (zero-shot)
Cardiac amyloidosis (zero-shot)
LV systolic dysfunction (zero-shot)
LV dilation (zero-shot)
LV hypertrophy (zero-shot)
Embedding