Cedars-Sinai Medical Center (Smidt Heart Institute) / Ouyang Lab
Vision-language foundation model fine-tuned from CLIP on more than one million private echocardiogram video-report pairs, enabling zero-shot cardiac function assessment, device identification, and image/text retrieval without task-specific training. Combines a ConvNeXt-Base video encoder with a GPT-2-style text encoder under contrastive pretraining. Training data is private, but model weights and code are public. Developed by Cedars-Sinai's Ouyang lab.
Architecture
Hybrid
ConvNeXt-Base CLIP-style vision encoder + GPT-2 BPE text tokenizer/transformer; contrastive image-text pretraining; EchoCLIP-R variant uses a custom long-context report tokenizer
Framework
PyTorch
Added to catalog
2026-07-10
Research use only
License for model weights only. Associated code may be licensed seperately, check code source for specific terms.
10,030 deidentified apical-4-chamber echo videos from Stanford Health Care. Source reports age and sex breakdowns.
LVEF (zero-shot)
Embedding