CVAI Catalog

·

View Catalog

EchoPrime

Cedars-Sinai Medical Center (Smidt Heart Institute) / Ouyang Lab

Echocardiography video

Filter catalog by Modality:
Echocardiography

Clinical text

Filter catalog by Modality:
Text & EHR

Multimodal

Filter catalog by Modality:
Multimodal

General Purpose / Multi-task

Filter catalog by Disease / Trait:
General / Foundation

Retrieval

Filter catalog by Task Type:
Representation Learning

Hybrid

Filter catalog by Architecture:
Hybrid / Multi-branch

PyTorch

Filter catalog by Framework:
PyTorch

Research use only

Filter catalog by License:
Non-commercial / Research-only

Vision-language foundation model that interprets an entire transthoracic echocardiogram study rather than a single view or video: it classifies the view type of every clip, applies view-informed attention across the full study, and generates or retrieves comprehensive study-level interpretations in English or Italian. Pretrained on a private Cedars-Sinai corpus of 12 million echo video-report pairs. Developed by the Smidt Heart Institute and Stanford's Ouyang lab.

memory Specifications

category

Architecture

Hybrid

Multi-video, view-informed vision-language model: per-view video encoders with contrastive text alignment, plus view-classification and anatomic-attention pooling across a full echocardiographic study

code

Framework

PyTorch

calendar_month

Added to catalog

2026-07-10

gavel License

Research use only

License for model weights only. Associated code may be licensed seperately, check code source for specific terms.

description Publication

database Training & evaluation data

Cedars-Sinai Echo Video-Report Corpus (EchoPrime)

train

USA

12 million video-report pairs; unique patient count not stated in source paper

science Capabilities & performance

Study-level echocardiogram interpretation retrieved/generated from a view-informed embedding of the full study

Retrieval General Purpose / Multi-task