University of Toronto (McIntosh Lab)
Distills knowledge from EchoCLIP, a vision-language echocardiography model, into ECG embeddings, aiming to improve how well ECG signals alone can predict echo-derived measures of cardiac function. Combines a 1D ECG encoder with a BioBERT text encoder under a probabilistic cross-modal embedding objective that captures uncertainty. Published at MICCAI 2025 by the University of Toronto's McIntosh Lab.
Architecture
Hybrid
1D ECG encoder + BioBERT text encoder with Probabilistic Cross-Modal Embeddings (PCME++); knowledge-distilled from the EchoCLIP vision-language teacher model
Framework
PyTorch
Added to catalog
2026-07-10
CC BY-NC-ND 4.0
License for model weights only. Associated code may be licensed seperately, check code source for specific terms.
EchoCLIP paired ECG / echocardiogram-report cohort (distillation source)
Paired ECG and echocardiogram-report data used for cross-modal knowledge distillation from EchoCLIP; the underlying source cohort is not fully disclosed.
Embedding