V-JEPA: Learning Video Representations by Feature Prediction
JEPAComputer VisionDeep LearningVideoPython
Follow V-JEPA ViT-L/16 from video tubelets and feature prediction to pretrained encoder features, temporal tests and a small action-classification experiment.
Read the article
