Research Paper Collection

Video Understanding2024

LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Haoning Wu

Benchmark designed such that it requires to view more frames to increase performace. Retreival: find related details in longer sequence of video foota

Updated Sep 30, 2026
Video Understanding2024

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Yi Wang

First teach model how to view an video. Then teach model what the video means. Then teach model how to use language to answer video's questions. 3 Tra

Updated Sep 30, 2026
Video Understanding2019

TSM: Temporal Shift Module for Efficient Video Understanding

Ji Lin

让2D CNN看到时间信息 2D CNN + Temporal Shift = spatiotemporal CNN 3D CNN is very computationaly expensive. Conv(T, H, W) with T = Time Normal CNN:: Conv(H, W

Updated Sep 30, 2026
Video Understanding2019

SlowFast Networks for Video Recognition

Kaiming He

2 Pathways: Slow pathway: Whats in the frame Fast pathway: How are objects moving Video Recognition. Input: huamn waving hand Output: action classific

Updated Sep 30, 2026
Robot Policy2026

Reinforcement Learning for Flow-Matching Policies with Density Transport (RLDT)

Antonio Loquercio

Notes coming soon.

Updated Sep 29, 2026
Robot Policy2023

Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control (GRIF)

Vivek Myers

Connect language instruction with visual state change (start → goal) If we do language → goal image: - Instruction: "Move the pan to the left" - there

Updated Sep 28, 2026
Robot Policy2021

MBOLD - Model-Based Visual Planning with Self-Supervised Functional Distances

Stephen Tian

Notes coming soon.

Updated Sep 28, 2026
Reinforcement Learning2026

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

Perry Dong

With a pretrained VLA, using limited data, using RL to finetune the policy to make it more reliable. Given a pretrained VLA + RL-based error correctio

Updated Sep 28, 2026
Image Generation2024

One-step Diffusion with Distribution Matching Distillation

Tianwei Yin

Diffusion is good quality, but inference is too slow. DMD takes 100+ steps diffusion model, distill it to one forward pass generator. DMD is $x = G_\t

Updated Sep 27, 2026
Image Editing2025

NP Edit: Learning an Image Editing Model without Image Editing Pairs

Nupur Kumari

This is an instruction-based image editing model. Input: Input image + editing instruction Output: Edited Image Problem: for image editing model, the

Updated Sep 26, 2026
Robot Policy2023

MUTEX: Learning Unified Policies from Multimodal Task Specifications

Rutav Shah

Give multiple modality input to the same representation space, output unified robot policy. Multimodal information is used mainly to train a better ta

Updated Sep 26, 2026
Robot Policy2024

GENIMA: Generative Image as Action Models

Mohit Shridhar

Input: at timesetp t: \[O_t= \{I_t^{front}, I_t^{left}, I_t^{right}, I_t^{wrist} \} \]4 RGB camera views Language instruction Output: action policy ou

Updated Sep 26, 2026
Robot Policy2023

VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Jason Ma

VIP is a fine tuned Resnet-50 encoder, when you give it current image, it outputs a latent vector $z_t$. When you give it a goal image, it outputs a l

Updated Sep 26, 2026
Audio Generation2021

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Wei-Ning Hsu

Speech Representation Learning: turn raw sound into numerical features that make useful information easier to extract. - what sounds were spoken Same

Updated Sep 23, 2026
Audio Generation2020

Deep Audio Priors Emerge From Harmonic Convolutional Networks

Zhoutong Zhang

1. Given noisy audio track, it will output the track with less / without noise 2. Given a track with multiple insturments i.e. flute, piano, violin. I

Updated Sep 23, 2026
Image Generation2024

Compositional Image Decomposition with Diffusion Models

Yilun Du

从图片中检测compositional concepts

Updated Sep 21, 2026
Image Generation2023

ControlNet: Adding Conditional Control to Text-to-Image Diffusion Models

Lvmin Zhang,

提供一个骨架图image+文字指定prompt,然后生成一个好image Input: prompt, control image, random noise Output: edited image

Updated Sep 21, 2026
Image Generation2023

InstructPix2Pix: Learning to Follow Image Editing Instructions

Tim Brooks

修图师傅:之前需要specify新的图片长什么样,现在可以直接通过语言prompt去描述怎么改这张图片 学习条件: 图片 I: 一张红色汽车 指令 c:把汽车变成蓝色,背景不变 图片 Y: 蓝汽车 Input: Instruction that tells the model How to edit

Updated Sep 21, 2026
Reading view