Research Paper Collection
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Benchmark designed such that it requires to view more frames to increase performace. Retreival: find related details in longer sequence of video foota
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
First teach model how to view an video. Then teach model what the video means. Then teach model how to use language to answer video's questions. 3 Tra
TSM: Temporal Shift Module for Efficient Video Understanding
让2D CNN看到时间信息 2D CNN + Temporal Shift = spatiotemporal CNN 3D CNN is very computationaly expensive. Conv(T, H, W) with T = Time Normal CNN:: Conv(H, W
SlowFast Networks for Video Recognition
2 Pathways: Slow pathway: Whats in the frame Fast pathway: How are objects moving Video Recognition. Input: huamn waving hand Output: action classific
Reinforcement Learning for Flow-Matching Policies with Density Transport (RLDT)
Notes coming soon.
Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control (GRIF)
Connect language instruction with visual state change (start → goal) If we do language → goal image: - Instruction: "Move the pan to the left" - there
MBOLD - Model-Based Visual Planning with Self-Supervised Functional Distances
Notes coming soon.
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
With a pretrained VLA, using limited data, using RL to finetune the policy to make it more reliable. Given a pretrained VLA + RL-based error correctio
One-step Diffusion with Distribution Matching Distillation
Diffusion is good quality, but inference is too slow. DMD takes 100+ steps diffusion model, distill it to one forward pass generator. DMD is $x = G_\t
NP Edit: Learning an Image Editing Model without Image Editing Pairs
This is an instruction-based image editing model. Input: Input image + editing instruction Output: Edited Image Problem: for image editing model, the
MUTEX: Learning Unified Policies from Multimodal Task Specifications
Give multiple modality input to the same representation space, output unified robot policy. Multimodal information is used mainly to train a better ta
GENIMA: Generative Image as Action Models
Input: at timesetp t: \[O_t= \{I_t^{front}, I_t^{left}, I_t^{right}, I_t^{wrist} \} \]4 RGB camera views Language instruction Output: action policy ou
VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
VIP is a fine tuned Resnet-50 encoder, when you give it current image, it outputs a latent vector $z_t$. When you give it a goal image, it outputs a l
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Speech Representation Learning: turn raw sound into numerical features that make useful information easier to extract. - what sounds were spoken Same
Deep Audio Priors Emerge From Harmonic Convolutional Networks
1. Given noisy audio track, it will output the track with less / without noise 2. Given a track with multiple insturments i.e. flute, piano, violin. I
Compositional Image Decomposition with Diffusion Models
从图片中检测compositional concepts
ControlNet: Adding Conditional Control to Text-to-Image Diffusion Models
提供一个骨架图image+文字指定prompt,然后生成一个好image Input: prompt, control image, random noise Output: edited image
InstructPix2Pix: Learning to Follow Image Editing Instructions
修图师傅:之前需要specify新的图片长什么样,现在可以直接通过语言prompt去描述怎么改这张图片 学习条件: 图片 I: 一张红色汽车 指令 c:把汽车变成蓝色,背景不变 图片 Y: 蓝汽车 Input: Instruction that tells the model How to edit