ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Donghao Zhou*, Haoyang He*, Fan Zhang, Hao Yang†, Guisheng Liu, Xin Gao, Zhongwei Wan, Xingyuan Bu, Jie Wang, Qiangpeng Yang, Shilei Wen§, Chi-Wing Fu, Pheng-Ann Heng§
The Chinese University of Hong Kong · Zhejiang University · ByteDance · The Ohio State University
* Equal contribution · † Project lead · § Corresponding authors
Overview

Think before you edit

Real-world video editing instructions can be implicit, indirect, and causality-heavy. ThinkV2V empowers an MLLM to first reason over the source video and instruction, then turns the refined intent into robust DiT conditions for generating the edited video.

Overview of ThinkV2V
Motivation

Implicit instructions need reasoning

Existing video editors often work well for direct prompts, but struggle when the true editing goal must be derived from temporal dynamics, object relations, or causal context.

Core Idea

MLLM as an explicit thinker

Instead of using the MLLM as a static semantic encoder, ThinkV2V activates Chain-of-Thought reasoning and produces a refined prompt before generation.

Outcome

Reasoning becomes executable

MLLM hidden states, original text embeddings, and source-video VAE features jointly condition a DiT generator for faithful editing execution.

01 / Method

From MLLM thinking to DiT editing

ThinkV2V builds on a practical MLLM-to-DiT architecture and incorporates dedicated training and inference strategies that elicit reasoning under complex editing instructions.

Pipeline of ThinkV2V
Explicit thinking

MLLM for intent-reasoning

The MLLM reads the source video and instruction, reasons about the hidden intent, and passes features through the connector to align them with the DiT conditioning space.

Editing execution

Multi-condition DiT

Reasoning features, original prompt embeddings, and source-video VAE features jointly condition the DiT generator for faithful editing execution.

Training

Progressive curriculum training

Training progresses from stable basic editing to high-resolution adaptation and reasoning-intensive cases, gradually strengthening reliable execution.

Inference

Inference-time thinking scaling

Serial refinement plus best-of-N selection lets the MLLM revisit candidate prompts and select the one best aligned with the original request.

02 / Dataset & Benchmark

Reasoning-oriented training and evaluation

We also introduce ThinkV2V-150K and ThinkV2V-Bench for the community, serving as a reasoning-oriented training set and benchmark, respectively.

150K ThinkV2V-150K pairs
308 ThinkV2V-Bench samples
5 reasoning-intensive edit categories
720p default editing resolution
Dataset and benchmark statistics
03 / Quantitative Comparison

Competitive performance across complex and standard editing

Results are reported on ThinkV2V-Bench and OpenVE-Bench using Seed-1.6-VL and Gemini-2.5-Pro judges, covering diverse editing categories and instruction complexities.

Method #Param. Global Style Bg. Change Local Change Local Remove Local Add Overall
VACE14B1.331.171.411.001.051.19
OmniVideo1.3B1.921.401.911.961.871.82
ReCo1.3B2.251.622.503.132.392.26
InsVIE2B2.081.181.481.041.241.40
Lucy-Edit5B1.692.082.841.092.212.01
ICVE13B1.952.042.912.052.122.23
DITTO14B3.281.932.121.001.822.02
OpenVE-Edit5B2.242.492.551.142.052.09
ThinkV2V5B3.132.673.072.302.422.72
Method#Param.Global StyleBg. ChangeLocal ChangeLocal RemoveLocal AddOverall
VACE14B1.771.911.952.071.481.83
OmniVideo1.3B2.481.181.251.561.231.52
ReCo1.3B2.581.602.032.712.422.27
InsVIE2B2.901.181.461.381.291.63
Lucy-Edit5B2.842.013.121.872.392.46
ICVE13B2.901.992.982.232.362.52
DITTO14B3.831.361.971.391.572.01
OpenVE-Edit5B3.062.052.061.321.902.08
ThinkV2V5B3.742.392.892.032.022.61
Method#Param.Global StyleBg. ChangeLocal ChangeLocal RemoveLocal AddOverall
VACE14B1.411.161.431.001.051.21
OmniVideo1.3B1.021.001.001.001.001.00
ReCo1.3B2.681.692.672.462.422.38
InsVIE2B2.251.231.601.001.231.46
Lucy-Edit5B2.172.203.301.032.372.21
ICVE13B2.351.862.912.682.272.41
DITTO14B3.702.232.281.002.082.26
OpenVE-Edit5B3.112.723.191.422.412.57
ThinkV2V5B3.412.493.492.832.432.93
Method#Param.Global StyleBg. ChangeLocal ChangeLocal RemoveLocal AddOverall
VACE14B1.491.552.071.461.261.57
OmniVideo1.3B1.111.181.141.141.361.19
ReCo1.3B2.691.642.092.711.952.20
InsVIE2B2.201.061.481.361.171.45
Lucy-Edit5B2.271.573.201.752.302.22
ICVE13B2.221.622.572.511.972.18
DITTO14B4.011.682.031.531.412.13
OpenVE-Edit5B3.162.362.981.852.152.50
ThinkV2V5B3.812.123.482.391.852.72
Citation

ThinkV2V

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Paper ↗
BibTeX
@article{zhou2026thinkv2v,
  title = {ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing},
  author = {Zhou, Donghao and He, Haoyang and Zhang, Fan and Yang, Hao and Liu, Guisheng and Gao, Xin and Wan, Zhongwei and Bu, Xingyuan and Wang, Jie and Yang, Qiangpeng and Wen, Shilei and Fu, Chi-Wing and Heng, Pheng-Ann},
  year = {2026}
}