TL;DR: We propose a novel method to scale up the data for 3D-free video re-shooting.
Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods while requiring minimal paired data.
Initial Viewpoint & Camera Motion Control
TARS jointly controls the target initial viewpoint and camera trajectory, allowing the camera to begin from a specified viewing angle before following the requested motion. For each example, the left is the reference and the right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
Portrait: First-Person to Third-Person
TARS transforms a portrait video from a first-person observation into a third-person viewpoint while preserving identity, scene context, and temporal motion. For each example, the left is the reference and the right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
Portrait: Third-Person to First-Person
A second portrait setting demonstrates the same third-person-to-first-person transition under a different scene and motion pattern. For each example, the left is the reference and the right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
Animal: Third-Person to First-Person
TARS changes an animal video from third-person to first-person, maintaining plausible subject motion and a coherent new perspective. For each example, the left is the reference and the right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
Robotic Arm Scene
TARS supports viewpoint changes in embodied robotic-arm scenes, preserving the task-relevant scene layout and the arm's temporal motion while moving to a new perspective. For each example, the left is the reference and the right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
Typical Camera Re-shooting
TARS controls both the initial viewing angle and camera motion across human and embodied robotic-arm scenes. The examples below demonstrate side-view, rear-view, and robotic-camera re-shooting; in each video, the left is the reference and the right is the TARS result.
Perspective Switching
TARS enables coherent first-person and third-person conversion across portrait and animal scenes, including large viewpoint changes beyond the source view.
More Cases Exploration
Explore additional TARS results for initial viewpoint and camera motion control, portrait perspective switching, and animal perspective switching.
More: Initial Viewpoint & Camera Motion Control
Additional examples with jointly controlled initial viewpoint and camera motion. Left is the reference, and right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
More: Portrait: First-Person to Third-Person
Additional portrait first-person-to-third-person transitions. Left is the reference, and right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
More: Portrait: Third-Person to First-Person
Additional portrait third-person-to-first-person transitions in different scenes and motion patterns. Left is the reference, and right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
More: Animal: Third-Person to First-Person
Additional animal third-person-to-first-person transitions. Left is the reference, and right is the TARS result.
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
TARS case slot Add the replacement media in case/
Qualitative Comparisons
Qualitative evaluations. TARS accurately re-shoots source videos under novel camera trajectories and viewpoints. Compared with CamClone, TrajCrafter, and SD 2.0, TARS better preserves visual appearance and temporal action dynamics while following the requested camera transformation.
Pipeline
Timestep-aware self-supervised learning framework.Top: High-noise diffusion timesteps primarily learn global structure, viewpoint, and camera motion, while mid- and low-noise timesteps refine texture details. Bottom: Guided by this observation, we construct self-supervised training pairs by temporally splitting each video into two clips, enabling large-scale learning of camera transitions from unlabeled videos.
Acknowledgement
We thank the Kling team and all annotation contributors whose work made this project possible.
Responsible Use Statement
The examples on this page are intended solely to communicate research results on controllable video re-shooting. We will use only materials that are authorized for research demonstration and will update or remove any example when a valid concern is raised.
BibTeX
@article{liu2026tars,
title={TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting},
author={Liu, Jiwen and Li, Shujuan and Li, Xiaohan and Meng, Zijie and Liu, Xinyue and Xu, Yulong and Zhou, Yan and Zhang, Guoxin},
year={2026}
}