TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

Jiwen Liu, Shujuan Li, Xiaohan Li, Zijie Meng,
Xinyue Liu, Yulong Xu, Yan Zhou, Guoxin Zhang
Abstract

TL;DR: We propose a novel method to scale up the data for 3D-free video re-shooting.

Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods while requiring minimal paired data.

Initial Viewpoint & Camera Motion Control
TARS jointly controls the target initial viewpoint and camera trajectory, allowing the camera to begin from a specified viewing angle before following the requested motion. For each example, the left is the reference and the right is the TARS result.
Portrait: First-Person to Third-Person
TARS transforms a portrait video from a first-person observation into a third-person viewpoint while preserving identity, scene context, and temporal motion. For each example, the left is the reference and the right is the TARS result.
Portrait: Third-Person to First-Person
A second portrait setting demonstrates the same third-person-to-first-person transition under a different scene and motion pattern. For each example, the left is the reference and the right is the TARS result.
Animal: Third-Person to First-Person
TARS changes an animal video from third-person to first-person, maintaining plausible subject motion and a coherent new perspective. For each example, the left is the reference and the right is the TARS result.
Robotic Arm Scene
TARS supports viewpoint changes in embodied robotic-arm scenes, preserving the task-relevant scene layout and the arm's temporal motion while moving to a new perspective. For each example, the left is the reference and the right is the TARS result.

More Cases Exploration

Explore additional TARS results for initial viewpoint and camera motion control, portrait perspective switching, and animal perspective switching.

More: Initial Viewpoint & Camera Motion Control
Additional examples with jointly controlled initial viewpoint and camera motion. Left is the reference, and right is the TARS result.
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
More: Portrait: First-Person to Third-Person
Additional portrait first-person-to-third-person transitions. Left is the reference, and right is the TARS result.
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
More: Portrait: Third-Person to First-Person
Additional portrait third-person-to-first-person transitions in different scenes and motion patterns. Left is the reference, and right is the TARS result.
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
More: Animal: Third-Person to First-Person
Additional animal third-person-to-first-person transitions. Left is the reference, and right is the TARS result.
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
TARS case slot
Add the replacement media in case/
Qualitative Comparisons
Qualitative comparison between TARS and prior video re-shooting methods
Qualitative evaluations. TARS accurately re-shoots source videos under novel camera trajectories and viewpoints. Compared with CamClone, TrajCrafter, and SD 2.0, TARS better preserves visual appearance and temporal action dynamics while following the requested camera transformation.
Pipeline
TARS timestep-aware self-supervised learning framework
Timestep-aware self-supervised learning framework. Top: High-noise diffusion timesteps primarily learn global structure, viewpoint, and camera motion, while mid- and low-noise timesteps refine texture details. Bottom: Guided by this observation, we construct self-supervised training pairs by temporally splitting each video into two clips, enabling large-scale learning of camera transitions from unlabeled videos.
Acknowledgement
We thank the Kling team and all annotation contributors whose work made this project possible.
Responsible Use Statement
The examples on this page are intended solely to communicate research results on controllable video re-shooting. We will use only materials that are authorized for research demonstration and will update or remove any example when a valid concern is raised.
BibTeX

@article{liu2026tars,
  title={TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting},
  author={Liu, Jiwen and Li, Shujuan and Li, Xiaohan and Meng, Zijie and Liu, Xinyue and Xu, Yulong and Zhou, Yan and Zhang, Guoxin},
  year={2026}
}