📝 Publications
A full publication list is available on my Google Scholar page.
(*: Equal contribution; †: Corresponding authors.)
🎬 Video Generation, World Model & Multimodal Model

[arXiv 2026 Kling Team] OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, et al.
- We propose OmniDirector, a general framework for multi-shot camera cloning that operates without the need for cross-paired data, significantly advancing camera control in video synthesis.

[SCA 2026] ParaScale: Scale-Calibrated Camera-Motion Transfer via a Gauge-Invariant Parallax Number
Zijie Meng.
[ResearchPod] (CCF-B Top-tier in Computer Animation)
- We introduce the Parallax Number ($\Pi$), a dimensionless, gauge-invariant descriptor that enables scale-faithful camera-motion transfer across disparate scenes (e.g., from galaxy sweeps to desk nudges).
- ParaScale is a plug-and-play module that re-realizes camera moves against target depths without retraining, cutting Parallax Consistency Error (PCE) by over 3x while preserving cinematic rotation.

[arXiv 2026] ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation
Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan†.
[Code] (Internal)
- We propose ARGUS, a novel framework for subject-preserving video generation using stacked multi-view identity mosaic injection, ensuring high fidelity and temporal consistency.
- This work was conducted during my internship at Kuaishou Kling, focusing on controllable identity injection in foundation video models.

[ICASSP 2026 oral] Make a Game: A Novel Paradigm for Interactive Game Rendering
Zijie Meng, Jinming Che, Bingcai Wei, Xixin Cao†.
[Award] First Prize in PKU Challenge Cup
- We introduce a novel paradigm for interactive game rendering using unified tokens and lightweight plugins, enhancing controllability in video generation.
- Successfully generalized to complex interactive game scenarios, providing a bridge between generative AI and real-time game engines.

[arXiv 2026] OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation
Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan Yuan, Miao Zhang†.
- We present OmniDrive, an LLM-choreographed world model for multi-view driving video generation, enabling coordinated multi-agent simulation.
- Introduced a unified latent co-compression mechanism to ensure spatial-temporal consistency across multiple camera views in complex driving environments.
Dataset

[NeurIPS 2026] 3d-rad: A Comprehensive 3d Radiology Med-vqa Dataset with Multi-temporal Analysis and Diverse Diagnostic Tasks
Xiaotang Gai, Jiaxiang Liu, Yichen Li, Zijie Meng, Jian Wu, Zuozhu Liu.
- We introduce 3D-RAD, the most comprehensive 3D radiology dataset for Medical VQA, supporting multi-temporal analysis and diverse clinical diagnostic tasks.
🎨 Image-Generation & Restoration & Segmentation

[Science China Info. Sci. 2025] Orpaint: A Zero-Shot Inpainting Model for Oracle Bone Inscription Rubbings with Visual Mamba Block
Zijie Meng, Yuanze Zeng, Xiang Chang, Tianshuo Xu, Fei Chao†, Xixin Cao, Changjing Shang, Qiang Shen.
[Journal] JCR-Q1, CCF-A
- We propose Orpaint, the first zero-shot inpainting model specifically designed for Oracle Bone Inscription (甲骨文) restoration.
- By integrating the Visual Mamba Block into the Diffusion denoising network, we achieve significantly faster inference and better structural restoration for damaged ancient rubbings.

[ACM MM 2025] Robust Single Image Sand Removal by Leveraging Uncertainty-aware SAM Priors and Prompt Learning with Refined Perceptual Loss
Bingcai Wei, Hui Liu, Chuang Qian, Zijian Li, Wangyu Wu, Zijie Meng.
CCF-A Conference
- We address the challenging task of sand-dust image restoration by leveraging uncertainty-aware SAM (Segment Anything Model) priors and prompt learning.
- My contribution focused on the Llama3 fine-tuning for generating refined perceptual instructions.

[ICME 2026 Spotlight] Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment
Zijie Meng.
CCF-B Conference | GitHub
- I propose a multi-scale two-stream vision-language alignment framework that decouples semantic understanding from distortion perception for robust AI-generated image quality assessment.
- My contribution focused on the overall framework design and vision-language alignment strategy.

[MICCAI 2025] SynPo: Boosting Training-Free Few-Shot Medical Segmentation via High-Quality Negative Prompts
Yufei Liu, Haoke Xiao, J Chai, Y Zhang, R Wang, Zijie Meng, Zhiming Luo†.
CCF-B Conference / Medical AI Top Conference
- We propose SynPo, which boosts training-free medical image segmentation by utilizing high-quality negative prompts to refine few-shot boundary detection.