📝 Publications

A full publication list is available on my Google Scholar page.

(*: Equal contribution; †: Corresponding authors.)

🎬 Video Generation, World Model & Multimodal Model

arXiv 2026
OmniDirector

[arXiv 2026 Kling Team] OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, et al.

  • We propose OmniDirector, a general framework for multi-shot camera cloning that operates without the need for cross-paired data, significantly advancing camera control in video synthesis.
SCA 2026
ParaScale

[SCA 2026] ParaScale: Scale-Calibrated Camera-Motion Transfer via a Gauge-Invariant Parallax Number
Zijie Meng.
[ResearchPod] (CCF-B Top-tier in Computer Animation)

  • We introduce the Parallax Number ($\Pi$), a dimensionless, gauge-invariant descriptor that enables scale-faithful camera-motion transfer across disparate scenes (e.g., from galaxy sweeps to desk nudges).
  • ParaScale is a plug-and-play module that re-realizes camera moves against target depths without retraining, cutting Parallax Consistency Error (PCE) by over 3x while preserving cinematic rotation.
arXiv 2026
ARGUS

[arXiv 2026] ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation
Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan†.
[Code] (Internal)

  • We propose ARGUS, a novel framework for subject-preserving video generation using stacked multi-view identity mosaic injection, ensuring high fidelity and temporal consistency.
  • This work was conducted during my internship at Kuaishou Kling, focusing on controllable identity injection in foundation video models.
ICASSP 2026 Oral
Make A Game

[ICASSP 2026 oral] Make a Game: A Novel Paradigm for Interactive Game Rendering
Zijie Meng, Jinming Che, Bingcai Wei, Xixin Cao†.
[Award] First Prize in PKU Challenge Cup

  • We introduce a novel paradigm for interactive game rendering using unified tokens and lightweight plugins, enhancing controllability in video generation.
  • Successfully generalized to complex interactive game scenarios, providing a bridge between generative AI and real-time game engines.
arXiv 2026
OmniDrive

[arXiv 2026] OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation
Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan Yuan, Miao Zhang†.

  • We present OmniDrive, an LLM-choreographed world model for multi-view driving video generation, enabling coordinated multi-agent simulation.
  • Introduced a unified latent co-compression mechanism to ensure spatial-temporal consistency across multiple camera views in complex driving environments.

Dataset

NeurIPS 2026
3D-RAD

[NeurIPS 2026] 3d-rad: A Comprehensive 3d Radiology Med-vqa Dataset with Multi-temporal Analysis and Diverse Diagnostic Tasks
Xiaotang Gai, Jiaxiang Liu, Yichen Li, Zijie Meng, Jian Wu, Zuozhu Liu.

  • We introduce 3D-RAD, the most comprehensive 3D radiology dataset for Medical VQA, supporting multi-temporal analysis and diverse clinical diagnostic tasks.

🎨 Image-Generation & Restoration & Segmentation

SCIS (CCF-A)
Orpaint

[Science China Info. Sci. 2025] Orpaint: A Zero-Shot Inpainting Model for Oracle Bone Inscription Rubbings with Visual Mamba Block
Zijie Meng, Yuanze Zeng, Xiang Chang, Tianshuo Xu, Fei Chao†, Xixin Cao, Changjing Shang, Qiang Shen.
[Journal] JCR-Q1, CCF-A

  • We propose Orpaint, the first zero-shot inpainting model specifically designed for Oracle Bone Inscription (甲骨文) restoration.
  • By integrating the Visual Mamba Block into the Diffusion denoising network, we achieve significantly faster inference and better structural restoration for damaged ancient rubbings.
ACM MM 2025
Sand Removal

[ACM MM 2025] Robust Single Image Sand Removal by Leveraging Uncertainty-aware SAM Priors and Prompt Learning with Refined Perceptual Loss
Bingcai Wei, Hui Liu, Chuang Qian, Zijian Li, Wangyu Wu, Zijie Meng.
CCF-A Conference

  • We address the challenging task of sand-dust image restoration by leveraging uncertainty-aware SAM (Segment Anything Model) priors and prompt learning.
  • My contribution focused on the Llama3 fine-tuning for generating refined perceptual instructions.
ICME 2026 Spotlight
MST-CLIPIQA

[ICME 2026 Spotlight] Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment
Zijie Meng.
CCF-B Conference | GitHub

  • I propose a multi-scale two-stream vision-language alignment framework that decouples semantic understanding from distortion perception for robust AI-generated image quality assessment.
  • My contribution focused on the overall framework design and vision-language alignment strategy.
MICCAI 2025
SynPo

[MICCAI 2025] SynPo: Boosting Training-Free Few-Shot Medical Segmentation via High-Quality Negative Prompts
Yufei Liu, Haoke Xiao, J Chai, Y Zhang, R Wang, Zijie Meng, Zhiming Luo†.
CCF-B Conference / Medical AI Top Conference

  • We propose SynPo, which boosts training-free medical image segmentation by utilizing high-quality negative prompts to refine few-shot boundary detection.