TKCAM: Text and Keyframe to
Camera Trajectory Generation

1The University of Hong Kong 2The Hong Kong University of Science and Technology 3Macau University of Science and Technology 4Texas A&M University

† Corresponding authors

NeurIPS 2026 (Poster)

arXiv:2610.11105

An example prompt and two keyframes produce a TKCAM camera trajectory that guides video synthesis.
Text and visual keyframes guide a camera path, which can serve as a motion prior for video synthesis.

Abstract

Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a text- and keyframe-conditioned camera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Fréchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization.

Method

TKCAM discretizes continuous 12D camera trajectories into hierarchical motion tokens. A masked base transformer predicts the main trajectory from text and timestamped RGB keyframes; a residual transformer refines finer motion details.

TKCAM framework with residual vector quantization, a masked base transformer, and a masked residual transformer conditioned on text and sparse visual keyframes.
Overview of TKCAM. Text and RGB keyframe embeddings condition both stages through cross-attention.

RealEstate10K-Cap Dataset

We curate approximately 25K annotated camera trajectories, 5 million frames, and 50 hours of indoor footage. Captions combine pose-derived camera motion with visual scene details.

Motion category distribution in RealEstate10K-Cap compared with DataDoP and E.T.
Motion distribution across camera trajectory datasets. RealEstate10K-Cap includes a high proportion of compound motions.

Results

On a mixed-domain benchmark, TKCAM improves trajectory realism and text-motion alignment. Sparse visual keyframes provide an additional gain.

Selected quantitative results from the preprint
ConditionMethodFID ↓Matching ↑R@10 ↑
Text only · 1,500 samples
TextGenDoP (fine-tuned)0.5930.0827.80%
TextTKCAM0.5290.1009.60%
Visual keyframes + text · 994 samples
RGB-D + textGenDoP (fine-tuned)0.6160.10911.07%
Sparse RGB + textTKCAM0.5030.12912.67%

FID is measured on Universal CLaTr features. R@10 is retrieval on each track's evaluation set. Full results and experimental details are in the paper.

Qualitative comparison of TKCAM camera paths with ground truth, GenDoP, E.T., Director3D, and CCD.
Qualitative camera trajectory comparison. TKCAM follows compound directorial prompts across these examples.

Citation

If you find TKCAM useful, please cite the preprint.

@misc{yang2026tkcam,
  title={TKCAM: Text and Keyframe to Camera Trajectory Generation},
  author={Haozhe Yang and Zhiyang Dou and Zekai Gu and Cheng Lin and Wenping Wang and Yuan Liu and Taku Komura},
  year={2026},
  eprint={2610.11105},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.11105}
}