TKCAM: Text and Keyframe to
Camera Trajectory Generation
NeurIPS 2026 (Poster)
Abstract
Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a text- and keyframe-conditioned camera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Fréchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization.
Method
TKCAM discretizes continuous 12D camera trajectories into hierarchical motion tokens. A masked base transformer predicts the main trajectory from text and timestamped RGB keyframes; a residual transformer refines finer motion details.
RealEstate10K-Cap Dataset
We curate approximately 25K annotated camera trajectories, 5 million frames, and 50 hours of indoor footage. Captions combine pose-derived camera motion with visual scene details.
Results
On a mixed-domain benchmark, TKCAM improves trajectory realism and text-motion alignment. Sparse visual keyframes provide an additional gain.
| Condition | Method | FID ↓ | Matching ↑ | R@10 ↑ |
|---|---|---|---|---|
| Text only · 1,500 samples | ||||
| Text | GenDoP (fine-tuned) | 0.593 | 0.082 | 7.80% |
| Text | TKCAM | 0.529 | 0.100 | 9.60% |
| Visual keyframes + text · 994 samples | ||||
| RGB-D + text | GenDoP (fine-tuned) | 0.616 | 0.109 | 11.07% |
| Sparse RGB + text | TKCAM | 0.503 | 0.129 | 12.67% |
FID is measured on Universal CLaTr features. R@10 is retrieval on each track's evaluation set. Full results and experimental details are in the paper.
Citation
If you find TKCAM useful, please cite the preprint.
@misc{yang2026tkcam,
title={TKCAM: Text and Keyframe to Camera Trajectory Generation},
author={Haozhe Yang and Zhiyang Dou and Zekai Gu and Cheng Lin and Wenping Wang and Yuan Liu and Taku Komura},
year={2026},
eprint={2610.11105},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.11105}
}