Researchers have developed OmniCam, a new autoregressive model designed to generate camera trajectories for video generation, scene reconstruction, and robotic perception. The model utilizes a geometry-grounded pose token learning approach, incorporating a panoramic point-cloud encoder, hybrid absolute-rotation and relative-translation tokenization, and separate geometric and semantic conditioning streams with a 3D target anchor. A new dataset, OmniCaT, comprising over 267,000 trajectories, was also constructed to evaluate OmniCam, which demonstrated significant reductions in trajectory errors and collision rates compared to existing methods. AI
IMPACT This model could improve the realism and control of AI-generated video and enhance robotic perception systems.
RANK_REASON This is a research paper detailing a new model and dataset. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →