Building multimodal models for spatial reasoning

Researchers from UC San Diego and Lambda present an approach for 3D spatial reasoning in multimodal models. The technique uses a target-affinity token to balance local task-dependent focus with global spatial awareness, addressing challenges in how models interpret physical environments and object relationships.

Cover image for Building multimodal models for spatial reasoning

Researchers report that their CVP model beats the Video‑3D‑LLM baseline on every benchmark they tested. On ScanQA the CIDEr score rises from 102.1 to 107.1, and exact‑match accuracy on SQA3D climbs from 58.6 to 62.3. Visual‑grounding performance improves, with Acc@0.25 on ScanRefer moving from 58.1 to 62.0 and F1@0.25 on Multi3DRefer from 58.0 to 60.2. Dense captioning also benefits, with CIDEr on Scan2Cap increasing from 83.8 to 90.5. The team evaluated CVP across five 3D vision‑language benchmarks that test question answering, visual grounding and dense captioning. Each benchmark uses established metrics such as CIDEr, exact‑match accuracy, Acc@0.25 and F1@0.25, allowing direct comparison with the prior Video‑3D‑LLM system. Training involved joint optimization on five datasets covering those tasks, and the reported numbers come from held‑out test splits. The paper does not provide insight into how the model would perform on scenes that differ substantially from the training distributions, such as outdoor environments or highly cluttered rooms. It also leaves unclear whether the modest 6 × 6 allocentric grid would scale to larger spaces without losing accuracy. Finally, the reported gains are averaged across benchmarks, so the variability of improvement for any single query remains uncertain.