ROMO-S: Multimodal Spatial Understanding

Studying generated surround views for vision-language spatial understanding from a front camera and 360-degree LiDAR.

Status: Ongoing research, September 2026.

I am investigating whether generated surround views can help a vision-language model interpret the environment around a mobile robot. ROMO-S combines an observed front-camera image with 360-degree LiDAR geometry to infer visual context outside the camera’s field of view.

The central question is whether these generated views add useful spatial information compared with bird’s-eye-view (BEV) representations when both approaches receive the same sensor inputs.

  • Research focus: camera-LiDAR fusion, surround-view generation, and vision-language spatial understanding.
  • Approach: compare BEV-based and generated-view-assisted representations under matched sensor inputs and evaluation conditions.
  • Evaluation focus: object presence and relative location around the robot, including missing objects and unsupported descriptions.
  • Current scope: model development and comparative evaluation. Improvements in navigation or planning remain to be established.