Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3583 storiesTowards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds
arXiv:2505.14366v2 Announce Type: replace-cross Abstract: We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspec…
FIRMGrasp: A Friction-Informed Risk Margin for Robust Grasp Synthesis
arXiv:2607.25049v1 Announce Type: new Abstract: Classical grasp quality metrics assume a single deterministic friction coefficient, so they cannot predict wheth…
Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design
arXiv:2607.25798v1 Announce Type: new Abstract: An often overlooked factor of robot manipulation performance is the embodiment of the robot itself. Motivated by…
Steeringless Drifting: Differential-Torque Control of a Four-Wheel Independently Driven Vehicle
arXiv:2607.24863v1 Announce Type: new Abstract: Control methods for emerging vehicle chassis architectures are important for autonomous driving near handling li…
Diff2DGS: Reliable Reconstruction of Occluded Surgical Scenes via 2D Gaussian Splatting
arXiv:2602.18314v2 Announce Type: replace-cross Abstract: Real-time reconstruction of deformable surgical scenes is vital for advancing robotic surgery, improvi…
Track-Leakage-Free Hold-Out Self-Validation for Photogrammetric Reconstruction: Protocol, Sensitivity, and Limits
arXiv:2607.24852v1 Announce Type: cross Abstract: Automated photogrammetric inspection emits metric measurements from a 3D reconstruction whose own correctness …
Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming
arXiv:2607.25047v1 Announce Type: cross Abstract: Extended Reality (XR) is increasingly used in human-robot interaction to communicate robot intent, planned mot…
DC-WAM: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models
arXiv:2607.25918v1 Announce Type: new Abstract: World-Action Models (WAMs) augment robot policies with future visual prediction, but it remains unclear what the…
S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information
arXiv:2607.26047v1 Announce Type: new Abstract: Acoustic information provides rich cues about object location, material properties, and changes caused by contac…
When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning
arXiv:2607.25593v1 Announce Type: new Abstract: Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator confi…
Decompose and Reorganize: Planning with Primitives and Visuomotor Policies Learned from Demonstrations
arXiv:2607.25397v1 Announce Type: new Abstract: Successfully automating dexterous, long-horizon robotic manipulation requires frameworks capable of both high-le…
$\pi\mathbf{R}^2$: Reactive Real-time Flow Policies
arXiv:2607.26055v1 Announce Type: new Abstract: Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pret…
Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Critic Controller
arXiv:2607.25728v1 Announce Type: new Abstract: This paper presents a cooperative indoor UAV guidance framework that combines a shared voxel-map world model wit…
CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
arXiv:2601.02295v2 Announce Type: replace Abstract: Current work on robot failure detection and correction typically operates in a post hoc manner, analyzing er…
HOME: Robust Hough-space Matching Method for Structured and Textureless Videos
arXiv:2607.25389v1 Announce Type: cross Abstract: Visual front-ends for robotic localization typically rely on point-based features such as Oriented FAST and Ro…
SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation
arXiv:2607.25219v1 Announce Type: new Abstract: Social navigation has progressed from simplified 2D environments toward a more general vision-based setting, in …
Tri-Manual Visuomotor Imitation Learning of Robot Policies
arXiv:2607.25731v1 Announce Type: new Abstract: Bimanual teleoperation provides an effective way to collect robot demonstrations, but it assumes that the operat…
INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models
arXiv:2607.26056v1 Announce Type: new Abstract: Forward latent world models predict how actions change a scene, but recover actions for a desired change only th…
Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
arXiv:2607.25337v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rath…
Belief-Aware Influence and Trust (BAIT): Shaping Human Belief During Repeated Human-Robot Interaction
arXiv:2607.25327v1 Announce Type: new Abstract: Repeated human-robot interaction (HRI) requires proactively accounting for humans who continually adapt to evolv…
Input Shaping for Point-to-Point Motion with a Continuum Robot Arm
arXiv:2607.25071v1 Announce Type: new Abstract: A cable-driven continuum robot arm is an underactuated mechanism and may suffer residual vibration at the end of…
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
arXiv:2607.25895v1 Announce Type: new Abstract: Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and…
Tripody: An Overconstrained 3-SPR-like Parallel Robot for High-Reach Construction Tasks
arXiv:2607.25781v1 Announce Type: new Abstract: Many ceiling construction tasks still rely on heavy serial manipulators that are difficult to deploy in cluttere…
VisualPatchWorld: Code World Models as Latent Structured Representations for Planning
arXiv:2607.25236v1 Announce Type: cross Abstract: Different research lines use the term world model in different ways, yet they share a common aim: to capture h…
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
arXiv:2607.25912v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existin…
Pictura: Perspective-View Self-Play at Scale for Driving
arXiv:2607.26005v1 Announce Type: cross Abstract: Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been m…
Nanobot Algorithms for Treatment of Diffuse Cancer
arXiv:2509.06893v2 Announce Type: replace-cross Abstract: Motile nanosized particles, or "nanobots", promise more effective and less toxic targeted drug deliver…
Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling
arXiv:2602.08058v4 Announce Type: replace-cross Abstract: In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- w…
Nautilus: From One Prompt to Plug-and-Play Robot Learning
arXiv:2605.11665v2 Announce Type: replace Abstract: Robot learning research is fragmented across policy families, benchmark suites, and real robots; each implem…
A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models
arXiv:2607.25516v1 Announce Type: new Abstract: Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructio…