Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3583 storiesFrom Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching
arXiv:2607.26817v1 Announce Type: new Abstract: Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching ego…
Task and Skill Planning: Hierarchical Robot Planning with Black-Box Skills
arXiv:2504.17901v3 Announce Type: replace Abstract: Task and motion planning (TAMP) is a well-established approach for solving long-horizon robot planning probl…
Vision-TL-Action: Neuro-Symbolic Trajectory Generation from Visual Observations and Temporal Logic
arXiv:2607.26770v1 Announce Type: new Abstract: Temporal logic (TL) provides a compositional language for the formulation of long horizon robotic tasks, but exi…
Self-Configurable Mesh-Networks for Scalable Distributed Submodular Bandit Optimization
arXiv:2602.19366v2 Announce Type: replace-cross Abstract: We study how to scale distributed bandit submodular coordination under realistic communication constra…
SG-CoT: An Ambiguity-Aware Robotic Planning Framework using Scene Graph Representations
arXiv:2603.18271v3 Announce Type: replace Abstract: Ambiguity poses a major challenge to large language models (LLMs) used as robotic planners. In this letter, …
HumanCLAW: Can Vision-Language Models Act Through a Body?
arXiv:2607.27180v1 Announce Type: cross Abstract: Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome o…
ContactFlow: A video action conditioning that transfers across embodiments
arXiv:2607.26579v1 Announce Type: new Abstract: World models offer a promising route toward robot planning by enabling agents to imagine and verify the conseque…
PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
arXiv:2606.00537v2 Announce Type: replace Abstract: Recent vision-language-action and diffusion-based robot policies often use action chunking, where each polic…
ActSWM: Action-Sensitive World Models for Long-Horizon Planning in Open-World Games
arXiv:2607.26712v1 Announce Type: new Abstract: Latent world models support efficient model-predictive control by optimizing future control sequences in latent …
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
arXiv:2607.26789v1 Announce Type: new Abstract: Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action…
Enfold: Folding World-Generator Computation into Predictive Representations for Efficient Embodied Control
arXiv:2607.26657v1 Announce Type: new Abstract: World generative models are typically used through what they produce: a rendered future, a video-conditioned act…
Reeling It In: Flexible Needle Pick Up via Thread Manipulation for Autonomous Suturing
arXiv:2607.26337v1 Announce Type: new Abstract: Suture-needle pickup is necessary for autonomous suturing, as a needle can be unexpectedly dropped or strategica…
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception
arXiv:2607.26985v1 Announce Type: new Abstract: Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow …
Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment
arXiv:2607.26802v1 Announce Type: new Abstract: This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and effic…
Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
arXiv:2607.26121v1 Announce Type: new Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and…
HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework
arXiv:2607.26283v1 Announce Type: cross Abstract: Collaborative Perception (CP) improves autonomous systems' awareness of their surroundings by sharing sensor d…
RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
arXiv:2607.26991v1 Announce Type: new Abstract: Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance…
Learning Implicit Causal World Models from Multi-Agent Demonstrations
arXiv:2607.26336v1 Announce Type: cross Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often con…
Sensor-Placement-Agnostic Sonomyography: Toward Continuous High-Dimensional Control by Users with Tetraplegia
arXiv:2607.26401v1 Announce Type: cross Abstract: Sonomyography (SMG) enables continuous device control via ultrasound-measured muscle deformation signals, but …
TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans
arXiv:2603.09971v2 Announce Type: replace Abstract: We present TiPToP, a modular manipulation system that integrates pretrained foundation models with a GPU-acc…
BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories
arXiv:2607.26914v1 Announce Type: new Abstract: Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing em…
Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots
arXiv:2607.26567v1 Announce Type: new Abstract: Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the …
Self-Adaptive Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret
arXiv:2607.26370v1 Announce Type: new Abstract: We propose a self-adaptive online learning for control method for tracking unknown target dynamics. The target d…
Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
arXiv:2607.26148v1 Announce Type: new Abstract: Autonomous embodied agents must sustain a long decision-making loop that involves perceiving, acting, verifying,…
GBPP: Grasp-Aware Base Placement Prediction for Robots via Two-Stage Learning
arXiv:2509.11594v3 Announce Type: replace Abstract: GBPP is a fast learning based scorer that selects a robot base pose for grasping from a single RGB-D snapsho…
PolygMap: A Perceptive Locomotion Framework for Humanoid Robot Stair Climbing
arXiv:2510.12346v2 Announce Type: replace Abstract: Recently, biped robot walking technology has been significantly developed, mainly in the context of a bland …
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
arXiv:2607.27017v1 Announce Type: cross Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize …
FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets
arXiv:2607.26423v1 Announce Type: cross Abstract: As autonomous drone deployments scale from individual units to coordinated swarms, the human operator's role s…
Semi-Decentralized Multi-Spacecraft Collision Avoidance under Communication Constraints
arXiv:2607.26570v1 Announce Type: new Abstract: Current spacecraft collision-avoidance operations rely on intermittent ground-station contacts, requiring operat…
Multi-Objective Compliance-Integrated Coevolution For Simulated And Real-World Deployment Of Multi-Robot Marine Autonomy
arXiv:2607.26279v1 Announce Type: new Abstract: Collaborative robots are well-suited to maritime missions that benefit from coordination, such as the exploratio…