Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 storiesBody-Grounded Replanning for Physically Adaptive Manipulation
arXiv:2609.30024v1 Announce Type: new Abstract: Manipulation requires not only reasoning about the external environment, but also about the robot's physical con…
M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis
arXiv:2609.30056v1 Announce Type: new Abstract: Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generat…
Real-Time Force Regulation for Whole-Hand Dexterous Grasping
arXiv:2609.30082v1 Announce Type: new Abstract: Robust dexterous grasping requires maintaining physical stability despite contacts interactively evolving across…
Self-Adaptive VLA for Robust Robot Deployment
arXiv:2609.30092v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models demonstrate impressive capabilities in robotic manipulation, their mem…
Faster Visuomotor Policy Learning on Action Manifolds via Riemannian MeanFlow
arXiv:2609.30127v1 Announce Type: new Abstract: Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based o…
Training-free Behavior Cloning
arXiv:2609.30134v1 Announce Type: new Abstract: Neural behavior cloning compresses demonstrations into large models, making individual actions difficult to trac…
Contact as a Decision Variable: Capability-Tradeoff Contact Selection for Legged Loco-Manipulation
arXiv:2609.30140v1 Announce Type: new Abstract: In this paper, we study the joint selection of an environmental support contact and a whole-body configuration f…
ReVAMP: Vector-Accelerated Motion Planning for Kinematically-Constrained Systems via Reparameterization
arXiv:2609.30213v1 Announce Type: new Abstract: Robots often must satisfy one or more constraints during motion planning for real-world tasks. When such constra…
Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage
arXiv:2609.30214v1 Announce Type: new Abstract: We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-…
Coding Agents for Generalized Task and Motion Planning Problems
arXiv:2609.30233v1 Announce Type: new Abstract: Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states…
Rolling-WAM: World Action Models with Rolling Imagination
arXiv:2609.30247v1 Announce Type: new Abstract: World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. Howe…
RAPID: Robot Agentic Programming from Demonstrations
arXiv:2609.30249v1 Announce Type: new Abstract: Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their pote…
DeltaWAM: Delta World Action Models for Bimanual Manipulation
arXiv:2609.28811v1 Announce Type: cross Abstract: World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control…
Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models
arXiv:2609.28865v1 Announce Type: cross Abstract: Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains…
A Procedure for Classifying Attachments and Affective Social Bonds in Human-Robot Dyads
arXiv:2609.29063v1 Announce Type: cross Abstract: Human-robot interaction (HRI) claims that people form attachments and social bonds with artificial agents, yet…
UpDown-SC: Gravity-Canonicalized Dual-Envelope Scan Context for Indoor LiDAR Place Recognition
arXiv:2609.29118v1 Announce Type: cross Abstract: LiDAR place recognition is a key front end for loop closure and global relocalization, yet indoor retrieval re…
Visual Representation and History Modeling for Navigation World Models
arXiv:2609.29555v1 Announce Type: cross Abstract: Navigation World Models (NWMs) predict action-conditioned visual futures for planning. Two practical challenge…
Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge
arXiv:2609.29648v1 Announce Type: cross Abstract: Video object detection on edge devices runs computationally expensive detectors over long frame streams, causi…
System Identification of an Octocopter in Hover using Full-Harmonic Orthogonal Multisine Inputs
arXiv:2609.29832v1 Announce Type: cross Abstract: A new method for multi-input flight maneuver design for system identification is presented. The method consist…
Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding
arXiv:2609.29835v1 Announce Type: cross Abstract: LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomo…
SplatLabel: Pseudo-Labelling through 4D Gaussian Splatting
arXiv:2609.29836v1 Announce Type: cross Abstract: While 2D Vision Foundation Models offer a pathway to automate 3D semantic pseudo-labelling, translating these …
Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation
arXiv:2609.29934v1 Announce Type: cross Abstract: Does progress on spatial reasoning benchmarks translate into better navigation? Existing benchmarks test isola…
Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think
arXiv:2609.30036v1 Announce Type: cross Abstract: Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goa…
Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures
arXiv:2609.30187v1 Announce Type: cross Abstract: Ego-Exo4D is a large-scale dataset providing synchronized egocentric and multi-view exocentric video, a rich r…
TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
arXiv:2609.30222v1 Announce Type: cross Abstract: Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points…
AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control
arXiv:2609.30264v1 Announce Type: cross Abstract: Latent world models are typically trained to predict factual transitions, whereas model predictive control (MP…
Synthetic Enclosed Echoes: A New Dataset to Mitigate the Gap Between Simulated and Real-World Sonar Data
arXiv:2505.15465v2 Announce Type: replace Abstract: This paper introduces Synthetic Enclosed Echoes (SEE), a novel dataset designed to enhance robot perception …
Physics-Guided Residual Reinforcement Learning for Humanoid Narrow-Path Traversal
arXiv:2508.20661v5 Announce Type: replace Abstract: Traversing narrow paths is challenging for humanoid robots due to the sparse and safety-critical footholds r…
Novelty Adaptation Through Hybrid Large Language Model (LLM)-Symbolic Planning and LLM-guided Reinforcement Learning
arXiv:2603.11351v2 Announce Type: replace Abstract: In dynamic open-world environments, autonomous agents often encounter novelties that hinder their ability to…
Coordinate-Independent Robot Model Identification
arXiv:2603.14656v2 Announce Type: replace Abstract: Robot model identification is commonly performed by least-squares regression on inverse dynamics, but existi…