Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3542 storiesSelective Agentic Recovery for UAV Autonomy with a Persistent Mission Runtime
arXiv:2606.14219v1 Announce Type: new Abstract: Agentic AI can support unmanned aerial vehicle (UAV) autonomy by providing high-level recovery reasoning when lo…
$\mu_0$: A Scalable 3D Interaction-Trace World Model
arXiv:2606.13769v1 Announce Type: new Abstract: World models that capture how actions induce physical change enable scalable robot learning without reliance on …
ParkourFormer: Integrating Predictive Supervision and Sequence Modeling into Parkour Locomotion
arXiv:2605.25782v3 Announce Type: replace Abstract: Humanoid parkour requires locomotion policies to coordinate whole-body dynamics across rapidly changing terr…
Efficient Domain-Adaptive Policy Learning via Kernel Representation with Application to Quadrotor Control under Non-Stationary Disturbances
arXiv:2606.13842v1 Announce Type: new Abstract: We present an algorithm for efficient domain-adaptive policy learning via kernel representations. Learning domai…
GAIT: Legged Robot Proprioceptive State Estimation with Attention over Inertial-Leg Tokens
arXiv:2606.14160v1 Announce Type: new Abstract: In this paper, we propose a method that applies Inertial-Leg (IL) tokenization to an attention-based network for…
EqCollide: Equivariant and Collision-Aware Deformable Objects Neural Simulator
arXiv:2506.05797v2 Announce Type: replace-cross Abstract: Simulating collisions of deformable objects is a fundamental yet challenging task due to the complexit…
Robustness without Wrinkles: Parallel Simulation and Robust MPC for Certified Deformable Manipulation
arXiv:2606.14188v1 Announce Type: new Abstract: We present CORD-SLS, a real-time control method for safe deformable object manipulation, with a focus on ropes a…
BIM-Loc: BIM-Integrated Discrepancy-Aware LiDAR-based Indoor Localization
arXiv:2606.14237v1 Announce Type: new Abstract: Accurate and robust localization is a fundamental requirement for service and inspection robots, particularly in…
When and How Severely: Scenario-Specific Safety Envelopes for Driving VLAs
arXiv:2606.14238v1 Announce Type: new Abstract: Safety certification of Vision-Language-Action (VLA) driving planners under ISO 21448 (SOTIF) rests on an Operat…
SyLink Hand: A Synergy-Inspired Linkage-Driven Anthropomorphic Hand for Human-Like Dexterity
arXiv:2606.14250v1 Announce Type: new Abstract: Designing anthropomorphic robotic hands that balance functional dexterity with mechanical simplicity remains a s…
Planning with the Views via Scene Self-Exploration
arXiv:2605.29563v2 Announce Type: replace-cross Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this c…
Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera
arXiv:2606.14535v1 Announce Type: new Abstract: Recent visual imitation learning systems have widely adopted multi-camera setups with wrist-mounted cameras as t…
ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation
arXiv:2606.13877v1 Announce Type: new Abstract: Contact-rich manipulation requires world models to reason over complex contact dynamics from multimodal sensory …
Occupancy-Grounded Room Segmentation for Hierarchical 3D Scene Graphs
arXiv:2606.13727v1 Announce Type: new Abstract: Hierarchical 3D scene graphs (3DSGs) for indoor robots organize geometric and semantic information across spatia…
An Attention-based Model for Robust Forecasting with Missing Modality
arXiv:2606.13970v1 Announce Type: new Abstract: Learning with missing modalities is a fundamental challenge in multimodal robot learning, as real-world robotic …
Guided Diffusion with Distilled Vision-Language Reliability for Aerial Navigation
arXiv:2606.13883v1 Announce Type: new Abstract: Autonomous UAV navigation is conventionally solved by pipelines that separate perception, mapping, and planning …
FlowMo-WM: A World Model with Object Momentum and Hidden Ambient Drift
arXiv:2606.13817v1 Announce Type: new Abstract: World models in robot learning predict future states from visual observations and actions, enabling agents to re…
Lifted Schr\"odinger Bridges for Gaussian Mixture Endpoints: Projection Gaps and Path-Space Obstructions
arXiv:2605.24795v2 Announce Type: replace-cross Abstract: We study stochastic density control between Gaussian-mixture endpoint distributions under Brownian pri…
Sensitivity Shaping for Latent Modeling
arXiv:2606.14585v1 Announce Type: new Abstract: Generative dynamics models enable planning in challenging robotic systems, but safe deployment requires reliably…
Impedance MPC with Disturbance Estimation for Dexterous Hand Control
arXiv:2606.14606v1 Announce Type: new Abstract: Dexterous hands must simultaneously track precise finger trajectories and maintain safe, compliant contact -- ob…
Digital Twin Driven Textile Classification and Foreign Object Recognition in Automated Sorting Systems
arXiv:2603.05230v2 Announce Type: replace-cross Abstract: The increasing demand for sustainable textile recycling requires robust automation solutions capable o…
EgoGuide: Egocentric Guidance for Efficient Robot-Free Demonstration Collection and Learning
arXiv:2606.14665v1 Announce Type: new Abstract: Robot learning from real-world demonstrations is currently constrained by data scaling. Universal Manipulation I…
Encoder Winners Do Not Reliably Transfer Across VLA Backbone Scale: A Frozen-Backbone Grafting Diagnostic
arXiv:2606.14153v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies typically inherit their vision encoder from upstream VLM releases, but i…
Micro-Swarm Locomotion Optimization in Dynamic Flow using Multi-Objective Multi-Agent Reinforcement Learning
arXiv:2605.25025v2 Announce Type: replace Abstract: Coordinating micro-robotic swarms in realistic, time-dependent fluid environments remains a major challenge …
Output-Level Regularization Eliminates the Seed Lottery in Single-GPU VLA Fine-Tuning
arXiv:2606.13856v1 Announce Type: new Abstract: Fine-tuning a vision-language-action model (VLA-JEPA) on a single GPU should be simple: load a pretrained checkp…
Low-Burden LLM-Based Preference Learning: Personalizing Assistive Robots from Natural Language Feedback for Users with Paralysis
arXiv:2604.01463v2 Announce Type: replace Abstract: Physically Assistive Robots require personalized behaviors to ensure user safety and comfort. However, tradi…
A Unified Control Architecture for Macro-Micro Manipulation using a Active Remote Center of Compliance for Manufacturing Applications
arXiv:2602.01948v2 Announce Type: replace Abstract: Macro-micro manipulators combine a macro manipulator with a large workspace, such as an industrial robot, wi…
Optimality-Preserving Decomposition for Scalable QAOA in Natural-Language-Guided Multi-Drone Assignment
arXiv:2606.14252v1 Announce Type: new Abstract: As multi-drone fleets scale, zone assignment rapidly evolves into an intractable NP-hard combinatorial problem t…
AERMANI-PLACE: Language Guided Object Placement with Aerial Manipulators
arXiv:2606.14531v1 Announce Type: new Abstract: Object placement is a fundamental component of aerial manipulation tasks, yet existing systems typically require…
PhysVLA: Towards Physically-Grounded VLA for Embodied Robotic Manipulation
arXiv:2606.13886v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models excel at mapping visual inputs and natural language instructions directly to…