Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6171 storiesClosing the Motion Execution Gap: From Semantic Motion Task Constraints to Kinematic Control
arXiv:2605.12053v2 Announce Type: replace Abstract: This paper addresses the Motion Execution Gap, the disconnect between high-level symbolic task descriptions …
EKF-Based Depth Camera and Deep Learning Fusion for UAV-Person Distance Estimation and Following in SAR Operations
arXiv:2602.20958v2 Announce Type: replace Abstract: Vision-based Unmanned Aerial Vehicles (UAVs) frameworks aid human search tasks by detecting and recognizing …
Phase-Based Multi-Gait Learning for a Salamander-Like Robot
arXiv:2511.08299v2 Announce Type: replace Abstract: Salamander-like robots are designed inspired by the skeletal structure of their biological counterparts. How…
The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning
arXiv:2505.03296v2 Announce Type: replace Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy repre…
Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments
arXiv:2606.12339v1 Announce Type: cross Abstract: Sound source distance estimation (SDE) is a critical capability in human-robot interaction. An inappropriate i…
Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
arXiv:2606.12217v1 Announce Type: cross Abstract: World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to …
DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace
arXiv:2606.11687v1 Announce Type: cross Abstract: Unmanned Aerial Vehicle (UAV) threats have emerged as a defining security challenge of the 21st century. This …
FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning
arXiv:2606.12406v1 Announce Type: new Abstract: Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to th…
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
arXiv:2605.03065v2 Announce Type: replace-cross Abstract: Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged a…
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
arXiv:2605.00321v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) policies often fail under distribution shift, suggesting that decisions may dep…
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
arXiv:2604.20348v2 Announce Type: replace Abstract: Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Co…
Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution
arXiv:2604.07833v4 Announce Type: replace Abstract: Embodied Agents are evolving from passive reasoning systems into active executors that interact with tools, …
Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning
arXiv:2511.14427v4 Announce Type: replace Abstract: Effective contact-rich manipulation requires robots to synergistically leverage vision, force, and proprioce…
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
arXiv:2511.13207v2 Announce Type: replace Abstract: Object navigation in unseen indoor environments requires agents to perform semantic search under partial obs…
Cross-Modal Benchmarking for Robotic Perception in Natural Environments
arXiv:2606.11563v1 Announce Type: cross Abstract: Natural environments present a complex challenge to robotics perception systems. Current models, particularly …
Energy-Conserved Neural Pipelines: Attenuating Error Propagation in Modular Neural Networks via Physical Conservation Constraints
arXiv:2606.11341v1 Announce Type: cross Abstract: Modular neural network pipelines suffer from error compounding: noise at any module boundary propagates and po…
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
arXiv:2606.12352v1 Announce Type: new Abstract: Multi-robot collaboration allows robots to efficiently take on a wide range of tasks, from moving a couch throug…
Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering
arXiv:2606.12299v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from …
DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems
arXiv:2606.12236v1 Announce Type: new Abstract: Many autonomous driving systems are increasingly incorporating foundation models to improve generalization and h…
AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents
arXiv:2606.12142v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) are increasingly used in inspection, search and rescue, environmental monitoring…
Bridging the Morphology Gap: Adapting VLA Models to Dexterous Manipulation via Intent-Conditioned Fine-Tuning
arXiv:2606.12109v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable zero-shot generalization in robotic manipulatio…
DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model
arXiv:2606.12105v1 Announce Type: new Abstract: Vision-language-action (VLA) models inherit a shared synchronous clock from vision-language pretraining, process…
KinematicRL: A Sim-to-Real Reinforcement Learning Framework For Social Navigation With Kinodynamic Feasibility
arXiv:2606.12042v1 Announce Type: new Abstract: Deep Reinforcement Learning (DRL) has shown promise for social navigation, yet its real-world deployment remains…
VICX: Generalizable Robot Manipulation via Video Generation and In-Context Operator Network
arXiv:2606.12028v1 Announce Type: new Abstract: Generalizable robot manipulation requires not only task-level reasoning over unseen scenes, but also reliable gr…
MPPI-based Informative Trajectory Planning for Search and Capture of Drifting Targets with ASVs
arXiv:2606.12019v1 Announce Type: new Abstract: Autonomous surface vehicles offer an efficient solution for environmental cleanup as well as search and rescue o…
Modular Anthropomorphic Hand Design via Multi-Parameter Finger Benchmarking and Selection
arXiv:2606.11826v1 Announce Type: new Abstract: Designing anthropomorphic dexterous robotic hands remains challenging as the design space straddles morphology, …
TacCoRL: Integrating Tactile Feedback into VLA via Simulation
arXiv:2606.11743v1 Announce Type: new Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, b…
Learning Object Manipulation from Scratch via Contrastive Interaction
arXiv:2606.11525v1 Announce Type: new Abstract: Contrastive Reinforcement Learning (CRL) has seen recent success in a wide variety of goal-conditioned robotics …
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
arXiv:2606.11324v1 Announce Type: new Abstract: We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied rea…
Model-based Optimization of Anguilliform Swimming Gaits for Soft Robotic Applications
arXiv:2606.11278v1 Announce Type: new Abstract: In this paper, we introduce the Soft Lamprey-Inspired Dual Environment Robot (SLIDER) and a proper modeling and …