Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3542 storiesLearning Ordinal Response Policies in Rank-Based Stochastic Prize-Collecting Games
arXiv:2510.24515v2 Announce Type: replace Abstract: The Team Orienteering Problem (TOP) generalizes many real-world multi-agent scheduling and routing tasks tha…
SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation
arXiv:2511.05203v3 Announce Type: replace Abstract: Today's autonomous agents, largely driven by foundation models (FMs), can understand natural language instru…
DynaRetarget: Dynamically-Feasible Retargeting using Sampling-Based Trajectory Optimization
arXiv:2602.06827v3 Announce Type: replace Abstract: In this paper, we introduce DynaRetarget, a complete pipeline for retargeting human motions to humanoid cont…
Consensus-based optimization (CBO): Towards Global Optimality in Robotics
arXiv:2602.06868v2 Announce Type: replace Abstract: Zero-order optimization has recently received significant attention for designing optimal trajectories and p…
Vision-Aided Relative State Estimation for Approach and Landing on a Moving Platform with Inertial Measurements
arXiv:2512.19245v2 Announce Type: replace-cross Abstract: This paper tackles the problem of estimating the relative position, orientation, and velocity between …
SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
arXiv:2605.12386v2 Announce Type: replace Abstract: Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee sa…
Continual Quadruped Robots Coordination via Semantic Skill Discovery
arXiv:2606.08102v2 Announce Type: replace Abstract: Multi-quadruped coordination has attracted increasing attention due to its enhanced payload capacity, broade…
GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation
arXiv:2606.08530v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world de…
TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation
arXiv:2606.09337v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have become a powerful framework for robotic manipulation, and recent st…
Vision-Language-Action Jump-Starting for Reinforcement Learning Robotic Agents
arXiv:2604.13733v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but …
ActionMap: Robot Policy Learning via Voxel Action Heatmap
arXiv:2606.06904v2 Announce Type: replace Abstract: Vision-language-action (VLA) models have advanced rapidly across backbones, training recipes, and data scale…
Closing the Motion Execution Gap: From Semantic Motion Task Constraints to Kinematic Control
arXiv:2605.12053v2 Announce Type: replace Abstract: This paper addresses the Motion Execution Gap, the disconnect between high-level symbolic task descriptions …
EKF-Based Depth Camera and Deep Learning Fusion for UAV-Person Distance Estimation and Following in SAR Operations
arXiv:2602.20958v2 Announce Type: replace Abstract: Vision-based Unmanned Aerial Vehicles (UAVs) frameworks aid human search tasks by detecting and recognizing …
Phase-Based Multi-Gait Learning for a Salamander-Like Robot
arXiv:2511.08299v2 Announce Type: replace Abstract: Salamander-like robots are designed inspired by the skeletal structure of their biological counterparts. How…
The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning
arXiv:2505.03296v2 Announce Type: replace Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy repre…
Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments
arXiv:2606.12339v1 Announce Type: cross Abstract: Sound source distance estimation (SDE) is a critical capability in human-robot interaction. An inappropriate i…
Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
arXiv:2606.12217v1 Announce Type: cross Abstract: World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to …
DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace
arXiv:2606.11687v1 Announce Type: cross Abstract: Unmanned Aerial Vehicle (UAV) threats have emerged as a defining security challenge of the 21st century. This …
FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning
arXiv:2606.12406v1 Announce Type: new Abstract: Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to th…
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
arXiv:2605.03065v2 Announce Type: replace-cross Abstract: Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged a…
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
arXiv:2605.00321v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) policies often fail under distribution shift, suggesting that decisions may dep…
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
arXiv:2604.20348v2 Announce Type: replace Abstract: Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Co…
Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution
arXiv:2604.07833v4 Announce Type: replace Abstract: Embodied Agents are evolving from passive reasoning systems into active executors that interact with tools, …
Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning
arXiv:2511.14427v4 Announce Type: replace Abstract: Effective contact-rich manipulation requires robots to synergistically leverage vision, force, and proprioce…
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
arXiv:2511.13207v2 Announce Type: replace Abstract: Object navigation in unseen indoor environments requires agents to perform semantic search under partial obs…
Cross-Modal Benchmarking for Robotic Perception in Natural Environments
arXiv:2606.11563v1 Announce Type: cross Abstract: Natural environments present a complex challenge to robotics perception systems. Current models, particularly …
Energy-Conserved Neural Pipelines: Attenuating Error Propagation in Modular Neural Networks via Physical Conservation Constraints
arXiv:2606.11341v1 Announce Type: cross Abstract: Modular neural network pipelines suffer from error compounding: noise at any module boundary propagates and po…
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
arXiv:2606.12352v1 Announce Type: new Abstract: Multi-robot collaboration allows robots to efficiently take on a wide range of tasks, from moving a couch throug…
Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering
arXiv:2606.12299v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from …
DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems
arXiv:2606.12236v1 Announce Type: new Abstract: Many autonomous driving systems are increasingly incorporating foundation models to improve generalization and h…