Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3542 storiesBeyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation
arXiv:2608.05999v1 Announce Type: new Abstract: Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leverag…
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
arXiv:2608.06374v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a singl…
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
arXiv:2603.25406v3 Announce Type: replace Abstract: Vision-Language-Action (VLA) models map visual observations and natural-language instructions to robot actio…
Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models
arXiv:2608.05903v1 Announce Type: cross Abstract: Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, trans…
Reinforcing Action Policies by Prophesying
arXiv:2511.20633v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) policies excel in aligning language, perception, and robot control. However, mo…
VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations
arXiv:2608.06210v1 Announce Type: new Abstract: Contact-rich manipulation requires precise tracking and mechanical compliance, where variable impedance control …
Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators
arXiv:2608.06219v1 Announce Type: new Abstract: Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in c…
JTA: Joint Testability Architecture for Scenario-Based Validation of Safety-Critical Software
arXiv:2608.05594v1 Announce Type: cross Abstract: Validation adequacy in safety-critical software depends on more than the system under test. Critical scenarios…
Failing Gracefully: Mitigating Impact of Inevitable Robot Failures
arXiv:2608.05313v1 Announce Type: new Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are …
JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment
arXiv:2608.05674v1 Announce Type: new Abstract: Robot data is scarce, so generalist policies need to learn from heterogeneous sources, including human egocentri…
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
arXiv:2608.06332v1 Announce Type: new Abstract: Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments r…
A System for Train Condition Monitoring and Structural Health Assessment of Rail Vehicles
arXiv:2608.05221v1 Announce Type: new Abstract: The ongoing digitalization of rail systems and the increasing use of artificial intelligence (AI) are fundamenta…
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
arXiv:2608.05369v1 Announce Type: new Abstract: Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs,…
Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations
arXiv:2608.05588v1 Announce Type: new Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that cont…
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
arXiv:2503.03556v3 Announce Type: replace-cross Abstract: Object affordance reasoning, the ability to infer object functionalities based on physical properties,…
Sliding Sensors: Configurable Confidence in State Estimation for Continuum Robots
arXiv:2608.05410v1 Announce Type: new Abstract: Continuum robots often operate in uncertain environments, where accurate state estimation is essential for safe …
TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions
arXiv:2608.05975v1 Announce Type: new Abstract: In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned…
Visual Grounding in Zero-Shot Vision-Language Control
arXiv:2608.06154v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not…
KILVO: Kinematic-Inertial-LiDAR-Visual Odometry with Robust Multimodal Adaptation for Humanoid Robots
arXiv:2608.05647v1 Announce Type: new Abstract: This article presents a kinematic-inertial-LiDAR-visual odometry for humanoid robots, called KILVO. Tailored to …
Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
arXiv:2605.04568v3 Announce Type: replace-cross Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, populati…
Multi-Robot Motions in Milliseconds: Vector-Accelerated Primitives for Sampling-Based Planning
arXiv:2604.23960v2 Announce Type: replace Abstract: In this paper, we extend the recent Vector-Accelerated Motion Planning (VAMP) framework to multi-robot motio…
CADRE: Dynamic Catching via Implicit Contact Descriptors and Task-Appropriate Recovery Affordances
arXiv:2510.14768v2 Announce Type: replace Abstract: Real-world dexterous manipulation often encounters unexpected errors and disturbances, which can lead to cat…
Curiosity-Diffuser: Curiosity Guide Diffusion Models for Reliability
arXiv:2503.14833v2 Announce Type: replace Abstract: One of the bottlenecks in robotic intelligence is the instability of neural network models. This leads to ri…
Zero-shot Sim2Real Transfer for Magnet-Based Tactile Sensor on Insertion Tasks
arXiv:2505.02915v2 Announce Type: replace Abstract: Tactile sensing is an important sensing modality for robot manipulation. Among different types of tactile se…
Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments
arXiv:2608.04933v1 Announce Type: new Abstract: Long-horizon embodied task requires agents to act under partial observability while preserving both scene belief…
SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration
arXiv:2608.04246v1 Announce Type: new Abstract: Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor…
Structured LLM Reasoning for Zero-Shot Human--Robot Coordination Under Hidden Goals
arXiv:2608.04309v1 Announce Type: new Abstract: We present a structured large-language-model (LLM) architecture for zero-shot human--robot coordination in a coo…
PRIMAL3: Pathfinding via Reinforcement and Imitation Multi-Agent Learning - Leveraging LaCAM3
arXiv:2608.04905v1 Announce Type: new Abstract: We present PRIMAL3, an ultra-large-scale learning-based framework for multi-agent pathfinding (MAPF) that integr…
GUARD: Grounding Uncertainty and Ablation-Based Risk Detection for Diffusion-Based VLAs
arXiv:2608.04510v1 Announce Type: new Abstract: Diffusion-based vision-language-action (VLA) policies can generate plausible actions even when their predictions…
Design and Flight of an Ion-propelled Micro Hovercraft Leveraging Ground Proximity Effects
arXiv:2608.04343v1 Announce Type: new Abstract: Electroaerodynamic propulsion is compelling for use in micro air vehicles due to its silent and solid-state natu…