Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3584 storiesReal-Time Visual Intelligence on Low-Cost UAVs: A Modular Approach for Tracking, Scanning, and Navigation
arXiv:2607.02298v1 Announce Type: new Abstract: Autonomous drones are rapidly transforming modern warfare and civil applications alike. This paper presents the …
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
arXiv:2603.22435v2 Announce Type: replace Abstract: "Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) me…
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
arXiv:2607.02466v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- t…
LIME: Learning Intent-aware Camera Motion from Egocentric Video
arXiv:2607.02417v1 Announce Type: new Abstract: Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded …
Controllable Sim Agents with Behavior Latents
arXiv:2607.02496v1 Announce Type: new Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpre…
Bridge-WA: Predicting Where and How the World Changes for Robotic Action
arXiv:2607.02195v1 Announce Type: new Abstract: General-purpose vision-language-action models benefit from large vision-language priors, but effective manipulat…
SPLC: Social Preference Learning for Crowd Robot Navigation
arXiv:2607.01925v1 Announce Type: new Abstract: Offline reinforcement learning (RL) holds significant potential for crowd robot navigation in human-robot coexis…
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
arXiv:2605.11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching…
Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
arXiv:2508.05941v2 Announce Type: replace Abstract: Visuomotor policies trained via behavior cloning are vulnerable to covariate shift, where small deviations f…
The Three Dimensions of ROS 2 Middleware
arXiv:2607.01304v1 Announce Type: new Abstract: ROS 2 (Robot Operating System 2) has emerged as the de facto standard for modern robot software development, wit…
WaveLander: A Generalizable Hierarchical Control Framework for UAV Landing on Wave-Disturbed Platforms via Reinforcement Learning
arXiv:2607.01281v1 Announce Type: new Abstract: Autonomous landing of unmanned aerial vehicles (UAVs) on wave-disturbed marine platforms remains challenging due…
Multi-Rate Nonlinear Model Predictive Control for Wall-Supported Bipedal Locomotion of Quadrupedal Robots
arXiv:2607.01574v1 Announce Type: new Abstract: This paper presents a novel layered planning and control framework based on multi-rate nonlinear model predictiv…
PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation
arXiv:2607.01803v1 Announce Type: cross Abstract: Recent advances in 3D content generation from text or images have achieved impressive results, yet view incons…
Distilling Collaborative Dynamics into Latent Space for Implicit Coordination in Decentralized Multi-Agent Manipulation
arXiv:2606.22982v2 Announce Type: replace Abstract: Multi-arm manipulation demands precise spatiotemporal coordination, yet many centralized approaches scale po…
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
arXiv:2512.22539v3 Announce Type: replace Abstract: While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remai…
Actuator Reality Shaping for Zero-Shot Sim-to-Real Robot Learning
arXiv:2607.02205v1 Announce Type: new Abstract: Sim-to-real transfer in robot learning is often limited by discrepancies between the ideal actuator dynamics ass…
Influence of Radial Basis Activation Functions on Intelligent Controller for Robotic Manipulators
arXiv:2607.02167v1 Announce Type: cross Abstract: This paper presents an intelligent control framework for trajectory tracking of robotic manipulators using rad…
BIFROST: Bridging Invariant Feature Representation for Observation-space Sim2Real Transfer
arXiv:2607.01410v1 Announce Type: new Abstract: Sim2real transfer for robot policy learning suffers due to mismatch between simulation and reality. Existing met…
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
arXiv:2603.18315v2 Announce Type: replace Abstract: Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision sign…
SE(2) Navigation Mesh
arXiv:2607.01454v1 Announce Type: new Abstract: Global navigation for ground robots in complex multi-level environments requires representations that accurately…
VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
arXiv:2607.01586v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different…
ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
arXiv:2603.28545v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models and world-action models have emerged as central paradigms for general-pu…
Episodic-to-Semantic Consolidation Without Identity Drift
arXiv:2607.01988v1 Announce Type: cross Abstract: Long-running adaptive intelligent agents face a structural tension between knowledge consolidation and informa…
DL-VINS-Factory: A Modular Framework for Learned Visual Front-Ends in Visual-Inertial SLAM
arXiv:2607.01757v1 Announce Type: cross Abstract: Deep-learning features excel in visual matching, yet their practical value in tightly coupled visual-inertial …
DL-SLAM: Enabling High-Fidelity Gaussian Splatting SLAM in Dynamic Environments based on Dual-Level Probability
arXiv:2607.01860v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled significant progress in dense dynamic Simultaneous …
SPOT: Spatio-Temporal Obstacle-free Trajectory Planning for UAVs in Unknown Dynamic Environments
arXiv:2602.01189v3 Announce Type: replace Abstract: We address the problem of reactive motion planning for quadrotors operating in unknown environments with dyn…
Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
arXiv:2512.18368v2 Announce Type: replace Abstract: Scaling imitation learning to diverse multi-task robot manipulation remains challenging due to suboptimal de…
CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation
arXiv:2607.02222v1 Announce Type: new Abstract: Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map cons…
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
arXiv:2604.19775v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, plan…
MetaTune: Adjoint-based Meta-tuning via Robotic Differentiable Dynamics
arXiv:2603.27313v2 Announce Type: replace Abstract: Disturbance observer-based control has shown promise in robustifying robotic systems against uncertainties. …