Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6273 storiesSeeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning
arXiv:2510.03885v4 Announce Type: replace Abstract: In this paper, we demonstrate that mobile manipulation policies utilizing a 3D latent map achieve stronger s…
PerFACT: Motion Policy with LLM-Powered Dataset Synthesis and Fusion Action-Chunking Transformers
arXiv:2512.03444v2 Announce Type: replace Abstract: Deep learning methods have significantly enhanced motion planning for robotic manipulators by leveraging pri…
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
arXiv:2512.22983v2 Announce Type: replace Abstract: Recent advances in vision, language, and multimodal learning have significantly accelerated progress in robo…
Hybrid System Planning using a Mixed-Integer ADMM Heuristic and Hybrid Zonotopes
arXiv:2602.17574v2 Announce Type: replace Abstract: Embedded optimization-based planning for hybrid systems is challenging due to the use of mixed-integer progr…
Grounding Robot Generalization in Training Data via Retrieval-Augmented VLMs
arXiv:2603.11426v3 Announce Type: replace Abstract: Recent work on robot manipulation has advanced policy generalization to novel scenarios. However, it is ofte…
Morphology-Conditioned World Model for Cross-Embodiment Quadrupedal Locomotion
arXiv:2604.08780v2 Announce Type: replace Abstract: World models promise a paradigm shift in robotics, where an agent learns the physics of its environment once…
SADP: Subgoal-Aware Diffusion Policy for Long-Horizon Manipulation Learned from Foundation Model Generated Demonstrations
arXiv:2605.16871v2 Announce Type: replace Abstract: Long-horizon robot manipulation requires policies to coordinate multiple intermediate subgoals and determine…
The functional and temporal roles of gaze evolve across the phases and constraints of multi-stage robot-mediated manipulation
arXiv:2606.21920v2 Announce Type: replace Abstract: Goal-directed eye movements are a fundamental component of visuomotor control, enabling humans to anticipate…
VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models
arXiv:2605.29605v2 Announce Type: replace Abstract: Task-success confidence estimation for Vision-Language-Action (VLA) models provides a crucial task-level sig…
Learning Versatile Humanoid Manipulation with Touch Dreaming
arXiv:2604.13015v3 Announce Type: replace Abstract: Humanoid robots promise general-purpose assistance, yet real-world humanoid loco-manipulation remains challe…
I-Perceive: A Foundation Model for Active Perception with Language Instructions
arXiv:2603.00600v2 Announce Type: replace Abstract: Active perception - the ability of a robot to proactively select viewpoints to acquire task-relevant informa…
MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery
arXiv:2602.12407v3 Announce Type: replace Abstract: Background: Robot-assisted minimally invasive surgery (RMIS) research increasingly relies on multimodal data…
On Minimum Aerial Photographs for Planar Region Coverage: Hardness and Approximation
arXiv:2512.18268v4 Announce Type: replace Abstract: Aerial photography with drones often requires covering a planar region with a limited number of images while…
Fully distributed and resilient source seeking for robot swarms
arXiv:2410.15921v3 Announce Type: replace Abstract: Existing source-seeking algorithms for robot swarms typically require either direct gradient measurements or…
EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints
arXiv:2608.15502v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inf…
FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge
arXiv:2608.15410v1 Announce Type: cross Abstract: Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests i…
UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection
arXiv:2608.15259v1 Announce Type: cross Abstract: Unmanned Aerial Vehicles (UAVs) play a crucial role in various scenarios ranging from disaster response to tra…
EgoTac: In-the-wild Tactile Prediction from Egocentric Vision
arXiv:2608.15060v1 Announce Type: cross Abstract: Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot lea…
NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving
arXiv:2608.14767v1 Announce Type: cross Abstract: Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Ex…
Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade
arXiv:2608.14650v1 Announce Type: cross Abstract: Existing adaptive-inference and world-action-model systems use cheap-stage outputs or predicted futures to all…
$\tau_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
arXiv:2608.16885v1 Announce Type: new Abstract: Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them co…
When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents
arXiv:2608.16806v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-…
Neurosymbolic Embodied Agents
arXiv:2608.16794v1 Announce Type: new Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as the…
Observation-Constrained Joint-Space Viewpoint Optimization for Robotic Inspection of Cylindrical Cavities
arXiv:2608.16442v1 Announce Type: new Abstract: Inspection is a core capability in many mobile robotics applications, including industrial facility monitoring, …
Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration
arXiv:2608.16229v1 Announce Type: new Abstract: Coordinated multi-agent exploration requires not only efficient individual coverage but also non-redundant cover…
Arm-Aware Guided Dexterous Grasp Generation with Arm-Agnostic Grasp Models
arXiv:2608.16351v1 Announce Type: new Abstract: Dexterous grasp generation that considers arm-related constraints is crucial in real-world scenarios involving a…
Readiness Barrier Functions: Forward-Invariant Control Authority for Overactuated Multirotor Allocation
arXiv:2608.16335v1 Announce Type: new Abstract: Allocation schemes that greedily maximize a readiness metric over the actuator fiber bundle of an overactuated m…
Marker-Constrained Pose-Graph Correction for Cross-Platform Georeferencing in GNSS-Denied Environments
arXiv:2608.16281v1 Announce Type: new Abstract: Autonomous operation in GNSS-denied environments requires heterogeneous mapping pipelines to maintain a consiste…
RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing
arXiv:2608.16195v1 Announce Type: new Abstract: Achieving human-level competitive intelligence and physical agility in humanoid robots remains a profound challe…
SparkVLA: Stop-Aware Hierarchical VLA with Adaptive Action Chunking for Long-Horizon Manipulation
arXiv:2608.16172v1 Announce Type: new Abstract: At every re-observation point in a hierarchical Vision-Language-Action (VLA) system, two interface decisions mus…