Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6273 storiesEvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
arXiv:2605.21862v2 Announce Type: replace Abstract: Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on…
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
arXiv:2604.09860v4 Announce Type: replace Abstract: The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based bench…
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
arXiv:2512.23649v4 Announce Type: replace Abstract: Humans learn locomotion through visual observation, interpreting visual content first before imitating actio…
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
arXiv:2510.10903v2 Announce Type: replace Abstract: Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer visi…
DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts
arXiv:2510.00358v2 Announce Type: replace Abstract: Soft snake robots offer remarkable flexibility and adaptability in complex environments, yet their control r…
Relay-Based Coordination for Energy-Efficient Multi-Robot Pickup and Delivery
arXiv:2509.14127v3 Announce Type: replace Abstract: We consider the problem of delivering multiple packages from a single depot to distinct goal locations using…
X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
arXiv:2608.16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding …
Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain
arXiv:2608.16164v1 Announce Type: cross Abstract: Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration…
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
arXiv:2504.06961v2 Announce Type: replace Abstract: 3D assembly tasks, such as furniture assembly and component fitting, play a crucial role in daily life and r…
Adaptive Bridge: A Proxy-Based Decoupling Layer for Mitigating DDS Backpressure in ROS 2
arXiv:2608.15380v1 Announce Type: cross Abstract: In ROS 2 systems using DDS, a single slow subscriber on a RELIABLE topic can cause backpressure that degrades …
FollowUpBot: An LLM-Based Conversational Robot for Automatic Postoperative Follow-up
arXiv:2507.15502v1 Announce Type: cross Abstract: Postoperative follow-up plays a crucial role in monitoring recovery and identifying complications. However, tr…
Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation
arXiv:2608.16843v1 Announce Type: new Abstract: Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied a…
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
arXiv:2608.16837v1 Announce Type: new Abstract: Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist visi…
Semantic- and Density-Aware Planning for Accessibility-Preserving Multi-Object Placement
arXiv:2608.16741v1 Announce Type: new Abstract: Long-term manipulation planning requires robots to reason not only about immediate task success but also about h…
Design Optimization for Large High-Force Soft Robot Manipulators Under Gravitational Loads
arXiv:2608.16728v1 Announce Type: new Abstract: Designing large soft robots capable of generating high forces for physical human-robot interaction remains a sig…
Throwing a Tight Spiral American Football by a Humanoid Robot
arXiv:2608.16642v1 Announce Type: new Abstract: Accurate throwing of the American football requires precise regulation of release conditions, where coupled line…
DPNet: Efficient Dead-End Prediction and Avoidance for Vision-Based UAV Navigation
arXiv:2608.16640v1 Announce Type: new Abstract: Vision-based Unmanned Aerial Vehicles (UAVs) often suffer from navigation failures in dead ends due to limited s…
Zetta $\zeta$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
arXiv:2608.16590v1 Announce Type: new Abstract: Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path ha…
Co-design of Neural and Muscle Network based on Embodied Perceptron Representation
arXiv:2608.16555v1 Announce Type: new Abstract: Recent advances in AI technologies have enabled the advanced design of complex control policies. In contrast, fo…
NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation
arXiv:2608.16503v1 Announce Type: new Abstract: Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance tra…
Cyclops: LiDAR as a Camera That Dreams in Color
arXiv:2608.16264v1 Announce Type: new Abstract: Conventionally, robotic perception relies heavily on cameras due to the rich semantic texture they provide. Howe…
Unified Condition-Action Modeling for Accurate One-Step Action Generation
arXiv:2608.16153v1 Announce Type: new Abstract: Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to chan…
ScenarioCharacterization: A Modular Toolkit for Characterizing Safety across Trajectory Datasets
arXiv:2608.16041v1 Announce Type: new Abstract: We introduce ScenarioCharacterization, an open-source framework for automated, dataset-agnostic profiling of dri…
Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots
arXiv:2608.16030v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction …
Revisiting Open-Loop Execution in Robotics: Toward Reactive, Higher-Performing Policies
arXiv:2608.15938v1 Announce Type: new Abstract: Action chunking --- the practice of predicting a sequence of actions and executing a prefix open-loop --- has em…
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
arXiv:2608.15875v1 Announce Type: new Abstract: Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstratin…
ViTaR: Visuo-Tactile Residual Adaptation for Foundation VLA Manipulation
arXiv:2608.15816v1 Announce Type: new Abstract: As Vision-Language-Action (VLA) models scale toward real-world deployment, contact-rich manipulation exposes a c…
Tac4Loco: Learning Spatiotemporal Plantar Pressure Representations for Humanoid Locomotion
arXiv:2608.15766v1 Announce Type: new Abstract: Humanoid robots are expected to traverse complex terrains, where the plantar support may vary dramatically due t…
Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation
arXiv:2608.15680v1 Announce Type: new Abstract: Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, sc…
Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification
arXiv:2608.15636v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but t…