Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6273 storiesAttune: A Self-Annotation Tool for Understanding Robot Operator Attention Profiles
arXiv:2608.12650v1 Announce Type: new Abstract: Deploying robot fleets in complex, real-world environments requires human operators to supervise multiple robots…
FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition
arXiv:2608.12683v1 Announce Type: new Abstract: Embodied agents must often identify and interact with objects based on their function rather than their identity…
SAP-Nav: Spatial Semantic Representation Meets Active Perception for Hierarchical Open-Vocabulary Object Navigation
arXiv:2608.12707v1 Announce Type: new Abstract: Hierarchical open-vocabulary object navigation (OVON) requires agents to follow free-form instructions that may …
Genetic Fuzzy System-Based Multi-Robot Coordination for Planetary Missions
arXiv:2608.12755v1 Announce Type: new Abstract: This paper proposes a decentralized approach for a multi-robot system (MRS) using a genetic fuzzy system to perf…
AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN
arXiv:2608.12835v1 Announce Type: new Abstract: Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, in…
ASPIRE-VINS: Adaptive Spline-based Visual-inertial Navigation System With Robust 3D Measurement Residuals
arXiv:2608.12840v1 Announce Type: new Abstract: Visual-inertial navigation systems estimate six-degree-of-freedom motion by fusing visual and inertial data. Mod…
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
arXiv:2608.12854v1 Announce Type: new Abstract: Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-en…
HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments
arXiv:2608.12860v1 Announce Type: new Abstract: Vision-Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: biped…
Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning
arXiv:2608.13026v1 Announce Type: new Abstract: Outcome-driven reinforcement learning offers a scalable way to post-train vision-language-action (VLA) policies …
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
arXiv:2608.13049v1 Announce Type: new Abstract: Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expen…
Semantic Radiance Fields as Simulators for Spatial Reasoning in Real-World Scenes
arXiv:2608.13095v1 Announce Type: new Abstract: Training and evaluating spatial reasoning in embodied agents requires diverse environments that are both geometr…
S2-HWM: Sparse Event-Structured Hierarchical World Model for Long-Horizon Surgical Robot Manipulation
arXiv:2608.13103v1 Announce Type: new Abstract: Long-horizon surgical robot manipulation is challenging because task rewards are sparse, while meaningful intera…
FAM-DQ: A Dual-Quadrotor-Based Fully Actuated Aerial Manipulator for High-Torque Interaction
arXiv:2608.13220v1 Announce Type: new Abstract: Aerial physical interaction requires aerial manipulation platforms to generate large interaction forces and torq…
Manufacturing Complex Airtight Soft Pneumatic Actuators for Soft Robotics: Process Evaluation and Optimization
arXiv:2608.13233v1 Announce Type: new Abstract: Manufacturing complex soft pneumatic actuators remains challenging because geometric fidelity, compliance, struc…
NestDex: Nested Policy Learning with Copilot Assisted Teleoperation for Dexterous Manipulation
arXiv:2608.13362v1 Announce Type: new Abstract: Dexterous manipulation promises substantially richer robot interaction with the physical world, but learning the…
FIRE-VLA: Failure-Informed Self-Evolution for Vision-Language-Action Models in Autonomous Driving
arXiv:2608.13395v1 Announce Type: new Abstract: Reinforcement learning improves autonomous-driving vision-language-action (VLA) models by evaluating trajectorie…
Capstan-driven Continuum Surgical Robot: Design, Modeling, and Perception
arXiv:2608.13396v1 Announce Type: new Abstract: Shape and force sensing have long been critical bottlenecks in the development of compact capstan-driven continu…
Deliberate Practice: Learning Robot Skills under a Budget
arXiv:2608.13415v1 Announce Type: new Abstract: We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tas…
ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models
arXiv:2608.13438v1 Announce Type: new Abstract: Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is esp…
Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement
arXiv:2608.13448v1 Announce Type: new Abstract: Social robots are expected to operate across diverse environments, where similar arrangements can imply differen…
Decoding Task Progress from VLA Representations
arXiv:2608.13474v1 Announce Type: new Abstract: Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation polic…
A Browser-Native Digital Test Range for Benchmarking 4D Ocean-Glider Planning Algorithms
arXiv:2608.13511v1 Announce Type: new Abstract: Repeated in-situ evaluation of ocean-glider planners requires scarce vehicles, operators, deployment and recover…
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
arXiv:2608.13555v1 Announce Type: new Abstract: Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees wi…
Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?
arXiv:2608.12515v1 Announce Type: cross Abstract: Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in hu…
Excitation-Supervised Closed-Loop Self-Calibration and Target Seeking for an Unknown-Pose Range-Bearing Relay
arXiv:2608.12528v1 Announce Type: cross Abstract: A vehicle seeking a hidden target through a range-bearing relay of unknown position and yaw must decide, onlin…
Entropy-Augmented Multi-Objective Policy Optimization in Multiagent Systems
arXiv:2608.12534v1 Announce Type: cross Abstract: Autonomous agent teams deployed in settings such as marine and extraterrestrial outposts must coordinate actio…
Do LLMs Beat Nash? Testing Decentralized Coordination in Self-Play Multi-Agent Games
arXiv:2608.12547v1 Announce Type: cross Abstract: Large language model agents deployed without a central controller are often assumed to require communication t…
Genetic Fuzzy System-based Control for Final Approach of Spacecraft Rendezvous and Proximity Operations
arXiv:2608.12760v1 Announce Type: cross Abstract: In-space servicing has been receiving great attention to extend the operation of spacecraft with defective com…
Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling
arXiv:2608.12917v1 Announce Type: cross Abstract: Developing effective robot navigation methods in crowded environments is essential for real-world applications…
EgoPHI: Estimating Contact and Force from Egocentric Vision
arXiv:2608.13014v1 Announce Type: cross Abstract: Understanding hand-object interaction from egocentric vision is essential for modeling how people physically e…