Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6273 storiesSocioGesture: Real-Time and Adaptive Social Gesture Perception for Human-Robot Interaction
arXiv:2609.04545v1 Announce Type: new Abstract: Robots interacting with people must recognize not only explicit commands, but also social cues such as invitatio…
NavArena: Automated Construction of Goal-Oriented Navigation Benchmarks from 3D Gaussian Splatting Reconstructions
arXiv:2609.04602v1 Announce Type: new Abstract: Fixed 3D Gaussian Splatting (3DGS) reconstructions provide realistic novel views but lack the traversability con…
ToPos: Automated Optimal Positioning on Topographic Manifolds using Constrained Geodesic Voronoi Decomposition
arXiv:2609.05084v1 Announce Type: new Abstract: Reliable autonomous mapping, environmental sampling, last-mile logistics, and infrastructure deployment depend o…
Sound-based Multi-Person 3D Pose Estimation
arXiv:2609.04902v1 Announce Type: cross Abstract: Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to esti…
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
arXiv:2604.18271v2 Announce Type: replace Abstract: As the world of agentic artificial intelligence applied to robotics evolves, the need for agents capable of …
Toward Context-Aware Exoskeleton Assistance: Integrating Computer Vision Payload Estimation with a Multi-Metric Optimization Space
arXiv:2508.06207v3 Announce Type: replace Abstract: Back-support exoskeletons mitigate musculoskeletal strain, yet current systems rely on reactive sensing and …
Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models
arXiv:2609.04381v1 Announce Type: cross Abstract: Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service …
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
arXiv:2609.05324v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. …
Morphology and actuation as inductive biases in robotic hand manipulation
arXiv:2609.05206v1 Announce Type: new Abstract: Robotic hands vary widely in anatomical fidelity and mechanical complexity, and these structural choices influen…
LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models
arXiv:2609.05178v1 Announce Type: new Abstract: Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in r…
A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning
arXiv:2609.05133v1 Announce Type: new Abstract: This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reasonin…
HaptiNet: Networked Haptic Robots Enable Physical Co-presence in Geographically-Unconstrained Rehabilitation
arXiv:2609.04799v1 Announce Type: new Abstract: Cooperative rehabilitation enhances engagement, task performance, and social-motor interaction, yet it demands p…
Continuous Cognitive Coverage for Autonomous Robots via Event-Dependent Cognitive Treatment and Learning
arXiv:2609.04770v1 Announce Type: new Abstract: Autonomous robots continuously encounter objects, changes, and situations, and every event admitted into cogniti…
AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision
arXiv:2609.04411v1 Announce Type: new Abstract: Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigation …
MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision
arXiv:2609.04958v1 Announce Type: cross Abstract: Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity …
CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation
arXiv:2609.05397v1 Announce Type: cross Abstract: Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-v…
Octopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts
arXiv:2605.09055v2 Announce Type: replace Abstract: Bringing a previously unintegrated device under the control of an AI agent still requires device-specific en…
YOLO with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection with trustworthy multimodal AI in computer vision perception
arXiv:2603.23037v2 Announce Type: replace-cross Abstract: The trustworthy object detection capabilities of a novel Kolmogorov-Arnold network framework are exami…
LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory
arXiv:2609.02350v2 Announce Type: replace-cross Abstract: Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in…
SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control
arXiv:2605.22894v3 Announce Type: replace-cross Abstract: Controlling physics-based humanoids from natural-language instructions is a critical step toward gener…
Out-of-Distribution Semantic Occupancy Prediction
arXiv:2506.21185v3 Announce Type: replace-cross Abstract: 3D semantic occupancy prediction is crucial for autonomous driving, providing a dense, semantically ri…
Multi-Robot Bearing-based Pose Estimation via Angle Rigidity
arXiv:2606.03931v2 Announce Type: replace Abstract: This letter proposes a novel distributed pose estimator for multi-robot systems evolving on $\mathrm{SE}(3)$…
Risk-Aware Optimal Control with Rulebooks
arXiv:2609.05199v1 Announce Type: cross Abstract: We consider safety-critical control problems involving multiple requirements with different priorities and unc…
Pack It My Way: Triadic Human-Robot Collaboration for Personalized Autonomous Packing
arXiv:2609.04620v1 Announce Type: new Abstract: Personalized autonomous packing requires robots to account for resident preferences that cannot be inferred from…
Open-Set 3D Scene Graphs for Field Robotics: An Outdoor Case Study
arXiv:2609.04607v1 Announce Type: new Abstract: Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounded,…
Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI
arXiv:2609.04552v1 Announce Type: new Abstract: Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with huma…
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
arXiv:2609.04355v1 Announce Type: new Abstract: Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demandin…
Scalable Edge-assisted Fusion and Path Prediction for Connected Autonomous Vehicles
arXiv:2609.04364v1 Announce Type: new Abstract: The planning algorithms inside an Autonomous Vehicle (AV) rely on information from on-board sensors whose line o…
Dressing in Motion: A Human Motion-Aware Diffusion Policy for Robot-Assisted Dressing
arXiv:2609.04759v1 Announce Type: new Abstract: Robotic dressing assistance is a promising solution for supporting older adults with physical impairments in dai…
Coupled Control and Wireless World Models for Resilient Remote Robotic Control
arXiv:2609.04851v1 Announce Type: new Abstract: Remote robotic systems operating over wireless networks must maintain reliable control despite limited communica…