Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6273 storiesKoopman Representation of Nonlinear Virtual Environments in Kinesthetic Haptic Systems
arXiv:2608.11461v1 Announce Type: new Abstract: Rendering haptic feedback with nonlinear virtual environments (VEs) is important in many applications that requi…
Repurposing RGB-based Foundation Model for Depth Estimation on Thermal Images Using Hierarchical Supervision
arXiv:2608.11564v1 Announce Type: cross Abstract: Depth estimation from thermal images is highly valuable for robotic applications in adverse conditions, such a…
Towards the Harness of Embodied Agents
arXiv:2608.11246v1 Announce Type: cross Abstract: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on …
LiDAR-based 3D Change Detection at City Scale
arXiv:2510.21112v3 Announce Type: replace-cross Abstract: High-definition 3D city maps enable city planning and change detection, which is essential for municip…
Utilizing Inpainting for Keypoint Detection for Vision-Based Control of Robotic Manipulators
arXiv:2604.13309v2 Announce Type: replace Abstract: We present a novel visual servoing framework for controlling a robotic manipulator in configuration space us…
RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
arXiv:2510.06710v3 Announce Type: replace Abstract: Recent studies have demonstrated the potential of reinforcement learning (RL) to improve the task performanc…
FEWT: Frequency-Enhanced Wavelet-based Transformer for Multimodal Wheeled Bimanual Manipulation
arXiv:2509.11109v4 Announce Type: replace Abstract: Embodied intelligence bridges the physical world and information spaces, with robots demonstrating immense p…
First-order friction models with bristle dynamics: lumped and distributed formulations
arXiv:2602.09429v3 Announce Type: replace-cross Abstract: Dynamic models, particularly rate-dependent models, have proven effective in capturing the key phenome…
RoboHarness: A Memory-Augmented Policy Harness for Vision-Language-Action Model Robustness via In-Context Adaptation
arXiv:2603.24060v3 Announce Type: replace Abstract: Despite the promise of Vision-Language-Action (VLA) models as generalist robotic controllers, their robustne…
A Generalized Theory of Load Distribution in Redundantly-actuated Robotic Systems
arXiv:2603.11431v2 Announce Type: replace Abstract: This paper presents a generalized theory which describes how applied loads are distributed within rigid bodi…
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
arXiv:2603.08131v2 Announce Type: replace Abstract: 3D Visual Grounding (3DVG) localizes objects from natural-language descriptions in 3D scenes and is fundamen…
Keep the Future, Drop the Rollout: RIFT for World Action Models
arXiv:2608.11521v1 Announce Type: new Abstract: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases d…
IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework
arXiv:2608.11597v1 Announce Type: new Abstract: As smart port infrastructures increasingly rely on autonomous maritime devices enabled by the Internet of Things…
ContactIPM: A Structure-Exploiting Interior-Point Solver for Contact-Implicit Trajectory Optimization
arXiv:2608.11731v1 Announce Type: new Abstract: Contact-implicit trajectory optimization avoids prescribing contact sequences, but yields mathematical programs …
RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation
arXiv:2608.11580v1 Announce Type: new Abstract: Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluatio…
Video2Track: From Real-World Interaction Videos to Steerable Adversarial Closed-Track Testing for Automated Driving Systems
arXiv:2608.11592v1 Announce Type: new Abstract: Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (A…
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
arXiv:2608.11671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often …
Policy-Induced Hand Priors in Humanoid Dual-Arm Manipulation: Diagnosing and Mitigating Initial-Pose Dependence
arXiv:2608.11769v1 Announce Type: new Abstract: Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial …
Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)
arXiv:2608.11229v1 Announce Type: cross Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align rob…
Self-Evolving Embodied Agents via Skill-Harness Evolution
arXiv:2608.11350v1 Announce Type: cross Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only…
Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards
arXiv:2608.11451v1 Announce Type: new Abstract: Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that…
World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
arXiv:2608.09730v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficie…
Scalable Multi-Agent Maze Traversal with Local Communication
arXiv:2608.11895v1 Announce Type: new Abstract: Cave networks, pipe systems, and similar maze-like environments pose significant challenges for multi-agent navi…
Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation
arXiv:2608.11870v1 Announce Type: new Abstract: In vision-based behavior cloning (BC), conventional image augmentations such as Random Crop and Color Jitter oft…
G0.5: One Autoregressive Stream for Robot Reasoning and Action
arXiv:2608.11739v1 Announce Type: new Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained…
Adaptation of Generalist Robot Policies with Minimal Data
arXiv:2608.11363v1 Announce Type: new Abstract: A central goal in robot learning is to move beyond task-specific human data collection toward robots that improv…
Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment
arXiv:2608.12198v1 Announce Type: new Abstract: Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches…
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
arXiv:2608.12122v1 Announce Type: new Abstract: Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the hi…
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
arXiv:2608.12063v1 Announce Type: new Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Lear…
DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements
arXiv:2608.11901v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) has progressively expanded from indoor to outdoor environments. However, ex…