Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
3583 storiesMyGO-Splat: Multi-Objective Closed-Loop Geometric Feedback for RGB-Only Gaussian SLAM
arXiv:2606.29738v1 Announce Type: new Abstract: Real-time monocular Simultaneous Localization and Mapping (SLAM) fundamentally suffers from scale ambiguity and …
VISTA-DZ: Visual Semantic Trajectory Adaptation for Personalized Dilemma Zone Prediction
arXiv:2606.29548v1 Announce Type: cross Abstract: Driver decision making in the dilemma zone at signalized intersections is safety critical, as vehicles approac…
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
arXiv:2510.03142v2 Announce Type: replace Abstract: Visual navigation policy is widely regarded as a promising direction, as it mimics humans by using egocentri…
RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis
arXiv:2606.28385v1 Announce Type: new Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. Ho…
Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients
arXiv:2606.28345v1 Announce Type: new Abstract: LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms…
Trajectory Optimization for Collision-Aware Redundant Robotic Multi-Axis Additive Manufacturing by Constrained Gradient Projection
arXiv:2606.29766v1 Announce Type: new Abstract: Redundant robotic multi-axis additive manufacturing (MAAM) enables support-free and conformal fabrication, but t…
RoamFlow: Reinforcement-Aligned One-Step Action MeanFlow Policy for Image-Goal Navigation
arXiv:2606.29934v1 Announce Type: new Abstract: Image-goal navigation is a key challenge in embodied robotics, where an agent must reach a target specified sole…
Cross-Spectral Stereo Inertial Odometry
arXiv:2606.29757v1 Announce Type: new Abstract: Standard stereo VIO focuses exclusively on the benefit of metric scale via single-spectrum baselines, often over…
Flying to Image-Specified Objects: 3D Quadrotor Navigation via Cross-Graph Memory and Viewpoint Planning
arXiv:2606.29917v1 Announce Type: new Abstract: Instance-Specific Image-Goal Navigation (InstanceImageNav) requires a robot to navigate toward the exact object …
Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model
arXiv:2606.29384v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA mo…
DRIVE-Nav: Directional Reasoning, Inspection, and Verification for Efficient Open-Vocabulary Navigation
arXiv:2603.28691v2 Announce Type: replace Abstract: Open-Vocabulary Object Navigation (OVON) requires an embodied agent to locate a language-specified target in…
Robotic Arm-Based Spectral Sensing for Strawberry Positioning and Non-Destructive Sweetness Measurement
arXiv:2606.28555v1 Announce Type: new Abstract: Accurate assessment of sweetness is essential for quality control in agriculture, yet conventional methods rely …
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
arXiv:2605.07306v2 Announce Type: replace Abstract: Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable…
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
arXiv:2512.13660v3 Announce Type: replace Abstract: Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it r…
SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance
arXiv:2606.30113v1 Announce Type: new Abstract: Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recove…
PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking
arXiv:2606.30476v1 Announce Type: cross Abstract: We introduce Point-supervised Multi-Object Tracking (PS-MOT) as a cost-effective alternative to traditional bo…
Demonstration-Free Robotic Control via LLM Agents
arXiv:2601.20334v2 Announce Type: replace Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong perf…
SCREP: Scene Coordinate Regression and Evidential Learning-based Perception-Aware Trajectory Generation
arXiv:2507.07467v3 Announce Type: replace Abstract: Autonomous flight in GPS-denied indoor spaces requires trajectories that keep visual-localization error tigh…
VibES: Induced Vibration for Persistent Event-Based Sensing
arXiv:2508.19094v3 Announce Type: replace-cross Abstract: Event cameras are a bio-inspired class of sensors that asynchronously measure per-pixel intensity chan…
Privacy-Preserving Decentralized Cooperative Localization with Range-Only Measurements: A Convex Optimization Based Approach
arXiv:2606.29673v1 Announce Type: new Abstract: Cooperative localization using range-based measurements is critical for multi-robot systems operating in GPS-den…
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
arXiv:2603.19235v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer…
Heterogeneous Tactile Transformer
arXiv:2606.29948v1 Announce Type: new Abstract: Tactile sensors are inherently heterogeneous: a model trained on one sensor cannot be directly used on another, …
HJ-SafeDMP: Hamilton-Jacobi Reachability-Guided Dynamic Movement Primitives for Provably Safe Robot Motion
arXiv:2606.28995v1 Announce Type: new Abstract: Robots deployed in safety-critical environments must execute motions that are simultaneously robust to disturban…
AERIS: Aerial-Edge Role-Driven Intelligence at Runtime via Orchestrated Language-Model Swarm
arXiv:2606.30151v1 Announce Type: new Abstract: Integrating large language models into robotic systems holds promise for enhancing autonomy, yet practical deplo…
Generation of Uncertainty-Aware High-Level Spatial Concepts in Factorized 3D Scene Graphs via Graph Neural Networks
arXiv:2409.11972v4 Announce Type: replace Abstract: Enabling robots to autonomously discover high-level spatial concepts (e.g., rooms and walls) from primitive …
Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation
arXiv:2606.29908v1 Announce Type: new Abstract: Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, deco…
Sphere-VIO: Fast and Robust Visual-Inertial Odometry via Unified Spherical Representation for Heterogeneous Multi-Camera Systems
arXiv:2606.29910v1 Announce Type: new Abstract: Multi-camera visual-inertial odometry (VIO) overcomes the inherent limitations of pure visual systems by expandi…
CSAR: Containerized System Architecture for Robotics
arXiv:2606.30293v1 Announce Type: new Abstract: Robotic applications increasingly rely on distributed computational infrastructures that combine embedded device…
An Overview of Formulae for the Higher-Order Kinematics of Lower-Pair Chains with Applications in Robotics and Mechanism Theory
arXiv:2309.05055v2 Announce Type: replace Abstract: The motions of mechanisms can be described in terms of screw coordinates by means of an exponential mapping.…
Vision-Language-Action Models: Experimental Insights from a Real-World UR5 Platform
arXiv:2606.30456v1 Announce Type: new Abstract: This project investigates whether recent Vision-Language-Action (VLA) models can be transferred from controlled …