Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
6273 storiesStability Control for Real World Testing in Autonomous Racing
arXiv:2608.17779v1 Announce Type: new Abstract: Controlling an autonomous vehicle at the limits of handling is a challenging task. Due to external influences, s…
OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects
arXiv:2608.17633v1 Announce Type: new Abstract: Integrating open-vocabulary perception into object-level 3D scene graphs is a double-edged sword. While vision-l…
Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision
arXiv:2608.17628v1 Announce Type: new Abstract: Developing robots capable of understanding and manipulating objects requires compact, interpretable, and general…
Physics-Informed Sliding-Window Particle Filtering for Tactile-Only In-Hand 6-DoF Object Pose Refinement
arXiv:2608.17601v1 Announce Type: new Abstract: This paper studies tactile-only 6-DoF pose refinement and belief maintenance for grasped objects in static and s…
LIBERO-VIFO: Benchmarking the Capability and Safety of Visual Cue Following in Vision-Language-Action Models
arXiv:2608.17600v1 Announce Type: new Abstract: Visual cues are increasingly adopted to guide robot learning, but whether Vision-Language-Action (VLA) models ca…
Scalix: Uncertainty-Aware Scale-Consistent Monocular SLAM
arXiv:2608.17553v1 Announce Type: new Abstract: Cameras are ubiquitous sensors in robotics due to their compact form factor and the perceptual richness captured…
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
arXiv:2608.17512v1 Announce Type: new Abstract: Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deplo…
EATR-Stereo: Embodiment-Aware Routing of Paired Stereo Evidence for Humanoid Vision-Language-Action Control
arXiv:2608.17453v1 Announce Type: new Abstract: Long-horizon humanoid vision--language--action (VLA) control with head-mounted stereo cameras requires visual in…
Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups
arXiv:2608.17423v1 Announce Type: new Abstract: GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PP…
Bi-Layer Ant Colony Optimization for Multi-Robot Task Allocation and Routing in Delivery Applications
arXiv:2608.17416v1 Announce Type: new Abstract: This paper addresses the multi-robot task allocation (MRTA) problem, which is essential for delivery and logisti…
ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback
arXiv:2608.17323v1 Announce Type: new Abstract: Robotic manipulation policies trained via imitation learning, such as Action Chunking with Transformers (ACT), c…
MANIGUARD: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation
arXiv:2608.17386v1 Announce Type: new Abstract: Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluatio…
Teach and Grow: An Agent-Centered Architecture for General Robot Learning
arXiv:2608.17209v1 Announce Type: new Abstract: End-to-end vision-language-action (VLA) and world-action models offer an elegant route to general-purpose roboti…
VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation
arXiv:2608.16978v1 Announce Type: new Abstract: Turning a frontier vision-language model into a robot policy usually means fine-tuning it to emit an action repr…
Symmetry-Breaking in Multi-Agent Navigation: Winding Number-Aware MPC with a Learned Topological Strategy
arXiv:2511.15239v3 Announce Type: replace Abstract: In decentralized multi-agent navigation, agents that independently compute their controls without communicat…
Parallel Branch Model Predictive Control on GPUs
arXiv:2506.13624v2 Announce Type: replace-cross Abstract: We present a GPU-based solver for trajectory planning problems using branch Model Predictive Control. …
Planning-aligned Token Compression for Long-Context Autonomous Driving
arXiv:2606.07464v2 Announce Type: replace Abstract: Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architec…
Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning
arXiv:2512.00074v4 Announce Type: replace Abstract: Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underpe…
A Diffusion-Refined Planner with Reinforcement Learning Priors for Confined-Space Parking
arXiv:2510.14000v2 Announce Type: replace Abstract: The growing demand for parking has increased the need for automated parking planning methods that can operat…
Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT
arXiv:2508.08748v2 Announce Type: replace Abstract: Robotic pick-and-place tasks in convenience stores pose challenges due to dense object arrangements, occlusi…
ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
arXiv:2406.01586v4 Announce Type: replace Abstract: Diffusion models have been verified to be effective in generating complex distributions from natural images …
Training with synthetic data for drone detection in thermal imagery
arXiv:2608.17799v1 Announce Type: cross Abstract: Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due t…
Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning
arXiv:2608.17347v1 Announce Type: cross Abstract: Repetition is a fundamental mechanism in human learning, where revisiting successful experiences strengthens m…
Hydra-0: Action Flow for Generalist World Modeling and Control
arXiv:2608.18077v1 Announce Type: new Abstract: We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pix…
CompCPZ: Preserving Multi-Modal Intent in Language-Guided Robot Manipulation
arXiv:2608.17717v1 Announce Type: new Abstract: A robot asked to "place the cup near the red plate or the blue plate" may reach the centroid between them and ap…
Force-Based Offset Estimation for Keyed Peg-in-Hole Assembly Using Local Gaussian Process Regression
arXiv:2608.17691v1 Announce Type: new Abstract: Key-keyway assembly tasks impose strict geometric constraints and are highly sensitive to grasp pose deviations …
tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots
arXiv:2608.17596v1 Announce Type: new Abstract: In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as c…
Reconfiguration-Complete Motion Primitives with Constructive Planning for Deformable Planar Modular Robots
arXiv:2608.17324v1 Announce Type: new Abstract: The continuously deformable geometry of modular robots makes it difficult to define a fixed representation for r…
FetchMan: Learning Visual Humanoid Loco-Manipulation Policies from Simulated Experiences
arXiv:2608.17027v1 Announce Type: new Abstract: Visual loco-manipulation policies that can generalize to novel scenes and objects have long been a goal of robot…
SpotlessGS: Relightable 3D Gaussian Splatting under Dynamic Illumination for Robotic Perception
arXiv:2608.14713v1 Announce Type: new Abstract: Robots operating in dark or poorly lit environments rely on onboard lights, which often produce uneven illuminat…