Robotics

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

Robos News Newsroom

Editorial Desk

2026-06-10 · 2 min read

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

Published June 10, 2026 · Category: Robotics

Overview

arXiv:2606.10495v1 Announce Type: new Abstract: Safe social navigation requires robots to distinguish people from ordinary obstacles and to react before danger becomes imminent. We show that pretrained Vision-Language-Action (VLA) models already encode pedestrian-object distinctions and future collision signals in their internal representations, but behavior cloning fails to translate these signals into socially appropriate actions. To address this mismatch, we propose SALSA, a two-stage annotation-free post-training framework: (1) social behavioral alignment bridges intermediate-layer social features to the action head and trains on counterfactual human-object scene pairs to break visual saliency shortcuts; (2) temporal safety alignment provides automatically generated future-risk supervision to enable anticipatory collision avoidance. On SCAND and real-world deployment, SALSA reduces near-collisions by 86.4% and improves social counterfactual accuracy from 53% to 93%, demonstrating that safer social navigation can be achieved by teaching VLA policies to act on representations they already possess. These results show that pretrained VLA policies can be adapted for safer social navigation by better aligning their latent representations with action generation.

Source

Originally published at arxiv.org.

Robos News Newsroom

Robos News covers markets, crypto and commodities for Asia & the Middle East — tier-1 desk research, AI-driven analysis, institutional-grade data. Tip our newsroom: [email protected]

Email the newsroom →

Disclaimer: This article is for informational purposes only and does not constitute investment advice. Data may be delayed up to 15 minutes. Past performance is not indicative of future results. Consult a licensed financial advisor before making investment decisions.

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

Overview

Source

Related Articles

Related Stories

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

Overview

Source

Related Articles

Related Stories

Human-robot interaction design retreat

Learning robust controllers that work across many partially observable environments

Robot Talk Episode 135 – Robot anatomy and design, with Chapa Sirithunge

Why companies don’t share AV crash data – and how they could

Cookie Preferences