OmniGuide

Abstract

Vision-language-action (VLA) models have shown great promise as generalist policies for a large range of relatively simple tasks. However, they demonstrate limited performance on more complex tasks, such as those requiring complex spatial or semantic understanding, manipulation in clutter, or precise manipulation. We propose OmniGuide, a flexible framework that improves VLA performance on such tasks by leveraging arbitrary sources of guidance, such as 3D foundation models, semantic-reasoning VLMs, and human pose models. We show how many kinds of guidance can be naturally expressed as differentiable energy functions with task-specific attractors and repellers located in 3D space, that influence the sampling of VLA actions. In this way, OmniGuide enables guidance sources with complementary task-relevant strengths to improve a VLA model's performance on challenging tasks. Extensive experiments in both simulation and real-world environments, across diverse sources of guidance, demonstrate that OmniGuide significantly enhances the performance of state-of-the-art generalist policies (e.g., π0.5, GR00T N1.6) across success (68.2%) and safety (86.5%) rates. Critically, our unified framework matches or surpasses the performance of prior methods designed to incorporate specific sources of guidance into VLA policies.

Method

Generalist robot policies (VLAs) are often "jacks-of-all-trades, masters of none." While they understand broad instructions, they often lack the "last-mile" precision needed for complex spatial reasoning or avoiding tight collisions. OmniGuide addresses this by providing inference-time guidance leveraging external sources of information, such as 3D foundation models, semantic-reasoning VLMs, and human pose models.

Results

We evaluate OmniGuide on a diverse set of tasks spanning three modalities of guidance against the baseline π0.5 in the real-world and GR00T N1.6 in simulation using the RoboCasa environments. See the results below! OmniGuide consistently improves base VLAs on all tasks.

OmniGuide Universal Guidance Fields for Enhancing Generalist Robot Policies

Abstract

Method

Three Modalities of Guidance

Results

BibTeX