PHYSICAL AI  /  AUTONOMOUS DRIVING

Grounding What Shapes The Plan

Rethinking Groundedness for Physical Intelligence in Autonomous Driving

GroundAct preserves the physical connection from reasoning to planning.

Minkyoung Cho1,2,†Zewei Zhou1,3,†Wenhao Ding1Shuhan Tan1Boyi Li1Yuxiao Chen1Yan Wang1Zheng Lian1Min-Hung Chen1Chaowei Xiao1,4Z. Morley Mao5Boris Ivanovic1Marco Pavone1,6Yulong Cao1
1NVIDIA2UMich3UCLA4JHU5USC6Stanford

† Work done during an internship at NVIDIA.

Entity As Grounding UnitToken-as-Reference As Grounding Mechanism

Make Grounding Consequential For Planning.

Interactive Illustration

01Start With The Physical Scene.

0:00

Schematic. Dots mark equal future-time intervals; refinement updates the whole proposal.

01 / THE QUESTION

Grounded Reasoning.
But Is The Plan Grounded?

Grounded reasoning does not by itself ensure desirable driving outcomes.

a

The Grounding-To-Plan Transfer Gap

b

Same Reasoning. Different Actions Required.

OUR CENTRAL QUESTION

What should groundedness mean when the output is an action?

↓
02 / THE PRINCIPLE

Keep The Physical Connection.

Driving unfolds through physical entities and their interactions. GroundAct keeps the states of reasoning-selected entities accessible through action generation, so their interactions can shape the plan.

Prior Interface

Token-as-Content

  • The token carries grounded information: causal statements, spatial representations, or predicted futures.
  • Planning is conditioned on this generated content.
  • The token itself does not provide an explicit address for retrieving the grounded physical state.

GroundAct

Token-as-Reference

  • Alongside generated content, reference tokens point to continuous entity states kept inside the model.
  • Reasoning selects the entities; subsequent reasoning and planning access the same states.
  • Their interactions with the evolving proposal guide the explicit plan-refinement pathway.
Read The Abstract

Driving models increasingly ground reasoning in causal relations, spatial structure, perceptual evidence, and predicted futures. These advances make reasoning more faithful to the driving scene, but leave a fundamental question unresolved: what should groundedness mean when the model ultimately outputs an action? Correctly grounded reasoning does not, by itself, ensure desirable driving outcomes. We introduce GroundAct, which starts from a simple premise: driving unfolds through physical entities and their interactions. Entities therefore become the unit of grounding; a lightweight reference token keeps each selected entity’s continuous state addressable through symbolic reasoning; and only the referenced entities’ interactions with the evolving proposal correct the plan. The result is an explicit path from what reasoning grounds to what the plan does, which we call grounded planning. To assess its practical value, we evaluate GroundAct in both open- and closed-loop settings. GroundAct shows strong open-loop planning across normal, out-of-distribution, and safety-critical scenarios, with closed-loop results extending this evidence to driving in simulation.

03 / THE METHOD

The Representation Changes.
The Physical Referent Does Not.

References resolve to continuous entity states shared by reasoning and planning. The action expert retains full scene context; the selected entities’ interactions guide an explicit refinement pathway.

GroundAct ArchitectureClick to enlarge ↗
04 / THE EVIDENCE

Grounding That Matters To Driving.

Driving performance and tests of how grounding shapes the plan.

Driving Performance

↑ Higher Is Better   ↓ Lower Is Better

Bold: Best  ·  C: Continuous  ·  D: Discrete  ·  † Re-implemented on the shared backbone.

Evaluation Details

Examining The Grounding Connection

Referencing, alignment, and selective action consequences.

Table 6 · Physical Grounding

Spatial Referencing

42.44F1@IoU
vs. 40.09 PETRv2

Geometry-aware referencing at BEV IoU ≥ 0.5.

Table 5 · Reasoning To Action

Reasoning–Action Alignment

68.44Reported F1
vs. 51.61 Alpamayo 1.5

Against predicted-trajectory meta-actions; continuous generation.

Table 3 · Selective Consequences

Selective Plan Correction

70.4%Safe → Preserved
2.2%Safe → Unsafe

Among initially safe plans, after correction.
Preserved: maximum deviation ≤ 3 m over 6.4 s.

Each measure retains its own evaluation protocol; these are not a combined score.

Grounding should shape
what the plan does.

BibTeX

@misc{cho2026groundact,
  title  = {Grounding What Shapes the Plan: Rethinking Groundedness
            for Physical Intelligence in Autonomous Driving},
  author = {Minkyoung Cho and Zewei Zhou and Wenhao Ding and Shuhan Tan
            and Boyi Li and Yuxiao Chen and Yan Wang and Zheng Lian
            and Min-Hung Chen and Chaowei Xiao and Z. Morley Mao
            and Boris Ivanovic and Marco Pavone and Yulong Cao},
  year   = {2026},
  note   = {Manuscript}
}

Draft citation; publication details will be added when available.