Vision-Language-Action Autonomous Driving Agent with Language-Based Memory

Kai Yan* · Xiangyu Chen · Yulong Cao · Alex Naumann · Peter Karkus · Yan Wang · Jef Packer ·
Alexander Schwing · Yuxiong Wang · Boris Ivanovic · Wenjie Luo · Marco Pavone

University of Illinois Urbana-Champaign · NVIDIA

*Corresponding author (kaiyan3@illinois.edu). Work done as an intern at NVIDIA.

Can we fully exploit the knowledge in a VLA's backbone by using language-based, in-episode memory that is brief, explainable, and portable?

Memory Helps Autonomous Driving

VLA agents cannot take too many frames as input due to the high token cost of images. Thus, the current image input can be ambiguous when the relevant evidence appeared only seconds earlier. Below shows example of how memory can help the agent make the right driving decision.

All-way-stop visual input with the approaching left car marked in red
All-way stop: who arrived first? When ego and the left car are both stopped, the current image input does not reveal arrival order. Memory is needed to decide who should yield.
Click to see how memory helps
Current general-driving visual input with the pedestrian marked in red
Occluded object: is there a pedestrian entering the road? A pedestrian becomes occluded by a crosswalk sign. Earlier observations make the pedestrian easier to track and clarify the intention to enter the road.
Click to see how memory helps

AD-Memo: Autonomous Driving with Language-Based Memory

AD-Memo is a streaming video agent switching between two modes: driving and VQA. We divide a video by visual input horizon t, and consider the n+1 "keyframes" each separated by t; in keyframe 1 to n, the agent is in driving mode, generates trajectories and language memory. The memory becomes part of the agent's future input. At keyframe n+1, the agent switches to VQA mode and answers question based on the memory.

Memory is emitted as an extension of Chain-of-Thought. A rolling window of previous entries becomes language input at the next keyframe, while the model continues to produce its 6.4-second driving trajectory.

AD-Memo combines earlier language memory with current visual input to make the right driving decision.
Language memory recovers temporal evidence that is absent from the current visual input. The same memory supports driving, scene-understanding VQA, and plug-and-play use by other models.

Driving and VQA Mode

The idea of appending VQA at the end of the video is that, by asking a question on critical objects that affects driving, the agent must remember every critical object to answer the question correctly as it does not know what the question will be during driving. When deployed, the VQA mode will never be activated, but the agent will still generate relevant memory. VQA is also a good way to evaluate the effectiveness of the agent's memory.

BriefTracks only the decision-relevant objects of every past frame.
ExplainableExposes what the agent retained in a human-readable way, fully exploiting VLM backbone reasoning capabilities.
PortableCan be passed to another model without architecture-specific memory modules.

Da Capo: Semi-Closed-Loop Reinforcement Learning

Supervised training never exposes the model to its own imperfect memory. Da Capo addresses this exposure bias without requiring an expensive, fully interactive driving simulator.

Open-loop in control, closed-loop in memory

Across K parallel rollouts, visual observations and past ego trajectories are replayed from ground truth. Each rollout nevertheless consumes the memories it generated at preceding steps.

Different Advantages for Heterogeneous Tokens

Driving tokens in open-loop control are conditionally independent from future rollouts, but memory tokens must receive reward signals from future; to address this heterogeneity, we propose Deviation-Adjusted Causal Advantage Policy Optimization (Da Capo) for better credit assignment.

  • Driving tokens: receive only stepwise advantage as a predicted trajectory does not alter later replayed states.
  • Memory tokens: receive future-dependent advantage from downstream driving and terminal VQA rewards.
  • Deviation adjustment: standard-deviation normalization balances heterogeneous reward scales.
Da Capo uses stepwise advantage for driving tokens and trajectory-wise advantage for memory and VQA
Da Capo improves credit assignment by applying only causally relevant reward components to each output block. See our paper for theoretical proofs on equivalence between the expectation of the gradient of Da Capo and that of trajectory-level GRPO.

How to Curate Driving-Critical Memory?

We curate two datasets: all-way stop and general driving. For the former, object of interest are the cars in the crossing and thus we can generate rule-based memory; for the latter, we need to identify objects that really affects driving decisions.

Original visual input for decision graph construction

Original front-camera observation.

Driving-relevant objects grounded in the visual input

Driving-relevant actors, map elements, and traffic rules are grounded in the image.

Decision graph connecting grounded evidence to the calibrated action

Evidence nodes connect to the calibrated ego action through causal, spatial, and regulatory relations.

Decision graphs keep memory relevant.

Describing every visible object would make language memory long and distracting. We curate a novel agentic pipeline that builds a decision graph, which identifies which grounded entities actually affect the expert action.

Memory is generated from action-connected nodes and then merged across the clip to preserve stable references and temporal continuity.

Experimental Results

The tables below reports the main results reported in the paper. ↓ indicates lower is better; ↑ indicates higher is better.

All-Way Stop

All-way stop dataset serves as a task-specific evaluation where we know memory are important and know what memory are needed.

Table 1. Results on 2,479 test clips and 50,533 keyframes. “Ref. mem” uses reference memory generated in the same way as SFT labels. ML. = Most Likely, SR = Success Rate, Δpos = Position Difference, Δdur = Duration Difference, Roll = Roll-through Rate, and MCQ = Multiple Choice Question Accuracy.

ModelminADE ↓Avg. ADE ↓ML. ADE ↓ Stop SR ↑Go SR ↑Δpos ↓Δdur ↓ Roll ↓MCQ Acc. ↑
Base Model1.3862.3512.39382.06%33.74%1.7251.87111.79%0%
Alpamayo 2 Super 34B0.9932.293N/A78.75%34.30%2.1821.48815.94%33.19%
No mem. (SFT only)1.1022.1852.31887.04%38.59%1.3581.5646.62%33.77%
CoT mem. (SFT only)1.0962.1712.29686.89%39.37%1.3771.5316.69%45.41%
AD-Memo (SFT only)1.0492.1212.25687.48%40.66%1.2571.4696.05%45.99%
CoT mem. (SFT + ref. mem)1.0982.1692.28686.95%39.45%1.3671.5216.61%45.24%
AD-Memo (SFT + ref. mem)0.9631.9632.08687.70%44.66%1.1851.2965.84%89.53%
No mem. (SFT + Da Capo)0.9441.9522.00588.92%43.76%1.1251.2485.21%35.32%
CoT mem. (SFT + Da Capo)0.9681.9061.94289.20%44.88%1.1571.1945.31%45.87%
AD-Memo (Ours)0.9511.8661.91589.26%45.01%1.0701.1835.26%51.66%

General Driving

General driving tests the agent's ability to generalize to diverse driving scenarios.

Table 2. Results on 9,112 test clips and 145,462 keyframes. Corner = Corner distance.

ModelminADE ↓Avg. ADE ↓ML. ADE ↓ minFDE ↓Avg. FDE ↓ML. FDE ↓Corner ↓MCQ Acc. ↑
Base Model1.0321.9812.0492.8535.9376.1941.0010%
Alpamayo 2 Super 34B0.9542.081N/A2.5566.141N/A0.90225.71%
No mem. (SFT only)1.0072.0332.0752.7166.1086.2790.96858.02%
CoT mem. (SFT only)1.0312.0592.1052.7876.1926.3750.99357.51%
AD-Memo (SFT only)1.0052.0112.0492.7056.0426.1960.96765.63%
CoT mem. (SFT + ref. mem)1.0342.0642.1052.7956.2076.3730.99557.36%
AD-Memo (SFT + ref. mem)1.0122.0142.0512.7346.0536.2080.97467.65%
No mem. (SFT + Da Capo)1.0901.8831.8913.0425.5835.6481.06257.43%
CoT mem. (SFT + Da Capo)1.1051.9071.9183.0715.6315.7071.07956.94%
AD-Memo (Ours)1.0911.8641.8693.0275.5125.5761.06665.51%

Memory Portability

AD-Memo is special for its portability of the memory: it can be used by other models in a plug-and-play, 0-shot manner.

Table 3. Accuracy of GPT-5.6 Luna on LingoQA and WaymoQA using AD-Memo's generated memory.

BenchmarkOur memory + last frameOnly last frameFull video
WaymoQA71.65%66.63%75%
LingoQA67%62.6%70%