Act2Answer: Does VLA Even Know the Basics?

We introduce Act2Answer - an embodied evaluation protocol that adapts VLM knowledge benchmarks to VLA by requiring action-based answer selection instead of text generation.

Act2Answer episodes

Fine-tuning VLMs on actions isn't enough.

The next generation of embodied agents must retain and use what they know about the world.

What we find

RQ1 · Simple primitives

Simple concepts like Color and Shape remain behaviorally accessible in most VLAs.

RQ2 · Richer semantics

On richer categories most VLAs fail and sit near chance.

RQ3 · The VLM–VLA gap

VLAs drop ~20–40 points vs their source VLMs across most knowledge domains.

RQ4 · Where knowledge goes

Answer-relevant signal is retained in the VLA backbone, but attenuates toward action generation.

RQ5 · VL supervision

VL co-training improves knowledge retention.

RQ6 · Downstream tuning

Extra SFT/RL action fine-tuning does not reliably help and can degrade performance.

Act2Answer Construction

Data curation pipeline

Main results

VLMs (top) stay green across categories; VLAs (bottom) collapse toward chance on richer semantics.

<46% 46–53% 54–59% 60–89% ≥90%
Model SocialPhysicalQuant. TemporalNormativeCulturalBiological
EmotionAttributeStateColorShapeSymmetryCounting TimePublic InfoTrafficCelebrityLiving World
VLM (action-free text probe)
InternVL3.5-8B95%68%64%100%89%69%52%99%85%75%99%91%
InternVL3.5-38B99%73%68%100%96%83%59%100%94%81%100%96%
Ovis2.5-9B89%69%69%100%98%83%59%99%88%85%100%97%
Qwen2.5-7B89%64%68%100%90%78%62%99%80%86%100%94%
Qwen2.5-32B99%69%69%100%93%83%61%99%85%86%100%96%
Qwen3-8B86%68%67%100%97%81%65%98%83%93%100%95%
Qwen3-32B92%67%66%100%98%83%59%100%87%90%100%86%
Prismatic-VLM-7B82%59%61%96%85%67%52%96%75%76%99%82%
PaliGemma-3B53%50%52%47%48%49%49%51%48%49%48%49%
VLA (embodied action selection)
OpenVLA48%51%49%89%64%45%48%49%49%46%50%52%
OpenVLA (SFT)41%45%44%82%53%38%46%37%47%46%42%45%
OpenVLA (RL)46%50%47%88%61%44%50%48%46%48%47%52%
SpatialVLA47%48%50%87%83%45%52%46%51%57%55%49%
π051%50%48%86%49%46%50%48%46%48%38%45%
Magma72%63%59%89%81%37%51%77%88%80%94%77%
Xiaomi-Robotics-R063%52%50%91%82%58%48%52%64%57%68%56%
InternVLA-M153%49%53%90%66%43%48%49%54%53%52%58%

Knowledge Is Present in VLA — But Not Recoverable in Actions

Validation Acc. (%)
Layer Index
ModelPrefixActionRetention
Magma75.2372.600.8702
Xiaomi-Robotics-R068.0464.980.8159
SpatialVLA65.7062.600.7808
OpenVLA68.7164.610.7697
SmolVLA63.1857.730.5809
π064.9955.400.3620

At each VLA layer, we train a linear probe to predict the correct answer from hidden representations. We probe both the VLM backbone ( Prefix) and the Action Part.

Where Does the Knowledge Go?

  • Answer-relevant information is linearly recoverable in the VLM backbone, but attenuates on the path to the Action Part.