We introduce Act2Answer - an embodied evaluation protocol that adapts VLM knowledge benchmarks to VLA by requiring action-based answer selection instead of text generation.
The next generation of embodied agents must retain and use what they know about the world.
Simple concepts like Color and Shape remain behaviorally accessible in most VLAs.
On richer categories most VLAs fail and sit near chance.
VLAs drop ~20–40 points vs their source VLMs across most knowledge domains.
Answer-relevant signal is retained in the VLA backbone, but attenuates toward action generation.
VL co-training improves knowledge retention.
Extra SFT/RL action fine-tuning does not reliably help and can degrade performance.
VLMs (top) stay green across categories; VLAs (bottom) collapse toward chance on richer semantics.
| Model | Social | Physical | Quant. | Temporal | Normative | Cultural | Biological | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Emotion | Attribute | State | Color | Shape | Symmetry | Counting | Time | Public Info | Traffic | Celebrity | Living World | |
| VLM (action-free text probe) | ||||||||||||
| InternVL3.5-8B | 95% | 68% | 64% | 100% | 89% | 69% | 52% | 99% | 85% | 75% | 99% | 91% |
| InternVL3.5-38B | 99% | 73% | 68% | 100% | 96% | 83% | 59% | 100% | 94% | 81% | 100% | 96% |
| Ovis2.5-9B | 89% | 69% | 69% | 100% | 98% | 83% | 59% | 99% | 88% | 85% | 100% | 97% |
| Qwen2.5-7B | 89% | 64% | 68% | 100% | 90% | 78% | 62% | 99% | 80% | 86% | 100% | 94% |
| Qwen2.5-32B | 99% | 69% | 69% | 100% | 93% | 83% | 61% | 99% | 85% | 86% | 100% | 96% |
| Qwen3-8B | 86% | 68% | 67% | 100% | 97% | 81% | 65% | 98% | 83% | 93% | 100% | 95% |
| Qwen3-32B | 92% | 67% | 66% | 100% | 98% | 83% | 59% | 100% | 87% | 90% | 100% | 86% |
| Prismatic-VLM-7B | 82% | 59% | 61% | 96% | 85% | 67% | 52% | 96% | 75% | 76% | 99% | 82% |
| PaliGemma-3B | 53% | 50% | 52% | 47% | 48% | 49% | 49% | 51% | 48% | 49% | 48% | 49% |
| VLA (embodied action selection) | ||||||||||||
| OpenVLA | 48% | 51% | 49% | 89% | 64% | 45% | 48% | 49% | 49% | 46% | 50% | 52% |
| OpenVLA (SFT) | 41% | 45% | 44% | 82% | 53% | 38% | 46% | 37% | 47% | 46% | 42% | 45% |
| OpenVLA (RL) | 46% | 50% | 47% | 88% | 61% | 44% | 50% | 48% | 46% | 48% | 47% | 52% |
| SpatialVLA | 47% | 48% | 50% | 87% | 83% | 45% | 52% | 46% | 51% | 57% | 55% | 49% |
| π0 | 51% | 50% | 48% | 86% | 49% | 46% | 50% | 48% | 46% | 48% | 38% | 45% |
| Magma | 72% | 63% | 59% | 89% | 81% | 37% | 51% | 77% | 88% | 80% | 94% | 77% |
| Xiaomi-Robotics-R0 | 63% | 52% | 50% | 91% | 82% | 58% | 48% | 52% | 64% | 57% | 68% | 56% |
| InternVLA-M1 | 53% | 49% | 53% | 90% | 66% | 43% | 48% | 49% | 54% | 53% | 52% | 58% |
| Model | Prefix | Action | Retention |
|---|---|---|---|
| Magma | 75.23 | 72.60 | 0.8702 |
| Xiaomi-Robotics-R0 | 68.04 | 64.98 | 0.8159 |
| SpatialVLA | 65.70 | 62.60 | 0.7808 |
| OpenVLA | 68.71 | 64.61 | 0.7697 |
| SmolVLA | 63.18 | 57.73 | 0.5809 |
| π0 | 64.99 | 55.40 | 0.3620 |
At each VLA layer, we train a linear probe to predict the correct answer from hidden representations. We probe both the VLM backbone ( Prefix) and the Action Part.