Cross-Platform VLA Scaling

How vision-language-action models scale from power-constrained edge SoCs to datacenter GPUs

Vision-language-action models are becoming the standard policy class for robotic control, but robots ship with edge SoCs while most VLA benchmarking happens on datacenter GPUs.

We evaluated five representative VLA models — state-of-the-art baselines and two newly proposed architectures — on the LIBERO benchmark across edge platforms under varying power budgets and across datacenter GPU configurations, measuring task accuracy jointly with latency, throughput, and peak memory.

Key findings:

  • Architectural choices (action tokenization, backbone size) strongly influence throughput and memory footprint, independent of raw hardware capability.
  • Power-constrained edge devices degrade non-linearly — yet well-chosen edge configurations can match or exceed older datacenter GPUs, challenging the assumption that robotic inference belongs in the cloud.
  • High-throughput VLA variants are achievable without significant accuracy loss.

Published at GLSVLSI 2026 (Taherin et al., 2026); the related VOTE trajectory-ensemble optimization (Lin et al., 2025) reaches 46 Hz throughput on edge platforms with public code.

References

2026

  1. GLSVLSI
    Amir Taherin, Juyi Lin, Arash Akbari, and 5 more authors
    In Proceedings of the Great Lakes Symposium on VLSI (GLSVLSI), 2026

2025

  1. arXiv
    Juyi Lin, Amir Taherin, Arash Akbari, and 11 more authors
    2025