Robotics & Embodied AI

Vision-language-action models from edge to cloud — and where world models take embodied intelligence next

Goal: understand and optimize the policies that let machines act in the physical world — across the full hardware spectrum robots actually ship with.

Cross-platform VLA scaling

Five representative vision-language-action models evaluated on the LIBERO benchmark from power-constrained edge SoCs to datacenter GPUs, measuring task accuracy jointly with latency, throughput, and peak memory. Architectural choices dominate throughput and memory; edge devices degrade non-linearly yet well-chosen edge configurations match or exceed older datacenter GPUs — challenging the assumption that robotic inference belongs in the cloud. Published at GLSVLSI 2026 (Taherin et al., 2026).

VOTE — efficient VLA optimization

A training framework that finetunes VLA models to emit far fewer action tokens, plus a voting-based ensemble over action trajectories: 39× faster inference than OpenVLA with 46 Hz throughput on edge platforms, at higher success rates (Lin et al., 2025).

World models

With collaborators, a unified perspective on world models organized by the cognitive functions each line of work innovates — memory, perception, language, reasoning, imagining, motivation, and metacognition — identifying motivation and metacognition as drastically under-researched and introducing epistemic world models for scientific discovery (Rupprecht* et al., 2026). From my systems perspective: predictive models of physical environments add sustained memory, latency, and energy pressure exactly where they can least be afforded — the edge.

References

2026

  1. GLSVLSI
    Amir Taherin, Juyi Lin, Arash Akbari, and 5 more authors
    In Proceedings of the Great Lakes Symposium on VLSI (GLSVLSI), 2026
  2. arXiv
    Timothy Rupprecht*, Pu Zhao*, Amir Taherin*, and 20 more authors
    2026

2025

  1. arXiv
    Juyi Lin, Amir Taherin, Arash Akbari, and 11 more authors
    2025