Vision-Language-Action Models

Vision-language-action (VLA) models are AI systems that combine visual perception, language understanding, and motor control in a single model, letting a robot interpret a scene, follow an instruction, and act on it directly. This hub covers VLA research, releases, and the labs and robots built on the approach.

Combine with: