HumanoidGPT · Part of The Humanoid Group

Vision-Language-Action Models

Vision-language-action models connect three things: what a robot sees, what it is asked to do and how it moves. They are a widely discussed route to humanoids that people can direct in plain language.

By Arjun Rao · Updated

How a VLA model works

A vision-language-action model, or VLA, takes camera images and a written or spoken instruction as input and produces actions as output, such as arm and hand movements. Many VLAs start from a model already trained on images and text, which gives them a broad sense of objects and language, then learn actions from robot demonstrations. The result is a policy that can, in principle, follow an instruction like "put the cup on the shelf" in a scene it has not seen before.

Fast and slow thinking

Some systems split the work in two. A slower, larger model interprets the scene and the instruction and decides what to do next, while a faster, smaller controller turns that plan into smooth, precise movement many times a second. Others use a single model end to end. Each design balances responsiveness, generality and computing demands differently. For buyers, the practical question is less about architecture and more about how the system behaves on your tasks, in your lighting, with your objects.

Strengths and limits

VLAs make robots easier to direct and can generalise to new objects and phrasing better than hand-written rules. They can also be unpredictable: a model may perform well in one setting and fail in another for reasons that are hard to diagnose. Safety therefore still depends on conventional layers such as speed limits, force limits, protective stops and well-designed workcells. Ask any maker how learned behaviour is bounded, monitored and updated over time.

Sources and further reading

Common questions

Can I talk to a humanoid running a VLA model?

VLA models are built to take instructions in natural language, typed or transcribed from speech. Holding a conversation is a different ability, so check what the robot actually supports.

Do VLA models learn on the job?

Usually not on their own. Most are trained offline and updated by the maker. Some systems support fine-tuning with new demonstrations, which is worth asking about.

Visit The Humanoid Group

People also search for vision language action model, VLA model, what is a VLA model, VLA robotics, vision language model for robots and AI robot that understands language.