HumanoidGPT · Part of The Humanoid Group
Vision-Language-Action Models
Vision-language-action models connect three things: what a robot sees, what it is asked to do and how it moves. They are a widely discussed route to humanoids that people can direct in plain language.
By Arjun Rao · Updated
How a VLA model works
A vision-language-action model, or VLA, takes camera images and a written or spoken instruction as input and produces actions as output, such as arm and hand movements. Many VLAs start from a model already trained on images and text, which gives them a broad sense of objects and language, then learn actions from robot demonstrations. The result is a policy that can, in principle, follow an instruction like "put the cup on the shelf" in a scene it has not seen before.
Fast and slow thinking
Some systems split the work in two. A slower, larger model interprets the scene and the instruction and decides what to do next, while a faster, smaller controller turns that plan into smooth, precise movement many times a second. Others use a single model end to end. Each design balances responsiveness, generality and computing demands differently. For buyers, the practical question is less about architecture and more about how the system behaves on your tasks, in your lighting, with your objects.
Strengths and limits
VLAs make robots easier to direct and can generalise to new objects and phrasing better than hand-written rules. They can also be unpredictable: a model may perform well in one setting and fail in another for reasons that are hard to diagnose. Safety therefore still depends on conventional layers such as speed limits, force limits, protective stops and well-designed workcells. Ask any maker how learned behaviour is bounded, monitored and updated over time.
Sources and further reading
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control — Google DeepMind (arXiv).
A vision-language model fine-tuned on robot trajectories to output actions as text tokens, carrying web knowledge into control. - OpenVLA: An Open-Source Vision-Language-Action Model — Stanford University, UC Berkeley and partners (arXiv).
An open 7B-parameter VLA trained on 970k robot demonstrations, built to be fine-tuned on consumer hardware. - Embodied large language models enable robots to complete complex tasks in unpredictable environments — Nature Machine Intelligence.
Peer-reviewed study pairing a large language model with force and vision feedback so a robot adapts mid-task while making coffee.
Common questions
Can I talk to a humanoid running a VLA model?
VLA models are built to take instructions in natural language, typed or transcribed from speech. Holding a conversation is a different ability, so check what the robot actually supports.
Do VLA models learn on the job?
Usually not on their own. Most are trained offline and updated by the maker. Some systems support fine-tuning with new demonstrations, which is worth asking about.
People also search for vision language action model, VLA model, what is a VLA model, VLA robotics, vision language model for robots and AI robot that understands language.