HumanoidGPT · Part of The Humanoid Group

LLMs as Task Planners

Ask a humanoid to tidy a kitchen and something has to turn that request into a sequence of actions the robot can perform. Large language models increasingly do that job. This guide explains how they plan and where they go wrong.

By Arjun Rao · Updated

From a request to a list of skills

In many LLM-based planners, the language model does not move the robot directly. It sits above a library of skills the robot already has, such as pick up, open, place or go to, and decides which to call and in what order. Given a request like 'I spilled my drink, can you help?', a planner might choose to find a sponge, pick it up, bring it over and wipe the table. The skills themselves run as separate learned or programmed controllers. This division of labour lets the language model contribute everyday knowledge about how chores are usually done, while the robot's own software handles grasping, balance and motion.

Grounding plans in what the robot can do

A language model knows a lot about kitchens in general but nothing about the robot standing in this one. Left unchecked, it may propose steps the robot cannot perform or that make no sense in the current scene. The SayCan project at Google tackled this by scoring each step the model suggested with value functions that estimate how likely the corresponding skill is to succeed from the robot's present state, so the chosen step is both useful and feasible. In the authors' framing, the robot serves as the language model's hands and eyes. Other systems ground plans with scene descriptions from a vision model, or by feeding back the result of each step before choosing the next.

Writing plans as code

Another approach asks the model to write a short program rather than a list of words. In Code as Policies, Google researchers prompted code-writing language models with examples and a set of perception and control functions, and the models composed new robot programs from fresh instructions. Code brings real advantages to planning: it can loop, test conditions, do geometry with maths libraries and turn vague words such as 'a bit faster' into numbers. It is also easier for an engineer to read before it runs. The weakness is shared with any generated code, since it can call functions wrongly or assume something false about the scene, so it needs limits and checks.

Knowing when to ask for help

Language models sound confident even when they are guessing, which is risky when their output moves a machine. If someone says 'put the bowl in the microwave' and two bowls sit on the counter, the planner should ask rather than pick one. KnowNo, from Princeton University and Google DeepMind, uses a statistical method called conformal prediction to measure a planner's uncertainty and request human help when the options are ambiguous, while giving formal guarantees on how often tasks are completed. For buyers, the practical points are how a robot's planner handles unclear instructions, whether it confirms before irreversible steps and whether its plans are limited to actions that have been tested.

Sources and further reading

Common questions

Can a general chatbot plan tasks for a humanoid robot?

It can suggest steps, but on its own it knows nothing about the robot's skills, surroundings or limits. Useful systems connect the language model to the robot's skill library, perception and safety checks, and the combination is tested on real tasks before anyone relies on it.

Does an LLM planner need an internet connection?

That depends on where the model runs. Larger models are often hosted remotely, while smaller ones can run on the robot's own computer. Ask the maker what happens to planning, and to the robot's behaviour, if the connection drops halfway through a task.

Visit The Humanoid Group

People also search for LLM task planning, large language model robot planning, language model robot actions, SayCan, Code as Policies and robot skill library.