HumanoidGPT · Part of The Humanoid Group

Robot brains are moving out of the lab.

Foundation models now let robots attempt unseen homes and ten-minute chores, while safety tests expose what still fails. Follow the progress, then match hardware with The Humanoid Group.

What moved in robot AI this year.

  • 89.4% success on RLBench's 18 simulated tasks, up from about 50% in 2022 (Stanford HAI, AI Index 2026, 2025)
  • 12.4% full-task success by the top team on BEHAVIOR-1K's 1,000 home tasks (eWeek, citing the Stanford AI Index, 2026)
  • 3–10× slower than humans: typical speed of robots in Epoch AI's 2026 review (Epoch AI, 2026)
  • 56% success of Figure's Helix 2.5 in 30 unseen homes (9% without Index) (TechRepublic, 2026)

Figure's Helix 2.5 is tested in 30 unseen homes and succeeds 56% of the time

· Milestone

Figure tested its Helix 2.5 robot policy in 30 Bay Area homes it had never seen, on tidying toys, folding towels and making beds. Across 420 attempts with no partial credit, it succeeded 56% of the time, versus 9% without Figure's Index pretraining. Tidying scored lowest at 40%.

Why it matters: A 44% miss rate shows why home-style pilots need measured success rates on your own tasks, not demo clips. Ask any vendor for trial counts, failure definitions and results from sites the robot has not seen.

Source: humanoid.guide

Safe-Stop paper teaches a humanoid to judge when stopping will not make it fall

· Research

Researchers posted a method that estimates whether a humanoid can still stop safely from its current motion, then falls back to another policy if it cannot. On a Unitree G1 in simulation it stopped successfully in 96.4% of 179,650 test starts, with 150 trials on hardware.

Why it matters: Emergency stops on legged robots are not automatic. Ask humanoid vendors how their stop behavior was tested, from which motions, and how often real-world trials disagreed with simulation.

Source: arXiv

Skild AI says its S1 model learns a 10-minute task from one human video

· Launch

Skild AI introduced S1, a robot model prompted with a single human video instead of retraining. It reports 66% per-step success on unseen tasks lasting up to 10 minutes, against 9% for language-prompted policies. The results come from Skild's own benchmarks, and S1 has no public weights or API yet.

Why it matters: A model that learns from one video could shorten setup for site-specific jobs, but today's evidence is vendor-run and the model is not yet public. Ask for independent trials before planning around it.

Source: Runtime Wire

Sources and further reading

Visit The Humanoid Group

People also search for AI humanoid robot, physical AI, embodied AI, AI for robots, AI that controls robots and how humanoid robots learn.