Skild AI Publishes Manipulation Results for S1 - Show It One Video and It Runs Tasks It Never Trained On, Company Says

Skild AI published manipulation results for its robotic foundation model S1. Tasks are specified by a video demonstration rather than language, and the company says one set of weights executes unseen tasks with no fine-tuning. It calls the demonstration on ten-minute tasks a first.

Robotics company Skild AI published manipulation results for its foundation model S1 on its blog1. The company describes S1 as designed from the outset to learn in context, and says a video is all it needs in order to attempt a task, whether that task is a quick one or a long one and whether or not it appeared in training1.

The distinguishing choice is how a task gets communicated: a video demonstration takes the place of a written or spoken instruction1. Because pre-training spans many different tasks, the company argues, the model has to work out what the demonstrator is trying to do, which it says leaves one set of weights able to handle tasks absent from training, with no additional tuning1. It states plainly that there is no fine-tuning and no post-training involved1. By its own account of the earlier approach, every new task meant collecting hours of data and tuning a specialist policy for it, and for tasks of moderate complexity the volume of post-training data had to run from tens into hundreds of hours recorded under the conditions where the robot would actually work1.

Skild AI presents this as a first: by its account, no robotics foundation model before S1 has demonstrated in-context learning on tasks that run as long as ten minutes and were absent from pre-training1. These are the company’s own claims, not third-party verification. It says a future post will cover how S1 is trained1.

Sources

  1. Introducing S1: In-Context Learning for Robotics - Skild AI official blog

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →