Published: 26 August 2026
Skild AI announced S1 on 25 August 2026, a robot foundation model that learns a new task from a single human video, reaching 66% success on tasks it had never seen against 9% for an equivalent language-prompted policy. Tasks run up to ten minutes and dozens of steps. There are no weights, no API and no paper.
What can Skild AI S1 do?
S1 takes one egocentric human video as its prompt and carries out the task with no fine-tuning and no post-training, so the weights never change. Skild shows four tasks that were absent from pretraining: potting a plant, cooking pancakes, making pour-over coffee and assembling a kit. The longest runs to ten minutes across dozens of manipulation steps.
This is a deliberate break from the vision-language-action approach, where the task is specified in words. Skild argues a video carries spatial and temporal detail that a sentence drops. On the plant-potting task, 11 minutes passed between showing the demonstration and the robot working on its own, and most of that was scene setup.
Skild AI S1 benchmarks: 66% against 9%
Skild ran a controlled comparison holding data, architecture and compute constant and varying only whether the policy was prompted by video or by language. At 100,000 hours of training data the in-context model scored 66% on unseen tasks against 9% for the language-conditioned one. One video demonstration was worth roughly 380 post-training episodes.
The more revealing number sits at the other end of the curve. At 1,000 hours the language policy wins, 53% against 43%. The advantage only appears at scale, which makes this a claim about data volume as much as about method. A post-trained policy still reaches 86%, but needs 2,000 teleoperated demonstrations to get there.
Can you actually run S1?
No. There are no downloadable weights, no public API, no licence and no pricing. Skild says it is deploying S1 with a limited set of industrial partners now and will widen access over the coming months, with a sign-up form for early access. Its named robot OEM partners are ABB Robotics, Universal Robots and MiR.
There is also no technical report. The 66% comes from an internal benchmark Skild has not released, with no trial counts, no error bars and no external baseline. Parameter count, total pretraining hours and the robot platforms used in evaluation are all unpublished. Everything here rests on one blog post and a set of demonstration videos.
How does S1 compare to Generalist GEN-1.5 and pi-0.5?
Skild names no competitor. Its only baseline is its own language-conditioned ablation, so any head-to-head reading is inference rather than measurement. The nearest comparison shipped a day earlier: Generalist AI announced GEN-1.5 on 24 August 2026, which learns new tasks from a single demonstration of 3 to 12 seconds.
Set against Physical Intelligence’s pi-0.5 and the wider VLA family, S1’s differentiator is task length rather than raw dexterity. Ten minutes and dozens of steps sits well beyond the horizon most published VLA work reports. Whether it holds outside four curated demonstrations is exactly what the missing paper would settle.
What this means
The SI model is worth watching rather than testing, because you cannot test it. For robotics teams the useful result is the crossover point: below roughly 1,000 hours of data, language prompting still wins, so anyone fine-tuning a modest VLA on their own fleet is not doing the wrong thing. The video-prompt approach only pays off at Skild’s data scale.
Treat the 66% as a directional signal, not a benchmark. One blog post, an internal evaluation and four hand-picked tasks is the weakest evidence tier in this field, and Skild raised $1.4 billion at a valuation above $14 billion in January, so it has every incentive to publish a strong number. Ask for the paper before you plan around it.
For more information visit the official S1 announcement on the Skild AI blog.