AI nelle imprese, con i numeri veriLinkedIn ↗
RassegnaAI

Lukas FerrazziAnalisi24 August 2026 · 5 min read

The robot that beat Bolt and then ran into the crash mats

In Beijing the 100 metres record fell on 22 August, in a race neither of the two protagonists finished on its feet. For anyone evaluating an artificial intelligence system, the instructive part is the second one.

Originally published in Italian · leggi in italiano

Nine point three nine

On 22 August, at the National Speed Skating Oval in Beijing, the venue built for the 2022 Winter Games, the second edition of the World Humanoid Robot Games opened. On the 100 metres track a robot from Beijing's humanoid robotics innovation centre won in 9.39 seconds, under the 9.58 with which Usain Bolt set the world record in Berlin in 2009. The race was covered by ANSA and by the Associated Press, which added the numbers of the event, 2,056 robots, 666 teams from sixteen countries, 51 disciplines and 1,301 races over five days.

Then there is the detail that dropped out of almost every headline. The winning robot lost its balance just past the finish line and ended up against the foam barriers. The runner-up, Lightning, built by smartphone maker Honor, had run 9.32 in a heat, closed the final in 9.47 and collapsed to the ground, carried off the field of play. The high jump record fell the same day, 2.88 metres against Javier Sotomayor's 2.45 from 1993, with a technique that looks nothing like an athlete's.

From twenty-one seconds to nine

Twelve months earlier, at the first edition of the Games, the 100 metres had been won in 21.50 seconds and the robot high jump record stood at 95 centimetres. Registered teams grew by 138 per cent in a year. Il Sole 24 Ore traced the history of Lightning, which in April had won the Beijing half marathon in 50 minutes and 26 seconds, and whose engineers then fitted legs ten centimetres longer, going from 95 centimetres to 1.05 metres, in order to run the short distance.

Anyone dismissing all this as folklore is wrong. Halving a time in twelve months is not done by a marketing department, it is done by balance control, actuators and software managing the stride, and those are the same components that will be needed in a warehouse or on an assembly line. The interesting question is not whether the progress is real, because it is, but what conditions were required to produce that number.

The conditions of the test

The lane is straight, the surface is known, there is no contact with opponents, the attempt that counts comes after many attempts and no rulebook asks the machine to stop safely after the finish line. In a warehouse none of this is guaranteed. The floor is dirty, the route changes, someone walks across it, and the performance that matters is not the peak but repeatability, with decent behaviour when something goes wrong.

The same gap is visible right now in the software market. On 23 August the Financial Times reported that, based on spending data from 70,000 American companies collected by payments company Ramp, Anthropic's flagship model settled at around 11 per cent of what those companies spend on the firm's tools, overtaken by a model released at the end of July that costs less. The reasons cited by analysts and investors are two, a price roughly double that of its direct competitor and the fact that earlier models are enough for most office work. Anthropic is heading towards a listing with fast growing revenue, so this is not a story of decline but of purchasing criteria. Companies are buying what gets to the end of the task at a predictable cost, not the record.

When I test a voice agent, the figure I write down is not the most impressive answer the system gave in trials, it is the share of conversations that reach the end without the person having to repeat the same information. In early versions that share sits well below the level of the demonstration, and the useful work of the following weeks consists entirely of pushing it up. A pilot exists precisely for this, to produce a number that did not exist before.

Three questions for a supplier come out of this, to be asked before signing. How many runs were needed to obtain the result shown, and what is the median value, not the best one. Out of a hundred runs, how many finish within the agreed tolerance, with the log of failed attempts attached and not summarised. What the system does when it is wrong, whether it stops and hands over to a person or carries on like the two robots in Beijing, which crossed the line and went into the barriers. If the supplier cannot answer the second question, that figure belongs in the contract as the object of the pilot, with a threshold and a date.

When the industry bought itself a stopwatch

In the late nineteen eighties workstation makers sold power in MIPS, millions of instructions per second, and each of them measured with whichever program suited it best. The number was so hard to compare that in engineers' slang MIPS became the acronym for Meaningless Indicator of Processor Speed. In 1988 Apollo, Hewlett-Packard, MIPS Computer Systems and Sun Microsystems founded SPEC, a consortium that wrote common tests and rules on how to publish their results. It did not solve everything, manufacturers soon learned to optimise for the test, but from that moment buyers had a shared yardstick instead of the vendor's word.

Humanoid robotics, with all its medals and cameras, is building itself that yardstick. In Beijing there were a hundred metres that were the same for everyone, a stopwatch and a judge who can see a robot leave its lane. The software companies are buying to put artificial intelligence to work, for now, does not even have the lane.