
Why self-taught robot brains still need to get into the mud.
Silicon Valley just wrote a $320 million check on the theory that robots can learn to work by watching video games.
General Intuition, spun out of the clip-sharing platform Medal, raised its Series A at a $2.3 billion valuation to train foundation models on billions of gameplay clips. Physical Intelligence, Skild, and a wave of world-model labs are running versions of the same play: feed a large model enough video and simulation, and physical intuition emerges. No fleet. No dirt. No decade of deployments. Just GPUs.
They are half right, but I would argue the “money” is in the other half.
The half they have right
Robot intelligence is commoditizing, and faster than most people expected. NVIDIA now ships its GR00T robot foundation models under Apache 2.0, free to use commercially, backed by Cosmos world models and the Isaac simulation stack. When the most capable robotics AI on the planet is a download, intelligence stops being a moat and starts being plumbing.
The foundation model labs sit one layer up from NVIDIA, and they have the same problem. If you are training on video games and public video, there is no reason others cannot do exactly the same thing. The next lab can gather similar footage and rent a bigger cluster. NVIDIA can simply extend GR00T upward into the same territory. A model built on data anyone can get is an advantage anyone can copy.
The half that requires mud
Now ask what these models still cannot see. A wet orchard row at dawn in fog. Dust kicked up over a sensor at harvest. Glare off irrigation line at 4pm. A worker stepping out from between vine rows exactly when a machine is moving. Mud that changes what traction means from one hour to the next. In a video game, none of this exists and ultimately will have to be learned in the field.
Their best argument is that game clips carry human intent, not just pixels. Fair. But intent in Fortnite is not intent in a wet almond orchard with a person stepping into the row.
We have run this experiment before. Autonomous driving spent fifteen years and well over $100 billion learning that simulation is necessary but nowhere near sufficient. Waymo drove tens of billions of simulated miles and still needed tens of millions of real ones before removing the driver. Tesla built its entire autonomy strategy on the premise that fleet-scale real-world data is the asset, not the model. John Deere trains on data from over 300 million engaged acres and calls the resulting flywheel its core moat. Carbon Robotics built its weed-recognition advantage one lasered plant at a time, in actual fields, in actual weather.
The roboticist Ken Goldberg estimates robotics faces a 100,000-year data gap relative to what language models had. You do not close that gap with Fortnite clips. You close it with machines doing paid work outdoors, near people, for years.
Watch what they do, not what they say
Here is the tell, and it comes straight from the self-taught camp itself. General Intuition launched Nerve, a marketplace that pays people for teleoperation and real-world action data. Its CEO says the company will prioritize customers based on who can supply real-world data to improve its model. Its demo robot, trained on 100 hours of gameplay, still needed eight minutes of real-world data collected on an actual street before it could handle a new environment, and whether that transfer holds in complex industrial settings has not been proven on any public benchmark.[1]
Where the moat actually sits
The chart at the top of this piece is the whole argument. The NVIDIA layer is a commodity today. The foundation model layer directly above it is commoditizing in real time. The two layers you cannot download are sitting on top: real-world data earned in the field, and the workflow layer where a customer pays for completed work rather than intelligence.
The physical AI gold rush is real. But the picks and shovels are not GPUs, and they are not gameplay clips. They are fleets out in the field getting beaten up and dirty.