Someone told me the excitement about egocentric data was over. It doesn’t work, he said, so people are giving up on it.
If you haven’t run into the idea: the pitch was that you could teach robots by watching people. Put a camera on someone’s head, record them cooking and folding laundry and fixing a bike, and train a robot on what comes back. The internet already holds millions of hours of people using their hands. The dream was to point a model at all of it and have it learn manipulation the way language models learned to write, by swallowing an ocean of examples.
The claim now is that the dream died. The footage didn’t teach the robots much, so the field is moving on.
That claim is both true and false, and which one you get depends entirely on what you thought the words meant. “Egocentric data” points at four different things. They are, right now, going in four different directions. So any one sentence about whether it works is wrong before it reaches the period.
Let me pull the four apart, because once you see them separately the argument mostly dissolves.
The first is raw first-person video and nothing else. Someone’s GoPro. Pixels of a person doing a task, with no record of what their joints did or how hard they gripped. Just the picture.
The second is a person wearing a rig that records the picture and the actions. A gripper you hold in your hand that logs its own pose. A glove that captures how your fingers move and what they touch. It looks like the first thing. It behaves like something completely different, and that difference is the whole essay.
The third is using human video only to warm a model up. You don’t ask it to learn the task from the video. You let it watch a lot of people first, so it has some sense of how the world moves, and then you teach it the actual skill on real robot data.
The fourth isn’t even first-person. It’s motion capture of a whole body, filmed from outside, retargeted onto a robot so a humanoid can learn to walk and balance. People lump it in because “learn from humans, build a humanoid” sounds like one idea. It isn’t.
Now ask the dying question of each.
The criticism lands squarely on the first one, and only the first one. And here’s the part nobody says out loud: it was never really in doubt. Raw video of a person is missing everything a robot needs in order to act. There’s no record of the commands that moved the hand, so the robot can’t copy them. There’s no force, so it can’t tell a firm grip from a gentle one. The body is the wrong shape. The camera is swinging around. And every clip is a success, because nobody uploads the twenty times they dropped the jar. You cannot learn to recover from a mistake by watching a stranger who never made one. All of this was true the day the first big egocentric dataset shipped. If “the hype is dying” means people stopped believing you could train a policy on raw video alone, then fine. But that’s a small, unsurprising thing to be right about, wearing the coat of a big one.
Because look what happens to the other three.
The distinction that actually matters isn’t human versus robot. It’s whether the data records what the person did or only what they saw. A video of your hands is a postcard. A glove that logs your grip is a demonstration. Sort by that line and the confusion clears in one move.
The demonstrations - the gloves and the handheld rigs - are the fastest-growing corner of the whole field. They give you usable actions at a fraction of the cost of anything else, and you can collect them anywhere, in parallel, by the thousands of hours. The companies built on this are not retreating. They’re raising money and hiring.1 The reason is simple. A camera on a cheap glove that also records grip is not “human video.” It’s a robot demonstration with the robot taken out. The one thing raw video was missing, the action, is right there, because your own wrist supplies it.
The warm-up use is now just standard practice. The clearest result came out of Physical Intelligence late last year: once you’ve trained a policy on enough varied robot data, letting it also watch human video roughly doubles how well it handles situations that showed up only in the video. 2 Read that slowly. The human footage isn’t the foundation. It’s what you pour over a foundation you already have, to reach corners you couldn’t afford to visit with a real robot. Below some threshold of robot data it does nothing. Above it, the model stops being able to tell the human and robot examples apart, and the free coverage kicks in.
So the score, honestly: one narrow use of one kind of data lost a belief it never should have held. The other three are climbing. Calling that “egocentric data is dying” is like standing in a city where three neighborhoods are booming and announcing the place is finished because the fourth one emptied out.
The fix is to throw away the phrase. It’s doing too much work, and it hides the only question worth asking. Does the data carry actions, or not? If it does, it’s a demonstration, and demonstrations are getting cheaper and more plentiful every month. If it doesn’t, it’s coverage - cheap, useful, worth having, and never the thing you learn the skill from.
I’m being vague about the numbers on purpose. The raises are real and large, but the eye-catching performance claims around these companies - “beats robot data,” “zero robot data” - are mostly their own, and nobody without a stake has reproduced them. Believe the direction, not the decimal places
Also their result, also not yet independently replicated. I trust the shape of it because the same idea keeps surfacing in separate places, but it’s one lab’s finding


