Measured evidence instead of momentum
Keep ThinkingAI · Markets · Society
All minds
Stanford professor, co-director of Stanford HAI, CEO of World Labs

Fei-Fei Li: The data pioneer who thinks language is not enough

She built the dataset that ignited the deep learning era, then spent a decade arguing that AI should be designed around human values. Now she is betting a company that the next frontier is not better text, but machines that understand space.

The core position

Intelligence is grounded in perception and action in the three-dimensional world, not in text alone. Machines that can see, reason about, and construct physical spaces will unlock capabilities that language models structurally cannot reach, and the whole enterprise must be built to serve human needs rather than abstract benchmarks.

The lab's read

Li is proof that the infrastructure behind a breakthrough can matter more than the algorithm. ImageNet showed that the field had been optimizing the wrong constraint, and her spatial intelligence thesis is the same argument applied again: the bottleneck is not scale of text but the absence of grounded perception. For readers tracking where capital and research effort flow next, her career is a reliable leading indicator.

From Chengdu to Caltech

Fei-Fei Li was born in Beijing in 1976 and grew up in Chengdu. At sixteen she joined her father in Parsippany, New Jersey, where the family ran a dry-cleaning shop and she worked weekends behind the counter while attending high school. The immigrant biography is not decoration in her case. She has described those years as the source of a conviction that technology should be judged by what it does for ordinary people.

She studied physics at Princeton, graduating in 1999, then earned her PhD in electrical engineering at Caltech in 2005 under Pietro Perona and Christof Koch. Her dissertation combined computational models with human psychophysics, asking how people categorize visual scenes almost instantly. That dual training, half machine learning and half human perception, set the pattern for everything after: she has always treated human cognition as the reference standard, not an afterthought.

The dataset that changed the field

In 2007, as a young assistant professor at Princeton, Li concluded that computer vision's real bottleneck was not algorithms but data. Inspired by an estimate that humans recognize tens of thousands of object categories, she set out to build a visual database at a scale colleagues considered impractical. The result was ImageNet, published as a CVPR paper in 2009: roughly 14 million images across more than 20,000 categories, labeled largely through Amazon Mechanical Turk.

The accompanying competition, the ImageNet Large Scale Visual Recognition Challenge, ran from 2010 to 2017. In 2012 a convolutional neural network called AlexNet, built by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, won it by a margin so large that the field reorganized itself around deep learning almost overnight. Li had not invented the neural network. She had built the condition under which it became undeniable.

The lesson she draws is methodological. Progress came from measuring the field against a hard, public, data-rich benchmark, and from betting on scale before scale was fashionable. It is the same bet she would later make about physical space.

The human-centered turn

From 2013 to 2018 Li directed the Stanford Artificial Intelligence Laboratory, then spent 2017 and 2018 at Google as a vice president and chief scientist of AI and machine learning for Google Cloud, where her team worked on making AI accessible to ordinary businesses, including the AutoML products. The stint ended amid the Project Maven controversy, when leaked internal emails showed her enthusiasm for the Pentagon cloud contract alongside warnings about how military AI would be perceived. She told the New York Times that working on anything that weaponizes AI is deeply against her principles, and returned to Stanford in the fall of 2018.

Back at Stanford she converted the episode into institution building. In 2017 she had co-founded AI4ALL, a nonprofit bringing underrepresented students into AI. In 2019 she became founding co-director of the Stanford Institute for Human-Centered AI, alongside former provost John Etchemendy, an institute built on the premise that AI research should be evaluated by its effect on the human condition. HAI's annual AI Index report has since become the field's standard statistical reference.

Her policy positions follow the same empirical instinct. At the Paris AI Action Summit in February 2025 she argued that AI governance should be based on science rather than science fiction, meaning measured evaluation of actual capabilities rather than speculation about either utopia or catastrophe. She has also called for a moonshot level of public investment in academic AI research, warning that the compute and talent gap between industry and universities is distorting what gets studied.

I believe in human-centered AI to benefit people in positive and benevolent ways. It is deeply against my principles to work on any project that I think is to weaponize AI.

Interview with the New York Times, 2018

Words are not worlds

Li's current intellectual position is a deliberate counterweight to the language model era. Her argument, laid out in a 2024 TED talk and developed in a November 2025 essay titled From Words to Worlds, is that language is a recent evolutionary overlay on a far older perceptual intelligence. Evolution spent hundreds of millions of years building systems that see, navigate, and manipulate three-dimensional space long before any creature produced a sentence. An AI trained only on text, she argues, is a wordsmith in the dark: fluent about the world without being situated in it.

The constructive proposal is spatial intelligence: world models that can perceive, generate, and reason about persistent 3D environments, connecting perception to action. This puts her in loose alignment with researchers like Yann LeCun who argue that language models are a detour from grounded understanding, though Li's framing is less a critique of LLMs than a claim about what must come after them. Robotics, scientific discovery, and medicine, in her telling, all wait on machines that understand space.

Our dreams of truly intelligent machines will not be complete without spatial intelligence.

From Words to Worlds, November 2025

World Labs, and the honest ledger

In 2024 Li took partial leave from Stanford and co-founded World Labs to build large world models, emerging from stealth that September with 230 million dollars in funding at a valuation above one billion. The company released a real-time generative model, RTFM, in October 2025, and in November 2025 launched Marble, its first commercial product: a model that turns text, images, or video into persistent, editable 3D environments exportable as meshes or Gaussian splats. In February 2026 World Labs raised a further one billion dollars, including 200 million from Autodesk, to push world models into professional 3D workflows.

The honest assessment has two sides. She has been demonstrably right before: the ImageNet bet on data looked eccentric in 2007 and defined a decade, and her insistence on human-centered framing, once dismissed as soft, now describes the mainstream of policy and procurement. The contested side is that the spatial thesis is unproven at scale, world models are a crowded race with Google and a field of startups, and running a venture-backed company sits in visible tension with her warnings about the concentration of AI resources in private hands. Whether history repeats will depend on whether Marble and its successors do for spatial reasoning what ImageNet did for recognition: make a capability undeniable.

What to take seriously

1

Infrastructure beats insight

Li's greatest contribution was not an algorithm but a dataset and a benchmark. The unglamorous work of building shared measurement can move a field further than any single model.

2

Ask what the training data cannot contain

Her spatial intelligence argument is a general method: identify what the dominant paradigm's data excludes. Text models exclude the physical world, which is why she expects the next breakthrough elsewhere.

3

Human-centered is a design constraint, not a slogan

Through HAI, AI4ALL, and her policy testimony, Li treats human welfare as an engineering requirement to be specified and measured, not a value to be invoked at launch events.

4

Governance should follow measurement

Her science rather than science fiction position cuts against both doom and boosterism. Regulation grounded in evaluated capability ages better than regulation grounded in narrative.

5

Watch where the conviction builder goes next

Li has twice committed years to a thesis before the field agreed: data scale in 2007, spatial intelligence now. Her career is a useful map of which unfashionable bets are worth tracking.

Sources & further reading

Keep Thinking

Independent analysis: no reselling, no vendor commissions. On the side, we help a small number of companies implement what we write about.

Work with the lab

More minds

Newsletter

New papers, when they are ready. One email per analysis, nothing else.