Skip to main content

AI's Blind Spot: Why New Science Still Eludes It

AI's Blind Spot: Why New Science Still Eludes It

At the 2026 Inclusion·Bund Conference, Princeton professor Wang Mengdi posed a question that's been nagging at researchers: how close is AI to actually discovering new science on its own?

Her answer might surprise you. Today's large models have swallowed vast amounts of human knowledge, yet they still stumble when it comes to genuine scientific breakthroughs. Why? Because they're trained to find the most likely answer, while real discoveries often lurk in the long tail of probability distributions.

The 'Simulated Life' Experiment

Wang's team, working with Princeton sociologist Xie Yu, ran a clever test. They asked models like ChatGPT and Claude to simulate a British person born in 1930 and generate their life story. Across marriage age, occupation, and gender, the models consistently overestimated common scenarios—and badly underestimated minority situations and long-tail distributions.

"The model learns the peak of a probability distribution," Wang explained, "but it may overlook or even ignore knowledge in the long tail."

This explains a puzzle: why AI rockets ahead in math and coding but hasn't produced original discoveries of the same scale in physics, chemistry, or biology.

The difference? Validators. Coding has compilers and unit tests. Math has formal proof systems like Lean. The model can try, fail, get feedback, and improve. In real science, verification is far messier.

"Ideas are cheap," Wang said bluntly. "You need to show me code or an experiment."

Scientific research involves complex equipment, human judgment, and cross-team collaboration. Many experiments can't be stably replicated. And more information doesn't necessarily mean more effective new information—it can actually lower the signal-to-noise ratio.

Building the Infrastructure for Discovery

So what's the key to AI becoming an autonomous discoverer? Wang argues it's not about making the model know more. It's about making the real world increasingly verifiable.

Her team is building exactly that. In quantum materials research, they've created an automated experimental platform that turned a graphene experiment—once taking researchers months—into something anyone can call via API. They're also developing LabOS, an intelligent operating system for research labs, where multimodal AI perceives the experimental environment through smart glasses, assists researchers, and connects software, AI models, robots, and real equipment.

Wang's conclusion is striking: what determines whether AI becomes a "discoverer" isn't the model's size, but whether there's infrastructure linking hypotheses, validation, traceability, and replication. Only when physical-world experiments can be systematically recorded, continuously validated, and repeatedly replicated can AI transform the "likelihood" learned in training into exploration of "possibility."

In other words, the future of AI-driven science might not be written in code—it might be built in the lab.

Key Points

  • Large models excel at finding likely answers but miss long-tail discoveries where true novelty lies.
  • Math and coding advance quickly because they have clear validators; real science doesn't.
  • More data can hurt by lowering the signal-to-noise ratio in research.
  • Infrastructure matters more than model size—connecting hypotheses, validation, and replication is key.
  • Wang's team is building automated labs and LabOS to make physical experiments verifiable and reproducible.