Back to David Shapiro

Nobody gets this right

David ShapiroJune 7, 202615m
In a Nutshell

The speaker argues that "world model" claims against language models rest on false distinctions, since multimodal omni models already integrate text, vision, audio, and sensor data, and accurate next-token or next-frame prediction is evidence of abstract world understanding. Animals with strong embodied sensorimotor loops lack general intelligence useful to humans, showing that physical data alone does not produce AGI. Existing video-language-action models, code-writing capabilities, and layered safety systems already bridge the gap between generative models and real-world deployment, making rebranding around JEPA-style world models a category error rather than a fundamental shift.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

The question of whether language models are world models has been asked repeatedly, partly due to recent news. Papers exist claiming language models are world models, with the debate centering on a matter of degree. Language models were already approximating mental maps of spatial relationships even before multimodality, though the approximation was not highly accurate.

The phrase "the world is not made of words" is commonly used in this discussion. However, physicists and engineers often state that the world is made of math. Language models have become increasingly capable at mathematics. If AI develops intuitive understanding of geometry and physical space, this constitutes a world model.

Most discussions focus on the felt sense of intuition when moving through three-dimensional space. The most useful functions of AGI are expected to be in abstract mathematics, where advanced physics and advanced biophysics occur. Physical world models with sensor data, feedback loops, and proprioception are extremely useful for robotics specifically.

Birds possess excellent proprioception yet are not considered generally intelligent in ways useful to humans. Dogs, cats, monkeys, and baboons demonstrate high spatial intelligence with tight sensor feedback loops and good physical control, but lack general intelligence that would be useful to humans.

Leor Alexander posted on Twitter using the example "the real world isn't made of words" in reaction to Yann LeCun. The first claim states that large language models predict the next word, are trained on text, and therefore understand language, but the real world isn't made of words. This claim is considered false because models are no longer trained solely on words. They are trained on multimodal data including audio, video, text, and images. The term "language model" has been a misnomer for more than a year. ChatGPT-4o introduced the "O" for omni, making "omni model" the more accurate term, though LLM is expected to persist.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.