Mental World Modeling hits 87.9 F1 on human-action benchmarks
Reporting from The Decoder highlights new research arguing that widely discussed generative world models emphasize physical dynamics while neglecting internal mental states. The Mental World Modeling approach reportedly encodes beliefs and intentions as explicit variables rather than treating action as physics alone. Benchmarking under a shared human-action evaluation protocol yields 87.9 F1 for the complete pipeline, compared with 98.5 for human participants measured the same way. The article positions that gap as practical evidence that belief-blind models mispredict human behavior. Available excerpts cut off before fully describing weaker language-model components, leaving part of the comparative picture unspecified in this packet. Operators should treat headline scores as tied to that benchmark protocol rather than as universal proof of generalization.
Mental World Modeling hits 87.9 F1 on human-action benchmarks
Current world models like Sora or Genie only simulate physics and ignore what people think, want, or feel. The full MWM pipeline reaches 87.9. Humans hit 98.5 under the same protocol.
Key takeaway
Mental World Modeling reportedly reaches 87.9 F1 on human-action benchmarks by modeling beliefs and intentions, while humans score 98.5 under the same protocol.
What happened
The Decoder reports that new research finds world models that ignore human beliefs predict incorrect actions, noting that current systems such as Sora and Genie simulate physics but not what people think, want, or feel.
The article describes a Mental World Modeling framework that adds mental variables like beliefs and intentions, and states the full MWM pipeline reaches 87.9 F1 on human-action benchmarks while humans score 98.5 under the same protocol.
Evidence
Current world models like Sora or Genie only simulate physics and ignore what people think, want, or feel.
The Decoder · attributed
Current world models like Sora or Genie only simulate physics and ignore what people think, want, or feel.
The Mental World Modeling framework adds mental variables like beliefs and intentions.
The Decoder · attributed
The new "Mental World Modeling" framework adds mental variables like beliefs and intentions.
The full MWM pipeline reaches 87.9 F1 on human-action benchmarks.
The Decoder · attributed
Mental World Modeling hits 87.9 F1 on human-action benchmarks
Humans score 98.5 under the same human-action benchmark protocol.
The Decoder · attributed
Humans hit 98.5 under the same protocol.
Why it matters
Physics-only world models may misforecast intent-driven human behavior in agents, games, and simulations that depend on predicting what people will do next.
Limits and uncertainties
The Decoder excerpt truncates before completing the claim about weaker language-model variants, so part of the comparative analysis is missing from this packet.
The packet does not include peer-review status, dataset names, or independent replication beyond The Decoder's reporting.
Practical implications
Builders of agent or simulation stacks should evaluate whether explicit belief and intention modeling is required for human-action prediction beyond physical dynamics.
Product and research claims should reference the specific human-action benchmark protocol, given the reported 87.9 F1 versus 98.5 human score on the same test.
What to watch
Full publication or follow-on reporting that completes the truncated weaker language-model comparison in The Decoder excerpt.
Independent replication or extension of the 87.9 F1 human-action benchmark results on tasks beyond those cited in this reporting.