GPAI.WIKI
wiki·

World model

Also called
learned simulator · forward model · dynamics model
Notation
\(p(s_{t+1} \given s_t, a_t)\)

A world model answers: if the state is \(s_t\) and I take action \(a_t\), what happens next? Having one converts action selection from trial in the environment into search inside the model, which is cheaper, safer, and — crucially — reusable across tasks that share dynamics.

Prediction space matters

Models that predict raw observations spend most of their capacity on detail that is irrelevant to control. Models that predict in a learned representation space avoid this but must be prevented from collapsing to a constant. This is the core tension addressed by R-001 and demonstrated in P-0003.

The exploitation problem

Any planner optimising against a learned model will find and exploit its errors. Compounding error over a rollout of length \(H\) means the effective planning horizon is bounded by model accuracy, not by compute. Mitigations — uncertainty penalties, ensembles, short rollouts with learned value bootstrapping — all trade horizon for reliability rather than removing the bound.

see also
referenced by1
reactions
no reactions yet
react on GitHub