Code Plans, Diffusion Renders: Open-Ended Generative World Modeling
En palabras de los autores
We introduce CoDeR, a new paradigm for world modeling. Unlike existing video world models that implicitly represent world dynamics through visual observations, our system explicitly constructs an executable world with code and employs video generation models for visual realization. Specifically, we coordinate five complementary roles to translate high-level concepts into structured world rules, executable dynamics, and perceptual observations. This design enables long-term memory, open-ended interactions, autonomous world evolution, and multi-agent scenarios, where multiple entities can act, interact, and evolve persistently beyond the current observation. Extensive experiments demonstrate that our framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings. Code and model weights will be made publicly available. Project Page: \href{https://becauseimbatman0.github.io/CoDeR}{CoDeR}.
Apareció: miércoles, 23 de septiembre. arXiv. Preprint, todavía sin revisión por pares.
Comentario de los autores: https://becauseimbatman0.github.io/CoDeR