One shared 3D state
Register reconstructed observations from all agents into a common coordinate system.
* These authors contributed equally; their order was determined by drawing lots.
† Corresponding author. #Project Lead.
Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics and cross-agent interaction in the real world. However, existing approaches commonly adopt implicit inter-agent communications via cross attention, which lack explicit geometry constraints and unified 3D state, thereby leading to poor multi-view consistency and struggling with recovering out-of-sight agents. In addition, most of them assume a static background, failing to represent uncontrolled background dynamics.
To address these problems, we propose Artemis: a geometry-grounded multi-agent world model with explicit memory sharing. An explicit 3D world map is reconstructed from multi-agent observations to enforce a unified 3D state across agents, offering high cross-view consistency. Specifically, an action-guided geometric injection module is developed to simultaneously render decomposed foreground-background control maps, which are then injected into a diffusion transformer through a designed GeoAdapter block. Compared to previous methods assuming static-only background, our GeoAdapter can also distinguish uncontrolled non-agent dynamics, which are conditioned on their own multi-frame history positions to provide consistent motion cues.
Keyframes selected from progressive video rollouts are used to progressively update the reconstructed 3D world maps. To effectively capture complex dynamic patterns, we curate a novel dataset sampled from the CARLA simulator called MA-CARLA, with scalable agent numbers, physical plausibility, and abundant interaction modes. Extensive experiments demonstrate the superiority of our proposed method in terms of visual fidelity and cross-view consistency in the generated videos. In addition, our Artemis can support simultaneous multi-modal rollouts with both 2D video and 3D point map maintenance, scale to scenarios beyond two agents, and flexibly switch between single-camera or multi-camera setting.
A shared 3D map grounds each agent's perspective. Action-guided controls shape the rollout, and newly observed geometry becomes memory for the next step.
Register reconstructed observations from all agents into a common coordinate system.
Inject action-conditioned, decomposed control maps through GeoAdapter; exchange information with cross-agent attention.
Fuse selected keyframes into the static world map and maintain historical observations of scene dynamics.
MA-CARLA is a synchronized multi-agent driving dataset rendered in CARLA. It captures diverse agent interactions across urban environments, weather conditions, and static or dynamic backgrounds.
Each sequence contains 256 synchronized timestamps, with RGB observations, motion annotations, and camera calibration. The dataset includes opposing encounters, co-directional travel, and turning and crossing interactions.
Explore synchronized driving rollouts and cross-agent consistency across agent counts and interaction patterns.
Two independently controlled agents observe the same evolving scene. Paired views highlight agent interactions and cross-agent consistency.
GT, Artemis (Ours), Solaris public, Solaris finetuned, and OASIS share the same initial observations.
GT, Artemis (Ours), Solaris public, Solaris finetuned, and OASIS share the same initial observations.
Scale the shared world beyond two agents. Three synchronized perspectives expose more complex interactions within one global state.
Top row: ground truth. Bottom row: Artemis (Ours). Columns show Agent 1, Agent 2, and Agent 3 in matching order.
Top row: ground truth. Bottom row: Artemis (Ours). Columns show Agent 1, Agent 2, and Agent 3 in matching order.
Side-by-side comparisons across opposing, following, and turning & crossing scenarios. Each row compares GT, Artemis, Solaris public, Solaris finetuned, and OASIS under matched initial observations and actions.
Each method shows Agent 1 above Agent 2. Play all synchronizes the five methods in this scenario.
Each method shows Agent 1 above Agent 2. Play all synchronizes the five methods in this scenario.
Each method shows Agent 1 above Agent 2. Play all synchronizes the five methods in this scenario.
| Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Subject Consistency ↑ | Background Consistency ↑ | Aesthetic Quality ↑ | Imaging Quality ↑ |
|---|---|---|---|---|---|---|---|
| OASIS | 13.459 | 0.500 | 0.530 | 0.828 | 0.921 | 0.440 | 0.434 |
| Solaris | 10.900 | 0.378 | 0.660 | 0.759 | 0.916 | 0.392 | 0.525 |
| Solaris (finetuned) | 16.015 | 0.522 | 0.414 | 0.896 | 0.953 | 0.502 | 0.581 |
| Artemis (Ours) | 17.743 | 0.566 | 0.309 | 0.902 | 0.949 | 0.536 | 0.568 |
↑ Higher is better. ↓ Lower is better. Bold indicates the best result in each column. Metrics reproduced from Table 1 of the supplied manuscript.