Geometry-grounded multi-agent world modeling

Artemis: Geometry-Grounded Multi-Agent
Driving World Models with Shared 3D State
and Progressive Memory Update

Sitian Shen2*Jiuming Liu1*#Mengmeng Liu3*Yian Wang4Michael Ying Yang5

Francesco Nex3Hao Cheng3Daniele De Martini2Ayush Tewari1†Per Ola Kristensson1

1 University of Cambridge2 University of Oxford3 University of Twente4 National University of Singapore5 University of Bath

* These authors contributed equally; their order was determined by drawing lots.
† Corresponding author. #Project Lead.

Paper · Coming soonExplore demonstrationsCode · Coming soon
MULTIPLE AGENTS. MULTIPLE VIEWS. ONE SHARED WORLD.
Shared 3D world stateCross-agent memoryScalable agents & camerasDynamic environments
Artemis jointly rolls out driving videos while maintaining an explicit 3D world state across agents and cameras.

Abstract

Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics and cross-agent interaction in the real world. However, existing approaches commonly adopt implicit inter-agent communications via cross attention, which lack explicit geometry constraints and unified 3D state, thereby leading to poor multi-view consistency and struggling with recovering out-of-sight agents. In addition, most of them assume a static background, failing to represent uncontrolled background dynamics.

To address these problems, we propose Artemis: a geometry-grounded multi-agent world model with explicit memory sharing. An explicit 3D world map is reconstructed from multi-agent observations to enforce a unified 3D state across agents, offering high cross-view consistency. Specifically, an action-guided geometric injection module is developed to simultaneously render decomposed foreground-background control maps, which are then injected into a diffusion transformer through a designed GeoAdapter block. Compared to previous methods assuming static-only background, our GeoAdapter can also distinguish uncontrolled non-agent dynamics, which are conditioned on their own multi-frame history positions to provide consistent motion cues.

Keyframes selected from progressive video rollouts are used to progressively update the reconstructed 3D world maps. To effectively capture complex dynamic patterns, we curate a novel dataset sampled from the CARLA simulator called MA-CARLA, with scalable agent numbers, physical plausibility, and abundant interaction modes. Extensive experiments demonstrate the superiority of our proposed method in terms of visual fidelity and cross-view consistency in the generated videos. In addition, our Artemis can support simultaneous multi-modal rollouts with both 2D video and 3D point map maintenance, scale to scenarios beyond two agents, and flexibly switch between single-camera or multi-camera setting.

Method overview

From shared geometry to synchronized futures

A shared 3D map grounds each agent's perspective. Action-guided controls shape the rollout, and newly observed geometry becomes memory for the next step.

Overall pipeline of Artemis. Initial observations are registered into a shared 3D map. Decomposed controls for static geometry, controlled agents, and uncontrolled dynamics guide the diffusion transformer through GeoAdapter. Keyframes from generated videos progressively update the world map.
01 / INITIALIZE

One shared 3D state

Register reconstructed observations from all agents into a common coordinate system.

02 / GENERATE

Geometry-guided rollouts

Inject action-conditioned, decomposed control maps through GeoAdapter; exchange information with cross-agent attention.

03 / UPDATE

Progressive memory

Fuse selected keyframes into the static world map and maintain historical observations of scene dynamics.

Synchronized multi-agent driving data

MA-CARLA Dataset

MA-CARLA is a synchronized multi-agent driving dataset rendered in CARLA. It captures diverse agent interactions across urban environments, weather conditions, and static or dynamic backgrounds.

11,283Rendered sequences
2,478Unique trajectories
8Towns
17Weather presets
MA-CARLA dataset. Representative driving samples and distributions of interaction modes and agent counts. Motion groups may overlap; chart percentages are computed over 17,790 group assignments.
6,020Two-agent sequences with static backgrounds
3,382Two-agent sequences with dynamic backgrounds
1,881Three-agent sequences with static backgrounds

Each sequence contains 256 synchronized timestamps, with RGB observations, motion annotations, and camera calibration. The dataset includes opposing encounters, co-directional travel, and turning and crossing interactions.

Visual demonstrations

A world shared across agents

Explore synchronized driving rollouts and cross-agent consistency across agent counts and interaction patterns.

Two agents, single camera

01 / 2A · 1C

Two independently controlled agents observe the same evolving scene. Paired views highlight agent interactions and cross-agent consistency.

Opposing interaction

2 agents · single camera
GT
Artemis (Ours)
Solaris public
Solaris finetuned
OASIS
Agent 1
Agent 2
0:00 / 0:10

GT, Artemis (Ours), Solaris public, Solaris finetuned, and OASIS share the same initial observations.

Following interaction

2 agents · single camera
GT
Artemis (Ours)
Solaris public
Solaris finetuned
OASIS
Agent 1
Agent 2
0:00 / 0:10

GT, Artemis (Ours), Solaris public, Solaris finetuned, and OASIS share the same initial observations.

Three agents, single camera

03 / 3A · 1C

Scale the shared world beyond two agents. Three synchronized perspectives expose more complex interactions within one global state.

Following

3 agents · single camera
Agent 1
Agent 2
Agent 3
GT
Artemis (Ours)
0:00 / 0:10

Top row: ground truth. Bottom row: Artemis (Ours). Columns show Agent 1, Agent 2, and Agent 3 in matching order.

Crossing

3 agents · single camera
Agent 1
Agent 2
Agent 3
GT
Artemis (Ours)
0:00 / 0:10

Top row: ground truth. Bottom row: Artemis (Ours). Columns show Agent 1, Agent 2, and Agent 3 in matching order.

05 / Comparisons

Comparison with other world models

Side-by-side comparisons across opposing, following, and turning & crossing scenarios. Each row compares GT, Artemis, Solaris public, Solaris finetuned, and OASIS under matched initial observations and actions.

Opposing

2 agents · single camera
GT
Artemis (Ours)
Solaris public
Solaris finetuned
OASIS
0:00 / 0:10

Each method shows Agent 1 above Agent 2. Play all synchronizes the five methods in this scenario.

Following

2 agents · single camera
GT
Artemis (Ours)
Solaris public
Solaris finetuned
OASIS
0:00 / 0:10

Each method shows Agent 1 above Agent 2. Play all synchronizes the five methods in this scenario.

Turning & crossing

2 agents · single camera
GT
Artemis (Ours)
Solaris public
Solaris finetuned
OASIS
0:00 / 0:10

Each method shows Agent 1 above Agent 2. Play all synchronizes the five methods in this scenario.

MA-CARLA: quantitative results from Table 1
MethodPSNR ↑SSIM ↑LPIPS ↓Subject Consistency ↑Background Consistency ↑Aesthetic Quality ↑Imaging Quality ↑
OASIS13.4590.5000.5300.8280.9210.4400.434
Solaris10.9000.3780.6600.7590.9160.3920.525
Solaris (finetuned)16.0150.5220.4140.8960.9530.5020.581
Artemis (Ours)17.7430.5660.3090.9020.9490.5360.568

↑ Higher is better. ↓ Lower is better. Bold indicates the best result in each column. Metrics reproduced from Table 1 of the supplied manuscript.

Enlarged Artemis paper figure