Abstract
To address this gap, we introduce COMRAD, a COoperative Multi-Agent Reinforcement LeArning benchmark suite in Doom, featuring a diverse set of challenging scenarios spanning role asymmetry, temporal synchronization, and spatial navigation.
To introduce within-scenario variability, we develop DoomGen, a procedural map generator that produces diverse layout configurations for every scenario. We integrate COMRAD with Sample Factory, a high-throughput asynchronous RL framework, and implement seven MARL baselines on top of it, reaching ~25K frames per second during training. Our experiments show that COMRAD poses significant challenges for current CTDE methods, establishing visual cooperative MARL as an important open frontier.
What We Introduce
13 Cooperative Scenarios
Agents coordinate solely from their local 3D first-person visual observations: no shared state, no explicit communication.
Easy - D1
Medium - D3
Medium - D3
Medium - D3
Hard - D4
Hard - D4
Hard - D4
Difficult - D5
Difficult - D5
Difficult - D5
Severe - D6
Very Difficult - D7
Very Difficult - D7
Methodology
High-throughput infrastructure, procedural generation, and curriculum learning form the backbone of COMRAD.
To make spatial memorization as difficult as possible, DoomGen avoids traditional orthogonal, grid-based room layouts. It constructs maps starting from a relaxed Voronoi tessellation, grouping irregular cells into semantically defined areas before translating the graph into valid Doom geometry.
This geometric irregularity broadens the distribution of boundary shapes, geometries, sightlines, and local visual cues, exposing algorithms to a distribution of spatial configurations rather than a single memorizable map, while still preserving task semantics.
Task-mechanical difficulty controls scenario parameters (enemy count, time limits, hazard intensity) that shape the coordination challenge within a fixed map topology. Spatial difficulty controls layout complexity produced by DoomGen, from compact single-room arenas to sprawling multi-wing labyrinths.
Scaling one axis while holding the other fixed isolates whether a method fails from task demands or spatial generalization, allowing sharper experimental conclusions across curriculum, transfer, and generalization studies.
Results
Performance of 7 established CTDE baselines across all 13 COMRAD scenarios under a 100M-step training budget.
| Scenario | IPPO | MAPPO | HAPPO | IDQN | VDN | QMIX | QPLEX-D | QPLEX-Q | Ceiling |
|---|---|---|---|---|---|---|---|---|---|
| Stag Hunt Arena | 17.55 ± 0.22 | 25.66 ± 0.32 | 2.58 ± 0.10 | 14.33 ± 0.18 | 17.96 ± 0.39 | 0.28 ± 0.06 | 1.49 ± 0.10 | 3.42 ± 0.20 | 36 |
| Rhythm Sync | 0.59 ± 0.05 | 0.65 ± 0.05 | 0.22 ± 0.02 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 1.00 |
| Foraging Commons | 2355.64 ± 30.33 | 2510.93 ± 31.18 | 2199.89 ± 39.31 | 2190.70 ± 32.66 | 2327.13 ± 42.84 | 2544.89 ± 46.77 | 1905.26 ± 23.66 | 2345.75 ± 39.38 | 21000 |
| Co-op Puzzle | 3.03 ± 0.07 | 3.88 ± 0.09 | 2.23 ± 0.14 | 0.03 ± 0.01 | 0.07 ± 0.02 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.07 ± 0.03 | 5 |
| Platform Chain | 2.00 ± 0.02 | 4.78 ± 0.19 | 4.67 ± 0.18 | 0.68 ± 0.06 | 1.12 ± 0.08 | 0.97 ± 0.02 | 3.40 ± 0.22 | 0.97 ± 0.03 | 47 |
| Armory Siege | 16.71 ± 2.90 | 19.87 ± 2.32 | -0.71 ± 1.07 | 10.97 ± 0.98 | 18.32 ± 1.46 | 23.12 ± 3.83 | 1.19 ± 0.89 | 14.13 ± 2.47 | 100 |
| Co-op Health Gathering | 1104.90 ± 117.30 | 553.96 ± 27.82 | 197.74 ± 7.98 | 164.20 ± 2.12 | 573.83 ± 27.28 | 166.43 ± 2.07 | 193.95 ± 5.10 | 267.78 ± 18.39 | 2100 |
| Lava Pit | 0.92 ± 0.01 | 1.00 ± 0.01 | 1.00 ± 0.01 | 0.92 ± 0.01 | 0.80 ± 0.05 | 0.97 ± 0.03 | 0.00 ± 0.01 | 1.00 ± 0.01 | 11 |
| Smart Enemies | 11.56 ± 0.86 | 14.23 ± 0.65 | 6.85 ± 0.38 | 3.30 ± 0.23 | 4.89 ± 0.38 | 0.64 ± 0.08 | 0.00 ± 0.01 | 0.65 ± 0.07 | 84 |
| Dumb Enemies | 12.44 ± 1.93 | 2.29 ± 0.20 | 1.03 ± 0.13 | 1.07 ± 0.11 | 2.76 ± 0.22 | 1.10 ± 0.15 | 0.00 ± 0.01 | 1.80 ± 0.16 | 90 |
| Ammo Carrier | 7052.93 ± 483.17 | 1362.73 ± 69.09 | 1340.99 ± 58.07 | 3836.13 ± 384.96 | 2542.51 ± 219.25 | 2379.46 ± 139.55 | 1530.98 ± 59.06 | 2080.97 ± 192.02 | 8400 |
| Stealth Labyrinth | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.21 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 1.00 |
| Lava Maze | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 0.00 ± 0.01 | 6 |
When Centralized Training Hurts
Visualizing spatial coordination and division of labor between agents in the Ammo Carrier scenario.
Why IPPO Outperforms HAPPO on Ammo Carrier
Objective: Ammo Carrier is a fixed-role asymmetric defense task. One immobilized defender holds the hub against enemies, while a mobile runner maintains a depot-to-hub resupply loop. The challenge lies in learning a stable, long-horizon logistics routine with delayed, role-specific payoffs.
IPPO (Left): Decentralized updates allow the runner to independently optimize its loop without cross-agent interference. IPPO concentrates movement along a recurring depot-to-hub corridor, successfully learning the required role-specialized logistics behavior.
HAPPO (Right): Centralized training creates counterproductive coordination pressure. The runner's gradient updates are influenced by the defender's value signal, causing it to abandon the logistics routine. This lead to episodic ammo starvation and significantly lower survival times.
Citation
@inproceedings{nguyen2026comrad,
title={{COMRAD}: A Benchmark for Embodied Cooperative Multi-Agent Reinforcement Learning},
author={Khoi H.B. Nguyen and Dimitar Zhivkov Zhekov and Tristan Tomilin},
booktitle={New Frontiers in Game-Theoretic Learning - NExT-Game},
year={2026},
url={https://openreview.net/forum?id=bXau6dlyV4}
}