NeurIPS 2026

MoSDOT

Multi-Agent Coordination
via Support-Preserving Distillation

Sangmin LeeYoungju NaChanmi LeeSung-eui Yoon†

KAIST

†Corresponding author

Teacher routing, then distillation. First a 2D toy with trained flow teachers: flow matching pairs noise with targets at random, MoSDOT by SDOT. Then the landmark diagnostic, replaying the paper figure's runs: joint teachers and the one-step students distilled from them. Red marks off-support samples, between valid modes.

TL;DR Better coordination starts with a better teacher: align source noise with joint-action modes, then distill into local one-step policies.

Video

MoSDOT in two minutes.

Why multimodal teachers fail under distillation, how source assignment fixes it, and what the agents do on the benchmarks.

Video

Overview. Narration is synthesized (Kokoro TTS); English captions are included. Table values are the paper's; the landmark animation replays the paper figure's runs, and the benchmark footage comes from policies re-trained with the paper configurations.

The problem

A coordination mode
belongs to the team.

Each agent alone has good options. Only some combinations work together, and those combinations are what the teacher must learn and the students must keep.

Two cars must pass a rock from opposite sides. Left: the joint action space; each car alone goes above or below, but only the opposite-side pairs, modes A and B, are valid; same-side pairs collide and the average drives both cars into the rock. Right: flow matching leaves samples between the modes, splitting each car's own noise reaches head-on pairs 48% of the time, and MoSDOT splits the noise diagonally into one cell per joint mode.
Example

Modes are pairs, not single choices. Either car may pass above or below the rock; only opposite sides get through. Flow matching leaves samples between the two modes, splitting each car's own noise reaches head-on pairs, and MoSDOT gives each joint mode its own noise cell, sized by how often the mode occurs in the data.

The idea

Good endpoints.
Better routes.

A centralized teacher can reach valid joint-action modes yet still assign nearby noise samples to conflicting behaviors. Local actors can inherit these artifacts during distillation, producing actions between valid modes.

MoSDOT organizes the source-to-mode assignment first. Capacity-matched transport gives the teacher cleaner targets and preserves more of the support that decentralized actors can recover.

A clean teacher does not remove the strict-product gap: independent local policies cannot represent every correlated joint distribution. We study shared randomness separately.

Flow Policy scatters source trajectories between modes; MoSDOT organizes the same space into capacity-matched regions. A radar plot compares four teacher diagnostics.
Fig. 1

Give every noise sample a coordinated destination. MoSDOT assigns source regions to replay-derived modes before teacher training, reducing off-support mass and improving routing consistency.

Landmark means over 4 seeds (Table 1). Benchmark averages as reported in Tables 2–3; benchmark evaluation uses 6 seeds.

The method

Align. Train. Distill.

A structured source assignment before teacher training.
One-step local actors at execution.

Replay joint actions become mode support and capacities. Conditional SDOT assigns source noise to modes. A centralized teacher is distilled into local actors, optionally with shared randomness.
Fig. 2

The MoSDOT framework. Transport operates on coordinated joint-action modes. Each distilled actor uses its own observation and noise; the shared-randomness variant additionally supplies a common signal h.

01

Build mode support

Summarize replay into representative joint actions yk(x), using identity support or joint-action quantization.

02

Match capacities

Assign each mode a target source mass bk(x) from its empirical replay frequency.

03

Align with SDOT

Partition source noise into capacity-matched Laguerre cells, each assigned to one joint-action mode.

04

Distill local actors

Train the teacher on aligned paths and distill its joint actions into decentralized one-step policies.

Benchmark results

Coordination across
different kinds of replay.

Strong gains on diverse continuous replay, with results that vary by discrete-action regime. Explore every baseline and dataset below.

MPE Simple Spread

Continuous actions · four dataset qualities

Paper · Table 3
MPE Simple Spread. Higher is better.
MethodExpert ↑Medium ↑Medium-Replay ↑Random ↑Average ↑
MATD3BC108.3 ± 3.329.3 ± 4.815.4 ± 5.69.8 ± 4.940.7
MACQL98.2 ± 5.234.1 ± 7.220.0 ± 8.424.0 ± 9.844.1
ICQ114.9 ± 2.647.9 ± 18.937.9 ± 12.334.4 ± 5.358.8
OMAR104.0 ± 3.429.3 ± 5.513.6 ± 5.76.3 ± 3.538.3
OMIGA80.8 ± 13.830.1 ± 16.95.4 ± 11.0−3.8 ± 12.328.1
MADiff95.0 ± 5.364.9 ± 7.730.3 ± 2.56.9 ± 3.149.2
MAC-Flow101.7 ± 10.980.1 ± 20.650.4 ± 33.231.1 ± 6.865.8
MoSDOT Ours104.40 ± 19.6695.97 ± 12.6954.12 ± 15.2975.92 ± 23.8182.6

MoSDOT leads on Medium, Medium-Replay, and Random; ICQ has the highest Expert mean.

SMACv1

Five scenarios, three dataset qualities.

All scenarios · reported average

SMACv1 published average. Higher is better.
MethodAverage ↑
BC12.2
MABCQ5.5
MACQL13.1
Diffusion BC13.0
MADiff13.8
DoF15.6
Flow BC13.4
MAC-Flow15.6
MoSDOT Ours16.7

3m

SMACv1 / 3m. Higher is better.
MethodGood ↑Medium ↑Poor ↑
BC16.0 ± 1.08.2 ± 0.84.4 ± 0.1
MABCQ3.7 ± 1.14.0 ± 1.03.4 ± 1.0
MACQL19.1 ± 0.113.7 ± 0.34.2 ± 0.1
Diffusion BC19.5 ± 0.513.3 ± 0.74.2 ± 0.2
MADiff19.3 ± 0.516.4 ± 2.610.3 ± 6.1
DoF19.8 ± 0.218.6 ± 1.210.9 ± 1.1
Flow BC20.0 ± 0.014.7 ± 1.54.5 ± 0.1
MAC-Flow19.8 ± 0.218.0 ± 3.210.6 ± 2.2
MoSDOT Ours20.0 ± 0.018.8 ± 1.516.4 ± 4.9

8m

SMACv1 / 8m. Higher is better.
MethodGood ↑Medium ↑Poor ↑
BC16.7 ± 0.410.7 ± 0.55.3 ± 0.1
MABCQ4.8 ± 0.65.6 ± 0.63.6 ± 0.8
MACQL18.9 ± 0.915.5 ± 1.57.5 ± 1.0
Diffusion BC19.4 ± 0.518.6 ± 0.64.8 ± 0.2
MADiff18.9 ± 1.116.8 ± 1.69.8 ± 0.9
DoF19.6 ± 0.318.6 ± 0.812.0 ± 1.2
Flow BC19.5 ± 0.218.2 ± 0.84.9 ± 0.1
MAC-Flow19.7 ± 0.319.4 ± 0.611.5 ± 0.8
MoSDOT Ours20.0 ± 0.020.0 ± 0.010.8 ± 0.8

2s3z

SMACv1 / 2s3z. Higher is better.
MethodGood ↑Medium ↑Poor ↑
BC18.2 ± 0.412.3 ± 0.76.7 ± 0.3
MABCQ7.7 ± 0.97.6 ± 0.76.6 ± 0.2
MACQL17.4 ± 0.315.6 ± 0.48.4 ± 0.8
Diffusion BC18.0 ± 1.013.4 ± 1.46.2 ± 1.2
MADiff15.9 ± 1.215.6 ± 0.38.5 ± 1.3
DoF18.5 ± 0.818.1 ± 0.910.0 ± 1.1
Flow BC19.5 ± 0.115.1 ± 2.06.9 ± 0.8
MAC-Flow19.5 ± 0.517.6 ± 0.68.5 ± 0.6
MoSDOT Ours20.1 ± 0.118.2 ± 1.09.4 ± 0.8

5m_vs_6m

SMACv1 / 5m_vs_6m. Higher is better.
MethodGood ↑Medium ↑Poor ↑
BC15.8 ± 3.612.4 ± 0.97.5 ± 0.2
MABCQ2.4 ± 0.43.8 ± 0.53.3 ± 0.5
MACQL16.2 ± 1.615.1 ± 2.910.5 ± 3.1
Diffusion BC16.8 ± 2.312.5 ± 2.18.0 ± 1.0
MADiff16.5 ± 2.815.2 ± 2.68.9 ± 1.3
DoF17.7 ± 1.116.2 ± 0.910.8 ± 0.3
Flow BC14.7 ± 2.112.8 ± 0.87.7 ± 0.8
MAC-Flow18.6 ± 3.515.6 ± 1.39.8 ± 2.1
MoSDOT Ours18.5 ± 1.917.9 ± 1.212.3 ± 0.9

2c_vs_64zg

SMACv1 / 2c_vs_64zg. Higher is better.
MethodGood ↑Medium ↑Poor ↑
BC17.5 ± 0.412.5 ± 0.39.7 ± 0.2
MABCQ10.1 ± 0.29.9 ± 0.29.0 ± 0.2
MACQL12.9 ± 0.211.6 ± 0.110.2 ± 0.1
Diffusion BC17.8 ± 1.310.5 ± 1.110.2 ± 2.3
MADiff14.7 ± 2.212.8 ± 1.210.8 ± 1.1
DoF16.1 ± 0.813.9 ± 0.911.5 ± 1.1
Flow BC18.0 ± 1.311.8 ± 2.610.0 ± 0.3
MAC-Flow19.1 ± 0.814.9 ± 4.111.4 ± 0.4
MoSDOT Ours20.2 ± 0.616.4 ± 1.011.4 ± 1.4

Reported average reward: MoSDOT 16.7; MAC-Flow and DoF 15.6. Select a scenario to inspect all means and uncertainties.

SMACv2

Discrete actions · replay datasets

Paper · Table 2
SMACv2. Higher is better.
Methodterran_5_vs_5 ↑zerg_5_vs_5 ↑terran_10_vs_10 ↑Average ↑
BC7.3 ± 1.06.8 ± 0.67.4 ± 0.57.2
MABCQ13.8 ± 4.410.3 ± 1.212.7 ± 2.012.3
MACQL11.8 ± 0.910.3 ± 3.411.8 ± 2.011.3
Diffusion BC9.3 ± 0.98.1 ± 1.75.5 ± 1.57.6
MADiff13.3 ± 1.810.2 ± 1.113.8 ± 1.312.4
DoF15.4 ± 1.312.0 ± 1.114.6 ± 1.114.0
Flow BC8.3 ± 1.94.6 ± 0.55.8 ± 1.76.2
MAC-Flow16.6 ± 4.39.8 ± 1.513.0 ± 4.713.1
MoSDOT Ours16.6 ± 4.39.7 ± 2.89.4 ± 2.611.9

DoF has the highest reported average (14.0), followed by MAC-Flow (13.1); MoSDOT reports 11.9.

Mean ± 2σ · 6 seeds · higher is better

Bold best mean Underlined second-best mean · ties included

Tables 2–3 in the paper. Average columns reproduce the reported values without recalculation. Download table data Swipe tables horizontally to see all columns.

MoSDOT in action.

MoSDOT's decentralized one-step actors on every benchmark. StarCraft II clips are rendered by the game engine itself during evaluation (MoSDOT red, built-in AI blue). The clips show successful episodes of policies re-trained with the paper configurations.

MPE Simple Spread

StarCraft II · SMACv1

StarCraft II · SMACv2

Looking closer

What survives distillation?

Controlled diagnostics separate teacher routing quality from the limits of independent local execution.

A

Cleaner teachers, cleaner local policies.

MoSDOT reduces fan mass between modes and retains the teacher’s ray structure after distillation.

Fig. 3

Teacher → student. Top: centralized teacher rollouts. Bottom: one-step factorized students with independent local noise. Black crosses mark the six anchor mode endpoints. Animated from the same runs as the paper figure: rollouts are drawn one at a time, and red marks rollouts whose first step leaves both 10° landmark cones (off-support), with a live count per panel.

Endpoint quality and routing consistency

Joint-teacher landmark diagnostics · Table 1 · mean over 4 seeds

Joint-teacher landmark diagnostics. Higher is better.
MethodFan-free ↑Route consistency ↑Success ↑Balance ↑
Flow matching0.8570.8560.9150.917
IMLE0.9910.8570.8480.761
Drifting0.9520.6530.9850.936
MoSDOT Ours0.9830.9720.9910.961

IMLE has the highest fan-free score. MoSDOT has the highest route consistency, success, and balance. Bold marks the best mean; underlining marks the second-best.

B

A clean teacher still has an execution limit.

On XOR, independent local sampling cannot preserve the correlated target support. A shared mode-selection signal lets local actors coordinate.

Four XOR panels: target support, MAC-Flow, MoSDOT without shared signal, and MoSDOT with shared signal. The last panel concentrates both teacher and student samples on the two valid opposite-sign tuples.
Fig. 6

Isolating the strict-product gap. (a) Target support. (b) MAC-Flow. (c) MoSDOT without a shared signal. (d) MoSDOT with a shared signal. Blue: factored student; pink: joint teacher.

More evidence: joint mode-tuple support
Bars compare teacher and student mass assigned to joint mode tuples for Flow, IMLE, Drift, and MoSDOT.
Fig. 7

Joint mode-tuple support before and after distillation. The original appendix figure reports fractions over 32,768 samples.

Reference

Cite this work.

Accepted to NeurIPS 2026.
Sangmin Lee, Youngju Na, Chanmi Lee, and Sung-eui Yoon · KAIST.

Read the manuscript
BibTeX · NeurIPS 2026
@inproceedings{lee2026mosdot,
  title = {Multi-Agent Coordination via
           Support-Preserving Distillation},
  author = {Sangmin Lee and Youngju Na and
            Chanmi Lee and Sung-eui Yoon},
  booktitle = {Advances in Neural Information
               Processing Systems},
  year = {2026}
}

Figure

Original PDF ↗