Robots that plan their own paths have a familiar failure mode: the plan is smart and late. Diffusion-style trajectory generators are great at multimodal behavior. They are also iterative. Many denoising steps per plan is fine for offline demos. It is rough when a pedestrian steps into your path now.
A Simon Fraser University paper for ICRA 2026 swaps the slow sampler for Implicit Maximum Likelihood Estimation (IMLE). One forward pass. Mode coverage stays. Latency drops hard enough for closed-loop model predictive control.
Project page and code: gmpc-imle.github.io. Paper: arXiv:2603.13733.
The problem with diffusion planners in the loop
Diffusion planners (think Diffuser-style UNets over trajectories) sample by iterative denoising. Guidance can add even more passes. On the authors’ locomotion setup, Diffuser sat around 1.33 Hz on CPU and 2.25 Hz on GPU (batch 64). That is not a comfortable replan rate when people move.
Cut the denoising steps and you get speed. You also get worse trajectories. Real-time control needs both.
What IMLE changes
IMLE trains a generator (f_\theta(z, c)) that maps noise and context straight to a trajectory. Training uses nearest-neighbor matching: every dataset trajectory should have at least one nearby generated sample. At inference you sample latents and generate in a single step.
For planning they add reward weighting so the generator prefers higher-return behavior, then plug samples into MPC / MPPI-style selection.
| Setting | Diffuser-class | IMLE (paper) |
|---|---|---|
| MuJoCo locomotion avg return | 77.5 | 78.47 |
| CPU sampling (Hz) | 1.33 | 32.87 |
| GPU sampling (Hz) | 2.25 | 53.52 |
| Maze2D CPU sampling (Hz) | 0.96 | 114.63 |
That is the “orders of magnitude faster, competitive quality” pitch, with numbers attached. AlphaSignal framed a digest-friendly 19x jump (4.3 to 83 plans/sec) and 38% less jerk in navigation. The project site’s controlled tables show 20x+ speedups on the Diffuser comparison. Either way, the qualitative win is the same: generative planning that can keep up with the world.
I also care that they stress-tested across open-loop Maze2D, locomotion offline RL, simulation pedestrians, and a physical robot. Papers that only win on one MuJoCo average are easier to overfit to a blog headline.

Real-time navigation, not just MuJoCo charts
The interesting validation is closed-loop human navigation. They compare against diffusion (CoBL / DDIM) and flow matching baselines in simulation, then deploy on a mobile robot among pedestrians.
Diffusion baselines struggle at responsive CPU rates (example cited: ~0.5 Hz with 50 DDIM steps). IMLE proposals warm-start MPPI, which improves constraint satisfaction while keeping sampling fast enough to replan. Onboard they report replanning up to 50 Hz on CPU with a small batch of trajectories.


Why builders outside robotics should still care
If you ship agents, you already feel this tradeoff: expressive generative models vs latency budgets. Diffusion (and long chain-of-thought) win quality contests. Production loops need single-digit-second or sub-second responses.
IMLE is one more data point that single-step generative models with the right training objective can keep multimodality without the iterative tax. Same story as people chasing distilled policies, speculative decoding, and better KV reuse for LLM agents.
Practical takeaways I’d steal:
- Measure plans per second under your real batch and hardware, not just offline success rate.
- Use generative models as proposal distributions for a controller, not as the whole controller.
- Prefer architectures that degrade gracefully when you need more samples in a fixed latency budget.
How I’d explain IMLE to a product team
Skip the loss equation in the first meeting. Say this:
- Diffusion drafts a path many times and slowly cleans it up.
- IMLE learns to jump to a good path in one shot, while still covering multiple plausible behaviors.
- The controller (MPC / MPPI) still picks and smooths. The generative model proposes.
That split is the reusable pattern. Generative models are great proposal engines. Classical control (or a simpler policy) remains the adult in the room for constraints.
| Layer | Job | Failure if missing |
|---|---|---|
| Generative proposer | Multimodal trajectory ideas | Jerky, myopic local optima |
| Cost / reward filter | Prefer safe, high-return rolls | Pretty collisions |
| Low-level tracking | Execute at high rate | Plan that never lands on hardware |
Limits to keep honest
Like other generative planners, IMLE still suffers when the optimal path sits far outside the training distribution. Augmentation helps. Adaptive mechanisms are future work. Pedestrian datasets also do not perfectly match robot dynamics.
JAX vs PyTorch throughput also differs in their notes (JAX higher on GPU in their reimplementation). If you are reproducing, pin the stack before you argue about Hz on Twitter.
Still, “competitive offline RL + real robot pedestrians + CPU-friendly Hz” is a stronger package than another slow simulation-only diffusion paper.
The software-agent rhyme
I spend most of my week on software agents, not mobile bases. The shared lesson is latency discipline:
- Prefer single-step or few-step generators when the loop is closed.
- Cache and reuse anything expensive (KV, embeddings, plans).
- Benchmark on your hardware and batch size, not the abstract.
If you are building applied AI systems where planning has to keep up with users, book a free discovery call.
