This robot planner skips diffusion’s slow drafts and plans in one shot

IMLE-based generative MPC hits competitive offline RL scores while sampling trajectories over 20x faster than Diffuser. Real robots can replan around people without waiting on denoising.

SaifullahSaifullah
5 min read
This robot planner skips diffusion’s slow drafts and plans in one shot

Robots that plan their own paths have a familiar failure mode: the plan is smart and late. Diffusion-style trajectory generators are great at multimodal behavior. They are also iterative. Many denoising steps per plan is fine for offline demos. It is rough when a pedestrian steps into your path now.

A Simon Fraser University paper for ICRA 2026 swaps the slow sampler for Implicit Maximum Likelihood Estimation (IMLE). One forward pass. Mode coverage stays. Latency drops hard enough for closed-loop model predictive control.

Project page and code: gmpc-imle.github.io. Paper: arXiv:2603.13733.

The problem with diffusion planners in the loop

Diffusion planners (think Diffuser-style UNets over trajectories) sample by iterative denoising. Guidance can add even more passes. On the authors’ locomotion setup, Diffuser sat around 1.33 Hz on CPU and 2.25 Hz on GPU (batch 64). That is not a comfortable replan rate when people move.

Cut the denoising steps and you get speed. You also get worse trajectories. Real-time control needs both.

What IMLE changes

IMLE trains a generator (f_\theta(z, c)) that maps noise and context straight to a trajectory. Training uses nearest-neighbor matching: every dataset trajectory should have at least one nearby generated sample. At inference you sample latents and generate in a single step.

For planning they add reward weighting so the generator prefers higher-return behavior, then plug samples into MPC / MPPI-style selection.

SettingDiffuser-classIMLE (paper)
MuJoCo locomotion avg return77.578.47
CPU sampling (Hz)1.3332.87
GPU sampling (Hz)2.2553.52
Maze2D CPU sampling (Hz)0.96114.63

That is the “orders of magnitude faster, competitive quality” pitch, with numbers attached. AlphaSignal framed a digest-friendly 19x jump (4.3 to 83 plans/sec) and 38% less jerk in navigation. The project site’s controlled tables show 20x+ speedups on the Diffuser comparison. Either way, the qualitative win is the same: generative planning that can keep up with the world.

I also care that they stress-tested across open-loop Maze2D, locomotion offline RL, simulation pedestrians, and a physical robot. Papers that only win on one MuJoCo average are easier to overfit to a blog headline.

Diagram comparing IMLE single-step trajectory generation with multi-step diffusion denoising

Real-time navigation, not just MuJoCo charts

The interesting validation is closed-loop human navigation. They compare against diffusion (CoBL / DDIM) and flow matching baselines in simulation, then deploy on a mobile robot among pedestrians.

Diffusion baselines struggle at responsive CPU rates (example cited: ~0.5 Hz with 50 DDIM steps). IMLE proposals warm-start MPPI, which improves constraint satisfaction while keeping sampling fast enough to replan. Onboard they report replanning up to 50 Hz on CPU with a small batch of trajectories.

IMLE planner project visual for real-time generative model predictive control
IMLE method figure showing single-step generative trajectory sampling

Why builders outside robotics should still care

If you ship agents, you already feel this tradeoff: expressive generative models vs latency budgets. Diffusion (and long chain-of-thought) win quality contests. Production loops need single-digit-second or sub-second responses.

IMLE is one more data point that single-step generative models with the right training objective can keep multimodality without the iterative tax. Same story as people chasing distilled policies, speculative decoding, and better KV reuse for LLM agents.

Practical takeaways I’d steal:

  • Measure plans per second under your real batch and hardware, not just offline success rate.
  • Use generative models as proposal distributions for a controller, not as the whole controller.
  • Prefer architectures that degrade gracefully when you need more samples in a fixed latency budget.

How I’d explain IMLE to a product team

Skip the loss equation in the first meeting. Say this:

  1. Diffusion drafts a path many times and slowly cleans it up.
  2. IMLE learns to jump to a good path in one shot, while still covering multiple plausible behaviors.
  3. The controller (MPC / MPPI) still picks and smooths. The generative model proposes.

That split is the reusable pattern. Generative models are great proposal engines. Classical control (or a simpler policy) remains the adult in the room for constraints.

LayerJobFailure if missing
Generative proposerMultimodal trajectory ideasJerky, myopic local optima
Cost / reward filterPrefer safe, high-return rollsPretty collisions
Low-level trackingExecute at high ratePlan that never lands on hardware

Limits to keep honest

Like other generative planners, IMLE still suffers when the optimal path sits far outside the training distribution. Augmentation helps. Adaptive mechanisms are future work. Pedestrian datasets also do not perfectly match robot dynamics.

JAX vs PyTorch throughput also differs in their notes (JAX higher on GPU in their reimplementation). If you are reproducing, pin the stack before you argue about Hz on Twitter.

Still, “competitive offline RL + real robot pedestrians + CPU-friendly Hz” is a stronger package than another slow simulation-only diffusion paper.

The software-agent rhyme

I spend most of my week on software agents, not mobile bases. The shared lesson is latency discipline:

  • Prefer single-step or few-step generators when the loop is closed.
  • Cache and reuse anything expensive (KV, embeddings, plans).
  • Benchmark on your hardware and batch size, not the abstract.

If you are building applied AI systems where planning has to keep up with users, book a free discovery call.

Share this post

Related posts