Robotics & Autonomous Systems

A robotics benchmark that cannot reproduce a run is not a benchmark. This pillar is small on purpose: the first thing worth building is a simulator whose trajectories are byte-identical given a seed, because without that property no claim about a dispatching policy can be distinguished from a claim about the random number generator.

Determinism first, policy second

The class of system under study is the Robotic Mobile Fulfillment System: inventory on movable pods, a fleet of autonomous mobile robots, human pick stations, and a controller deciding which robot serves which task. It is a well-defined operations research problem with a large literature and, until recently, very few open simulators that were both fast and reproducible.

Reproducibility is cheaper to get than it looks and is usually skipped anyway. It requires a seeded generator rather than a system-entropy one, canonical identifier-based tie-breaking everywhere two events could be ordered either way, and single-threaded execution in the deterministic path. Once it holds, the difference between two runs is attributable to the policy, which is the entire point of running the experiment.

What the reward measures is what you get

The second thread in this pillar is reward design. Throughput is the metric operators care about and a poor training signal: it is terminal, dominated by exogenous order arrivals, and saturates once station capacity rather than dispatching becomes the bottleneck. A policy trained on it learns very little about the cycle-time tail, which is the part that breaks a cut-off time.

The alternative is to use an exact decomposition of each completed task’s cycle time into components a dispatcher can and cannot influence, and to charge each decision for its share. That produces a dense, causally-aligned signal at the cost of an explicit modelling choice about responsibility — a trade we would rather argue about in the open than hide inside a constant.

What simulation does not settle

Practitioners routinely report a gap between a modelling study and the commissioned system, often a large one, and it is worth being precise about where it comes from. Part is scope: a simulator that models dispatching will not predict a throughput ceiling set by replenishment or the pack line. Part is distributional: service times that are drawn independently in simulation are correlated in a real building, by operator, shift hour and item class. Part is simply that layouts change between the study and the build.

None of that makes simulation useless; it makes absolute throughput prediction the wrong use of it. The defensible use is comparative — policy A against policy B, fleet size N against N+5, under identical seeded conditions — where the shared modelling error largely cancels. The articles in this pillar treat that as the design constraint rather than as a caveat at the end.

The second thing simulation does not settle is transfer. A policy trained on a simulated distribution meets a different one in deployment, and the standard mitigations — domain randomisation, conservative action masking, falling back to a heuristic when the state leaves the training envelope — are engineering answers rather than guarantees. Being explicit about which of those is in play is part of reporting a result honestly.