Reward Design for Warehouse Dispatching Agents

Why throughput is a poor training signal for RMFS dispatching policies, and how a five-bucket decomposition of cycle time gives per-decision credit instead.

A dispatching policy learns whatever the reward function measures. In a Robotic Mobile Fulfillment System the obvious reward — orders completed per hour — is the one metric that cannot be attributed to any individual decision. In waremax’s own experiments, policies trained on naive sparse or dense rewards plateau near the weakest heuristic baseline, at roughly 82–85% on-time attainment; a candidate-scoring policy paid for the delay its decision controls reaches roughly 97%, level with the strong heuristics. The more useful signal is already sitting in the simulator: the exact decomposition of each completed task’s cycle time into the components a dispatcher can and cannot control.

This note describes the reward-design problem as it appears in waremax, our open discrete-event simulator and reinforcement-learning benchmark for RMFS task allocation, and argues that the decomposition is the interesting part of the benchmark rather than the environment wrapper.

The decision the policy actually makes

An RMFS is a warehouse in which inventory lives on movable pods instead of fixed shelving. A fleet of autonomous mobile robots lifts a pod, carries it to a human pick station, waits while the operator takes what the order needs, and returns the pod to the floor. Amazon’s Kiva-derived system is the best-known instance; waremax models this class and nothing else — no AS/RS cranes, no conveyor sortation, no tugger trains, no pickers walking aisles.

Within that scope the controller is asked one question, repeatedly: given a task that has become assignable and the robots that could take it, which robot gets it? It is asked this question continuously through a simulated shift, and it is asked at irregular moments — whenever a robot finishes, or a new order makes a pod relevant. Nothing happens on a fixed tick.

That irregularity is not a detail. It means the natural formulation is a semi-Markov decision process: the interval between two consecutive decision points is itself a random variable, driven by travel distance, queue depth at the station, and congestion in the aisles. A reward defined per step is not comparable across decisions, because the steps have different durations. A reward defined per unit of simulated time is.

Why throughput fails as a training signal

Throughput is the right acceptance criterion and the wrong gradient. Three problems compound.

It is terminal. Orders completed per hour is known only at the end of an episode. Crediting it back to an assignment made much earlier in simulated time spreads one scalar across thousands of actions with no mechanism to say which of them helped.

It is dominated by arrivals. Order arrival is exogenous. A run that happened to receive a friendlier mix of orders scores higher regardless of policy. Unless arrivals are held fixed across comparisons, most of the variance in the reward has nothing to do with the thing being learned.

It saturates. Above a certain fleet size, throughput is bounded by station service capacity rather than by dispatching. The policy receives an identical reward for a wide range of behaviours, and the gradient goes quiet exactly in the operating regime where cycle-time tails still differ enormously.

The last point is the most damaging in practice. A warehouse operator does not care only about the mean; they care about the ninety-fifth percentile of order cycle time, because that is what breaks the cut-off for the evening van. A reward that is blind to the tail trains a policy that is blind to the tail.

Five buckets that sum to cycle time

The alternative is to use the decomposition the simulator can compute exactly. For every completed task, waremax attributes the elapsed cycle time to five buckets:

BucketWhat it measuresControlled by the dispatcher?
AssignmentTime the task waited before any robot was assigned to itDirectly
TravelTime the assigned robot spent moving to the pod and on to the stationPartly: travel to pickup depends on which robot was chosen
QueueTime spent waiting in line at the pick stationIndirectly, through station balance
CongestionTime lost to aisle contention and yieldingIndirectly, through path overlap
ServiceTime the human operator spent pickingNot at all

The buckets are exhaustive and non-overlapping: they sum to the task’s cycle time. That property is what makes them usable as a reward. A dispatching decision can be charged for the assignment wait it caused and the travel to pickup it selected, and held harmless for service time, which is drawn from a lognormal distribution and is entirely outside its control.

The harder question is what to do with queue and congestion, which a decision influences but does not control. waremax ships four reward modes so this can be tested rather than argued: sparse and dense as baselines, attribution, which exposes the full causal decomposition, and routed, which charges each decision only the controllable cost. The result reported with the benchmark is what we call a controllability principle: restricting the reward to the delay the agent controls (assignment wait and travel to pickup) is directionally better than additionally penalising congestion and station queue, with the comparison tested by Welch’s t across seeds.

The resulting signal is dense — it arrives per decision rather than once per episode — and it is causally aligned with the decision that produced it. Neither is true of throughput.

Reward is only half of it. The same experiments show a representation effect: a permutation-equivariant candidate-scoring policy, which scores each candidate robot with a shared network and selects with a masked softmax, reaches strong-heuristic attainment with a controllable-delay reward, while a flattened MLP plateaus near the weakest heuristic. Neither the representation nor the reward suffices alone.

A protocol for claiming an improvement

Dense rewards make it easier to learn something. They do not make it easier to know whether what was learned generalises. The benchmark discipline matters as much as the reward.

  1. Fix the seeds. waremax’s Rust core seeds a ChaCha8 generator from a u64 and applies canonical, identifier-based tie-breaking throughout, so the same seed and the same action sequence produce a byte-identical trajectory. Determinism is a tested property, not an aspiration.
  2. Hold the scenario constant. Topology, station layout, order arrival process and traffic parameters are declared in a scenario YAML. A comparison that changes the policy and the scenario has measured nothing.
  3. Run a seed sweep, not a run. Evaluate the baseline and the candidate on the same set of seeds. Report the distribution.
  4. Test the difference. Welch’s t-test across seeds is the minimum. It is unimpressive and it is what separates a result from an anecdote.
  5. Compare against real baselines. Nearest-robot, least-busy, round-robin, auction and workload-balanced allocators are the heuristics an integrator would actually ship, and waremax includes all of them. Be prepared for the answer: on waremax’s built-in scenarios, learned dispatching matches round-robin and nearest-robot but does not beat them, because those scenarios are bound by station capacity and destination contention, and state-blind round-robin is close to optimal. That is a finding, not a failure, and the simulator’s tunable load, fleet size, traffic capacity and congestion-aware routing exist to find regimes where dispatching has more leverage.
  6. Report the tail, not just the mean. Publish the cycle-time percentiles, and publish the attribution split so a reader can see where the improvement came from.

The RL interface itself is deliberately conventional: WaremaxAllocEnv exposes a Gymnasium environment with a dictionary observation and an action mask, wired so that maskable policy-gradient methods such as MaskablePPO can be applied without a custom wrapper. Masking matters here because the set of legal robot-task pairs changes at every decision point; without it, most of the action space is invalid most of the time and the agent spends its budget learning the constraint instead of the policy.

Limitations

Reward shaping of this kind buys density at the cost of bias. Charging a decision for the travel it selected implicitly assumes the shortest path is the right comparison; under heavy congestion it is not. The attribution split is exact, but the allocation of responsibility across the five buckets is a modelling choice, and different choices produce different policies — which is precisely what the reward-mode comparison measures.

The headline numbers are simulation results on waremax’s built-in scenarios, approximate as reported in the README and accompanying paper, with multi-seed data committed to the repository. They have not been validated in a physical warehouse.

The simulator’s scope is a second limitation. waremax models pod-to-person dispatching. If the throughput bottleneck in a real facility is replenishment, or induct, or the pack line, a dispatching policy tuned to perfection will not move the number that matters. Simulation results transfer only within the scope of what was simulated, and the honest use of a benchmark like this is policy comparison under identical conditions rather than absolute throughput prediction.

Finally, service time is modelled as a lognormal draw. Real pick times correlate with operator, shift hour and item class. Treating them as independent understates the variance a policy will meet in a live building.

Open Questions

How should congestion be priced at all? The delay is real and symmetric between two robots that block each other; charging both decisions doubles the penalty, charging neither leaves congestion unpriced. The controllability result argues for leaving it out of the dispatching reward, but that leaves congestion to the routing layer, and whether a joint dispatch-and-route policy should be paid for it is unresolved.

Does a tail-weighted reward generalise across load regimes? A policy trained at high load learns to protect the tail; the same policy at low load may be needlessly conservative. Whether one policy can span the range, or whether load-conditioned policies are required, is unresolved.

Can the attribution itself be learned? The five buckets are hand-specified. A decomposition induced from trajectories might identify delay sources a human did not name — or might overfit to the simulator’s own idiosyncrasies, which would be worse than useless.

Where does the pilot-versus-simulation gap actually come from? Practitioners routinely report a throughput surprise between a modelling study and the commissioned system. Narrowing that gap is more valuable than another percentage point on a benchmark, and it is not primarily an RL problem.

Conclusion

The environment wrapper is the easy half of an RL benchmark. The hard half is deciding what the agent is paid for, and in warehouse dispatching the default answer — throughput — is terminal, noisy and saturating. A five-bucket decomposition of cycle time gives a dense, causally-aligned alternative at the cost of an explicit modelling choice about responsibility, which is a trade we would rather make in the open than hide inside a reward constant. And the honest headline from waremax is a bounded one: the right reward and representation bring a learned policy level with strong heuristics on these scenarios, not past them.

waremax exists so that trade can be inspected. The determinism guarantee, the scenario YAML and the baseline policies are there to make a comparison reproducible by someone who did not run it. For the wider argument about deterministic simulation and where waremax sits against other tools, see waremax: Deterministic Warehouse-Robotics Simulation, the project page, and our research pillars.