Robotics
Reward Design for Warehouse Dispatching Agents
Why throughput is a poor training signal for RMFS dispatching policies, and how a five-bucket decomposition of cycle time gives per-decision credit instead.
RMFSreinforcement-learningreward-design