|
Dynamic Order Allocation Optimization in Multi-Supplier Supply Chains under Cost Uncertainty: A Comparative Study of Dynamic Programming and Reinforcement Learning
|
Leila Hosseini1 , Mohammad saber Fallah nezahd *1 , Vali Derhami2  |
1- Department of Industrial Engineering, Faculty of Engineering, Yazd University, Yazd, Iran 2- Department of Computer Engineering, Faculty of Engineering, Yazd University, Yazd, Iran |
|
|
Abstract: (17 Views) |
Order management in multi-source supply chains under stochastic cost fluctuations poses significant structural challenges. This study defines the problem as a strategic decision-making process under uncertainty, wherein historical cost data are available, yet due to structural market volatility and continuous price variations, the distribution derived from past data cannot accurately represent the future, and the decision-maker lacks access to a stable and deterministic distribution of future costs at the time of choice. In the dynamic programming approach, transition probability matrices are extracted from historical data and serve as an optimal model-based benchmark. In contrast, the Q-learning reinforcement learning approach learns the decision policy without explicit probabilistic distribution estimation, solely through interaction with a simulated environment built upon real data. Model validation in the polymer industry using real data across 47 planning horizons demonstrates that the proposed framework achieves an average cost saving of 0.23% (p-value < 0.001) even under baseline conditions with close supplier costs, equivalent to a reduction of over 73,000 monetary units per horizon. Furthermore, sensitivity analyses reveal that as the cost gap between suppliers widens, the saving potential increases substantially, reaching up to 40%. Although dynamic programming provides the optimal theoretical performance, the Q-learning method exhibits rapid adaptability and high stability against market shocks without requiring future distribution modeling. By bridging the gap between theoretical optimization and data-driven methods, this research provides a practical and robust foundation for implementing decision support systems in supply chains with incomplete information and unstable stochastic costs. |
|
| Keywords: Supply Chain, Dynamic Programming, Reinforcement Learning, Q-Learning, Markov Decision Process, Ordering Management |
|
|
|
Type of Study: Research |
Received: 2026/04/30 | Accepted: 2026/07/22 | Published: 2026/07/22
|
|
|
|
|
|