[Home ] [Archive]   [ فارسی ]  
:: Main :: About :: Current Issue :: Archive :: Search :: Submit :: Contact ::
Main Menu
Home::
Journal Information::
Articles archive::
For Authors::
For Reviewers::
Registration::
Contact us::
Site Facilities::
::
Social Network Membership
Linkedin
Researchgate
..
Indexing Databases
..
DOI
کلیک کنید
..
ِDOR
..
Search in website

Advanced Search
..
Receive site information
Enter your Email in the following box to receive the site news and information.
..
:: Volume 15, Issue 1 (4-2026) ::
ieijqp 2026, 15(1): 73-89 Back to browse issues page
Dynamic Order Allocation Optimization in Multi-Supplier Supply Chains under Cost Uncertainty: A Comparative Study of Dynamic Programming and Reinforcement Learning
Leila Hosseini1 , Mohammad saber Fallah nezahd *1 , Vali Derhami2
1- Department of Industrial Engineering, Faculty of Engineering, Yazd University, Yazd, Iran
2- Department of Computer Engineering, Faculty of Engineering, Yazd University, Yazd, Iran
Abstract:   (17 Views)
Order management in multi-source supply chains under stochastic cost fluctuations poses significant structural challenges. This study defines the problem as a strategic decision-making process under uncertainty, wherein historical cost data are available, yet due to structural market volatility and continuous price variations, the distribution derived from past data cannot accurately represent the future, and the decision-maker lacks access to a stable and deterministic distribution of future costs at the time of choice. In the dynamic programming approach, transition probability matrices are extracted from historical data and serve as an optimal model-based benchmark. In contrast, the Q-learning reinforcement learning approach learns the decision policy without explicit probabilistic distribution estimation, solely through interaction with a simulated environment built upon real data. Model validation in the polymer industry using real data across 47 planning horizons demonstrates that the proposed framework achieves an average cost saving of 0.23% (p-value < 0.001) even under baseline conditions with close supplier costs, equivalent to a reduction of over 73,000 monetary units per horizon. Furthermore, sensitivity analyses reveal that as the cost gap between suppliers widens, the saving potential increases substantially, reaching up to 40%. Although dynamic programming provides the optimal theoretical performance, the Q-learning method exhibits rapid adaptability and high stability against market shocks without requiring future distribution modeling. By bridging the gap between theoretical optimization and data-driven methods, this research provides a practical and robust foundation for implementing decision support systems in supply chains with incomplete information and unstable stochastic costs.
Keywords: Supply Chain, Dynamic Programming, Reinforcement Learning, Q-Learning, Markov Decision Process, Ordering Management
     
Type of Study: Research |
Received: 2026/04/30 | Accepted: 2026/07/22 | Published: 2026/07/22


XML   Persian Abstract   Print


Download citation:
BibTeX | RIS | EndNote | Medlars | ProCite | Reference Manager | RefWorks
Send citation to:

Hosseini L, Fallah nezahd M S, Derhami V. Dynamic Order Allocation Optimization in Multi-Supplier Supply Chains under Cost Uncertainty: A Comparative Study of Dynamic Programming and Reinforcement Learning. ieijqp 2026; 15 (1) :73-89
URL: http://ieijqp.ir/article-1-1060-en.html


Rights and permissions
Creative Commons License This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Volume 15, Issue 1 (4-2026) Back to browse issues page
نشریه علمی- پژوهشی کیفیت و بهره وری صنعت برق ایران Iranian Electric Industry Journal of Quality and Productivity
Persian site map - English site map - Created in 0.11 seconds with 40 queries by YEKTAWEB 4766