[Home ] [Archive]   [ فارسی ]  
:: Main :: About :: Current Issue :: Archive :: Search :: Submit :: Contact ::
Main Menu
Home::
Journal Information::
Articles archive::
For Authors::
For Reviewers::
Registration::
Contact us::
Site Facilities::
::
Social Network Membership
Linkedin
Researchgate
..
Indexing Databases
..
DOI
کلیک کنید
..
ِDOR
..
Search in website

Advanced Search
..
Receive site information
Enter your Email in the following box to receive the site news and information.
..
:: Volume 15, Issue 1 (4-2026) ::
ieijqp 2026, 15(1): 73-87 Back to browse issues page
Dynamic Order Allocation Optimization in Multi-Supplier Supply Chains under Cost Uncertainty: A Comparative Study of Dynamic Programming and Reinforcement Learning
Leila Hosseini1 , Mohammad saber Fallah nezahd *1 , Vali Derhami2
1- Department of Industrial Engineering, Faculty of Engineering, Yazd University, Yazd, Iran
2- Department of Computer Engineering, Faculty of Engineering, Yazd University, Yazd, Iran
Abstract:   (274 Views)
Order management in multi-source supply chains under stochastic cost fluctuations poses significant structural challenges. This study defines the problem as a strategic decision-making process under uncertainty, wherein historical cost data are available, yet due to structural market volatility and continuous price variations, the distribution derived from past data cannot accurately represent the future, and the decision-maker lacks access to a stable and deterministic distribution of future costs at the time of choice. In the dynamic programming approach, transition probability matrices are extracted from historical data and serve as an optimal model-based benchmark. In contrast, the Q-learning reinforcement learning approach learns the decision policy without explicit probabilistic distribution estimation, solely through interaction with a simulated environment built upon real data. Model validation in the polymer industry using real data across 47 planning horizons demonstrates that the proposed framework achieves an average cost saving of 0.23% (p-value < 0.001) even under baseline conditions with close supplier costs, equivalent to a reduction of over 73,000 monetary units per horizon. Furthermore, sensitivity analyses reveal that as the cost gap between suppliers widens, the saving potential increases substantially, reaching up to 40%. Although dynamic programming provides the optimal theoretical performance, the Q-learning method exhibits rapid adaptability and high stability against market shocks without requiring future distribution modeling. By bridging the gap between theoretical optimization and data-driven methods, this research provides a practical and robust foundation for implementing decision support systems in supply chains with incomplete information and unstable stochastic costs.
Keywords: Supply Chain, Dynamic Programming, Reinforcement Learning, Q-Learning, Markov Decision Process, Ordering Management
Full-Text [PDF 1299 kb]   (51 Downloads)    
Type of Study: Research |
Received: 2026/04/30 | Accepted: 2026/07/22 | Published: 2026/07/22
References
1. Azaron, A., Tang, O., & Tavakkoli-Moghaddam, R. (2009). Dynamic lot sizing problem with continuous-time Markovian production cost. International Journal of Production Economics, 120(2), 607-612. [DOI:10.1016/j.ijpe.2009.04.007]
2. Barnes-Schuster, D., Bassok, Y., & Anupindi, R. (2002). Coordination and flexibility in supply contracts with options. Manufacturing & Service Operations Management, 4(3), 171-207. [DOI:10.1287/msom.4.3.171.7754]
3. Bellman, R. (1957). A Markovian decision process. Journal of Mathematics and Mechanics, 679-684. [DOI:10.1512/iumj.1957.6.56038]
4. Bilsel, R. U., & Kumara, S. R. (2007). A reinforcement learning approach for dynamic supplier selection. In 2007 IEEE International Conference on Service Operations and Logistics, and Informatics (pp. 1-6). IEEE. [DOI:10.1109/SOLI.2007.4383959]
5. Chauhan, V. K., Mak, S., Parlikad, A. K., Alomari, M., Casassa, L., & Brintrup, A. (2023). Real-time large-scale supplier order assignments across two-tiers of a supply chain with penalty and dual-sourcing. Computers & Industrial Engineering, 176, 108928. [DOI:10.1016/j.cie.2022.108928]
6. Chen, Z., & Rossi, R. (2021). A dynamic ordering policy for a stochastic inventory problem with cash constraints. Omega, 102, 102378. [DOI:10.1016/j.omega.2020.102378]
7. Cheraghalipour, A., & Farsad, S. (2018). A bi-objective sustainable supplier selection and order allocation considering quantity discounts under disruption risks: A case study in plastic industry. Computers & Industrial Engineering, 118, 237-250. [DOI:10.1016/j.cie.2018.02.041]
8. Dabbagh, R., Haghshnas, A., & Joodat Nia, S. (2025). Identification and prioritization of power supply chain risks using an analytical approach with FMEA and ARAS (Case study: West Azerbaijan Electricity Distribution Company) (In Persian). Iranian Electric Industry Journal of Quality and Productivity, 14(1), 19-29.
9. Darezereshki, F., & Derhami, V. (2024). A reinforcement learning approach to determine when and how many stocks to buy in stock trading. [Unpublished manuscript].
10. De Farias, D. P., & Van Roy, B. (2003). The linear programming approach to approximate dynamic programming. Operations Research, 51(6), 850-865. [DOI:10.1287/opre.51.6.850.24925]
11. Dezfoulian, H., Sheikhi, M., & Samoui, P. (2025). Identification and ranking of power supply chain risks using crowdsourcing and ARAS multi-criteria decision-making method (In Persian). Iranian Electric Industry Journal of Quality and Productivity, 14(1), 38.
12. Geevers, K., Van Hezewijk, L., & Mes, M. R. (2024). Multi-echelon inventory optimization using deep reinforcement learning. Central European Journal of Operations Research, 32(3), 653-683. [DOI:10.1007/s10100-023-00872-2]
13. Hosseini, Z. S., Flapper, S. D., & Pirayesh, M. (2022). Sustainable supplier selection and order allocation under demand, supplier availability and supplier grading uncertainties. Computers & Industrial Engineering, 165, 107811. [DOI:10.1016/j.cie.2021.107811]
14. Kawakami, K., Kobayashi, H., & Nakata, K. (2021). Seasonal inventory management model for raw materials in steel industry. INFORMS Journal on Applied Analytics, 51(4), 312-324. [DOI:10.1287/inte.2021.1073]
15. Kharvi, S., & Pakkala, T. (2022). An optimal ordering inventory policy when the purchase price follows a discrete-time Markov chain. Communications in Statistics-Simulation and Computation, 1-18. [DOI:10.1080/03610918.2022.2158340]
16. Khodadadi, H., & Derhami, V. (2025). Employing chaos theory for exploration-exploitation balance in deep reinforcement learning. Tabriz Journal of Electrical Engineering, 55(1), 113-121.
17. Leys, C., Ley, C., Klein, O., Bernard, P., & Licata, L. (2013). Detecting outliers: Do not use standard deviation around the mean, use absolute deviation around the median. Journal of Experimental Social Psychology, 49(4), 764-766. [DOI:10.1016/j.jesp.2013.03.013]
18. Li, Q., Wu, X., & Cheung, K. L. (2009). Optimal policies for inventory systems with separate delivery-request and order-quantity decisions. Operations Research, 57(3), 626-636. [DOI:10.1287/opre.1090.0696]
19. Luo, S., Ahiska, S. S., Fang, S.-C., King, R. E., Warsing Jr., D. P., & Wu, S. (2021). An analysis of optimal ordering policies for a two-supplier system with disruption risk. Omega, 105, 102517. [DOI:10.1016/j.omega.2021.102517]
20. Maulidi, I., Radhiah, R., Hayati, C., & Apriliani, V. (2022). Optimal raw material inventory analysis using Markov decision process with policy iteration method. JTAM (Jurnal Teori dan Aplikasi Matematika), 6(3), 638-650. [DOI:10.31764/jtam.v6i3.8563]
21. Milani, H. (2024). A Markov decision processes model for stochastic inventory problem under uncertainty in supply. Available at SSRN 4833769. [DOI:10.2139/ssrn.4833769]
22. Mohamed, I. B., Klibi, W., Sadykov, R., Şen, H., & Vanderbeck, F. (2023). The two-echelon stochastic multi-period capacitated location-routing problem. European Journal of Operational Research, 306(2), 645-667. [DOI:10.1016/j.ejor.2022.07.022]
23. Nadi, F., Derhami, V., & Alamiyan Harandi, F. (2024). Value iteration based fuzzy reinforcement learning in target following robot (In Persian). Journal of Control, 18(2), 1-12.
24. Nayeri, S., Khoei, M. A., Rouhani-Tazangi, M. R., GhanavatiNejad, M., Rahmani, M., & Tirkolaee, E. B. (2023). A data-driven model for sustainable and resilient supplier selection and order allocation problem in a responsive supply chain: A case study of healthcare system. Engineering Applications of Artificial Intelligence, 124, 106511. [DOI:10.1016/j.engappai.2023.106511]
25. Piao, M., Zhang, D., Lu, H., & Li, R. (2023). A supply chain inventory management method for civil aircraft manufacturing based on multi-agent reinforcement learning. Applied Sciences, 13(13), 7510. [DOI:10.3390/app13137510]
26. Pontrandolfo, P., Gosavi, A., Okogbaa, O. G., & Das, T. K. (2002). Global supply chain management: A reinforcement learning approach. International Journal of Production Research, 40, 1299-1317. [DOI:10.1080/00207540110118640]
27. Powell, W. B., Shapiro, J. A., & Simão, H. P. (2002). An adaptive dynamic programming algorithm for the heterogeneous resource allocation problem. Transportation Science, 36(2), 231-249. [DOI:10.1287/trsc.36.2.231.561]
28. Ryzhov, I., Powell, W., & Frazier, P. (2008). The knowledge-gradient algorithm for online learning. [Submitted for publication].
29. Savvopoulou, A. (2024). Inventory management: Application at a packaging materials manufacturer. [Master's thesis].
30. Scarf, H. (1960). The optimality of (S, s) policies in the dynamic inventory problem. In Mathematical Methods in the Social Sciences.
31. Sulistyoningarum, R., Rosyidi, C. N., & Rochman, T. (2020). Supplier selection and order allocation of recycled plastic materials: A case study in a plastic manufacturing company. International Journal of Information and Management Sciences, 31(4), 315-330.
32. Sutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction. MIT Press. [DOI:10.1109/TNN.1998.712192]
33. Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.
34. Thompson, S., Nunez, M., Garfinkel, R., & Dean, M. D. (2009). OR practice-efficient short-term allocation and reallocation of patients to floors of a hospital during demand surges. Operations Research, 57(2), 261-273. [DOI:10.1287/opre.1080.0584]
35. Van Roy, B., Bertsekas, D. P., Lee, Y., & Tsitsiklis, J. N. (1997). A neuro-dynamic programming approach to retailer inventory management. In Proceedings of the 36th IEEE Conference on Decision and Control (Vol. 4, pp. 4052-4057). IEEE. [DOI:10.1109/CDC.1997.652501]
36. Wagelmans, A., Van Hoesel, S., & Kolen, A. (1992). Economic lot sizing: An O(n log n) algorithm that runs in linear time in the Wagner-Whitin case. Operations Research, 40(1), S145-S156. [DOI:10.1287/opre.40.1.S145]
37. Wagner, H. M., & Whitin, T. M. (1958). Dynamic version of the economic lot size model. Management Science, 5(1), 89-96. [DOI:10.1287/mnsc.5.1.89]
38. Wagner, H. M., & Whitin, T. M. (2004). Dynamic version of the economic lot size model. Management Science, 50(12), 1770-1774. [DOI:10.1287/mnsc.1040.0262]
39. Watkins, C. J., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3), 279-292. [DOI:10.1007/BF00992698]
40. Wu, J., Su, L., Wang, G., & Yang, Y. (2024). Approximated dynamic programming for production and inventory planning problem in cold rolling process of steel production. Mathematics, 12(24), 3922. [DOI:10.3390/math12243922]
41. Wulan, Q. (2023). Order scheduling optimization in manufacturing enterprises based on MDP and dynamic programming. Scientific Reports, 13(1), 9783. [DOI:10.1038/s41598-023-36976-7]
42. Zhang, H., Li, N., & Lin, J. (2024). Modeling the decision and coordination mechanism of power battery closed-loop supply chain using Markov decision processes. Sustainability, 16(11), 4329. [DOI:10.3390/su16114329]
43. Zhao, G., & Sun, R. (2010). Application of multi-agent reinforcement learning to supply chain ordering management. In 2010 Sixth International Conference on Natural Computation (Vol. 7, pp. 3830-3834). IEEE. [DOI:10.1109/ICNC.2010.5582551]
44. Zhou, Y., Guo, K., Yu, C., & Zhang, Z. (2024). Optimization of multi-echelon spare parts inventory systems using multi-agent deep reinforcement learning. Applied Mathematical Modelling, 125, 827-844. [DOI:10.1016/j.apm.2023.10.039]


XML   Persian Abstract   Print


Download citation:
BibTeX | RIS | EndNote | Medlars | ProCite | Reference Manager | RefWorks
Send citation to:

Hosseini L, Fallah nezahd M S, Derhami V. Dynamic Order Allocation Optimization in Multi-Supplier Supply Chains under Cost Uncertainty: A Comparative Study of Dynamic Programming and Reinforcement Learning. ieijqp 2026; 15 (1) :73-87
URL: http://ieijqp.ir/article-1-1060-en.html


Rights and permissions
Creative Commons License This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Volume 15, Issue 1 (4-2026) Back to browse issues page
نشریه علمی- پژوهشی کیفیت و بهره وری صنعت برق ایران Iranian Electric Industry Journal of Quality and Productivity
Persian site map - English site map - Created in 0.14 seconds with 40 queries by YEKTAWEB 4766