Autors: Gergov, I. G., Tsochev, G. R., Aleksieva-Petrova, A. P.
Title: Multi-Agent Reinforcement Learning for Decentralized Computational Resource Negotiation in a Tokenized Double-Auction Market: A Reproducibility-First Benchmark
Keywords: computational resource allocation, double auction, inequality, multi-agent reinforcement learning, reproducibility, social welfare, tokenized market

Abstract: Decentralized compute markets require autonomous agents to negotiate heterogeneous resources under budget constraints, stochastic supply, and strategic interaction. We present Agora-RL, a reproducibility-first benchmark for repeated negotiation of GPU, memory, and bandwidth through token-denominated double auctions. The study asks two questions: how standard MARL baselines rank when reward, social welfare, and inequality are evaluated jointly; and whether a transparent benchmark protocol can make such comparisons auditable. PPO, MAPPO, MADDPG, and IQL are evaluated with matched 300-episode training budgets, 30 deterministic evaluation episodes, and 12 random seeds. Using percentile-bootstrap 95% confidence intervals, MAPPO achieves the highest reward (0.0140 [0.0124, 0.0154]) and social welfare (0.0952 [0.0854, 0.1045]), whereas IQL yields the lowest Gini coefficient (0.4477 [0.4360, 0.4613]). Secondary diagnostics show that reward leadership does not imply fairness, equilibrium closeness, or communication robustness. The contribution is an empirical benchmark and audit protocol rather than a new auction theorem or blockchain settlement layer.

References

  1. Schulman J. Wolski F. Dhariwal P. Radford A. Klimov O. Proximal Policy Optimization Algorithms arXiv 2017 10.48550/arXiv.1707.06347 1707.06347
  2. Yu C. Velu A. Vinitsky E. Gao J. Wang Y. Bayen A. Wu Y. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games Adv. Neural Inf. Process. Syst. 2022 35 24611 24624
  3. Lowe R. Wu Y. Tamar A. Harb J. Abbeel P. Mordatch I. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments Adv. Neural Inf. Process. Syst. 2017 30 6379 6390
  4. Tan M. Multi-Agent Reinforcement Learning: Independent versus Cooperative Agents Proceedings of the Tenth International Conference on Machine Learning, Amherst, MA, USA, 27–29 June 1993 Morgan Kaufmann San Francisco, CA, USA 1993 330 337
  5. Bettini M. Prorok A. Moens V. BenchMARL: Benchmarking Multi-Agent Reinforcement Learning J. Mach. Learn. Res. 2024 25 10557 10566
  6. Nardini M. Helmer S. El Ioini N. Pahl C. A Blockchain-Based Decentralized Electronic Marketplace for Computing Resources SN Comput. Sci. 2020 1 251 10.1007/s42979-020-00243-7
  7. Li Q. Jia X. Huang C. A Truthful Dynamic Combinatorial Double Auction Model for Cloud Resource Allocation J. Cloud Comput. 2023 12 106 10.1186/s13677-023-00479-7
  8. Weerasinghe N. Porambage P. Braeken A. Liyanage M. Ylianttila M. TokenNet: A Novel Tokenized Resource Marketplace for 6G Network Slicing IEEE Trans. Netw. Sci. Eng. 2025 12 4697 4713 10.1109/TNSE.2025.3574634
  9. Vickrey W. Counterspeculation, Auctions, and Competitive Sealed Tenders J. Financ. 1961 16 8 37 10.1111/j.1540-6261.1961.tb02789.x
  10. Clarke E.H. Multipart Pricing of Public Goods Public Choice 1971 11 17 33 10.1007/BF01726210
  11. Groves T. Incentives in Teams Econometrica 1973 41 617 631 10.2307/1914085
  12. Hu S. Zhong Y. Gao M. Wang W. Dong H. Liang X. Li Z. Chang X. Yang Y. MARLlib: A Scalable and Efficient Multi-agent Reinforcement Learning Library J. Mach. Learn. Res. 2023 24 14911 14933
  13. Tafsiri S.A. Yousefi S. Combinatorial Double Auction-Based Resource Allocation Mechanism in Cloud Computing Market J. Syst. Softw. 2018 137 322 334 10.1016/j.jss.2017.11.044
  14. Kumar D. Baranwal G. Raza Z. Vidyarthi D.P. A Truthful Combinatorial Double Auction-Based Marketplace Mechanism for Cloud Computing J. Syst. Softw. 2018 140 91 108 10.1016/j.jss.2018.03.003
  15. Singhal R. Singhal A. A Feedback-Based Combinatorial Fair Economical Double Auction Resource Allocation Model for Cloud Computing Future Gener. Comput. Syst. 2021 115 780 797 10.1016/j.future.2020.09.022
  16. Ma X. Xu D. Wolter K. Blockchain-Enabled Feedback-Based Combinatorial Double Auction for Cloud Markets Future Gener. Comput. Syst. 2022 127 225 239 10.1016/j.future.2021.09.009
  17. Esposito A. D’Angelo S. Casuccio D. Semantic Technologies Applied to Blockchain for Cloud Computing Services Marketplace Implementation Int. J. Semant. Web Inf. Syst. 2025 21 1 37 10.4018/IJSWIS.387386
  18. d’Eon G. Newman N. Leyton-Brown K. Understanding Iterative Combinatorial Auction Designs via Multi-Agent Reinforcement Learning Proceedings of the 25th ACM Conference on Economics and Computation (EC ‘24) New Haven, CT, USA 8–11 July 2024 1102 1130
  19. Tang X. Yu H. Competitive-Cooperative Multi-Agent Reinforcement Learning for Auction-Based Federated Learning Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence Macao, China 19–25 August 2023 4262 4270
  20. Wilson R. Incentive Efficiency of Double Auctions Econometrica 1985 53 1101 1115 10.2307/1911013
  21. Glosten L.R. Milgrom P.R. Bid, Ask and Transaction Prices in a Specialist Market with Heterogeneously Informed Traders J. Financ. Econ. 1985 14 71 100 10.1016/0304-405X(85)90044-3
  22. Samvelyan M. Rashid T. de Witt C.S. Farquhar G. Nardelli N. Rudner T.G.J. Hung C.-M. Torr P.H.S. Foerster J. Whiteson S. The StarCraft Multi-Agent Challenge Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems Montreal, QC, Canada 13–17 May 2019 2186 2188
  23. Kurach K. Raichuk A. Stańczyk P. Zajac M. Bachem O. Espeholt L. Riquelme C. Vincent D. Michalski M. Bousquet O. et al. Google Research Football: A Novel Reinforcement Learning Environment Proceedings of the AAAI Conference on Artificial Intelligence New York, NY, USA 7–12 February 2020 Volume 34 4501 4510
  24. Zheng S. Trott A. Srinivasa S. Naik N. Gruesbeck M. Parkes D.C. Socher R. The AI Economist: Taxation Policy Design via Two-Level Deep Multiagent Reinforcement Learning Sci. Adv. 2022 8 eabk2607 10.1126/sciadv.abk2607
  25. Terry J.K. Black B. Grammel N. Jayakumar M. Hari A. Sullivan R. Santos L. Perez R. Horsch C. Dieffendahl C. et al. PettingZoo: Gym for Multi-Agent Reinforcement Learning Adv. Neural Inf. Process. Syst. 2021 34 15032 15043
  26. Cheng L. Li M. Tan C. Huang P. Zhang M. Sun R. Computational Game-Theoretic Models for Adaptive Urban Energy Systems: A Comprehensive Review of Algorithms, Strategies, and Engineering Applications Arch. Comput. Methods Eng. 2026 33 2037 2114 10.1007/s11831-025-10364-y

Issue

Mathematics, vol. 14, 2026, Albania, https://doi.org/10.3390/math14111828

Вид: статия в списание, публикация в издание с импакт фактор, публикация в реферирано издание, индексирана в Scopus