Autors: Gergov, I. G., Tsochev, G. R., Aleksieva-Petrova, A. P. Title: Multi-Agent Reinforcement Learning for Decentralized Computational Resource Negotiation in a Tokenized Double-Auction Market: A Reproducibility-First Benchmark Keywords: computational resource allocation, double auction, inequality, multi-agent reinforcement learning, reproducibility, social welfare, tokenized marketAbstract: Decentralized compute markets require autonomous agents to negotiate heterogeneous resources under budget constraints, stochastic supply, and strategic interaction. We present Agora-RL, a reproducibility-first benchmark for repeated negotiation of GPU, memory, and bandwidth through token-denominated double auctions. The study asks two questions: how standard MARL baselines rank when reward, social welfare, and inequality are evaluated jointly; and whether a transparent benchmark protocol can make such comparisons auditable. PPO, MAPPO, MADDPG, and IQL are evaluated with matched 300-episode training budgets, 30 deterministic evaluation episodes, and 12 random seeds. Using percentile-bootstrap 95% confidence intervals, MAPPO achieves the highest reward (0.0140 [0.0124, 0.0154]) and social welfare (0.0952 [0.0854, 0.1045]), whereas IQL yields the lowest Gini coefficient (0.4477 [0.4360, 0.4613]). Secondary diagnostics show that reward leadership does not imply fairness, equilibrium closeness, or communication robustness. The contribution is an empirical benchmark and audit protocol rather than a new auction theorem or blockchain settlement layer. References - Schulman J. Wolski F. Dhariwal P. Radford A. Klimov O. Proximal Policy Optimization Algorithms arXiv 2017 10.48550/arXiv.1707.06347 1707.06347
- Yu C. Velu A. Vinitsky E. Gao J. Wang Y. Bayen A. Wu Y. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games Adv. Neural Inf. Process. Syst. 2022 35 24611 24624
- Lowe R. Wu Y. Tamar A. Harb J. Abbeel P. Mordatch I. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments Adv. Neural Inf. Process. Syst. 2017 30 6379 6390
- Tan M. Multi-Agent Reinforcement Learning: Independent versus Cooperative Agents Proceedings of the Tenth International Conference on Machine Learning, Amherst, MA, USA, 27–29 June 1993 Morgan Kaufmann San Francisco, CA, USA 1993 330 337
- Bettini M. Prorok A. Moens V. BenchMARL: Benchmarking Multi-Agent Reinforcement Learning J. Mach. Learn. Res. 2024 25 10557 10566
- Nardini M. Helmer S. El Ioini N. Pahl C. A Blockchain-Based Decentralized Electronic Marketplace for Computing Resources SN Comput. Sci. 2020 1 251 10.1007/s42979-020-00243-7
- Li Q. Jia X. Huang C. A Truthful Dynamic Combinatorial Double Auction Model for Cloud Resource Allocation J. Cloud Comput. 2023 12 106 10.1186/s13677-023-00479-7
- Weerasinghe N. Porambage P. Braeken A. Liyanage M. Ylianttila M. TokenNet: A Novel Tokenized Resource Marketplace for 6G Network Slicing IEEE Trans. Netw. Sci. Eng. 2025 12 4697 4713 10.1109/TNSE.2025.3574634
- Vickrey W. Counterspeculation, Auctions, and Competitive Sealed Tenders J. Financ. 1961 16 8 37 10.1111/j.1540-6261.1961.tb02789.x
- Clarke E.H. Multipart Pricing of Public Goods Public Choice 1971 11 17 33 10.1007/BF01726210
- Groves T. Incentives in Teams Econometrica 1973 41 617 631 10.2307/1914085
- Hu S. Zhong Y. Gao M. Wang W. Dong H. Liang X. Li Z. Chang X. Yang Y. MARLlib: A Scalable and Efficient Multi-agent Reinforcement Learning Library J. Mach. Learn. Res. 2023 24 14911 14933
- Tafsiri S.A. Yousefi S. Combinatorial Double Auction-Based Resource Allocation Mechanism in Cloud Computing Market J. Syst. Softw. 2018 137 322 334 10.1016/j.jss.2017.11.044
- Kumar D. Baranwal G. Raza Z. Vidyarthi D.P. A Truthful Combinatorial Double Auction-Based Marketplace Mechanism for Cloud Computing J. Syst. Softw. 2018 140 91 108 10.1016/j.jss.2018.03.003
- Singhal R. Singhal A. A Feedback-Based Combinatorial Fair Economical Double Auction Resource Allocation Model for Cloud Computing Future Gener. Comput. Syst. 2021 115 780 797 10.1016/j.future.2020.09.022
- Ma X. Xu D. Wolter K. Blockchain-Enabled Feedback-Based Combinatorial Double Auction for Cloud Markets Future Gener. Comput. Syst. 2022 127 225 239 10.1016/j.future.2021.09.009
- Esposito A. D’Angelo S. Casuccio D. Semantic Technologies Applied to Blockchain for Cloud Computing Services Marketplace Implementation Int. J. Semant. Web Inf. Syst. 2025 21 1 37 10.4018/IJSWIS.387386
- d’Eon G. Newman N. Leyton-Brown K. Understanding Iterative Combinatorial Auction Designs via Multi-Agent Reinforcement Learning Proceedings of the 25th ACM Conference on Economics and Computation (EC ‘24) New Haven, CT, USA 8–11 July 2024 1102 1130
- Tang X. Yu H. Competitive-Cooperative Multi-Agent Reinforcement Learning for Auction-Based Federated Learning Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence Macao, China 19–25 August 2023 4262 4270
- Wilson R. Incentive Efficiency of Double Auctions Econometrica 1985 53 1101 1115 10.2307/1911013
- Glosten L.R. Milgrom P.R. Bid, Ask and Transaction Prices in a Specialist Market with Heterogeneously Informed Traders J. Financ. Econ. 1985 14 71 100 10.1016/0304-405X(85)90044-3
- Samvelyan M. Rashid T. de Witt C.S. Farquhar G. Nardelli N. Rudner T.G.J. Hung C.-M. Torr P.H.S. Foerster J. Whiteson S. The StarCraft Multi-Agent Challenge Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems Montreal, QC, Canada 13–17 May 2019 2186 2188
- Kurach K. Raichuk A. Stańczyk P. Zajac M. Bachem O. Espeholt L. Riquelme C. Vincent D. Michalski M. Bousquet O. et al. Google Research Football: A Novel Reinforcement Learning Environment Proceedings of the AAAI Conference on Artificial Intelligence New York, NY, USA 7–12 February 2020 Volume 34 4501 4510
- Zheng S. Trott A. Srinivasa S. Naik N. Gruesbeck M. Parkes D.C. Socher R. The AI Economist: Taxation Policy Design via Two-Level Deep Multiagent Reinforcement Learning Sci. Adv. 2022 8 eabk2607 10.1126/sciadv.abk2607
- Terry J.K. Black B. Grammel N. Jayakumar M. Hari A. Sullivan R. Santos L. Perez R. Horsch C. Dieffendahl C. et al. PettingZoo: Gym for Multi-Agent Reinforcement Learning Adv. Neural Inf. Process. Syst. 2021 34 15032 15043
- Cheng L. Li M. Tan C. Huang P. Zhang M. Sun R. Computational Game-Theoretic Models for Adaptive Urban Energy Systems: A Comprehensive Review of Algorithms, Strategies, and Engineering Applications Arch. Comput. Methods Eng. 2026 33 2037 2114 10.1007/s11831-025-10364-y
Issue
| Mathematics, vol. 14, 2026, Albania, https://doi.org/10.3390/math14111828 |
|