[1]许荣飞,苏志远,麻付强,等.基于强化学习的负载感知CPU资源分配和管理方法[J].计算机技术与发展,2025,(10):81-88.[doi:10.20165/j.cnki.ISSN1673-629X.2025.0142]
 XU Rong-fei,SU Zhi-yuan,MA Fu-qiang,et al.Workload-aware CPU Resource Allocation and Management Based on Reinforcement Learning[J].,2025,(10):81-88.[doi:10.20165/j.cnki.ISSN1673-629X.2025.0142]
点击复制

基于强化学习的负载感知CPU资源分配和管理方法()

《计算机技术与发展》[ISSN:1006-6977/CN:61-1281/TN]

卷:
期数:
2025年10期
页码:
81-88
栏目:
软件技术与工程
出版日期:
2025-10-10

文章信息/Info

Title:
Workload-aware CPU Resource Allocation and Management Based on Reinforcement Learning
文章编号:
1673-629X(2025)10-0081-08
作者:
许荣飞12苏志远12麻付强123吴保锡12
1. 浪潮电子信息产业股份有限公司,山东济南 250101;
2. 浪潮集团有限公司,山东济南 250101;
3. 济南浪潮数据技术有限公司,山东济南 250101
Author(s):
XU Rong-fei12SU Zhi-yuan12MA Fu-qiang123WU Bao-xi12
1. Inspur Electronic Information Industry Co. ,Ltd. ,Jinan 250101,China;
2. Inspur Group Co. ,Ltd. ,Jinan 250101,China;
3. Jinan Inspur Data Technology Co. ,Ltd. ,Jinan 250101,China
关键词:
负载感知强化学习多核系统CPU资源分配管理绑核调频
Keywords:
workload-awarereinforcement learningmulti-core systemCPU resource allocation and managementbind core and frequency adjust
分类号:
TP309
DOI:
10.20165/j.cnki.ISSN1673-629X.2025.0142
摘要:
随着CPU核的数量增多,合理分配CPU核对于降低系统功耗具有重要意义,如何根据系统运行时的负载情况进行精准的CPU资源分配和管理是一个关键的问题。现在处理器设计提供了很多对功耗优化的机制(比如动态电压频率调整DVFS),但是要让这些机制发挥作用,只有芯片的支持是不够的,还需要软硬协同设计。当前缺乏基于软件来最大化利用这些硬件机制的手段。近年来,机器学习在各个领域展现出巨大的潜力,很多基于机器学习的研究工作应运而生。其中,强化学习具有较强的自适应性,适用于动态感知系统环境并进行资源管理。因此,该文提出了一种基于强化学习的负载感知CPU资源分配和管理方法——RLWAM。该方法提出基于最小原则根据系统中运行时的任务负载进行CPU资源分配和管理,基于强化学习提出了面向上述场景的Q-Learning算法,包括面向任务和系统的状态建模方式、面向绑核、调频和资源整合的动作空间和激励函数,从而帮助系统进一步降低功耗。最后,通过在真实平台上从单类型任务上的绑核调频和多类型任务上的资源整合两个场景对该方法进行实验验证,结果表明该方法具有显著的有效性和可扩展性。
Abstract:
With the increase of the number of CPU cores,rational allocation of CPU cores is of great significance to reduce the power con-sumption for system. How to accurately allocate and manager CPU resources according to the workload of a system at runtime is a key problem. Although the current processor design provides a lot of power optimization (such as dynamic voltage frequency adjustment DVFS),it is not enough to make these mechanisms work only relying on the chips. The integrated design of software and hardware is needed. However,currently there is no way to maximally adopt these hardware mechanisms through software. In recent years,the machine learning technique has shown great potential in various fields,and a lot of research work based on machine learning has emerged. Among them,reinforcement learning is suitable for system environment awareness and system resource management because of its strong adaptability. Therefore,a load-aware CPU resource allocation and management method based on reinforcement learning (RLWAM) is proposed. This method proposes to allocate and manage CPU resources based on the minimum principle according to the task workload at runtime in system. Based on reinforcement learning,a Q-Learning algorithm for the above scenarios is proposed. It includes the state modeling for task and system,the action space for binding core,adjusting frequency and integrate resource,to further reduce the power consumption. Finally,the proposed method is verified by using two scenarios on a real platform:binding core and adjusting frequency for single type tasks and integrating resource for multi-type tasks. The experimental results show the effectiveness and scalability of the proposed method.

相似文献/References:

[1]冯林 李琛 孙焘.Robocup半场防守中的一种强化学习算法[J].计算机技术与发展,2008,(01):59.
 FENG Lin,LI Chen,SUN Tao.A Reinforcement Learning Method for Robocup Soccer Half Field Defense[J].,2008,(10):59.
[2]汤萍萍 王红兵.基于强化学习的Web服务组合[J].计算机技术与发展,2008,(03):142.
 TANG Ping-ping,WANG Hong-bing.Web Service Composition Based on Reinforcement -Learning[J].,2008,(10):142.
[3]王朝晖 孙惠萍.图像检索中IRRL模型研究[J].计算机技术与发展,2008,(12):35.
 WANG Zhao-hui,SUN Hui-ping.Research of IRRL Model in Image Retrieval[J].,2008,(10):35.
[4]林联明 王浩 王一雄.基于神经网络的Sarsa强化学习算法[J].计算机技术与发展,2006,(01):30.
 LIN Lian-ming,WANG Hao,WANG Yi-xiong.Sarsa Reinforcement Learning Algorithm Based on Neural Networks[J].,2006,(10):30.
[5]冯鸣夏,伍卫国,邸德海.基于负载感知和QoS的多中心作业调度算法[J].计算机技术与发展,2018,28(12):1.[doi:10.3969/j.issn.1673-629X.2018.12.001]
 FENG Mingxia,WU Weiguo,DI Dehai.A Job Scheduling Algorithm in Multi-computing Centers Based on Load-aware and QoS[J].,2018,28(10):1.[doi:10.3969/j.issn.1673-629X.2018.12.001]
[6]农汉琦,孙蕴琪,黄 洁,等.基于机器学习的认知无线网络优化策略[J].计算机技术与发展,2020,30(05):125.[doi:10. 3969 / j. issn. 1673-629X. 2020. 05. 024]
 NONG Han-qi,SUN Yun-qi,HUANG Jie,et al.Optimization Strategy of Cognitive Radio Network Based on Machine Learning[J].,2020,30(10):125.[doi:10. 3969 / j. issn. 1673-629X. 2020. 05. 024]
[7]雷 莹,许道云.一种合作 Markov 决策系统[J].计算机技术与发展,2020,30(12):8.[doi:10. 3969 / j. issn. 1673-629X. 2020. 12. 002]
 LEI Ying,XU Dao-yun.A Cooperation Markov Decision Process System[J].,2020,30(10):8.[doi:10. 3969 / j. issn. 1673-629X. 2020. 12. 002]
[8]彭云建,梁 进.基于探索-利用权衡优化的 Q 学习路径规划[J].计算机技术与发展,2022,32(04):1.[doi:10. 3969 / j. issn. 1673-629X. 2022. 04. 001]
 PENG Yun-jian,LIANG Jin.Q-learning Path Planning Based on Exploration / Exploitation Tradeoff Optimization[J].,2022,32(10):1.[doi:10. 3969 / j. issn. 1673-629X. 2022. 04. 001]
[9]乔 通,周 洲,程 鑫,等.基于 Q-学习的底盘测功机自适应 PID 控制模型[J].计算机技术与发展,2022,32(05):117.[doi:10. 3969 / j. issn. 1673-629X. 2022. 05. 020]
 QIAO Tong,ZHOU Zhou,CHENG Xin,et al.Adaptive PID Control Model of Chassis Dynamometer Based on Q-Learning[J].,2022,32(10):117.[doi:10. 3969 / j. issn. 1673-629X. 2022. 05. 020]
[10]魏竞毅,赖 俊,陈希亮.基于互信息的智能博弈对抗分层强化学习研究[J].计算机技术与发展,2022,32(09):142.[doi:10. 3969 / j. issn. 1673-629X. 2022. 09. 022]
 WEI Jing-yi,LAI Jun,CHEN Xi-liang.Research on Hierarchical Reinforcement Learning of Intelligent Game Confrontation Based on Mutual Information[J].,2022,32(10):142.[doi:10. 3969 / j. issn. 1673-629X. 2022. 09. 022]

更新日期/Last Update: 2025-10-10