Skip to main navigation menu Skip to main content Skip to site footer

Reinforcement Learning Based Adaptive Scheduling Method for Throughput Improvement in Distributed Computing Environments

Abstract

To address the declining scheduling efficiency caused by the continuous growth of task size, increasingly heterogeneous resource types, and dynamic changes in system state in distributed computing environments, this paper proposes a reinforcement learning-based adaptive scheduling method aimed at improving task throughput. This method models the distributed task scheduling process as a continuous decision-making process. By jointly representing task queue state, node resource occupancy, communication load information, and task attribute characteristics, an input representation reflecting the system's operating state is constructed. Based on this, a state refinement mechanism is introduced to enhance the policy learning's ability to perceive complex environmental changes. To improve the rationality of task allocation, a task node compatibility scoring module is further designed. This module comprehensively evaluates candidate allocation relationships from the perspectives of computational adaptability, communication cost, and load balancing requirements, enabling scheduling decisions to more accurately match task needs with node capabilities. Simultaneously, a reward function is constructed around the throughput improvement objective, integrating task completion efficiency, waiting time control, load balancing constraints, and service stability into the optimization process. This guides the scheduling strategy to achieve continuous improvement in overall system performance in the long run. This method achieves coordinated coupling of system state modeling, compatibility matching evaluation, throughput-oriented reward shaping, and policy optimization updates within a unified framework, transforming the scheduling process from a traditional rule-driven approach to an adaptive optimization process oriented towards dynamic environments. Comparative results show that the proposed method achieves superior overall performance in terms of throughput, resource utilization, task waiting control, and quality of service assurance, providing an effective technical solution for intelligent task scheduling in complex distributed computing environments.

pdf