Optimal voltage control for distribution networks considering dynamic correction by policy safety layer
DOI:10.19783/j.cnki.pspc.260188
Key Words:distribution network  voltage optimization control  deep reinforcement learning  model predictive control  policy safety layer
Author NameAffiliation
LI Xiaolu 1. Faculty of Electrical Engineering, Shanghai University of Electric Power, Shanghai 200090, China
2. State Grid Shanghai Electric Power Research Institute, Shanghai 200437, China 
DU Benjie 1. Faculty of Electrical Engineering, Shanghai University of Electric Power, Shanghai 200090, China
2. State Grid Shanghai Electric Power Research Institute, Shanghai 200437, China 
SU Haotian 1. Faculty of Electrical Engineering, Shanghai University of Electric Power, Shanghai 200090, China
2. State Grid Shanghai Electric Power Research Institute, Shanghai 200437, China 
LIU Jinsong 1. Faculty of Electrical Engineering, Shanghai University of Electric Power, Shanghai 200090, China
2. State Grid Shanghai Electric Power Research Institute, Shanghai 200437, China 
LIN Shunfu 1. Faculty of Electrical Engineering, Shanghai University of Electric Power, Shanghai 200090, China
2. State Grid Shanghai Electric Power Research Institute, Shanghai 200437, China 
Hits: 9
Download times: 1
Abstract:The high penetration of distributed energy resources has significantly increased the risk of voltage limit violations in distribution networks. Voltage optimization control methods based on deep reinforcement learning (DRL) struggle to ensure constraint satisfaction during online execution and lack adaptability to network topology changes. Additionally, model predictive control (MPC) incurs high computational costs due to full-horizon optimization. To address these challenges, an optimal voltage control method for distribution networks using deep reinforcement learning, considering dynamic correction by a policy safety layer is proposed. First, the spatial coupling features of the distribution network topology extracted by the graph attention network (GAT) are integrated into the state space of the deep deterministic policy gradient (DDPG) algorithm for offline training, thereby enhancing adaptability to frequent topology changes. Then, MPC is embedded into the DDPG framework as a policy safety layer. A value network is introduced to approximate the terminal value term, replacing the long-horizon terminal cost optimization in MPC, while the problem is then transformed into a control sequence sampling distribution using the cross-entropy method (CEM). Finally, the proposed method is validated based on an improved IEEE 123-bus system. Simulation results show that the proposed method not only ensures control performance and constraint feasibility but also exhibits excellent optimization control effects and robustness under uncertainties, topology changes, and measurement inaccuracies.
View Full Text  View/Add Comment  Download reader