考虑策略安全层动态修正的配电网电压优化控制
CSTR:
作者:
作者单位:

1. 上海电力大学电气工程学部,上海 200090;2. 国网上海市电力公司电力科学研究院,上海 200437

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金项目资助 (No. 52577119)


Optimal voltage control for distribution networks considering dynamic correction by policy safety layer
Author:
Affiliation:

1. Faculty of Electrical Engineering, Shanghai University of Electric Power, Shanghai 200090, China; 2. State Grid Shanghai Electric Power Research Institute, Shanghai 200437, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    分布式能源高比例接入加剧了配电网的电压越限风险,基于深度强化学习的电压优化控制方法难以保证执行阶段的约束可行性且对拓扑变化适应性不足,而模型预测控制 (model predictive control, MPC) 方法全时域计算开销较大。为此,提出一种考虑策略安全层动态修正的配电网深度强化学习电压优化控制方法。首先,将图注意力网络提取的配电网拓扑空间耦合特征纳入深度确定性策略梯度 (deep deterministic policy gradient, DDPG) 的状态空间,进行离线训练,以适应配电网拓扑结构的频繁变化。然后,将 MPC 作为策略安全层嵌入 DDPG 框架,引入价值网络作为终端价值项近似替代长时域 MPC 的修正求解,并采用交叉熵方法对控制序列进行采样分布。最后,基于改进 IEEE123 节点系统进行方法验证。结果表明,所提方法不仅能够保证控制性能和约束可行性,而且对不确定性、拓扑变化及量测不准确均具有良好的控制效果和鲁棒性。

    Abstract:

    The high penetration of distributed energy resources has significantly increased the risk of voltage limit violations in distribution networks. Voltage optimization control methods based on deep reinforcement learning (DRL) struggle to ensure constraint satisfaction during online execution and lack adaptability to network topology changes. Additionally, model predictive control (MPC) incurs high computational costs due to full-horizon optimization. To address these challenges, an optimal voltage control method for distribution networks using deep reinforcement learning, considering dynamic correction by a policy safety layer is proposed. First, the spatial coupling features of the distribution network topology extracted by the graph attention network (GAT) are integrated into the state space of the deep deterministic policy gradient (DDPG) algorithm for offline training, thereby enhancing adaptability to frequent topology changes. Then, MPC is embedded into the DDPG framework as a policy safety layer. A value network is introduced to approximate the terminal value term, replacing the long-horizon terminal cost optimization in MPC, while the problem is then transformed into a control sequence sampling distribution using the cross-entropy method (CEM). Finally, the proposed method is validated based on an improved IEEE 123-bus system. Simulation results show that the proposed method not only ensures control performance and constraint feasibility but also exhibits excellent optimization control effects and robustness under uncertainties, topology changes, and measurement inaccuracies.

    参考文献
    相似文献
    引证文献
引用本文

李晓露,杜本杰,苏昊天,等.考虑策略安全层动态修正的配电网电压优化控制[J].电力系统保护与控制,2026,54(16):93-106.[LI Xiaolu, DU Benjie, SU Haotian, et al. Optimal voltage control for distribution networks considering dynamic correction by policy safety layer[J]. Power System Protection and Control,2026,V54(16):93-106]

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-03-08
  • 最后修改日期:2026-05-28
  • 录用日期:
  • 在线发布日期: 2026-08-14
  • 出版日期:
文章二维码
关闭
关闭