Exploring reward strategies for wind turbine pitch control by reinforcement learning

In this work, a pitch controller of a wind turbine (WT) inspired by reinforcement learning (RL) is designed and implemented. The control system consists of a state estimator, a reward strategy, a policy table, and a policy update algorithm. Novel reward strategies related to the energy deviation fro...

Descripción completa

Detalles Bibliográficos
Autores: Sierra-García, Jesús Enrique, Santos Peñas, Matilde
Tipo de recurso: artículo
Fecha de publicación:2020
País:España
Institución:Universidad Complutense de Madrid (UCM)
Repositorio:Docta Complutense
Idioma:inglés
OAI Identifier:oai:docta.ucm.es:20.500.14352/112247
Acceso en línea:https://hdl.handle.net/20.500.14352/112247
Access Level:acceso abierto
Palabra clave:Intelligent control
Pitch control
Wind turbines
Wind energy
Reinforcement learning
Reward strategies
Inteligencia artificial (Informática)
1203.04 Inteligencia Artificial
id ES_d134e7da2f68ea434e2a1d58a35342ce
oai_identifier_str oai:docta.ucm.es:20.500.14352/112247
network_acronym_str ES
network_name_str España
spelling Exploring reward strategies for wind turbine pitch control by reinforcement learning Sierra-García, Jesús Enrique Santos Peñas, Matilde Intelligent control Pitch control Wind turbines Wind energy Reinforcement learning Reward strategies Inteligencia artificial (Informática) 1203.04 Inteligencia Artificial In this work, a pitch controller of a wind turbine (WT) inspired by reinforcement learning (RL) is designed and implemented. The control system consists of a state estimator, a reward strategy, a policy table, and a policy update algorithm. Novel reward strategies related to the energy deviation from the rated power are defined. They are designed to improve the efficiency of the WT. Two new categories of reward strategies are proposed: “only positive” (O-P) and “positive-negative” (P-N) rewards. The relationship of these categories with the exploration-exploitation dilemma, the use of ϵ-greedy methods and the learning convergence are also introduced and linked to the WT control problem. In addition, an extensive analysis of the influence of the different rewards in the controller performance and in the learning speed is carried out. The controller is compared with a proportional-integral-derivative (PID) regulator for the same small wind turbine, obtaining better results. The simulations show how the P-N rewards improve the performance of the controller, stabilize the output power around the rated power, and reduce the error over time. MDPI https://hdl.handle.net/20.500.14352/112247
title Exploring reward strategies for wind turbine pitch control by reinforcement learning
spellingShingle Exploring reward strategies for wind turbine pitch control by reinforcement learning
Sierra-García, Jesús Enrique
Intelligent control
Pitch control
Wind turbines
Wind energy
Reinforcement learning
Reward strategies
Inteligencia artificial (Informática)
1203.04 Inteligencia Artificial
title_short Exploring reward strategies for wind turbine pitch control by reinforcement learning
title_full Exploring reward strategies for wind turbine pitch control by reinforcement learning
title_fullStr Exploring reward strategies for wind turbine pitch control by reinforcement learning
title_full_unstemmed Exploring reward strategies for wind turbine pitch control by reinforcement learning
title_sort Exploring reward strategies for wind turbine pitch control by reinforcement learning
author Sierra-García, Jesús Enrique
author_facet Sierra-García, Jesús Enrique
Santos Peñas, Matilde
author_role author
author2 Santos Peñas, Matilde
author2_role author
topic Intelligent control
Pitch control
Wind turbines
Wind energy
Reinforcement learning
Reward strategies
Inteligencia artificial (Informática)
1203.04 Inteligencia Artificial
topic_facet Intelligent control
Pitch control
Wind turbines
Wind energy
Reinforcement learning
Reward strategies
Inteligencia artificial (Informática)
1203.04 Inteligencia Artificial
description In this work, a pitch controller of a wind turbine (WT) inspired by reinforcement learning (RL) is designed and implemented. The control system consists of a state estimator, a reward strategy, a policy table, and a policy update algorithm. Novel reward strategies related to the energy deviation from the rated power are defined. They are designed to improve the efficiency of the WT. Two new categories of reward strategies are proposed: “only positive” (O-P) and “positive-negative” (P-N) rewards. The relationship of these categories with the exploration-exploitation dilemma, the use of ϵ-greedy methods and the learning convergence are also introduced and linked to the WT control problem. In addition, an extensive analysis of the influence of the different rewards in the controller performance and in the learning speed is carried out. The controller is compared with a proportional-integral-derivative (PID) regulator for the same small wind turbine, obtaining better results. The simulations show how the P-N rewards improve the performance of the controller, stabilize the output power around the rated power, and reduce the error over time.
publishDate 2020
format article
url https://hdl.handle.net/20.500.14352/112247
language eng
eu_rights_str_mv openAccess
publisher MDPI
institution Universidad Complutense de Madrid (UCM)
collection Docta Complutense
reponame_str Docta Complutense
instname_str Universidad Complutense de Madrid (UCM)
_version_ 1878740335313551360
publishDateSort 2020
author_browse Santos Peñas, Matilde
Sierra-García, Jesús Enrique
publisherStr MDPI
score 6.8972664