Multicore workflow characterisation methodology for payloads running in the ALICE Grid

For LHC Run3 the ALICE experiment software stack has been completely refactored, incorporating support for multicore job execution. Whereas in both LHC Run 1 and 2 the Grid jobs were single-process and made use of a single CPU core, the new multicore jobs spawn multiple processes and threads within...

Descripción completa

Detalles Bibliográficos
Autores: Bertran Ferrer, Marta, Grigoras, Costin, Badia Sala, Rosa Maria|||0000-0003-2941-5499
Tipo de recurso: artículo
Fecha de publicación:2024
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2117/409320
Acceso en línea:https://hdl.handle.net/2117/409320
https://dx.doi.org/10.1051/epjconf/202429504005
Access Level:acceso abierto
Palabra clave:Resource allocation
Computational grids (Computer systems)
Assignació de recursos
Computació distribuïda
Àrees temàtiques de la UPC::Informàtica::Arquitectura de computadors::Arquitectures distribuïdes
id ES_c717cdbb11b6a3a84ff4f5fa0da42847
oai_identifier_str oai:upcommons.upc.edu:2117/409320
network_acronym_str ES
network_name_str España
spelling Multicore workflow characterisation methodology for payloads running in the ALICE Grid Bertran Ferrer, Marta Grigoras, Costin Badia Sala, Rosa Maria|||0000-0003-2941-5499 Resource allocation Computational grids (Computer systems) Assignació de recursos Computació distribuïda Àrees temàtiques de la UPC::Informàtica::Arquitectura de computadors::Arquitectures distribuïdes For LHC Run3 the ALICE experiment software stack has been completely refactored, incorporating support for multicore job execution. Whereas in both LHC Run 1 and 2 the Grid jobs were single-process and made use of a single CPU core, the new multicore jobs spawn multiple processes and threads within the payload. Some of these multicore jobs deploy a high amount of shortlived processes, in the order of more than a dozen per second. The overhead of starting so many processes impacts the overall CPU utilization of the payloads, in particular its System component. Furthermore, the short-lived processes were not correctly accounted for by the monitoring system of the experiment. This paper presents the developed new methodology for supervising the payload execution. We also present a black box analysis of the new multicore experiment software framework tracing the used resources and system function calls issued by MonteCarlo simulation jobs. Multiple sources of overhead in the lifecycle of processes and threads have thus been identified. This paper describes how the source of each was traced and what solutions were implemented to address them. These improvements have impacted the resource consumption and the overall turnaround time of these payloads with a notable 35% reduction in execution time for a reference production job. We also introduce how this methodology will be used to further improve the efficiency of our experiment software and what other optimization venues are currently under research. Peer Reviewed EPJ Web of Conferences https://hdl.handle.net/2117/409320 https://dx.doi.org/10.1051/epjconf/202429504005
title Multicore workflow characterisation methodology for payloads running in the ALICE Grid
spellingShingle Multicore workflow characterisation methodology for payloads running in the ALICE Grid
Bertran Ferrer, Marta
Resource allocation
Computational grids (Computer systems)
Assignació de recursos
Computació distribuïda
Àrees temàtiques de la UPC::Informàtica::Arquitectura de computadors::Arquitectures distribuïdes
title_short Multicore workflow characterisation methodology for payloads running in the ALICE Grid
title_full Multicore workflow characterisation methodology for payloads running in the ALICE Grid
title_fullStr Multicore workflow characterisation methodology for payloads running in the ALICE Grid
title_full_unstemmed Multicore workflow characterisation methodology for payloads running in the ALICE Grid
title_sort Multicore workflow characterisation methodology for payloads running in the ALICE Grid
author Bertran Ferrer, Marta
author_facet Bertran Ferrer, Marta
Grigoras, Costin
Badia Sala, Rosa Maria|||0000-0003-2941-5499
author_role author
author2 Grigoras, Costin
Badia Sala, Rosa Maria|||0000-0003-2941-5499
author2_role author
author
topic Resource allocation
Computational grids (Computer systems)
Assignació de recursos
Computació distribuïda
Àrees temàtiques de la UPC::Informàtica::Arquitectura de computadors::Arquitectures distribuïdes
topic_facet Resource allocation
Computational grids (Computer systems)
Assignació de recursos
Computació distribuïda
Àrees temàtiques de la UPC::Informàtica::Arquitectura de computadors::Arquitectures distribuïdes
description For LHC Run3 the ALICE experiment software stack has been completely refactored, incorporating support for multicore job execution. Whereas in both LHC Run 1 and 2 the Grid jobs were single-process and made use of a single CPU core, the new multicore jobs spawn multiple processes and threads within the payload. Some of these multicore jobs deploy a high amount of shortlived processes, in the order of more than a dozen per second. The overhead of starting so many processes impacts the overall CPU utilization of the payloads, in particular its System component. Furthermore, the short-lived processes were not correctly accounted for by the monitoring system of the experiment. This paper presents the developed new methodology for supervising the payload execution. We also present a black box analysis of the new multicore experiment software framework tracing the used resources and system function calls issued by MonteCarlo simulation jobs. Multiple sources of overhead in the lifecycle of processes and threads have thus been identified. This paper describes how the source of each was traced and what solutions were implemented to address them. These improvements have impacted the resource consumption and the overall turnaround time of these payloads with a notable 35% reduction in execution time for a reference production job. We also introduce how this methodology will be used to further improve the efficiency of our experiment software and what other optimization venues are currently under research.
publishDate 2024
format article
url https://hdl.handle.net/2117/409320
https://dx.doi.org/10.1051/epjconf/202429504005
language eng
eu_rights_str_mv openAccess
publisher EPJ Web of Conferences
institution Universitat Politècnica de Catalunya (UPC)
collection UPCommons. Portal del coneixement obert de la UPC
reponame_str UPCommons. Portal del coneixement obert de la UPC
instname_str Universitat Politècnica de Catalunya (UPC)
_version_ 1878440570622312448
publishDateSort 2024
author_browse Badia Sala, Rosa Maria|||0000-0003-2941-5499
Bertran Ferrer, Marta
Grigoras, Costin
publisherStr EPJ Web of Conferences
score 6,924472