Methods of Data Popularity Evaluation in the ATLAS Experiment at the LHC

Thomas Beermann; Olga Chuchuk; Alessandro Di Girolamo; Maria Grigorieva; Alexei Klimentov; Mario Lassnig; Markus Schulz; Andrea Sciaba; Eugeny Tretyakov

doi:10.1051/epjconf/202125102013

EPJ

a
b
c
d
e
ap
st
h
plus
ds
pv
ti
qt
am
n

Proceedings

Open Access

EPJ Web of Conferences 251, 02013 (2021)
https://doi.org/10.1051/epjconf/202125102013

Methods of Data Popularity Evaluation in the ATLAS Experiment at the LHC

Thomas Beermann⁵, Olga Chuchuk⁶^,7^*, Alessandro Di Girolamo⁷, Maria Grigorieva¹^,2^**, Alexei Klimentov³, Mario Lassnig⁷, Markus Schulz⁷, Andrea Sciaba⁷^*** and Eugeny Tretyakov²^,4

¹ Lomonosov Moscow State University, Russian Federation
² Plekhanov Russian University of Economics, Russian Federation
³ Brookhaven National Laboratory, USA
⁴ National Research Nuclear University MEPhI, Russian Federation
⁵ Bergische Universitaet Wuppertal, Germany
⁶ Université Côte d’Azur, France
⁷ CERN, Geneva, Switzerland

^* e-mail: olga.chuchuk@cern.ch
^** e-mail: maria.grigorieva@cern.ch
^*** e-mail: andrea.sciaba@cern.ch

Published online: 23 August 2021

Abstract

The ATLAS Experiment at the LHC generates petabytes of data that is distributed among 160 computing sites all over the world and is processed continuously by various central production and user analysis tasks. The popularity of data is typically measured as the number of accesses and plays an important role in resolving data management issues: deleting, replicating, moving between tapes, disks and caches. These data management procedures were still carried out in a semi-manual mode and now we have focused our efforts on automating it, making use of the historical knowledge about existing data management strategies. In this study we describe sources of information about data popularity and demonstrate their consistency. Based on the calculated popularity measurements, various distributions were obtained. Auxiliary information about replication and task processing allowed us to evaluate the correspondence between the number of tasks with popular data executed per site and the number of replicas per site. We also examine the popularity of user analysis data that is much less predictable than in the central production and requires more indicators than just the number of accesses.

© The Authors, published by EDP Sciences, 2021

This is an Open Access article distributed under the terms of the Creative Commons Attribution License 4.0, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.