https://doi.org/10.1051/epjconf/201921404030
Dynamic and on demand data streams
1
INFN Sezione di Perugia,
I-06100 Perugia,
Italy
2
Università degli Studi di Perugia, I-06100 Perugia,
Italy
* Corresponding author: matteo.duranti@infn.it
** Corresponding author: valerio.formato@infn.it
*** Corresponding author: valerio.vagelli@unipg.it
Published online: 17 September 2019
Replicability and efficiency of data processing on the same data samples are a major challenge for the analysis of data produced by HEP experiments. High level data analyzed by end-users are typically produced as a subset of the whole experiment data sample to study interesting selection of data (streams). For standard applications, streams may be eventually copied from servers and analyzed on local computing centers or user machine clients. The creation of streams as copy of a subset of the original data results in redundant information stored in filesystems and may be not efficient: if the definition of streams changes, it may force a reprocessing of the low-level files with consequent impact on the data analysis efficiency. We propose an approach based on a database of lookup tables intended for dynamic and on-demand definition of data streams. This enables the end-users, as the data analysis strategy evolves, to explore different definitions of streams with minimal cost in computing resources. We also present a prototype demonstration application of this database for the analysis of the AMS-02 experiment data.
© The Authors, published by EDP Sciences, 2019
This is an Open Access article distributed under the terms of the Creative Commons Attribution License 4.0, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.