A highly efficient time-series database approach for monitoring infrastructures

(English) The rising interest in extracting value from data has led to a broad proliferation of monitoring infrastructures, most notably composed by sensors, intended to collect this new oil. Thus, gathering data has become fundamental for a great number of applications, such as predictive maintenan...

ver descrição completa

Detalhes bibliográficos
Autor: García Calatrava, Carlos
Formato: tesis doctoral
Estado:Versión publicada
Fecha de publicación:2022
País:España
Recursos:CBUC, CESCA
Repositorio:TDR. Tesis Doctorales en Red
OAI Identifier:oai:www.tdx.cat:10803/692068
Acesso em linha:http://hdl.handle.net/10803/692068
https://dx.doi.org/10.5821/dissertation-2117-413897
Access Level:acceso abierto
Palavra-chave:Àrees temàtiques de la UPC::Informàtica
004
Descrição
Resumo:(English) The rising interest in extracting value from data has led to a broad proliferation of monitoring infrastructures, most notably composed by sensors, intended to collect this new oil. Thus, gathering data has become fundamental for a great number of applications, such as predictive maintenance techniques or anomaly detection algorithms. However, before data can be refined into insights and knowledge, it has to be efficiently stored and prepared for its later retrieval. While General-purpose database management systems, such as Relational Database Management Systems, have been historically capable of managing a wide range of scenarios, they were found inefficient, or even unsuitable, in handling the Velocity and Volume of nowadays large Infrastructures. Aiming to address the specific challenges of Monitoring Infrastructures, specialized systems like Time-Series Database Management Systems arose, becoming the fastest-growing database category since 2019. However, as each monitoring infrastructure has its own particularities, choosing the best fitting candidate solution became fairly laborious. In consequence, implementing efficient solutions involving Time-Series databases became an arduous task, not only in terms of investing in the most appropriate software and hardware infrastructure, but also in terms of finding expert personnel able to keep track and master those rapidly evolving technologies. In order to mitigate these problems, this research proposes a highly efficient Time-Series database approach for monitoring Infrastructures, aimed at providing the best balance between performance and resource consumption, while enabling its deployment in general purpose document-oriented databases, relieving experts from having to learn yet-another database solution from scratch. More precisely, our research provides the three following main contributions: (1) A foundation data model for time-series data over document-oriented databases, aimed at obtaining the best properties from both schema-full and schema-less approximations. (2) A technique for efficiently integrating several contiguous data models into a single time-series data store, creating a data-flow pattern named Cascading Polyglot Persistence. This technique makes it possible to adapt the database to the nature and progression of time-series data along time, as it is tailored to the expected operations to be performed according to the data aging, empowering further performance while limiting resource consumption. (3) A holistic scalability strategy for time-series databases following Cascading Polyglot Persistence, aimed at further maximizing the benefits of our polyglot approach when deploying it in a cluster fashion. In order to evaluate the performance of our approach, we materialize it on top of MongoDB, the most popular NoSQL database, which further facilitates its adoption. In addition, we benchmark it against two alternative solutions: InfluxDB, the most popular time-series database, and MongoDB itself. Our results show that our approach is able to retrieve historical data up to more than 10 times faster than MongoDB, while also globally outperforming InfluxDB. In addition, it has shown to be able to ingest streams of real-time data two times faster than both MongoDB and InfluxDB, while requesting the same disk space as InfluxDB. Regarding its ad hoc scalability approach, it has shown to greatly reduce the number of needed machines, with respect to traditional approaches, while offering a scalability efficiency up to 85%. These outstanding outcomes pave the way towards NagareDB, our time-series database, aimed at integrating all these approaches, providing them as an out-of-the-box solution.