Evaluating the Performance and Scalability o fan Apache Spark and Hadoop Cluster in a Low-Cost Environment

This article presents the design, configuration, and implementation of a distributed computing cluster using Apache Spark and Hadoop on Ubuntu Server 24.04.1 LTS. The architecture consists of a master node and multiple slave nodes connected to a local network via Ethernet. The installation, configur...

Descripción completa

Detalles Bibliográficos
Autores: Cruz Tumba, Natalie, Lancho Ramos, Alex, Leon Hurtado, Henry, Quispe Merma, Rafael Ricardo, Luque Ochoa, Evelyn Naida
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2025
País:Perú
Institución:Universidad Nacional Micaela Bastidas de Apurímac
Repositorio:UNMB-Riqchary
Idioma:español
OAI Identifier:oai:ojs2.68.168.220.155:article/218
Acceso en línea:https://revistas.unamba.edu.pe/index.php/riqchary/article/view/218
Access Level:acceso abierto
Palabra clave:Apache Spark
data processing
Hadoop
distributed cluster
distributed computing
low-cost cluster
apache spark
procesamiento de datos
hadoop
clúster distribuido
computación distribuida
clúster de bajo costo
Descripción
Sumario:This article presents the design, configuration, and implementation of a distributed computing cluster using Apache Spark and Hadoop on Ubuntu Server 24.04.1 LTS. The architecture consists of a master node and multiple slave nodes connected to a local network via Ethernet. The installation, configuration, and performance testing process with PySpark are detailed. The results demonstrate that, while a local configuration is more efficient for small datasets (<100 MB), the distributed cluster offers significant improvements for data volumes greater than 1 GB, validating its scalability and viability for resource-constrained educational and research environments.