The Computational Efficiency of Task-Based Parallelism of Spark MLlib


ÖZTÜRK M. M., Nejkovic V.

12th International Conference on Electrical, Electronic and Computing Engineering, IcETRAN 2025, Cacak, Sırbistan, 9 - 12 Haziran 2025, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/icetran66854.2025.11114169
  • Basıldığı Şehir: Cacak
  • Basıldığı Ülke: Sırbistan
  • Anahtar Kelimeler: Apache Spark, MLlib, parallelization
  • Süleyman Demirel Üniversitesi Adresli: Evet

Özet

Apache Spark is a big data processing framework that provides various tools to practitioners. MLlib is one of the most well-known Spark libraries, presenting various machine learning algorithms to process large-scale data. Existing studies generally focus on the configurable hyperparameters of Spark to increase general performance. Resilient distributed dataset (RDD), which is provided by Spark, provides a data-based parallelism to increase the big data processing performance. To this end, the manager of Spark is responsible for the distribution of tasks among various workers. However, different from existing studies, this work investigates whether a second task-based parallelization has the potential to increase the computational performance of Spark. To that end, a parallelization algorithm is developed by utilizing the parallel library of R. The experiment enriched with binary classification datasets shows that each algorithm of MLlib does not reduce computation time remarkably. Further, there is a moderate speedup of 1.4. In particular, linear models are prone to give promising results in task-based parallelization.