Clustering Analysis Of Spotify Most Streamed Songs Using The K-Means Algorithm Based On Popularity

Authors

  • Arko Fernanda Wibawa Internet Engineering Technologi, Politeknik Negeri Lampung
  • Juwita Valentiya Internet Engineering Technologi, Politeknik Negeri Lampung

DOI:

https://doi.org/10.25181/rt.v5i1.4713

Keywords:

Spotify, KMeans, streaming popularity, logarithmic transformation, song segmentation

Abstract

The rapid growth of music streaming platforms has created very large catalogs, making popularity patterns difficult to understand using a single indicator. Total streams reflect cumulative success, but they do not always represent current listening momentum. This situation motivates the need for song segmentation based on more informative popularity patterns to support decision-making for streaming platforms, artists, and labels. This study applied a data mining approach using K-Means clustering to group Spotify most-streamed songs based on streaming popularity indicators. The main contribution was a segmentation framework that combined total streams, daily streams, and a daily-to-total streams ratio to better capture current momentum. The method included data cleaning, missing value imputation, logarithmic transformation to reduce skewness, feature engineering of a ratio variable, feature standardization, K-Means training, cluster number selection using the elbow method and Silhouette Score, and evaluation using Inertia, Silhouette Score, the Calinski–Harabasz Index, and the Davies–Bouldin Index. The final model with k = 4 achieved an Inertia of 2673.011 and a Silhouette Score of 0.364835 and produced four interpretable segments. Cluster 0 represented super-trending songs with the highest daily-to-total ratio, cluster 1 represented legacy popular songs with low daily activity, cluster 2 represented mega hits with extremely high total streams and still strong daily activity, and cluster 3 represented consistently performing songs with stable daily streams. These segments provided practical insights for promotion prioritization, playlist curation, and trend interpretation.

Downloads

Download data is not yet available.

References

S. Mukhopadhyay, A. Kumar, D. Parashar, and M. Singh, “Enhanced Music Recommendation Systems: A Comparative Study of Content-Based Filtering and K-Means Clustering Approaches,” RIA, vol. 38, no. 1, pp. 365–376, Feb. 2024, doi: 10.18280/ria.380138.

A. Parthasarathy, “Music Recommendation Using Machine Learning Algorithms,” vol. 11, no. 11, 2023. doi: http://doi.one/10.1729/Journal.36826.

N. Rohman and A. Wibowo, “Clustering Of Popular Spotify Songs In 2023 Using K-Means Method And Silhouette Coefficient,” pilar, vol. 20, no. 1, pp. 18–24, Apr. 2024, doi: 10.33480/pilar.v20i1.4937.

B. Daga, H. Kadam, S. Shrungare, S. Fernandes, and A. Johsnon, “Music Recommendation System,” IJCTT, vol. 71, no. 5, pp. 26–36, May 2023, doi: 10.14445/22312803/IJCTT-V71I5P105.

M. Gagolewski, M. Bartoszuk, and A. Cena, “Are Cluster Validity Measures (In)valid?,” Information Sciences, vol. 581, pp. 620–636, Dec. 2021, doi: 10.1016/j.ins.2021.10.004.

H. Magadum, H. K. Azad, H. Patel, and R. H. R, “Music Recommendation Using Dynamic Feedback And Content-Based Filtering,” Multimedia Tools and Applications, vol. 83, no. 32, pp. 77469–77488, Sep. 2024, doi: 10.1007/s11042-024-18636-8.

R. Kumar and Rakesh, “Music Recommendation System Using Machine Learning,” in 2022 4th International Conference on Advances in Computing, Communication Control and Networking (ICAC3N), Greater Noida, India: IEEE, Dec. 2022, pp. 572–576. doi: 10.1109/ICAC3N56670.2022.10074362.

L. E. Ekemeyong Awong and T. Zielinska, “Comparative Analysis of the Clustering Quality in Self-Organizing Maps for Human Posture Classification,” Sensors, vol. 23, no. 18, p. 7925, Sep. 2023, doi: 10.3390/s23187925.

M. Garanayak, S. K. Nayak, Sangeetha K., T. Choudhury, and Shitharth S., “Content and Popularity-Based Music Recommendation System:,” International Journal of Information System Modeling and Design, vol. 13, no. 7, pp. 1–14, Dec. 2022, doi: 10.4018/ijismd.315027.

L. Liu et al., “Personalized Music Recommendation Algorithm Based On Machine Learning,” Multimedia Systems, vol. 31, no. 3, p. 166, Jun. 2025, doi: 10.1007/s00530-025-01749-x.

M. G. Galety, R. Thiagarajan, R. Sangeetha, L. K. B. Vignesh, S. Arun, and R. Krishnamoorthy, “Personalized Music Recommendation model based on Machine Learning,” in 2022 8th International Conference on Smart Structures and Systems (ICSSS), Chennai, India: IEEE, Apr. 2022, pp. 1–6. doi: 10.1109/ICSSS54381.2022.9782288.

J. S. Gulmatico, J. A. B. Susa, M. A. F. Malbog, A. Acoba, M. D. Nipas, and J. N. Mindoro, “SpotiPred: A Machine Learning Approach Prediction of Spotify Music Popularity by Audio Features,” in 2022 Second International Conference on Power, Control and Computing Technologies (ICPC2T), Raipur, India: IEEE, Mar. 2022, pp. 1–5. doi: 10.1109/ICPC2T53885.2022.9776765.

J. T. Anthony, G. E. Christian, V. Evanlim, H. Lucky, and D. Suhartono, “The Utilization of Content Based Filtering for Spotify Music Recommendation,” in 2022 International Conference on Informatics Electrical and Electronics (ICIEE), Yogyakarta, Indonesia: IEEE, Oct. 2022, pp. 1–4. doi: 10.1109/ICIEE55596.2022.10010097.

M. Alizade, R. Kheni, S. Price, B. C. Sousa, D. L. Cote, and R. Neamtu, “A Comparative Study of Clustering Methods for Nanoindentation Mapping Data,” Integr Mater Manuf Innov, vol. 13, no. 2, pp. 526–540, Jun. 2024, doi: 10.1007/s40192-024-00349-3.

B. A. Hassan, N. B. Tayfor, A. A. Hassan, A. M. Ahmed, T. A. Rashid, and N. N. Abdalla, “From A-to-Z review of clustering validation indices,” Neurocomputing, vol. 601, p. 128198, Oct. 2024, doi: 10.1016/j.neucom.2024.128198.

A. Singh and S. Gupta, “Recommendation System Algorithms For Music Therapy,” in 2023 13th International Conference on Cloud Computing, Data Science & Engineering (Confluence), Noida, India: IEEE, Jan. 2023, pp. 138–143. doi: 10.1109/Confluence56041.2023.10048894.

Downloads

Published

2026-02-28

Issue

Section

Articles