Subject Area Classification of Journal Articles Based on Metadata Using Bag of Words and Naïve Bayes

https://doi.org/10.56741/IISTR.esl.002041

Authors

Keywords:

Data Analysis, Feature Extraction, Performance Evaluation, Scientific Document Classification, Supervised Machine Learning

Abstract

The rapid growth of scientific publications poses challenges in grouping journal articles based on subject area, especially when using metadata such as titles, abstracts, and keywords. However, differences in feature representation and classification algorithms often result in varying performance, requiring comparative studies to determine the optimal model combination. This study compares four combinations of subject area classification models, namely TF-IDF + Naïve Bayes, TF-IDF + Support Vector Machine, Bag-of-Words + Support Vector Machine, and Bag-of-Words + Naïve Bayes. The research process included text preprocessing, feature extraction, and testing using an 80% training and 20% testing data split scheme in five scenarios. The evaluation was performed using confusion matrices, accuracy, precision, recall, and F1-score. The experimental results showed variations in performance between models, with an average F1-score of 0.8103 for TF-IDF + Naïve Bayes, 0.8494 for TF-IDF + Support Vector Machine, 0.8297 for Bag-of-Words + Support Vector Machine, and 0.8335 for Bag-of-Words + Naïve Bayes as the best performance. These findings indicate that a word frequency-based approach combined with Naïve Bayes is effective for classifying journal article subject areas based on metadata, although challenges remain in subject areas with semantic proximity.

Downloads

Download data is not yet available.

Author Biographies

Ainunna’imah, Universitas Ahmad Dahlan

is a master's student in the Master's Program of Informatics, Faculty of Industrial Technology, Universitas Ahmad Dahlan, Yogyakarta, Indonesia. The author's academic interests include artificial intelligence, data science, information systems, and emerging digital technologies. Currently engaged in research on technology-driven innovation and intelligent systems, the author actively contributes to academic projects and scholarly activities. (email: 2408048020@webmail.uad.ac.id).

Herman Yuliansyah, Indonesia

is a lecturer in the Department of Informatics, Faculty of Industrial Technology, Universitas Ahmad Dahlan, Yogyakarta, Indonesia. His academic and research interests include artificial intelligence, software engineering, data analytics, information systems, and digital innovation. He is actively involved in teaching, research, and community engagement activities, contributing to the advancement of information technology and its applications in education, industry, and society. Email: herman.yuliansyah@tif.uad.ac.id.

Imam Riadi, Universitas Ahmad Dahlan

is a Professor and lecturer in the Department of Information Systems, Faculty of Applied Science and Technology, Universitas Ahmad Dahlan, Yogyakarta, Indonesia. His research interests encompass information systems, cybersecurity, digital forensics, artificial intelligence, and data governance. He has published extensively in reputable national and international journals and actively contributes to research, teaching, and community service initiatives that advance information technology and digital transformation. (email: imam.riadi@is.uad.ac.id).

References

W. E. Savage and A. J. Olejniczak, “More journal articles and fewer books: Publication practices in the social sciences in the 2010’s,” PLoS One, vol. 17, no. 2, pp. 1–16, 2022, doi: 10.1371/journal.pone.0263410. DOI: https://doi.org/10.1371/journal.pone.0263410

J. J. Mallett, “The resilience of scientific publication : From elite ancient academies to open access,” Learn. Publ., vol. 34, no. 1, pp. 49–56, 2021, doi: 10.1002/leap.1366. DOI: https://doi.org/10.1002/leap.1366

M. A. Hanson, P. G. Barreiro, P. Crosetto, and D. Brockington, “The strain on scientific publishing,” Quant. Sci. Stud., vol. 5, no. 4, pp. 823–843, Nov. 2024, doi: 10.1162/qss_a_00327. DOI: https://doi.org/10.1162/qss_a_00327

M. Thelwall and P. Sud, “Scopus 1900–2020: Growth in articles, abstracts, countries, fields, and journals,” Quant. Sci. Stud., vol. 3, no. 1, pp. 37–50, Apr. 2022, doi: 10.1162/qss_a_00177. DOI: https://doi.org/10.1162/qss_a_00177

O. Fedoruk and S. Lecturer, “Society. Document. Communication,” Soc. Doc. Commun., vol. 10, no. 1, pp. 55–64, 2025, doi: 10.69587/sdc/1.2025.55. DOI: https://doi.org/10.69587/sdc/1.2025.55

M. Aria, C. Cuccurullo, L. D. Aniello, and M. Misuraca, “Comparative science mapping : a novel conceptual structure analysis with metadata,” Scientometrics, vol. 129, no. 11, pp. 7055–7081, 2024, doi: 10.1007/s11192-024-05161-6. DOI: https://doi.org/10.1007/s11192-024-05161-6

P. Pottier et al., “Title, abstract and keywords: a practical guide to maximize the visibility and impact of academic papers,” Proc. R. Soc. B Biol. Sci., vol. 291, no. 2027, p. 20241222, Jul. 2024, doi: 10.1098/rspb.2024.1222. DOI: https://doi.org/10.1098/rspb.2024.1222

G. Mustafa, M. Usman, M. Afzal, A. Shahid, and A. Koubâa, “A Comprehensive Evaluation of Metadata-Based Features to Classify Research Paper’s Topics,” IEEE Access, vol. 9, pp. 133500–133509, 2021, doi: 10.1109/access.2021.3115148. DOI: https://doi.org/10.1109/ACCESS.2021.3115148

W. I. Al-Obaydy, H. Hashim, Y. A. Najm, and A. Jalal, “Document classification using term frequency-inverse document frequency and K-means clustering,” Indones. J. Electr. Eng. Comput. Sci., 2022, doi: 10.11591/ijeecs.v27.i3.pp1517-1524. DOI: https://doi.org/10.11591/ijeecs.v27.i3.pp1517-1524

R. C. Morales-hernández and D. Becerra-Alonso, “A Comparison of Multi-Label Text Classification Models in Research Articles Labeled With Sustainable Development Goals,” IEEE Access, vol. 10, no. November, 2022. DOI: https://doi.org/10.1109/ACCESS.2022.3223094

S. Rahman, H. K. Shanto, U. A. Koana, and S. M. Danish, “Automated Research Article Classification and Recommendation Using NLP and Machine Learning,” arXiv, vol. 1, 2025. DOI: https://doi.org/10.1109/FLLM67465.2025.11391003

A. M. Chaid, Z. A. Abdulrazzaq, R. N. Sadoon, and M. A. Aljabery, “Comparative Analysis of Innovative Machine Learning Algorithms : Advancements in Natural Language Processing,” J. Inf. Syst. Eng. Manag., vol. 10, 2025. DOI: https://doi.org/10.52783/jisem.v10i14s.2373

N. Susetyo, F. Putri, A. Prasetya, H. Ar, and A. Nur, “Classification of Engineering Journals Quartile using Various Supervised Learning Models,” Ilk. J. Ilm., vol. 15, no. 1, pp. 101–106, 2023. DOI: https://doi.org/10.33096/ilkom.v15i1.1483.101-106

R. Guns, “Fine-grained classification of social science journal articles using textual data : A comparison of supervised machine learning approaches,” Quant. Sci. Stud., vol. 2, no. 1, 2021, doi: 10.1162/qss. DOI: https://doi.org/10.1162/qss_a_00106

M. Lu, L. Tang, and X. Zhou, “An ensemble approach for research article classi fi cation : a case study in arti fi cial intelligence,” PeerJ Comput. Sci., pp. 1–22, 2024, doi: 10.7717/peerj-cs.2521. DOI: https://doi.org/10.7717/peerj-cs.2521

L. P. F. Garcia, J. P. Mena-Chalco, R. L. Grando, W. M. C. da Silva, F. P. Guimarães, and V. de A. Jorge, “Hierarchical Article Classification: A Multi-Level Framework for Organizing Scholarly Literature,” IEEE Access, vol. 13, no. May, pp. 102589–102601, 2025, doi: 10.1109/ACCESS.2025.3579232. DOI: https://doi.org/10.1109/ACCESS.2025.3579232

T. Gupta, M. Zaki, N. M. A. Krishnan, and Mausam, “MatSciBERT: A materials domain language model for text mining and information extraction,” npj Comput. Mater., vol. 8, no. 1, p. 102, 2022, doi: 10.1038/s41524-022-00784-w. DOI: https://doi.org/10.1038/s41524-022-00784-w

S. Aum and S. Choe, “srBERT : automatic article classification model for systematic review using BERT,” Syst. Rev., pp. 1–8, 2021, doi: 10.1186/s13643-021-01763-w. DOI: https://doi.org/10.1186/s13643-021-01763-w

C. Arhiliuc, R. Guns, W. Daelemans, and T. C. E. Engels, “Journal article classification using abstracts: a comparison of classical and transformer-based machine learning methods,” Scientometrics, vol. 130, no. 1, pp. 313–342, 2025, doi: 10.1007/s11192-024-05217-7. DOI: https://doi.org/10.1007/s11192-024-05217-7

F. Aguilar-canto, C. Macias, A. E. Juárez, M. A. Cardoso-moreno, and H. Calvo, “Quartile Prediction and Journal Recommendation Using Deep Learning Models for Artificial Intelligence Articles,” J. Scientometr. Res., vol. 14, no. 1, pp. 373–382, 2025, doi: 10.5530/jscires.20251460. DOI: https://doi.org/10.5530/jscires.20251460

M. Sciences, “Exploring Preprocessing Techniques for Natural Language Text : A Comprehensive Study sing Python Code,” Int. J. Eng. Technol. Manag. Sci., vol. 5, no. 7, pp. 390–399, 2023, doi: 10.46647/ijetms.2023.v07i05.047. DOI: https://doi.org/10.46647/ijetms.2023.v07i05.047

W. Costa and G. V. Pedrosa, A Textual Representation Based on Bag-of-Concepts and Thesaurus for Legal Information Retrieval. 2022. doi: 10.5753/kdmile.2022.227779. DOI: https://doi.org/10.5753/kdmile.2022.227779

D. Sugiarto, E. Utami, and A. Yaqin, “Perbandingan Kinerja Model TF-IDF dan BOW untuk Klasifikasi Opini Publik Tentang Kebijakan BLT Minyak Goreng,” vol. 12, no. 3, pp. 272–277, 2022. DOI: https://doi.org/10.25105/jti.v12i3.15669

K. Juluru, H.-H. Shih, K. N. Keshava Murthy, and P. Elnajjar, “Bag-of-Words Technique in Natural Language Processing: A Primer for Radiologists,” RadioGraphics, vol. 41, no. 5, pp. 1420–1426, Aug. 2021, doi: 10.1148/rg.2021210025. DOI: https://doi.org/10.1148/rg.2021210025

Z. Jiang, B. Gao, Y. He, Y. Han, P. Doyle, and Q. Zhu, “Text Classification Using Novel Term Weighting Scheme-Based Improved TF-IDF for Internet Media Reports,” Math. Probl. Eng., vol. 2021, no. ii, 2021, doi: 10.1155/2021/6619088. DOI: https://doi.org/10.1155/2021/6619088

R. Blanquero, E. Carrizosa, and P. Ramírez-cobo, “Computers and Operations Research Variable selection for Naïve Bayes classification,” Comput. Oper. Res., vol. 135, p. 105456, 2021, doi: 10.1016/j.cor.2021.105456. DOI: https://doi.org/10.1016/j.cor.2021.105456

T. Simbolon, A. Prasetya, I. Ari, E. Zaeni, and A. Ritahani, “Text classification of traditional and national songs using naïve bayes algorithm,” Sci. Inf. Technol. Lett., vol. 3, no. 2, pp. 59–72, 2022. DOI: https://doi.org/10.31763/sitech.v3i2.1215

D. M. Abdullah and A. M. Abdulazeez, “Machine Learning Applications based on SVM Classification : A Review,” Qubahan Acad. J., pp. 81–90, 2021, doi: 10.48161/Issn.2709-8206. DOI: https://doi.org/10.48161/qaj.v1n2a50

G. D. Bispo et al., “Applied Sciences Automatic Literature Mapping Selection : Classification of Papers on Industry Productivity,” Appl. Sci., vol. 14, no. 9, p. 3679, 2024, doi: https://doi.org/10.3390/app14093679. DOI: https://doi.org/10.3390/app14093679

A. Lyutov, Y. Uygun, and M. T. Hütt, “Machine learning misclassification networks reveal a citation advantage of interdisciplinary publications only in high-impact journals,” Sci. Rep., vol. 14, no. 1, pp. 1–11, 2024, doi: 10.1038/s41598-024-72364-5. DOI: https://doi.org/10.1038/s41598-024-72364-5

Published

2026-06-13

How to Cite

Ainunna’imah, Yuliansyah, H., & Riadi, I. (2026). Subject Area Classification of Journal Articles Based on Metadata Using Bag of Words and Naïve Bayes. Engineering Science Letter, 5(02), 80–88. https://doi.org/10.56741/IISTR.esl.002041

Plaudit