Subject Area Classification of Journal Articles Based on Metadata Using Bag of Words and Naïve Bayes
Keywords:
Data Analysis, Feature Extraction, Performance Evaluation, Scientific Document Classification, Supervised Machine LearningAbstract
The rapid growth of scientific publications poses challenges in grouping journal articles based on subject area, especially when using metadata such as titles, abstracts, and keywords. However, differences in feature representation and classification algorithms often result in varying performance, requiring comparative studies to determine the optimal model combination. This study compares four combinations of subject area classification models, namely TF-IDF + Naïve Bayes, TF-IDF + Support Vector Machine, Bag-of-Words + Support Vector Machine, and Bag-of-Words + Naïve Bayes. The research process included text preprocessing, feature extraction, and testing using an 80% training and 20% testing data split scheme in five scenarios. The evaluation was performed using confusion matrices, accuracy, precision, recall, and F1-score. The experimental results showed variations in performance between models, with an average F1-score of 0.8103 for TF-IDF + Naïve Bayes, 0.8494 for TF-IDF + Support Vector Machine, 0.8297 for Bag-of-Words + Support Vector Machine, and 0.8335 for Bag-of-Words + Naïve Bayes as the best performance. These findings indicate that a word frequency-based approach combined with Naïve Bayes is effective for classifying journal article subject areas based on metadata, although challenges remain in subject areas with semantic proximity.
Downloads
References
W. E. Savage and A. J. Olejniczak, “More journal articles and fewer books: Publication practices in the social sciences in the 2010’s,” PLoS One, vol. 17, no. 2, pp. 1–16, 2022, doi: 10.1371/journal.pone.0263410. DOI: https://doi.org/10.1371/journal.pone.0263410
J. J. Mallett, “The resilience of scientific publication : From elite ancient academies to open access,” Learn. Publ., vol. 34, no. 1, pp. 49–56, 2021, doi: 10.1002/leap.1366. DOI: https://doi.org/10.1002/leap.1366
M. A. Hanson, P. G. Barreiro, P. Crosetto, and D. Brockington, “The strain on scientific publishing,” Quant. Sci. Stud., vol. 5, no. 4, pp. 823–843, Nov. 2024, doi: 10.1162/qss_a_00327. DOI: https://doi.org/10.1162/qss_a_00327
M. Thelwall and P. Sud, “Scopus 1900–2020: Growth in articles, abstracts, countries, fields, and journals,” Quant. Sci. Stud., vol. 3, no. 1, pp. 37–50, Apr. 2022, doi: 10.1162/qss_a_00177. DOI: https://doi.org/10.1162/qss_a_00177
O. Fedoruk and S. Lecturer, “Society. Document. Communication,” Soc. Doc. Commun., vol. 10, no. 1, pp. 55–64, 2025, doi: 10.69587/sdc/1.2025.55. DOI: https://doi.org/10.69587/sdc/1.2025.55
M. Aria, C. Cuccurullo, L. D. Aniello, and M. Misuraca, “Comparative science mapping : a novel conceptual structure analysis with metadata,” Scientometrics, vol. 129, no. 11, pp. 7055–7081, 2024, doi: 10.1007/s11192-024-05161-6. DOI: https://doi.org/10.1007/s11192-024-05161-6
P. Pottier et al., “Title, abstract and keywords: a practical guide to maximize the visibility and impact of academic papers,” Proc. R. Soc. B Biol. Sci., vol. 291, no. 2027, p. 20241222, Jul. 2024, doi: 10.1098/rspb.2024.1222. DOI: https://doi.org/10.1098/rspb.2024.1222
G. Mustafa, M. Usman, M. Afzal, A. Shahid, and A. Koubâa, “A Comprehensive Evaluation of Metadata-Based Features to Classify Research Paper’s Topics,” IEEE Access, vol. 9, pp. 133500–133509, 2021, doi: 10.1109/access.2021.3115148. DOI: https://doi.org/10.1109/ACCESS.2021.3115148
W. I. Al-Obaydy, H. Hashim, Y. A. Najm, and A. Jalal, “Document classification using term frequency-inverse document frequency and K-means clustering,” Indones. J. Electr. Eng. Comput. Sci., 2022, doi: 10.11591/ijeecs.v27.i3.pp1517-1524. DOI: https://doi.org/10.11591/ijeecs.v27.i3.pp1517-1524
R. C. Morales-hernández and D. Becerra-Alonso, “A Comparison of Multi-Label Text Classification Models in Research Articles Labeled With Sustainable Development Goals,” IEEE Access, vol. 10, no. November, 2022. DOI: https://doi.org/10.1109/ACCESS.2022.3223094
S. Rahman, H. K. Shanto, U. A. Koana, and S. M. Danish, “Automated Research Article Classification and Recommendation Using NLP and Machine Learning,” arXiv, vol. 1, 2025. DOI: https://doi.org/10.1109/FLLM67465.2025.11391003
A. M. Chaid, Z. A. Abdulrazzaq, R. N. Sadoon, and M. A. Aljabery, “Comparative Analysis of Innovative Machine Learning Algorithms : Advancements in Natural Language Processing,” J. Inf. Syst. Eng. Manag., vol. 10, 2025. DOI: https://doi.org/10.52783/jisem.v10i14s.2373
N. Susetyo, F. Putri, A. Prasetya, H. Ar, and A. Nur, “Classification of Engineering Journals Quartile using Various Supervised Learning Models,” Ilk. J. Ilm., vol. 15, no. 1, pp. 101–106, 2023. DOI: https://doi.org/10.33096/ilkom.v15i1.1483.101-106
R. Guns, “Fine-grained classification of social science journal articles using textual data : A comparison of supervised machine learning approaches,” Quant. Sci. Stud., vol. 2, no. 1, 2021, doi: 10.1162/qss. DOI: https://doi.org/10.1162/qss_a_00106
M. Lu, L. Tang, and X. Zhou, “An ensemble approach for research article classi fi cation : a case study in arti fi cial intelligence,” PeerJ Comput. Sci., pp. 1–22, 2024, doi: 10.7717/peerj-cs.2521. DOI: https://doi.org/10.7717/peerj-cs.2521
L. P. F. Garcia, J. P. Mena-Chalco, R. L. Grando, W. M. C. da Silva, F. P. Guimarães, and V. de A. Jorge, “Hierarchical Article Classification: A Multi-Level Framework for Organizing Scholarly Literature,” IEEE Access, vol. 13, no. May, pp. 102589–102601, 2025, doi: 10.1109/ACCESS.2025.3579232. DOI: https://doi.org/10.1109/ACCESS.2025.3579232
T. Gupta, M. Zaki, N. M. A. Krishnan, and Mausam, “MatSciBERT: A materials domain language model for text mining and information extraction,” npj Comput. Mater., vol. 8, no. 1, p. 102, 2022, doi: 10.1038/s41524-022-00784-w. DOI: https://doi.org/10.1038/s41524-022-00784-w
S. Aum and S. Choe, “srBERT : automatic article classification model for systematic review using BERT,” Syst. Rev., pp. 1–8, 2021, doi: 10.1186/s13643-021-01763-w. DOI: https://doi.org/10.1186/s13643-021-01763-w
C. Arhiliuc, R. Guns, W. Daelemans, and T. C. E. Engels, “Journal article classification using abstracts: a comparison of classical and transformer-based machine learning methods,” Scientometrics, vol. 130, no. 1, pp. 313–342, 2025, doi: 10.1007/s11192-024-05217-7. DOI: https://doi.org/10.1007/s11192-024-05217-7
F. Aguilar-canto, C. Macias, A. E. Juárez, M. A. Cardoso-moreno, and H. Calvo, “Quartile Prediction and Journal Recommendation Using Deep Learning Models for Artificial Intelligence Articles,” J. Scientometr. Res., vol. 14, no. 1, pp. 373–382, 2025, doi: 10.5530/jscires.20251460. DOI: https://doi.org/10.5530/jscires.20251460
M. Sciences, “Exploring Preprocessing Techniques for Natural Language Text : A Comprehensive Study sing Python Code,” Int. J. Eng. Technol. Manag. Sci., vol. 5, no. 7, pp. 390–399, 2023, doi: 10.46647/ijetms.2023.v07i05.047. DOI: https://doi.org/10.46647/ijetms.2023.v07i05.047
W. Costa and G. V. Pedrosa, A Textual Representation Based on Bag-of-Concepts and Thesaurus for Legal Information Retrieval. 2022. doi: 10.5753/kdmile.2022.227779. DOI: https://doi.org/10.5753/kdmile.2022.227779
D. Sugiarto, E. Utami, and A. Yaqin, “Perbandingan Kinerja Model TF-IDF dan BOW untuk Klasifikasi Opini Publik Tentang Kebijakan BLT Minyak Goreng,” vol. 12, no. 3, pp. 272–277, 2022. DOI: https://doi.org/10.25105/jti.v12i3.15669
K. Juluru, H.-H. Shih, K. N. Keshava Murthy, and P. Elnajjar, “Bag-of-Words Technique in Natural Language Processing: A Primer for Radiologists,” RadioGraphics, vol. 41, no. 5, pp. 1420–1426, Aug. 2021, doi: 10.1148/rg.2021210025. DOI: https://doi.org/10.1148/rg.2021210025
Z. Jiang, B. Gao, Y. He, Y. Han, P. Doyle, and Q. Zhu, “Text Classification Using Novel Term Weighting Scheme-Based Improved TF-IDF for Internet Media Reports,” Math. Probl. Eng., vol. 2021, no. ii, 2021, doi: 10.1155/2021/6619088. DOI: https://doi.org/10.1155/2021/6619088
R. Blanquero, E. Carrizosa, and P. Ramírez-cobo, “Computers and Operations Research Variable selection for Naïve Bayes classification,” Comput. Oper. Res., vol. 135, p. 105456, 2021, doi: 10.1016/j.cor.2021.105456. DOI: https://doi.org/10.1016/j.cor.2021.105456
T. Simbolon, A. Prasetya, I. Ari, E. Zaeni, and A. Ritahani, “Text classification of traditional and national songs using naïve bayes algorithm,” Sci. Inf. Technol. Lett., vol. 3, no. 2, pp. 59–72, 2022. DOI: https://doi.org/10.31763/sitech.v3i2.1215
D. M. Abdullah and A. M. Abdulazeez, “Machine Learning Applications based on SVM Classification : A Review,” Qubahan Acad. J., pp. 81–90, 2021, doi: 10.48161/Issn.2709-8206. DOI: https://doi.org/10.48161/qaj.v1n2a50
G. D. Bispo et al., “Applied Sciences Automatic Literature Mapping Selection : Classification of Papers on Industry Productivity,” Appl. Sci., vol. 14, no. 9, p. 3679, 2024, doi: https://doi.org/10.3390/app14093679. DOI: https://doi.org/10.3390/app14093679
A. Lyutov, Y. Uygun, and M. T. Hütt, “Machine learning misclassification networks reveal a citation advantage of interdisciplinary publications only in high-impact journals,” Sci. Rep., vol. 14, no. 1, pp. 1–11, 2024, doi: 10.1038/s41598-024-72364-5. DOI: https://doi.org/10.1038/s41598-024-72364-5
Published
How to Cite
Issue
Section
Categories
Copyright (c) 2026 Ainunna’imah, Herman Yuliansyah, Imam Riadi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.












