Improving BiLSTM News Articles Classification with Frozen DistilBERT Embedding
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.5175Keywords:
DistilBERT, BiLSTM, text classification, embeddings, news categorizationAbstract
News classification is critical for organizing digital content and enhancing user engagement. Static embeddings such as GloVe, however, often fail to capture dynamic contextual relationships, limiting classifier performance. This study investigates the integration of frozen DistilBERT embeddings used strictly as feature extractors into a BiLSTM news‐classification pipeline to harness context‐aware representations while preserving computational efficiency. Experiments were conduccted on the Fancyzhx/ag_news dataset, which contains 127,700 samples, using a controlled 70/30 train–test split, comparing three BiLSTM variants: baseline (no pretrained embeddings), GloVe‐BiLSTM, and DistilBERT‐BiLSTM. Model architectures and hyperparameters are held constant to ensure a fair evaluation. Experimental results demonstrate that DistilBERT‐BiLSTM achieves 93.2 % accuracy, outperforming GloVe‐BiLSTM by 1.3% and the baseline by 2.5%. UMAP visualizations reveal more distinct semantic clusters with DistilBERT embeddings, and token‐level heatmaps confirm sharper intra‐sentence focus on domain‐specific terms. These findings contribute to informatics research by demonstrating the value of frozen transformer embeddings for lightweight yet high-performing text classification, with practical applications in information retrieval, content moderation, and real-time news analytics.
Downloads
References
V. Dogra et al., “A Complete Process of Text Classification System Using State-of-the-Art NLP Models,” Comput. Intell. Neurosci., vol. 2022, no. 1, p. 1883698, 2022, doi: 10.1155/2022/1883698.
M. Kayakuş and F. Y. Açıkgöz, “Classification of News Texts by Categories Using Machine Learning Methods,” Alphanumeric J., vol. 10, no. 2, Art. no. 2, Dec. 2022, doi: 10.17093/alphanumeric.1149753.
J. Shobana and M. Murali, “An Improved Self Attention Mechanism Based on Optimized BERT-BiLSTM Model for Accurate Polarity Prediction,” Comput. J., vol. 66, no. 5, pp. 1279–1294, May 2023, doi: 10.1093/comjnl/bxac013.
R. Zamith and O. Westlund, “Digital Journalism and Epistemologies of News Production,” in Oxford Research Encyclopedia of Communication, 2022. doi: 10.1093/acrefore/9780190228613.013.84.
F. M. Simon, “Escape Me If You Can: How AI Reshapes News Organisations’ Dependency on Platform Companies,” Digit. Journal., vol. 12, no. 2, pp. 149–170, Feb. 2024, doi: 10.1080/21670811.2023.2287464.
H. Allam, L. Makubvure, B. Gyamfi, K. N. Graham, and K. Akinwolere, “Text Classification: How Machine Learning Is Revolutionizing Text Categorization,” Information, vol. 16, no. 2, Art. no. 2, Feb. 2025, doi: 10.3390/info16020130.
X. Li, Y. Lei, and S. Ji, “BERT- and BiLSTM-Based Sentiment Analysis of Online Chinese Buzzwords,” Future Internet, vol. 14, no. 11, Art. no. 11, Nov. 2022, doi: 10.3390/fi14110332.
B. Liu, J. Chen, R. Wang, J. Huang, Y. Luo, and J. Wei, “Optimizing News Text Classification with Bi-LSTM and Attention Mechanism for Efficient Data Processing,” Sep. 23, 2024, arXiv: arXiv:2409.15576. doi: 10.48550/arXiv.2409.15576.
U. B. Mahadevaswamy and P. Swathi, “Sentiment Analysis using Bidirectional LSTM Network,” Procedia Comput. Sci., vol. 218, pp. 45–56, Jan. 2023, doi: 10.1016/j.procs.2022.12.400.
A. Glenn, P. LaCasse, and B. Cox, “Emotion classification of Indonesian Tweets using Bidirectional LSTM,” Neural Comput. Appl., vol. 35, no. 13, pp. 9567–9578, May 2023, doi: 10.1007/s00521-022-08186-1.
H. Padalko, V. Chomko, and D. Chumachenko, “A novel approach to fake news classification using LSTM-based deep learning models,” Front. Big Data, vol. 6, Jan. 2024, doi: 10.3389/fdata.2023.1320800.
W. Wang, “Textual Information Classification of Campus Network Public Opinion Based on BILSTM and ARIMA,” Wirel. Commun. Mob. Comput., vol. 2022, no. 1, p. 8323083, 2022, doi: 10.1155/2022/8323083.
J. Jasmir, W. Riyadi, S. R. Agustini, Y. Arvita, D. Meisak, and L. Aryani, “Bidirectional Long Short-Term Memory and Word Embedding Feature for Improvement Classification of Cancer Clinical Trial Document,” J. RESTI Rekayasa Sist. Dan Teknol. Inf., vol. 6, no. 4, Art. no. 4, Aug. 2022, doi: 10.29207/resti.v6i4.4005.
Q. Liu, M. J. Kusner, and P. Blunsom, “A Survey on Contextual Embeddings,” Apr. 13, 2020, arXiv: arXiv:2003.07278. doi: 10.48550/arXiv.2003.07278.
S. K. Rongali, “Natural Language Processing (NLP) in Artificial Intelligence,” World J. Adv. Res. Rev., vol. 25, no. 1, pp. 1931–1935, 2025, doi: 10.30574/wjarr.2025.25.1.0275.
Z. Gou and Y. Li, “Integrating BERT Embeddings and BiLSTM for Emotion Analysis of Dialogue,” Comput. Intell. Neurosci., vol. 2023, no. 1, p. 6618452, 2023, doi: 10.1155/2023/6618452.
Y. Yuan, S. Lv, Z. Bao, and K. Li, “A Joint Model for Text Classification with BERT-BiLSTM and GCN,” in Proceedings of the 2022 5th International Conference on Artificial Intelligence and Pattern Recognition, in AIPR ’22. New York, NY, USA: Association for Computing Machinery, May 2023, pp. 180–186. doi: 10.1145/3573942.3573970.
J. Ive, “Natural Language Processing: A Machine Learning Perspective by Yue Zhang and Zhiyang Teng,” Comput. Linguist., vol. 48, no. 1, pp. 233–235, Apr. 2022, doi: 10.1162/coli_r_00423.
X. Wen and W. Li, “Time Series Prediction Based on LSTM-Attention-LSTM Model,” IEEE Access, vol. 11, pp. 48322–48331, 2023, doi: 10.1109/ACCESS.2023.3276628.
S. Akpatsa et al., “Online News Sentiment Classification Using DistilBERT,” J. Quantum Comput., vol. 4, no. 1, pp. 1–11, 2022, doi: 10.32604/jqc.2022.026658.
Y. Jing and Y. Liu, “A Population-based Plagiarism Detection using DistilBERT-Generated Word Embedding,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 8, 2023, doi: 10.14569/IJACSA.2023.0140868.
T. Hegde, K. S. Sanjay, S. M. Thomas, R. Kambhammettu, M. Anand Kumar, and S. Ramanna, “Impact of Vector Embeddings on the Performance of Tolerance Near Sets-based Sentiment Classifier for Text Classification,” Procedia Comput. Sci., vol. 225, pp. 645–654, Jan. 2023, doi: 10.1016/j.procs.2023.10.050.
A. Qazi, R. H. Goudar, R. Patil, G. S. Hukkeri, and D. Kulkarni, “Leveraging BERT, DistilBERT, and TinyBERT for Rumor Detection,” IEEE Access, vol. 13, pp. 72918–72929, 2025, doi: 10.1109/ACCESS.2025.3563301.
A. Abbas, M. Lee, N. Shanavas, and V. Kovatchev, “Clinical concept annotation with contextual word embedding in active transfer learning environment,” Digit. Health, vol. 10, p. 20552076241308987, Sep. 2024, doi: 10.1177/20552076241308987.
Q. H. Nguyen et al., “Influence of Data Splitting on Performance of Machine Learning Models in Prediction of Shear Strength of Soil,” Math. Probl. Eng., vol. 2021, no. 1, p. 4832864, 2021, doi: 10.1155/2021/4832864.
A. Gasparetto, M. Marcuzzo, A. Zangari, and A. Albarelli, “A Survey on Text Classification Algorithms: From Text to Predictions,” Information, vol. 13, no. 2, Art. no. 2, Feb. 2022, doi: 10.3390/info13020083.
E. C. Garrido-Merchan, R. Gozalo-Brizuela, and S. Gonzalez-Carvajal, “Comparing BERT Against Traditional Machine Learning Models in Text Classification,” J. Comput. Cogn. Eng., vol. 2, no. 4, Art. no. 4, 2023, doi: 10.47852/bonviewJCCE3202838.
A. Kurniasih and L. P. Manik, “On the Role of Text Preprocessing in BERT Embedding-based DNNs for Classifying Informal Texts,” Int. J. Adv. Comput. Sci. Appl. IJACSA, vol. 13, no. 6, Art. no. 6, 30 2022, doi: 10.14569/IJACSA.2022.01306109.
S. J. Mielke et al., “Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP,” Dec. 20, 2021, arXiv: arXiv:2112.10508. doi: 10.48550/arXiv.2112.10508.
P. M. Subhash, K. C.r., D. Gupta, and V. kanjirangat, “Indo-Aryan Dialect Identification Using Deep Learning Ensemble Model,” Procedia Comput. Sci., vol. 235, pp. 2886–2896, Jan. 2024, doi: 10.1016/j.procs.2024.04.273.
X. Song, A. Salcianu, Y. Song, D. Dopson, and D. Zhou, “Fast WordPiece Tokenization,” Oct. 05, 2021, arXiv: arXiv:2012.15524. doi: 10.48550/arXiv.2012.15524.
J. Camacho-Collados and M. T. Pilehvar, “Embeddings in Natural Language Processing,” in Proceedings of the 28th International Conference on Computational Linguistics: Tutorial Abstracts, L. Specia and D. Beck, Eds., Barcelona, Spain (Online): International Committee for Computational Linguistics, Dec. 2020, pp. 10–15. doi: 10.18653/v1/2020.coling-tutorials.2.
M. Kamyab, G. Liu, and M. Adjeisah, “Attention-Based CNN and Bi-LSTM Model Based on TF-IDF and GloVe Word Embedding for Sentiment Analysis,” Appl. Sci., vol. 11, no. 23, Art. no. 23, Jan. 2021, doi: 10.3390/app112311255.
A. Bahaa, A. E.-R. Kamal, H. Fahmy, and A. S. Ghoneim, “DB-CBIL: A DistilBert-Based Transformer Hybrid Model Using CNN and BiLSTM for Software Vulnerability Detection,” IEEE Access, vol. 12, pp. 64446–64460, 2024, doi: 10.1109/ACCESS.2024.3396410.
D. Li, Y. Li, and Z. Zhang, “Automatic Extractive Summarization using GAN Boosted by DistilBERT Word Embedding and Transductive Learning,” Int. J. Adv. Comput. Sci. Appl. IJACSA, vol. 14, no. 11, Art. no. 11, 2023, doi: 10.14569/IJACSA.2023.0141107.
Md. A. Istiake Sunny, M. M. S. Maswood, and A. G. Alharbi, “Deep Learning-Based Stock Price Prediction Using LSTM and Bi-Directional LSTM Model,” in 2020 2nd Novel Intelligent and Leading Emerging Sciences Conference (NILES), Oct. 2020, pp. 87–92. doi: 10.1109/NILES50944.2020.9257950.
B. Jang, M. Kim, G. Harerimana, S. Kang, and J. W. Kim, “Bi-LSTM Model to Increase Accuracy in Text Classification: Combining Word2vec CNN and Attention Mechanism,” Appl. Sci., vol. 10, no. 17, Art. no. 17, Jan. 2020, doi: 10.3390/app10175841.
S. Selva Birunda and R. Kanniga Devi, “A Review on Word Embedding Techniques for Text Classification,” in Innovative Data Communication Technologies and Application, J. S. Raj, A. M. Iliyasu, R. Bestak, and Z. A. Baig, Eds., Singapore: Springer, 2021, pp. 267–281. doi: 10.1007/978-981-15-9651-3_23.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Hans Fanolo Kristian Daeli, Jasmir Jasmir, Nurhadi Nurhadi

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



