Enhancing Big Five Personality Prediction on Indonesian Tweets: A Comparative Study of LSTM-GRU Hybrid and IndoBERT/IndoXLNet with Augmentation Techniques
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.5444Keywords:
Big Five Personality, Data Augmentation, Feature Expansion, IndoBERT, IndoXLNet, LSTM-GRUAbstract
Over the past few years, researches on Big Five personality prediction model focusing on Indonesian tweets has mostly faced data scarcity and imbalance problem. This study aims to build Big Five personality prediction models for the low-resource language setting using Indonesian social media texts in platform X, and conduct a comparative performance analysis using deep learning model (hybrid BiLSTM-BiGRU) and transferred learning models (IndoBERT and IndoXLNet) with the optimization using various data augmentation and feature expansion techniques to overcome data scarcity and imbalance. We carried out data collection of 11746 tweets from hybrid sources (BFI questionnaire and keyword-based crawling), preprocessing using tokenization and stemming, extraction using Word2Vec and FastText, data augmentation (using SYN-REP, BACK-TRANS, RAND-INS, and SMOTE variants), Top-N feature expansion (with N = 1, 5, and 10), hyperparameter tuning using Optuna (for learning rate, batch size, dropout rate, and weight decay), and evaluating the model's accuracy, precision, recall, and F1-score. The result shows that IndoBERTweet+BACK-TRANS+Top-10 achieved 84% accuracy and an F1-score of 81.5%, surpassing IndoXLNet+BACKTRANS+Top-5 (82%) and LSTM+GRU+FastText+KMeansSMOTE+Top-5 (78%), with the BACK-TRANS technique improves overall F1 result by 13%. These findings bolster computational psychology informatics for low-resource NLP focusing on Indonesian language, as well as aiding mental health screening with 18% better generalization in imbalanced data, and enabling scalable trait inference in human resource and mental health applications.
Downloads
References
W. Mahnaz and D. S. Kiran, “Big Five Personality Traits and Social Network Sites Preferences: The Mediating Role of Academic Achievement in Educational Outcomes of Secondary School Students,” Social Science Review Archives, vol. 2, no. 2, pp. 1353–1370, Dec. 2024, doi: 10.70670/sra.v2i2.187.
L. R. Goldberg, “An alternative ‘description of personality’: The Big-Five factor structure,” J. Pers. Soc. Psychol., vol. 59, no. 6, pp. 1216–1229, Feb. 1990, doi: https://doi.org/10.1037/0022-3514.59.6.1216.
N. Natasya, I. Ummah, and S. Widowati, “Comparative Performance Analysis of Big Five Personality Prediction on X Platform Using BERT and XLNet Models,” in 2025 International Conference on Information and Communication Technology (ICoICT), 2025, pp. 1–7. doi: 10.1109/ICoICT66265.2025.11193100.
Z. Chi, H. Huang, L. Liu, Y. Bai, X. Gao, and X.-L. Mao, “Can Pretrained English Language Models Benefit Non-English NLP Systems in Low-Resource Scenarios?,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 32, pp. 1061–1074, 2024, doi: 10.1109/TASLP.2023.3267618.
A. M. Dos Santos et al., “Words Similarities on Personalities: A Language-Based Generalization Approach for Personality Factors Recognition,” IEEE Access, vol. 11, pp. 29823–29836, 2023, doi: 10.1109/ACCESS.2023.3261339.
A. Arora and E. Turcan, “Evaluating the Effectiveness of Data Augmentation for Emotion Classification in Low-Resource Settings,” 2024. [Online]. Available: https://arxiv.org/abs/2406.05190
J. M. Tapia-Téllez and H. J. Escalante, “Data Augmentation with Transformers for Text Classification,” in Advances in Computational Intelligence, L. Martínez-Villaseñor, O. Herrera-Alcántara, H. Ponce, and F. A. Castro-Espinoza, Eds., Cham: Springer International Publishing, 2020, pp. 247–259. doi: 10.1007/978-3-030-60887-3_22.
S. R. Nudin, R. Gernowo, M. Somantri, and A. Wibowo, “Multi Task Classification Using Deep Learning Approaches for Big Five Personality Traits Prediction: A Review,” in 2024 11th International Conference on Information Technology, Computer, and Electrical Engineering (ICITACEE), 2024, pp. 37–42. doi: 10.1109/ICITACEE62763.2024.10762804.
M. C. Aparna and M. N. Nachappa, “Personality Traits Classification from Text Using Deep Learning Model,” in 2022 IEEE Conference on Interdisciplinary Approaches in Technology and Management for Social Innovation (IATMSI), 2022, pp. 1–6. doi: 10.1109/IATMSI56455.2022.10119306.
V. Ong et al., “Personality prediction based on Twitter information in Bahasa Indonesia,” in 2017 Federated Conference on Computer Science and Information Systems (FedCSIS), 2017, pp. 367–372. doi: 10.15439/2017F359.
N. H. Jeremy, C. Prasetyo, and and D. Suhartono, “Identifying Personality Traits for Indonesian User from Twitter Dataset,” The International Journal of Fuzzy Logic and Intelligent Systems, vol. 19, no. 4, pp. 283–289, Dec. 2019, doi: 10.5391/IJFIS.2019.19.4.283.
A. R. Julianda and W. Maharani, “Personality Detection on Reddit Using DistilBERT,” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 7, no. 5, pp. 1140–1146, Oct. 2023, doi: 10.29207/resti.v7i5.5236.
E. B. Setiawan, D. H. Widyantoro, and K. Surendro, “Feature Expansion for Sentiment Analysis in Twitter,” in 2018 5th International Conference on Electrical Engineering, Computer Science and Informatics (EECSI), 2018, pp. 509–513. doi: 10.1109/EECSI.2018.8752851.
A. R. Royyan and E. B. Setiawan, “Feature Expansion Word2Vec for Sentiment Analysis of Public Policy in Twitter,” Jurnal RESTI: Rekayasa Sistem dan Teknologi Informasi, vol. 6, no. 1, pp. 78–84, Feb. 2022, doi: 10.29207/resti.v6i1.3525.
I. W. A. Widiarta and E. B. Setiawan, “Depression Detection on Social Media X Using Hybrid Deep Learning CNN-BiGRU with Attention Mechanism and FastText Feature Expansion,” Jurnal Ilmiah Teknik Elektro Komputer dan Informatika, vol. 11, no. 2, pp. 290–304, Jun. 2025, [Online]. Available: https://journal.uad.ac.id/index.php/JITEKI/article/view/30687
L. Ayash, A. Algarni, and O. Alqahtani, “Advancements in feature selection and extraction methods for text mining: a review,” Discover Applied Sciences, vol. 7, no. 8, p. 914, 2025, doi: 10.1007/s42452-025-07587-w.
S. Sonkar, A. E. Waters, and R. G. Baraniuk, “Attention Word Embedding,” 2020. [Online]. Available: https://arxiv.org/abs/2006.00988
Y. Gunawan, J. C. Young, and A. Rusli, “FastText Word Embedding and Random Forest Classifier for User Feedback Sentiment Classification in Bahasa Indonesia,” Ultimatics : Jurnal Teknik Informatika, vol. 13, no. 2, pp. 101–107, Jan. 2022, doi: 10.31937/ti.v13i2.2124.
L. D. Cahya, A. Luthfiarta, J. I. T. Krisna, S. Winarno, and A. Nugraha, “Improving Multi-label Classification Performance on Imbalanced Datasets Through SMOTE Technique and Data Augmentation Using IndoBERT Model,” Jurnal Nasional Teknologi dan Sistem Informasi, vol. 9, no. 3, pp. 290–298, Jan. 2024, doi: 10.25077/TEKNOSI.v9i3.2023.290-298.
Z. Shi and A. Lipani, “Rethink the Effectiveness of Text Data Augmentation: An Empirical Analysis,” 2023. [Online]. Available: https://arxiv.org/abs/2306.07664
A. Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” 2019. [Online]. Available: https://arxiv.org/abs/1912.01703
W. Jian et al., “SA-Bi-LSTM: Self Attention With Bi-Directional LSTM-Based Intelligent Model for Accurate Fake News Detection to Ensured Information Integrity on Social Media Platforms,” IEEE Access, vol. 12, pp. 48436–48452, 2024, doi: 10.1109/ACCESS.2024.3382832.
Q. Li, L. Chen, C. Wang, X. Zhang, and J. Zhao, “Daily Load Prediction Study on SCASL and BiGRU models,” J. Phys. Conf. Ser., vol. 2591, no. 1, p. 012050, Sep. 2023, doi: 10.1088/1742-6596/2591/1/012050.
D. Lewy and J. Mańdziuk, “AttentionMix: Data augmentation method that relies on BERT attention mechanism,” 2023. [Online]. Available: https://arxiv.org/abs/2309.11104
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V Le, “XLNet: Generalized Autoregressive Pretraining for Language Understanding,” 2020. [Online]. Available: https://arxiv.org/abs/1906.08237
M. Khaldi, Z. Alilat, H. Bendoubba, and N. Mahammed, “Hyperparameter Optimization for Malicious URL Detection: Leveraging Optuna and Random Search in Machine Learning and Deep Learning Models,” INFORMATICA: International Journal of Computing and Informatics, vol. 49, no. 27, pp. 57–68, May 2025, doi: 10.31449/inf.v49i27.9106.
F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” 2020. [Online]. Available: https://arxiv.org/abs/2011.00677
F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” 2021. [Online]. Available: https://arxiv.org/abs/2109.04607
H. Q. Abonizio, E. C. Paraiso, and S. Barbon, “Toward Text Data Augmentation for Sentiment Analysis,” IEEE Transactions on Artificial Intelligence, vol. 3, no. 5, pp. 657–668, 2022, doi: 10.1109/TAI.2021.3114390.
M. Bayer, M.-A. Kaufhold, B. Buchhold, M. Keller, J. Dallmeyer, and C. Reuter, “Data augmentation in natural language processing: a novel text generation approach for long and short text classifiers,” International Journal of Machine Learning and Cybernetics, vol. 14, no. 1, pp. 135–150, 2023, doi: 10.1007/s13042-022-01553-3.
H. Lucky, Roslynlia, and D. Suhartono, “Towards Classification of Personality Prediction Model: A Combination of BERT Word Embedding and MLSMOTE,” in 2021 1st International Conference on Computer Science and Artificial Intelligence (ICCSAI), 2021, pp. 346–350. doi: 10.1109/ICCSAI53272.2021.9609750.
G. A. Pradnyana, W. Anggraeni, E. M. Yuniarno, and M. H. Purnomo, “Fine-Tuning IndoBERT Model for Big Five Personality Prediction from Indonesian Social Media,” in 2023 International Seminar on Intelligent Technology and Its Applications (ISITIA), 2023, pp. 93–98. doi: 10.1109/ISITIA59021.2023.10221074.
T. Pratama and Suharjito, “IndoXLNet: Pre-Trained Language Model for Bahasa Indonesia,” International Journal of Engineering Trends and Technology, vol. 70, no. 5, pp. 367–381, May 2022, doi: 10.14445/22315381/IJETT-V70I5P240.
B. Li, Y. Hou, and W. Che, “Data augmentation approaches in natural language processing: A survey,” AI Open, vol. 3, pp. 71–90, 2022, doi: https://doi.org/10.1016/j.aiopen.2022.03.001.
G. Chao, J. Liu, M. Wang, and D. Chu, “Data augmentation for sentiment classification with semantic preservation and diversity,” Know.-Based Syst., vol. 280, no. C, Nov. 2023, doi: 10.1016/j.knosys.2023.111038.
R. Xiang, E. Chersoni, Q. Lu, C.-R. Huang, W. Li, and Y. Long, “Lexical data augmentation for sentiment analysis,” J. Assoc. Inf. Sci. Technol., vol. 72, no. 11, pp. 1432–1447, 2021, doi: https://doi.org/10.1002/asi.24493.
M. Soleimani and H. B. Kashani, “Frozen or Fine-tuned? Analyzing Deep Learning Models and Training Strategies for Optimizing Big Five Personality Traits Prediction from Text,” in 2024 10th International Conference on Artificial Intelligence and Robotics (QICAR), 2024, pp. 11–16. doi: 10.1109/QICAR61538.2024.10496606.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Izzatul Ummah, Fitriyani, Novia Natasya

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



