Semantic-Preserving Hybrid Tokenization With Codebert For Robust Webshell Detection In Hypertext Preprocessor Source Code
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.6121Keywords:
CodeBERT, Cybersecurity, PHP Webshell Detection, Semantic Tokenization, Shortcut LearningAbstract
Transformer-based code models can achieve high accuracy in PHP(Hypertext Preprocessor) webshell detection, but high in-distribution scores do not necessarily indicate reliable behavior under difficult benign cases. This study investigates shortcut learning in CodeBERT-based PHP webshell classification and proposes a semantic-preserving hybrid representation that combines PHP structural tokens with explicit security-relevant tokens. After curation and SHA-256 exact deduplication, the corpus consists of 6,430 PHP files, including 3,685 benign files and 2,745 malicious webshell samples. The proposed tokenizer normalizes remote code execution sinks, input sources, obfuscation functions, and file operations into explicit tokens such as SINK_RCE_SYSTEM, INPUT_POST, OBF_BASE64, and FILE_INCLUDE. CodeBERT is fine-tuned under repository-aware and cross-partition robustness settings and evaluated using global metrics, confusion matrices, false-positive rates by benign bucket, and attention-based interpretation. Hybrid Tokens achieved an accuracy of 0.9902 and an F1-score of 0.9886 in repository-aware testing, while maintaining an accuracy of 0.9695 and an F1-score of 0.9629 in cross-partition robustness testing. The results indicate that preserving security semantics reduces reliance on superficial artifacts while retaining behavior-critical cues for distinguishing procedural benign scripts from webshells. These findings reframe PHP webshell detection as a robustness problem rather than a pure accuracy-maximization task, and show that security-aware input representation is a practical lever for reducing shortcut dependence. The study contributes to robust code intelligence, security-oriented representation learning, and more reliable evaluation of machine-learning detectors for cybersecurity.
Downloads
References
A. Hannousse, M. C. Nait-Hamoud, and S. Yahiouche, “A deep learner model for multi-language webshell detection,” Int. J. Inf. Secur., vol. 22, no. 1, pp. 47–61, 2023, doi: 10.1007/s10207-022-00615-5.
G. Y. Wang, H. J. Ko, C. P. Chiang, and W. J. Wang, “WebShell Detection Based on CodeBERT and Deep Learning Model,” in ACM International Conference Proceeding Series, Association for Computing Machinery, May 2024, pp. 484–489. doi: 10.1145/3670105.3670190.
Z. Wang, H. Wang, S. Yuan, and Z. Tian, “Incremental learning research for webshell detection,” Journal of King Saud University - Computer and Information Sciences, vol. 37, no. 8, Oct. 2025, doi: 10.1007/s44443-025-00235-8.
Z. Qing-peng, C. Jiang-li, and W. Jian-sheng, “A Webshell detection method based on feature fusion and federated learning,” Int. J. Inf. Secur., vol. 24, no. 4, p. 161, 2025, doi: 10.1007/s10207-025-01078-0.
B. Xie, Q. Li, and Y. Wang, “PHP-based malicious webshell detection based on abstract syntax tree simplification and explicit duration recurrent networks,” Comput. Secur., vol. 146, p. 104049, 2024, doi: https://doi.org/10.1016/j.cose.2024.104049.
C. Dong and D. Li, “AST-DF: A New Webshell Detection Method Based on Abstract Syntax Tree and Deep Forest,” Electronics (Switzerland), vol. 13, no. 8, Apr. 2024, doi: 10.3390/electronics13081482.
Z. Feng et al., “CodeBERT: A Pre-Trained Model for Programming and Natural Languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu, Eds., Online: Association for Computational Linguistics, Nov. 2020, pp. 1536–1547. doi: 10.18653/v1/2020.findings-emnlp.139.
D. Guo et al., “GraphCodeBERT: Pre-training Code Representations with Data Flow,” ArXiv, vol. abs/2009.08366, 2020, [Online]. Available: https://api.semanticscholar.org/CorpusID:221761146
Y. Wang, W. Wang, S. Joty, and S. C. H. Hoi, “CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, and S. W. Yih, Eds., Online and Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 8696–8708. doi: 10.18653/v1/2021.emnlp-main.685.
D. Guo, S. Lu, N. Duan, Y. Wang, M. Zhou, and J. Yin, “UniXcoder: Unified Cross-Modal Pre-training for Code Representation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio, Eds., Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 7212–7225. doi: 10.18653/v1/2022.acl-long.499.
R. Geirhos et al., “Shortcut learning in deep neural networks,” Nat. Mach. Intell., vol. 2, no. 11, pp. 665–673, Nov. 2020, doi: 10.1038/s42256-020-00257-z.
Y. Zhou, R. Tang, Z. Yao, and Z. Zhu, “Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models,” 2024. [Online]. Available: https://arxiv.org/abs/2409.17455
S. Chakraborty, R. Krishna, Y. Ding, and B. Ray, “Deep Learning based Vulnerability Detection: Are We There Yet?,” Sep. 2020, [Online]. Available: http://arxiv.org/abs/2009.07235
P. W. Koh et al., “WILDS: A Benchmark of in-the-Wild Distribution Shifts,” Jul. 2021, [Online]. Available: http://arxiv.org/abs/2012.07421
D. Hendrycks et al., “The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization,” Jul. 2021, [Online]. Available: http://arxiv.org/abs/2006.16241
H. Hanif and S. Maffeis, “VulBERTa: Simplified Source Code Pre-Training for Vulnerability Detection,” May 2022, doi: 10.1109/IJCNN55064.2022.9892280.
M. Fu and C. Tantithamthavorn, “LineVul: A Transformer-based Line-Level Vulnerability Prediction,” 2022, doi: 10.1145/3524842.
S. Lu et al., “CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation,” Mar. 2021, [Online]. Available: http://arxiv.org/abs/2102.04664
Y. Xia, H. Shao, and X. Deng, “VulCoBERT: A CodeBERT-Based System for Source Code Vulnerability Detection,” in ACM International Conference Proceeding Series, Association for Computing Machinery, May 2024, pp. 249–252. doi: 10.1145/3665348.3665391.
M. Minderer et al., “Revisiting the Calibration of Modern Neural Networks,” Oct. 2021, [Online]. Available: http://arxiv.org/abs/2106.07998
T. An, X. Shui, and H. Gao, “Deep Learning Based Webshell Detection Coping with Long Text and Lexical Ambiguity,” in Information and Communications Security, C. Alcaraz, L. Chen, S. Li, and P. Samarati, Eds., Cham: Springer International Publishing, 2022, pp. 438–457.
P. Feng et al., “GlareShell: Graph learning-based PHP webshell detection for web server of industrial internet,” Comput. Netw., vol. 245, no. C, May 2024, doi: 10.1016/j.comnet.2024.110406.
Y. Zhang, H. Kang, and Q. Wang, “MMFDetect: Webshell Evasion Detect Method Based on Multimodal Feature Fusion,” Electronics (Switzerland), vol. 14, no. 3, Feb. 2025, doi: 10.3390/electronics14030416.
Z. Wang and F. Li, “GraphCodeBERT-TCN: A Hybrid Deep Learning Framework for PHP WebShell Detection,” 2025 International Conference on Advanced Computing and Intelligent Robotics Applications (ACIRA), pp. 442–446, 2025, [Online]. Available: https://api.semanticscholar.org/CorpusID:284924222
B. Cheng, Y. Guo, Y. Ren, G. Yang, and G. Xu, “MSDetector: A Static PHP Webshell Detection System Based on Deep-Learning,” in Theoretical Aspects of Software Engineering: 16th International Symposium, TASE 2022, Cluj-Napoca, Romania, July 8–10, 2022, Proceedings, Berlin, Heidelberg: Springer-Verlag, 2022, pp. 155–172. doi: 10.1007/978-3-031-10363-6_11.
X. Du, M. Wen, Z. Wei, S. Wang, and H. Jin, “An Extensive Study on Adversarial Attack against Pre-trained Models of Code,” in ESEC/FSE 2023 - Proceedings of the 31st ACM Joint Meeting European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Association for Computing Machinery, Inc, Nov. 2023, pp. 489–501. doi: 10.1145/3611643.3616356.
A. Jha and C. K. Reddy, “CodeAttack: Code-based Adversarial Attacks for Pre-Trained Programming Language Models,” 2023. [Online]. Available: www.aaai.org
Z. Yang, J. Shi, J. He, and D. Lo, “Natural Attack for Pre-trained Models of Code,” in Proceedings - International Conference on Software Engineering, IEEE Computer Society, Jul. 2022, pp. 1482–1493. doi: 10.1145/3510003.3510146.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Muhammad Kevin Adli Pratama, Muh Ghazy Daffa Sampe, Muhammad Ariando Ferdian, Putri Tendry Zahrany, Alif Rifa'i, Anindita Septiarini, Joan Angelina Widians, Akhmad Irsyad

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



