张
内蒙古大学
毓
内蒙古大学
杨
内蒙古大学
涛
内蒙古大学
姬杨帆
内蒙古大学
顾知义
Abstract
目的是验证仅利用邮件头信息检测垃圾邮件和钓鱼邮件的可行性,并比较监督与无监督模型性能。方法 基于 TREC 2007 和 Monkey.org 数据集,从 76707 封邮件中提取缺失字段、字符串模式、Received 头一致性、域名匹配等 94 项特征,构建 10 种监督模型、单类支持向量机和孤立森林,以准确率、精确率、召回率、F1 值和 AUC 评价性能。结果 垃圾邮件任务中,堆叠集成准确率和 F1 值分别为 .9971 和 .9978,随机森林分别为 .9966 和 .9975;多种监督模型在钓鱼邮件任务中的准确率达到 1.0000。无监督模型中,孤立森林检测垃圾邮件的准确率为 .7592,单类支持向量机检测钓鱼邮件的准确率为 .7403。域名一致性和关键头字段缺失是主要判别特征。结论 邮件头特征能够有效支持异常邮件检测,监督模型整体优于无监督模型,但仍需通过跨来源、跨时间数据验证泛化能力。
Keywords
- 异常邮件检测;邮件头;机器学习;集成学习;无监督学习
Preview
References
- [1] Anti-Phishing Working Group. Phishing Activity Trends Report, 4th Quarter 2022[EB/OL]. https://docs.apwg. org/reports/apwg_trends_report_q4_2022.pdf, 2023-05-10.
- [2] Beaman, C. and Isah, H. (2022) Anomaly Detection in Emails Using Machine Learning and Header Information[EB/OL]. https://arxiv.org/abs/2203.10408, 2022-03-19.
- [3] Cormack, G.V. (2007) TREC 2007 Spam Track Overview. In: Voorhees, E.M. and Buckland, L.P., Eds.,
- Proceedings of the Sixteenth Text REtrieval Conference, National Institute of Standards and Technology, Gaithersburg.
- [4] Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., et al. (2011) Scikit-Learn: Machine Learning in Python. Journal of Machine Learning Research, 12,2825-2830.
- [5] Breiman, L. (2001) Random Forests. Machine Learning, 45, 5-32.
- [6] Wolpert, D.H. (1992) Stacked Generalization. Neural Networks, 5, 241-259.
- [7] Schölkopf, B., Platt, J.C., Shawe-Taylor, J.C., Smola, A.J. and Williamson, R.C. (2001) Estimating the Support of a HighDimensional Distribution. Neural Computation, 13, 1443-1471.
- [8] Liu, F.T., Ting, K.M. and Zhou, Z.H. (2008) Isolation Forest. Proceedings of the 2008 IEEE International Conference on Data Mining, IEEE, Pisa, 413-422.