Adversarial Machine Learning: Detecting and Mitigating Evasion Attacks on Intrusion Detection Systems
Main Article Content
Abstract
Network Intrusion Detection Systems (NIDS) increasingly rely on machine learning (ML) to identify malicious traffic that signature-based tools cannot recognize. This dependence, however, introduces a new attack surface: adversarial evasion, in which an attacker crafts subtly perturbed network traffic that is misclassified as benign while preserving its malicious functionality. This paper presents a systematic study of evasion attacks against ML-based NIDS and proposes a layered detection-and-mitigation framework that combines adversarial training, input transformation defences, ensemble disagreement analysis, and statistical drift monitoring. We formalize the threat model across white-box, gray-box, and black-box attacker knowledge levels; implement and evaluate four representative evasion techniques (Fast Gradient Sign Method, Projected Gradient Descent, a genetic-algorithm-based black-box attack, and a GAN-based traffic synthesizer) against a Random Forest, a Multi-Layer Perceptron, and an XGBoost classifier trained on the CIC-IDS2017 and NSL-KDD datasets; and measure the effectiveness of our proposed defences. Our results show that undefended ML-based NIDS suffer detection-rate degradation of 38–71% under evasion attacks, while the proposed layered defence recovers between 61% and 84% of the lost detection accuracy with an acceptable false-positive overhead of under 3 percentage points. We conclude with a discussion of the fundamental trade-offs between robustness, accuracy, and computational cost, and outline directions for future research, including certified robustness for tabular network-flow data and adversarial robust federated intrusion detection
Downloads
Article Details
Section

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.