Details
- Supervisors
- Faculty
- Degree label
- Abstract
- The dead internet theory states that the majority of internet traffic is generated by bots instead of humans. With the advent of AI, this theory has become reality. Myriads of bots are sent accross the web to crawl and scrape to feed AI datasets. This has become a massive problem for website owners who see their content stolen and server load increase drastically. Bot detection and protection has been a topic of research as long as bot traffic has existed but no single miracle solution has yet been found. From Honeypots and Captchas to Fingerprinting and obfuscation, each known method has advantages and disadvantages when it comes to their efficiency, impact on normal users and uses in specific use-cases. In this paper, we propose a single-request on-the-fly bot detection method using machine learning with a K-Nearest Neighbors model and a Random Forest model trained on HTTP requests from Honeypots and logs from the University of Louvain-la-Neuve. Our objective for the machine learning is to outperform rule-based methods with the goal of detecting bots to reduce server load on the University's network. To that end, we start by reviewing the research that has been performed until now in order to gauge the efficiency of existing methods and understand what goes into designing one. After that, we build Honeypots to catch bot traffic, perform an analysis on the logs to design a rule-based benchmark and finally develop and test multiple machine learning detection system options. In the end, we are able to filter bots from humans from a single HTTP request with better accuracy than rule-based systems but holes in the training data still has an impact on our results.