Before delving into the primary datasets, it's essential to grasp the significance of cybersecurity and why these datasets play a critical role in safeguarding our digital realm. In our interconnected world, cybersecurity threats pose substantial risks to individuals, enterprises, and governments. With the surge in cybercrimes, ranging from data breaches to cyberattacks, having access to trustworthy and current cybersecurity datasets is paramount. These datasets empower us to detect and thwart potential threats effectively.
The cybersecurity field is vast, encompassing a wide range of topics and challenges. The community provides a selection of datasets designed for specific research and analysis within the realm of cybersecurity. Here are some key datasets used in cybersecurity, along with brief descriptions and links (as for my study):
The DARPA Intrusion Detection Evaluation datasets were collected as part of the 1998 and 1999 DARPA intrusion detection evaluations. These datasets contain a variety of network traffic data for evaluating intrusion detection systems.
- 1998 DARPA INTRUSION DETECTION EVALUATION DATASET
- 1999 DARPA INTRUSION DETECTION EVALUATION DATASET
- 2000 DARPA INTRUSION DETECTION SCENARIO SPECIFIC DATASETS
The KDD Cup 1999 dataset is one of the earliest datasets used for intrusion detection research. It contains data from the DARPA Intrusion Detection Evaluation experiment.
The NSL-KDD dataset is a modified version of the well-known KDD Cup 1999 dataset, addressing issues such as redundancy and balance. The new dataset is reduced to the unique values and balanced representation of the different types of the described attacks.
- NSL-KDD Dataset
- Shortcut to downloads
- Kaggle version
- More from the Canadian Institute for Cybersecurity
The CTU-13 dataset is particularly notable for its comprehensive representation of botnet traffic, allowing researchers to analyze and develop detection methods for botnet-related activities. CTU-13 features traffic data related to a total of 7 distinct botnets, featuring a mix of both real-world botnet and user-simulated traffic in a university-like environment.
- CTU-13 Dataset
- Kaggle version
- More from Stratosphere IPS Datasets
The ISCXIDS2012 dataset consists of network traffic data, including both normal network traffic and various types of simulated and real-world cyberattacks.
- ISCXIDS2012 Dataset
- Shortcut to downloads
- More datasets from ISCX
The Canadian Institute for Cybersecurity Intrusion Detection Systems (CICIDS2017) dataset contains network traffic data specific to machine learning for intrusion detection system (IDS) research, describing various attack scenarios, such as DoS, DDoS, and port scanning.
The colaboorative project between the Communications Security Establishment (CSE) and the Canadian Institute for Cybersecurity (CIC) resulted in a comprehensive dataset, describing various attack scenarios, such as DoS, DDoS, and port scanning, that can be used for machine learning intrusion detection system (IDS) research.
aws s3 sync s3://cse-cic-ids2018/dir/ ./localdir
The colaboorative project between the Communications Security Establishment (CSE) and the Canadian Institute for Cybersecurity (CIC) resulted in a comprehensive dataset, describing various attack scenarios, such as DoS, DDoS, and port scanning, that can be used for machine learning intrusion detection system (IDS) research.
- CIDDS - COBURG INTRUSION DETECTION DATA SETS
- Github
- Shortcut to download CIDDS-001
- Shortcut to download CIDDS-002
- CIDDS-001 Kaggle version
- CIDDS-002 Kaggle version
The Kyoto (Kyoto 2006+) dataset, also known as the Kyoto University Honeypot Dataset, is a collection of network traffic data specifically designed for cybersecurity research, with a focus on honeypot-based intrusion detection and analysis.
- Kyoto 2006+ Dataset
- Kyoto Data
- Statistical analysis of honeypot data and building of Kyoto 2006+ dataset for NIDS evaluation
The Hornet datasets consist of a collection of data sets created to explore the potential influence of geographic factors on the occurrence of network attacks. This data was gathered during April and May 2021 from eight identically configured honeypot servers strategically positioned in various regions spanning North America, Europe, and Asia.
- Hornet: Network Dataset of Geographically Placed Honeypots
- Downlaod Hornet 7 Dataset
- Downlaod Hornet 15 Dataset
- Downlaod Hornet 40 Dataset
The UNSW-NB15 dataset is another network intrusion detection dataset containing diverse network traffic data. It includes normal traffic and a wide range of attacks, making it suitable for evaluating intrusion detection systems.
The MAWILab dataset consists of real-world traffic data collected from the MAWI (Monitoring and Analysis of Internet Wide-Area Network Traffic) project. It is used for anomaly detection and network traffic analysis.
This dataset is used for malware classification tasks. It contains a large collection of files, each labeled as benign or malicious, making it suitable for machine learning-based malware detection.
The AWID project aims to offer robust tools, methodologies, and datasets to help researchers create advanced security solutions for present and future wireless networks. The AWID2 dataset includes a large packet set (F) and a smaller one (R), focusing on WEP-based infrastructure with over 150 different attributes. AWID3 targets WPA2 Enterprise, 802.11w, and Wi-Fi 5, featuring multi-layer and contemporary attacks like Krack and Kr00k.
The H23Q dataset is a extensive, labeled 802.3 corpus features traces of ten different attacks targeting HTTP/2, HTTP/3, and QUIC services, including modern attacks specific to HTTP/3. The dataset, which is 30 GB in size, is accessible in both pcap and CSV formats.
The MTA-KDD'19 is a curated dataset designed for training and evaluating machine learning algorithms in malware traffic analysis. It was developed from extensive online network traffic databases, emphasizing relevant features while minimizing size and noise through cleaning and preprocessing. This dataset is versatile, not tailored to any particular application, and can be automatically updated to remain current.
This dataset is constructed using real traffic and contemporary attacks, sourced from various NetFlow v9 collectors positioned strategically within the network of a Spanish ISP. It consists of two distinct sets of data, each pre-divided on a weekly basis.
The Contagio Mobile Mini-Dump is an extensive collection of mobile malware samples, essential for cybersecurity research and the development of mobile security solutions. It provides a diverse range of real-world malware examples, aiding in understanding and combating mobile threats.
The MC2 - Computer Networking Operations, part of the VAST Challenge 2011, presents a scenario involving detailed analysis of intricate computer network operations. Participants are tasked with sifting through extensive data to identify unusual patterns, potential cyber threats, and anomalies in network behavior. This challenge simulates a real-world situation the dataset may be used for developing purposes.
The Software Assurance Reference Dataset (SARD) is a growing collection of test programs with documented weaknesses. Test cases vary from small synthetic programs to large applications. The programs are in C, C++, Java, PHP, and C#, and cover over 150 classes of weaknesses.
The Draper dataset consists of the source code of 1.27 million functions mined from open source software, labelled by static analysis for potential vulnerabilities. This dataset can be used for developing vulnerability prediction models.
DEFCON regularly releases cybersecurity-focused challenge datasets, engaging participants with diverse machine learning security problems. Their DEFCON30 event, hosted by the AI Village, exemplified this with a unique competition on Kaggle.
- DEFCON30
- DEFCON8
- DEFCON10
- DEFCON11
The CDX 2009 dataset, provided by the Cyber Research Center at West Point, captures network traffic and system logs from the 2009 Cyber Defense Exercise. It's a valuable resource for cybersecurity research, offering real-world data for studying network attacks, defense strategies, and system vulnerabilities. This dataset is especially useful for developing and testing intrusion detection systems and other security tools.
The dataset from the University of Twente focuses on flow-based intrusion detection in high-speed networks (1-10 Gbps). It was built on a honeypot connected directly to the Internet, ensuring data relevance. This dataset is ideal for tuning, training, and evaluating intrusion detection systems.
The ASNM-CDX-2009 dataset, part of the Advanced Security Network Metrics & CDX 2009 initiative, features ASNM features extracted from tcpdump captures of both malicious and legitimate TCP communications. It focuses on network services susceptible to buffer overflow attacks, providing a rich dataset for network traffic analysis.
The ADFA IDS Datasets, available through UNSW Research, are tailored for system call-based Host Intrusion Detection Systems (HIDS) evaluation. These datasets encompass both Linux and Windows environments, providing a comprehensive resource for developing and assessing HIDS capabilities.
This dataset consists of full-header network traffic from a medium-sized location, excluding payload. The dataset underwent extensive anonymization to eliminate any details that could reveal individual IP identities. Nevertheless, the LBNL consist of 100 hours of activity specifying the traces of packet header for identifying malicious traffic.
The UMass Network Datasets, available on the UMass Trace Repository website, are a comprehensive collection of network-related data, designed for a wide array of research applications. These datasets include detailed network traffic traces from a variety of sources and settings, offering valuable insights for studies in network security, performance analysis, and protocol development.
VirusTotal - Malware samples Phishtank - Phishing samples Kaggle - Challenges and public datasets Scispace - Find research papers relevant to your study
To download the files mentioned above, you may access the provided URL directly, or just call the below command.
wget --mirror -np -nH --cut-dirs=1 -P /path/to/save/directory <URL>
# Example
wget --mirror -np -nH --cut-dirs=1 -P Kyoto2006+ -N -r -A '*' https://www.takakura.com/Kyoto_data/new_data201704/
# `-r``: Recursively download files.
# `-A '*'`: Download files with any extension.
# `-N`: Download only files that are newer.Please note that the availability and specifics of these datasets may change over time, and it's important to review the dataset documentation and terms of use before using them for research or analysis.
In cybersecurity, the design, development, and implementation of effective Intrusion Detection Systems (IDS) are important for safeguarding IT&C infrastructures from unauthorized access, data breaches, and various forms of malicious activities. The selection of an appropriate ML/DL algorithm plays a essential role in ensuring the security and integrity of protected systems.
But before we can dive in the development of a new-edge algorithm, we shoud have the appropriate data, that needs to be studied and analysed in order to undestant the reality and challenges of our ML problem. In accordance with this paradigm, we chosed to study the early created datasets designed for IDS systems in order to derive leasons learn for feature dataset development.
This experiment aims to comprehensively evaluate the performance of different ML and DL algorithms on a variety of datasets, encompassing a wide range of network traffic scenarios. The datasets used for this analysis include well-known benchmark datasets such as KDD, NSL-KDD, CTU-13, ISCXIDS2012, CIC-IDS2017, CSE-CIC-IDS2018, CIDDS-001/CIDDS-002, and Kyoto 2006+. Each dataset represents a distinct set of challenges and characteristics, making this evaluation both diverse and insightful.
The experiment is divided into three main phases:
- Data Acquisition and Preprocessing:
- In this phase, we acquire the selected datasets from reputable sources, ensuring the integrity and accuracy of the data.
- Data preprocessing tasks include handling missing values, selecting the most relevant features using feature selection techniques, normalizing the data, and, if necessary, performing feature engineering to enhance the dataset's suitability for machine learning.
- Algorithm Evaluation:
- We evaluate the performance of a range of ML/DL algorithms on each dataset. The chosen algorithms include baseline methods like ZeroRule and OneRule, traditional machine learning approaches like Naive Bayes and Random Forest, as well as some of the most used anomaly detection deep learning algorithms.
- Cross-validation is applied to ensure the robustness of our results. Performance metrics such as precision, variance, and Mean Absolute Error (MAE) are calculated for each algorithm and dataset.
- Results and Insights:
- The results of this evaluation provide valuable insights into the strengths and weaknesses of different IDS algorithms under various conditions.
- We analyze the performance of algorithms on both the original datasets and balanced datasets to address the challenge of class imbalance in intrusion detection.
- Observations and additional details regarding the algorithms' performance are documented, providing a comprehensive overview of their behavior.
By conducting this experiment, we aim to contribute to the understanding of cyber domain dataset generation. The findings will assist in making informed decisions when developing a cybersecurity AI application, by deriving necesary steps and procedures in selecting the appropriate learning data.
The following Jupyter notebooks will provide a detailed walkthrough of the experiments, including code snippets, visualizations, and discussions of the results:
- KDD99-BM
- NSL-KDD-BM
- CTU-13-BM
- ISCXIDS2012-BM
- CIC-IDS2017-BM
- CSE-CIC-IDS2018-BM
- CIDDS-001-BM
- CIDDS-002-BM
- Kyoto2006+-BM
The results are also saved under the pickle files mentioned below:
- KDD99-BM
- NSL-KDD-BM
- CTU-13-BM
- ISCXIDS2012-BM
- CIC-IDS2017-BM
- CSE-CIC-IDS2018-BM
- CIDDS-001-BM
- CIDDS-002-BM
- Kyoto2006+-BM
We’d love your help to grow CY0P5_ML_Datasets!
If you find this project helpful, please share it with your friends, colleagues, or communities—on social media, newsletters, or blogs. Every mention helps us reach more people who might benefit from it.
You can also support the project directly and
—it means a lot and helps keep everything running!
Thank you for being part of our journey and for helping CY0P5_ML_Datasets thrive 🚀

