{
 "cells": [
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Imports assumed from previous context\n",
    "import numpy as np\n",
    "import pandas as pd\n",
    "import matplotlib.pyplot as plt\n",
    "import random\n",
    "\n",
    "# Note: This notebook requires the following libraries:\n",
    "!pip install numpy pandas matplotlib scikit-learn torch transformers xgboost shap imbalanced-learn scipy"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Machine Learning Essentials for Cybersecurity\n",
    "\n",
    "## Introduction\n",
    "Welcome to the core of modern intelligent defense systems. This chapter provides the essential machine learning (ML) knowledge that every cybersecurity professional needs. We will move past the buzzwords and build a solid foundation, starting with the clear distinctions between Artificial Intelligence, Machine Learning, and Deep Learning. You will learn about the main ML paradigms—supervised, unsupervised, semi-supervised, and reinforcement learning—and see how each one applies to real-world security problems like malware classification and anomaly detection.\n",
    "\n",
    "We will then explore the key families of ML models, from classical algorithms like Random Forest to advanced deep learning architectures like Convolutional and Recurrent Neural Networks. A critical focus of this chapter is on the practical aspects of building ML systems for security. This includes the crucial process of feature engineering, the art of selecting evaluation metrics that truly reflect a model's effectiveness in security contexts, and an honest look at the significant challenges you will face, such as imbalanced data and adversarial attacks. By the end of this chapter, you will have completed a hands-on project, building a basic threat classification pipeline from data loading to model evaluation, giving you a tangible and repeatable workflow for your own projects."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## AI, ML, and DL: Demystifying the Buzzwords\n",
    "\n",
    "To build intelligent defense systems, we must first understand the language of the field. The terms Artificial Intelligence (AI), Machine Learning (ML), and Deep Learning (DL) are often used interchangeably, but they represent distinct concepts with a clear hierarchical relationship.\n",
    "\n",
    "> **Figure:** The hierarchical relationship: DL is part of ML, and ML is part of AI.\n",
    "\n",
    "### Artificial Intelligence (AI)\n",
    "**Artificial Intelligence (AI)** is the broadest field. It is a branch of computer science dedicated to creating machines or systems that can perform tasks that normally require human intelligence. These tasks include capabilities like visual perception, speech recognition, decision-making, problem-solving, and understanding natural language. The ultimate goal of AI is to create general intelligence, but in practice, most current AI systems are narrow, meaning they are designed to perform a specific task very well.\n",
    "\n",
    "### Machine Learning (ML)\n",
    "**Machine Learning (ML)** is a subfield of AI. *It gives computers the ability to learn from data without being explicitly programmed for every possible scenario*. Instead of writing rules by hand, an ML algorithm analyzes a dataset to find patterns. It then uses these learned patterns to make predictions or decisions on new, unseen data. Over time, as it processes more data, the algorithm's performance on its specific task improves.\n",
    "\n",
    "### Deep Learning (DL)\n",
    "**Deep Learning (DL)** is a specialized subfield within machine learning. It uses a specific type of algorithm called an *artificial neural network*, which is inspired by the structure of the human brain. These networks have multiple layers of interconnected nodes, and the term \"deep\" refers to having many of these layers. Deep learning is especially powerful for finding very complex patterns in large datasets, particularly unstructured data like images, audio, and text.\n",
    "\n",
    "The relationship between these concepts is simple:\n",
    "* **AI** is the overall goal of creating intelligent machines.\n",
    "* **ML** is a primary method for achieving AI by learning from data.\n",
    "* **DL** is a powerful and popular type of ML that uses deep neural networks to achieve state-of-the-art results in many areas."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Core ML Paradigms: Your AI Toolset\n",
    "\n",
    "Machine learning is not a single technique but a collection of approaches, or paradigms, each suited for different types of problems and data. Understanding these core paradigms is fundamental to selecting the right tool for a cybersecurity task.\n",
    "\n",
    "## Supervised Learning\n",
    "Supervised learning is the most common ML paradigm. The algorithm learns from a dataset where each piece of input data is paired with a correct output label.\n",
    "\n",
    "Think of this like a teacher and a student. The teacher provides the student with practice exams (input data) along with the answer key (labels). The student's goal is to learn the logic needed to arrive at the correct answers so they can pass a future exam where the answers are hidden.\n",
    "\n",
    "> **Figure:** Supervised Machine Learning: Learning the mapping from Input ($X$) to Output ($y$).\n",
    "\n",
    "> **Figure:** Supervised Machine Learning System Overview.\n",
    "\n",
    "### Applications in Cybersecurity\n",
    "\n",
    "#### 1. Classification (Categorical Prediction)\n",
    "This is the task of predicting a discrete category.\n",
    "* **Binary Classification:** Two classes (e.g., \"Benign\" vs. \"Malware\", \"Spam\" vs. \"Ham\").\n",
    "* **Multiclass Classification:** More than two classes (e.g., identifying a malware family: \"Ransomware\", \"Trojan\", \"Spyware\", or \"Worm\").\n",
    "\n",
    "> **Figure:** ML Classification Overview.\n",
    "\n",
    "> **Figure:** Classification: Drawing a boundary to separate classes.\n",
    "\n",
    "The code below demonstrates a simple Phishing detector. The model looks at URL features to classify the URL as either Safe (0) or Phishing (1)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import numpy as np\n",
    "from sklearn.linear_model import LogisticRegression\n",
    "\n",
    "# --- CLASSIFICATION EXAMPLE: Phishing Detection ---\n",
    "\n",
    "# 1. Dataset (Features: [URL Length, Number of '@' symbols])\n",
    "# Long URLs and '@' symbols are suspicious.\n",
    "X_train = np.array([\n",
    "    [15, 0], [22, 0], [18, 0],  # Benign samples (Short, no '@')\n",
    "    [85, 1], [90, 2], [75, 1]   # Phishing samples (Long, has '@')\n",
    "])\n",
    "\n",
    "# Labels (0: Benign, 1: Phishing)\n",
    "y_train = np.array([0, 0, 0, 1, 1, 1])\n",
    "\n",
    "# 2. Train Model\n",
    "clf = LogisticRegression()\n",
    "clf.fit(X_train, y_train)\n",
    "\n",
    "# 3. Predict on new data\n",
    "new_email = np.array([[80, 1]]) # A long URL with an '@' symbol\n",
    "prediction = clf.predict(new_email)\n",
    "\n",
    "print(f\"Classification Prediction: {prediction[0]} (1=Phishing, 0=Benign)\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "#### 2. Regression (Numerical Prediction)\n",
    "This is the task of predicting a continuous numerical value. In cybersecurity, this is often used for risk assessment or forecasting.\n",
    "* **Examples:** Predicting a vulnerability's CVSS severity score (0.0 to 10.0), estimating the number of days until a patch is reverse-engineered, or predicting network throughput.\n",
    "\n",
    "> **Figure:** Regression Overview.\n",
    "\n",
    "> **Figure:** Regression: Fitting a line to predict numerical trends.\n",
    "\n",
    "The code below shows a model predicting a **Risk Score**. The model learns that as the number of unpatched vulnerabilities increases, the Risk Score goes up."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import numpy as np\n",
    "from sklearn.linear_model import LinearRegression\n",
    "\n",
    "# --- REGRESSION EXAMPLE: Risk Scoring ---\n",
    "\n",
    "# 1. Dataset (Feature: [Count of Open Critical Vulnerabilities])\n",
    "X_history = np.array([\n",
    "    [0], [1], [2], [5], [10]\n",
    "])\n",
    "\n",
    "# Labels (Target Risk Score: 0.0 to 100.0)\n",
    "# Notice the relationship: More vulnerabilities = Higher Score\n",
    "y_history = np.array([10.0, 25.0, 40.0, 65.0, 95.0])\n",
    "\n",
    "# 2. Train Model\n",
    "reg = LinearRegression()\n",
    "reg.fit(X_history, y_history)\n",
    "\n",
    "# 3. Predict Risk for a system with 7 open vulnerabilities\n",
    "new_system = np.array([[7]])\n",
    "predicted_score = reg.predict(new_system)\n",
    "\n",
    "print(f\"Regression Prediction (Risk Score): {predicted_score[0]:.2f}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Unsupervised Learning\n",
    "In contrast to supervised learning, unsupervised learning algorithms work with data that has no predefined labels.\n",
    "\n",
    "Think of this as a detective walking into a crime scene without knowing what crime occurred. They must look at the evidence, find connections, and identify oddities entirely on their own.\n",
    "\n",
    "> **Figure:** Unsupervised Learning: The model finds structure in unlabeled data.\n",
    "\n",
    "### Applications in Cybersecurity\n",
    "\n",
    "#### 1. Anomaly Detection\n",
    "This involves identifying data points that are significantly different from the norm. In security, this is the primary method for detecting *Zero-Day attacks* or *Insider Threats*, where the attack pattern is unknown, but the behavior is clearly \"abnormal.\"\n",
    "\n",
    "> **Figure:** Anomaly Detection: Identifying outliers that do not fit the standard distribution.\n",
    "\n",
    "The code below uses the *Isolation Forest* algorithm. It learns what \"normal\" login behavior looks like and flags an aggressive brute-force attempt as an anomaly."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.ensemble import IsolationForest\n",
    "import numpy as np\n",
    "\n",
    "# --- ANOMALY DETECTION: Insider Threat ---\n",
    "\n",
    "# 1. Training Data (Normal User Behavior)\n",
    "# Feature: [Number of File Accesses per Minute]\n",
    "# Most users access about 5-10 files.\n",
    "X_normal = np.array([[5], [6], [5], [7], [10], [8], [5]])\n",
    "\n",
    "# 2. Train Model\n",
    "# contamination=0.01 tells the model that anomalies are rare.\n",
    "iso = IsolationForest(contamination=0.01, random_state=42)\n",
    "iso.fit(X_normal)\n",
    "\n",
    "# 3. Live Data Monitoring\n",
    "# User A accesses 6 files (Normal)\n",
    "# User B accesses 1000 files (Data Exfiltration Attempt)\n",
    "X_live = np.array([[6], [1000]])\n",
    "\n",
    "predictions = iso.predict(X_live)\n",
    "# Output: 1 = Normal, -1 = Anomaly\n",
    "print(f\"Predictions: {predictions}\") \n",
    "# Expected: [1, -1] -> User B flagged as anomaly"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "#### 2. Clustering\n",
    "Clustering groups similar data points together based on their features. Security analysts use this to categorize new malware samples automatically or to group compromised hosts into botnets.\n",
    "\n",
    "> **Figure:** Clustering: Grouping unlabelled data based on similarity.\n",
    "\n",
    "The code below uses *K-Means* to separate network traffic into two distinct groups without being told what the groups are."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.cluster import KMeans\n",
    "import numpy as np\n",
    "\n",
    "# --- CLUSTERING: Traffic Categorization ---\n",
    "\n",
    "# 1. Unlabeled Data (Features: [Packet Size, Destination Port])\n",
    "# We have mixed traffic, but we don't know what is what.\n",
    "X_traffic = np.array([\n",
    "    [64, 53], [70, 53], [60, 53],    # Small packets on Port 53 (DNS)\n",
    "    [1500, 443], [1450, 443], [1500, 443] # Large packets on Port 443 (HTTPS)\n",
    "])\n",
    "\n",
    "# 2. Train Model\n",
    "# We ask the model to find 2 distinct groups (clusters)\n",
    "kmeans = KMeans(n_clusters=2, random_state=42)\n",
    "kmeans.fit(X_traffic)\n",
    "\n",
    "# 3. Analyze Results\n",
    "print(f\"Cluster Labels: {kmeans.labels_}\")\n",
    "# Expected: [0, 0, 0, 1, 1, 1] \n",
    "# The model successfully grouped DNS traffic separately from HTTPS traffic."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "#### 3. Dimensionality Reduction\n",
    "Security datasets can have hundreds of features (packet headers, flags, payload entropy), making them slow to process and hard to visualize. Dimensionality reduction compresses this data while keeping the important information.\n",
    "\n",
    "> **Figure:** Dimensionality Reduction: Compressing complex data into fewer dimensions.\n",
    "\n",
    "The code below uses *PCA (Principal Component Analysis)* to compress a dataset with 5 features down to just 2, making it possible to plot on a 2D graph."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.decomposition import PCA\n",
    "import numpy as np\n",
    "\n",
    "# --- DIMENSIONALITY REDUCTION: Log Compression ---\n",
    "\n",
    "# 1. High-Dimensional Data (5 Features per log entry)\n",
    "# Imagine features: [CPU, RAM, Disk, Net_In, Net_Out]\n",
    "X_logs = np.random.rand(10, 5) \n",
    "\n",
    "print(f\"Original Shape: {X_logs.shape} (10 samples, 5 features)\")\n",
    "\n",
    "# 2. Apply PCA\n",
    "# We want to compress 5 features down to 2 \"Principal Components\"\n",
    "pca = PCA(n_components=2)\n",
    "X_compressed = pca.fit_transform(X_logs)\n",
    "\n",
    "# 3. Result\n",
    "print(f\"Compressed Shape: {X_compressed.shape} (10 samples, 2 features)\")\n",
    "print(f\"Information Retained: {np.sum(pca.explained_variance_ratio_):.2%}\")\n",
    "# We typically retain ~95% of the information while removing noise."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Semi-supervised Learning\n",
    "Semi-supervised learning is a hybrid approach that uses a small amount of labeled data along with a large amount of unlabeled data.\n",
    "\n",
    "In cybersecurity, manually analyzing a file to confirm it is malware is expensive and slow (requires human experts). However, collecting raw, unanalyzed files is easy. Semi-supervised learning solves this by using the few known bad files to infer the nature of the unknown files based on their similarity.\n",
    "\n",
    "> **Figure:** Semi-supervised Learning: Propagating labels from known to unknown data.\n",
    "\n",
    "The code below demonstrates *Label Spreading*. We mark unknown files with `-1`. The model then looks at the geometry of the data and spreads the known labels (0 or 1) to the nearest unknown neighbors."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import numpy as np\n",
    "from sklearn.semi_supervised import LabelSpreading\n",
    "\n",
    "# --- SEMI-SUPERVISED EXAMPLE: Malware Label Propagation ---\n",
    "\n",
    "# 1. Dataset (Feature: [File Entropy, File Size])\n",
    "X = np.array([\n",
    "    [0.1, 100], [0.2, 120],  # Cluster A (Likely Benign)\n",
    "    [0.9, 900], [0.8, 950],  # Cluster B (Likely Malware)\n",
    "    [0.15, 110],             # New File 1 (Unknown)\n",
    "    [0.85, 920]              # New File 2 (Unknown)\n",
    "])\n",
    "\n",
    "# 2. Labels (0: Benign, 1: Malware, -1: Unknown/Unlabeled)\n",
    "# We only know the ground truth for the first 4 files.\n",
    "y_mixed = np.array([0, 0, 1, 1, -1, -1])\n",
    "\n",
    "# 3. Train Model (Label Spreading)\n",
    "# The model measures similarity to guess the -1 labels.\n",
    "model = LabelSpreading(kernel='knn', n_neighbors=2)\n",
    "model.fit(X, y_mixed)\n",
    "\n",
    "# 4. Results\n",
    "print(f\"Original Labels: {y_mixed}\")\n",
    "print(f\"Inferred Labels: {model.transduction_}\")\n",
    "# Expected: The last two files should become [0, 1] based on neighbors."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Reinforcement Learning (RL)\n",
    "In reinforcement learning, an \"agent\" learns to make decisions by performing actions in an environment to achieve a goal.\n",
    "\n",
    "Unlike other methods, the agent is not told *what* to do. It learns through trial and error. It takes an action, receives a *Reward* (positive points) or a *Penalty* (negative points), and updates its strategy to maximize rewards over time.\n",
    "\n",
    "> **Figure:** Reinforcement Learning: The Agent-Environment-Reward loop.\n",
    "\n",
    "The code below simulates a simple *Firewall Agent*. It learns whether to \"Block\" or \"Allow\" traffic from a specific IP range. At first, it guesses randomly. Over time, it learns that blocking malicious IPs yields a reward (preventing a breach), while blocking valid users yields a penalty (user complaint)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import numpy as np\n",
    "import random\n",
    "\n",
    "# --- RL EXAMPLE: Q-Learning Firewall Agent ---\n",
    "\n",
    "# 1. Setup Environment\n",
    "# Actions: 0 = Allow, 1 = Block\n",
    "# State: A specific IP Address (e.g., \"192.168.1.5\")\n",
    "q_table = { '192.168.1.5': [0.0, 0.0] } # [Value of Allow, Value of Block]\n",
    "\n",
    "def get_reward(action, is_malicious):\n",
    "    if is_malicious and action == 1: return +10  # Good: Blocked attack\n",
    "    if is_malicious and action == 0: return -50  # Bad: Allowed attack (Breach)\n",
    "    if not is_malicious and action == 1: return -10 # Bad: Blocked user (Complaint)\n",
    "    return +5 # Good: Allowed user\n",
    "\n",
    "# 2. Simulation Loop (Training)\n",
    "ip = '192.168.1.5'\n",
    "is_malicious = True  # Ground truth (Agent doesn't know this yet!)\n",
    "learning_rate = 0.5\n",
    "\n",
    "print(\"--- Training Start ---\")\n",
    "for episode in range(5):\n",
    "    # Step A: Choose Action (Exploration vs Exploitation)\n",
    "    # Simple logic: mostly pick the action with highest current value\n",
    "    current_values = q_table[ip]\n",
    "    if random.random() < 0.5: \n",
    "        action = random.choice([0, 1]) # Explore (Random)\n",
    "    else:\n",
    "        action = np.argmax(current_values) # Exploit (Best known)\n",
    "\n",
    "    # Step B: Act and get Reward\n",
    "    reward = get_reward(action, is_malicious)\n",
    "\n",
    "    # Step C: Update Q-Table (Learn)\n",
    "    # New Value = Old Value + Learning Rate * Reward\n",
    "    old_value = q_table[ip][action]\n",
    "    q_table[ip][action] = old_value + (learning_rate * reward)\n",
    "    \n",
    "    print(f\"Ep {episode+1}: Action {'Block' if action==1 else 'Allow'} -> Reward {reward}\")\n",
    "\n",
    "print(\"--- Training End ---\")\n",
    "print(f\"Final Q-Values for IP: {q_table[ip]}\")\n",
    "# Expected: Index 1 (Block) should have a much higher value than Index 0 (Allow)."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Supervised Learning Workflow\n",
    "A typical supervised learning project follows a structured workflow. The order of these steps is critical: performing them out of sequence (such as scaling data before splitting it) can lead to invalid results, a problem known as *data leakage*.\n",
    "\n",
    "> **Figure:** The Supervised Learning Workflow\n",
    "\n",
    "1. **Data Collection:** Gather a dataset where each example has a correct label. For instance, collect network flow records, with each record labeled as 'benign' or 'malicious'.\n",
    "\n",
    "2. **Data Splitting:** Immediately divide the raw dataset into two parts: a *training set* and a *testing set*. The model learns **only** from the training set. The testing set is locked away and used only for the final evaluation. A common split is 80% for training and 20% for testing."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import pandas as pd\n",
    "from sklearn.model_selection import train_test_split\n",
    "\n",
    "# 1. Data Collection (Mock Dataset)\n",
    "# Features: Duration (sec), Source Bytes, Destination Bytes\n",
    "data = {\n",
    "    'duration':  [0.1, 5.2, 0.2, 12.0, 0.1, 0.5, 4.5, 1.2, 0.1, 3.0],\n",
    "    'src_bytes': [64, 5000, 128, 9000, 64, 256, 4500, 500, 64, 6000],\n",
    "    'dst_bytes': [1000, 0, 500, 0, 800, 1200, 0, 300, 950, 0],\n",
    "    'label':     [0, 1, 0, 1, 0, 0, 1, 0, 0, 1]  # 0: Benign, 1: Malicious\n",
    "}\n",
    "df = pd.DataFrame(data)\n",
    "\n",
    "# 2. Data Splitting\n",
    "X = df.drop('label', axis=1)\n",
    "y = df['label']\n",
    "\n",
    "# stratify=y ensures both sets maintain the same ratio of attacks\n",
    "X_train, X_test, y_train, y_test = train_test_split(\n",
    "    X, y, test_size=0.2, random_state=42, stratify=y\n",
    ")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "3. **Feature Engineering:** Select, transform, or create numerical characteristics, called *features*, from the raw data. For example, rather than feeding raw packet logs into the model, you might calculate the \"average packet size\" or \"connection duration.\" This step often relies on cybersecurity domain knowledge.\n",
    "\n",
    "4. **Data Preprocessing:** Clean the training data and scale numerical features (e.g., forcing values between 0 and 1). Crucially, you must calculate the scaling parameters (like the mean) using *only* the training data, and then apply those same parameters to the test data."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.preprocessing import StandardScaler\n",
    "\n",
    "# 3. Feature Engineering\n",
    "# Create a new feature: Bytes per Second (Transfer Rate)\n",
    "# We apply the calculation to both sets independently\n",
    "X_train['bytes_per_sec'] = X_train['src_bytes'] / (X_train['duration'] + 0.01)\n",
    "X_test['bytes_per_sec'] = X_test['src_bytes'] / (X_test['duration'] + 0.01)\n",
    "\n",
    "# 4. Data Preprocessing (Scaling)\n",
    "scaler = StandardScaler()\n",
    "\n",
    "# CRITICAL: We 'fit' the scaler ONLY on the Training data\n",
    "X_train_scaled = scaler.fit_transform(X_train)\n",
    "\n",
    "# We use those learned parameters to transform the Test data\n",
    "X_test_scaled = scaler.transform(X_test)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "5. **Model Selection:** Choose a suitable machine learning algorithm. For example, a Random Forest classifier is often effective for structured data like network logs, while a Convolutional Neural Network (CNN) is better for unstructured data like malware binaries.\n",
    "\n",
    "6. **Hyperparameter Tuning:** Before the final evaluation, use a technique called *Cross-Validation* on the training set to find the best settings for the model (e.g., the number of trees in a Random Forest). This ensures the model is optimized without peeking at the test set.\n",
    "\n",
    "7. **Model Training:** Train the model using the best hyperparameters on the full training dataset. The algorithm adjusts its internal parameters to map the input features to the output labels."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.ensemble import RandomForestClassifier\n",
    "from sklearn.model_selection import GridSearchCV\n",
    "\n",
    "# 5. Model Selection (Random Forest)\n",
    "rf = RandomForestClassifier(random_state=42)\n",
    "\n",
    "# 6. Hyperparameter Tuning (using Cross-Validation on X_train)\n",
    "param_grid = {'n_estimators': [10, 50], 'max_depth': [3, 5]}\n",
    "grid_search = GridSearchCV(rf, param_grid, cv=2, scoring='f1')\n",
    "grid_search.fit(X_train_scaled, y_train)\n",
    "\n",
    "# 7. Model Training (Automatic via GridSearchCV)\n",
    "best_model = grid_search.best_estimator_\n",
    "print(f\"Best Parameters: {grid_search.best_params_}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "8. **Model Evaluation:** Finally, use the isolated *testing set* to assess performance. Use specific metrics like precision, recall, and the F1-score to measure how well the model detects threats in data it has never seen before."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.metrics import classification_report\n",
    "\n",
    "# 8. Model Evaluation\n",
    "# Predict using the locked-away X_test_scaled\n",
    "y_pred = best_model.predict(X_test_scaled)\n",
    "\n",
    "print(\"--- Final Test Set Evaluation ---\")\n",
    "print(classification_report(y_test, y_pred, target_names=['Benign', 'Malicious']))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "9. **Deployment and Monitoring:** Integrate the model into a live security system to make real-time predictions. Once deployed, the model must be continuously monitored for performance degradation caused by changing attack patterns (concept drift)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# 9. Deployment (Simulation)\n",
    "def analyze_packet(duration, src_bytes, dst_bytes):\n",
    "    # Create single-row DataFrame\n",
    "    live_data = pd.DataFrame([[duration, src_bytes, dst_bytes]], \n",
    "                           columns=['duration', 'src_bytes', 'dst_bytes'])\n",
    "    \n",
    "    # Repeat Feature Engineering\n",
    "    live_data['bytes_per_sec'] = live_data['src_bytes'] / (live_data['duration'] + 0.01)\n",
    "    \n",
    "    # Scale using the TRAINING scaler\n",
    "    live_scaled = scaler.transform(live_data)\n",
    "    \n",
    "    # Predict\n",
    "    prediction = best_model.predict(live_scaled)[0]\n",
    "    return \"ALERT: Malicious\" if prediction == 1 else \"Benign\"\n",
    "\n",
    "# Simulate an incoming packet\n",
    "print(analyze_packet(15.0, 10000, 0))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Key ML Model Families for Security\n",
    "\n",
    "Different machine learning models are suited for different types of security data and tasks. We can broadly categorize them into classical models, which work well on structured data (like logs), and deep learning models, which excel with unstructured data (like raw code or text).\n",
    "\n",
    "## Classical Machine Learning\n",
    "These models are the industry standard for working with tabular data. Before jumping to complex algorithms, security professionals often start here to establish a performance baseline.\n",
    "\n",
    "### Logistic Regression\n",
    "Despite the name, this is a classification algorithm. It estimates the probability that an event belongs to a specific class (e.g., \"75% chance this is spam\"). It serves as an excellent baseline model because it is fast, interpretable, and easy to implement.\n",
    "\n",
    "> **Figure:** Logistic Regression uses a sigmoid function to map predictions to probabilities between 0 and 1."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.linear_model import LogisticRegression\n",
    "\n",
    "# 1. Initialize\n",
    "clf = LogisticRegression()\n",
    "\n",
    "# 2. Train (Fit)\n",
    "# X_train: Matrix of features, y_train: Vector of labels\n",
    "print(\"Training Logistic Regression...\")\n",
    "clf.fit(X_train_scaled, y_train) # Using scaled data from previous section\n",
    "\n",
    "# 3. Use\n",
    "sample_data = X_train_scaled[0].reshape(1, -1) # Mock sample data\n",
    "print(\"Prediction:\", clf.predict(sample_data))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Support Vector Machines (SVMs)\n",
    "SVMs are powerful for binary classification. Imagine two groups of data points (benign and malicious) on a graph. The SVM tries to draw a straight line (or a flat plane in 3D) that separates these groups with as much empty space, or \"margin,\" between them as possible.\n",
    "\n",
    "> **Figure:** SVM maximizes the margin between the two classes."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.svm import SVC\n",
    "\n",
    "# 1. Initialize\n",
    "# Kernel='linear' creates a straight boundary; 'rbf' creates a curved one\n",
    "clf = SVC(kernel='linear', C=1.0)\n",
    "\n",
    "# 2. Train (Fit)\n",
    "# SVM training can be slow on very large datasets\n",
    "print(\"Training SVM...\")\n",
    "clf.fit(X_train_scaled, y_train)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Decision Trees\n",
    "These models use a flowchart-like structure to make decisions. Each internal node represents a \"test\" on a feature (e.g., \"Is the packet size > 100 bytes?\"), and each leaf node represents a final prediction. They are highly interpretable but prone to \"overfitting\" (memorizing the training data).\n",
    "\n",
    "> **Figure:** A decision tree makes predictions by following a path of logical tests."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.tree import DecisionTreeClassifier\n",
    "\n",
    "# 1. Initialize\n",
    "# max_depth=5 restricts the tree growth to prevent overfitting\n",
    "clf = DecisionTreeClassifier(max_depth=5)\n",
    "\n",
    "# 2. Train (Fit)\n",
    "print(\"Training Decision Tree...\")\n",
    "clf.fit(X_train, y_train)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Ensemble Methods\n",
    "These combine multiple models to improve stability and accuracy.\n",
    "* **Random Forest:** Builds hundreds of decision trees on random subsets of the data and averages their votes.\n",
    "* **Gradient Boosting Machines (GBMs):** Algorithms like XGBoost build trees sequentially, where each new tree tries to fix the errors of the previous one."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.ensemble import RandomForestClassifier\n",
    "from xgboost import XGBClassifier\n",
    "\n",
    "# 1. Initialize\n",
    "rf_model = RandomForestClassifier(n_estimators=100)\n",
    "xgb_model = XGBClassifier(n_estimators=100, learning_rate=0.1)\n",
    "\n",
    "# 2. Train (Fit)\n",
    "# Random Forest trains trees in parallel (fast)\n",
    "rf_model.fit(X_train, y_train)\n",
    "\n",
    "# XGBoost trains trees sequentially (slower but often more accurate)\n",
    "xgb_model.fit(X_train, y_train)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Deep Learning Fundamentals\n",
    "Deep learning models use \"neural networks\" to learn complex patterns directly from raw, unstructured data without manual feature engineering.\n",
    "\n",
    "### Neural Networks (NNs)\n",
    "Inspired by the human brain, these consist of layers of interconnected \"neurons.\" The network learns by adjusting the mathematical \"weights\" of connections between neurons to minimize error.\n",
    "\n",
    "> **Figure:** A Feed-Forward Neural Network."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import torch\n",
    "import torch.nn as nn\n",
    "import torch.optim as optim\n",
    "\n",
    "# 1. Define Network\n",
    "model = nn.Sequential(nn.Linear(10, 64), nn.ReLU(), nn.Linear(64, 2))\n",
    "\n",
    "# 2. Define Loss and Optimizer\n",
    "criterion = nn.CrossEntropyLoss()\n",
    "optimizer = optim.Adam(model.parameters(), lr=0.001)\n",
    "\n",
    "# 3. Training Step (Conceptual)\n",
    "# Dummy input: 5 samples, 10 features each\n",
    "inputs = torch.randn(5, 10) \n",
    "labels = torch.tensor([0, 1, 0, 1, 0])\n",
    "\n",
    "optimizer.zero_grad()           # Reset gradients\n",
    "outputs = model(inputs)         # Forward pass\n",
    "loss = criterion(outputs, labels) # Calculate error\n",
    "loss.backward()                 # Backward pass (calculate gradients)\n",
    "optimizer.step()                # Update weights\n",
    "print(f\"Loss: {loss.item():.4f}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Convolutional Neural Networks (CNNs)\n",
    "Originally designed for computer vision, CNNs scan data for local patterns. In security, we can convert a malware binary file into a grayscale image; the CNN then \"looks\" for visual textures corresponding to malicious code [**ref**].\n",
    "\n",
    "> **Figure:** CNNs use filters to detect local patterns in grid-like data."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# 1D Conv Layer for sequence data (e.g., packet bytes)\n",
    "conv_layer = nn.Conv1d(in_channels=1, out_channels=16, kernel_size=3)\n",
    "\n",
    "# Training Phase input preparation:\n",
    "# Input Shape: [Batch Size, Channels, Length]\n",
    "# e.g., 10 packets, 1 channel (raw byte), 1000 bytes long\n",
    "input_data = torch.randn(10, 1, 1000)\n",
    "\n",
    "# Forward Pass: The filter slides over the sequence\n",
    "feature_map = conv_layer(input_data)\n",
    "print(f\"Output shape: {feature_map.shape}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Recurrent Neural Networks (RNNs) & LSTMs\n",
    "These networks have a \"memory\" loop [**ref**], allowing them to process sequences. This makes them ideal for analyzing time-series data, such as the sequence of API calls made by a program.\n",
    "\n",
    "> **Figure:** RNNs maintain an internal state to process sequential data."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# LSTM Layer\n",
    "lstm_layer = nn.LSTM(input_size=10, hidden_size=20, batch_first=True)\n",
    "\n",
    "# Training Phase input preparation:\n",
    "# Input Shape: [Batch Size, Sequence Length, Features per Event]\n",
    "# e.g., 5 sessions, 100 events long, 10 features per event\n",
    "sequence_data = torch.randn(5, 100, 10)\n",
    "\n",
    "# Forward Pass\n",
    "# output contains features for every time step; hidden is the final state\n",
    "output, (hidden, cell) = lstm_layer(sequence_data)\n",
    "print(f\"Encoded Sequence Shape: {output.shape}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Transformer Architectures (e.g., LLMs)\n",
    "Transformers [**ref**] use a mechanism called **Attention**. This allows the model to dynamically focus on relevant parts of a long sequence—for example, linking a suspicious \"URL\" at the end of an email body to the word \"Urgent\" in the subject line—regardless of the distance between them.\n",
    "\n",
    "> **Figure:** Transformer Architectures."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Note: This code block requires the 'transformers' library and a valid dataset to run.\n",
    "try:\n",
    "    from transformers import BertForSequenceClassification, Trainer, TrainingArguments\n",
    "\n",
    "    # 1. Load Pre-trained Model\n",
    "    model = BertForSequenceClassification.from_pretrained('bert-base-uncased')\n",
    "\n",
    "    # 2. Define Training Arguments\n",
    "    args = TrainingArguments(output_dir=\"./results\", num_train_epochs=1)\n",
    "\n",
    "    # 3. Initialize Trainer (High-level API)\n",
    "    # trainer = Trainer(\n",
    "    #     model=model,\n",
    "    #     args=args,\n",
    "    #     train_dataset=train_dataset, # Assumes dataset is pre-loaded\n",
    "    #     eval_dataset=test_dataset\n",
    "    # )\n",
    "\n",
    "    # 4. Train\n",
    "    # trainer.train()\n",
    "    print(\"Transformers library loaded successfully. Setup required for execution.\")\n",
    "except ImportError:\n",
    "    print(\"Transformers library not found.\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Feature Engineering\n",
    "**Feature engineering** is the art of using domain knowledge to create numerical features that help algorithms learn. A good feature simplifies the problem for the model.\n",
    "\n",
    "### Common Security Features\n",
    "\n",
    "#### 1. Text to Numbers (NLP)\n",
    "Converting a URL or email body into numbers. Methods range from simple *Bag-of-Words* (counting word frequency) to *Embeddings*.\n",
    "\n",
    "Listing below shows the **Bag-of-Words** technique. The model learns that the word \"urgent\" appears often in phishing, while \"meeting\" appears in benign emails."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.feature_extraction.text import CountVectorizer\n",
    "import pandas as pd\n",
    "\n",
    "# --- TEXT EXAMPLE: Bag-of-Words for Email ---\n",
    "\n",
    "# 1. Raw Data (Email Subjects)\n",
    "emails = [\n",
    "    \"Urgent: Update your password now\",  # Phishing\n",
    "    \"Meeting agenda for tomorrow\",       # Benign\n",
    "    \"Win a free iPhone instantly\"        # Phishing\n",
    "]\n",
    "\n",
    "# 2. Transform to Numbers\n",
    "# CountVectorizer creates a column for every unique word\n",
    "vectorizer = CountVectorizer(stop_words='english')\n",
    "X = vectorizer.fit_transform(emails)\n",
    "\n",
    "# 3. View Result (The Feature Vector)\n",
    "features = vectorizer.get_feature_names_out()\n",
    "df_text = pd.DataFrame(X.toarray(), columns=features)\n",
    "\n",
    "print(df_text)\n",
    "# Result: A matrix where columns are words ('urgent', 'win', 'meeting') \n",
    "# and values are counts (1 or 0)."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "#### 2. Network and Temporal Features\n",
    "Aggregating packet data and extracting time-based patterns. Attackers often operate during off-hours to avoid detection, making \"Hour of Day\" a powerful feature.\n",
    "\n",
    "Listing below demonstrates extracting the hour from a timestamp and calculating bytes-per-second from flow data."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import pandas as pd\n",
    "\n",
    "# --- NETWORK/TEMPORAL EXAMPLE: Log Parsing ---\n",
    "\n",
    "# 1. Raw Log Data\n",
    "df = pd.DataFrame({\n",
    "    'timestamp': ['2026-01-15 03:15:00', '2026-01-15 14:30:00'],\n",
    "    'bytes_sent': [15000, 450],\n",
    "    'duration_sec': [5.0, 0.5]\n",
    "})\n",
    "\n",
    "# 2. Feature Engineering\n",
    "# Extract 'Hour' (Attackers love 3 AM)\n",
    "df['timestamp'] = pd.to_datetime(df['timestamp'])\n",
    "df['hour'] = df['timestamp'].dt.hour\n",
    "\n",
    "# Calculate Transfer Rate\n",
    "df['transfer_rate'] = df['bytes_sent'] / df['duration_sec']\n",
    "\n",
    "print(df[['hour', 'transfer_rate']])\n",
    "# Result: Numerical features ready for analysis."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "#### 3. Executable Analysis (Entropy)\n",
    "Extracting static properties from malware. A key metric is **Entropy** (randomness). Legitimate code has structure (low entropy), while packed or encrypted malware appears random (high entropy).\n",
    "\n",
    "Listing below calculates the Shannon Entropy of a file header."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import numpy as np\n",
    "from scipy.stats import entropy\n",
    "import collections\n",
    "\n",
    "# --- MALWARE EXAMPLE: Calculating File Entropy ---\n",
    "\n",
    "def calc_entropy(data_bytes):\n",
    "    # Count frequency of each byte (0-255)\n",
    "    counts = collections.Counter(data_bytes)\n",
    "    probs = [count / len(data_bytes) for count in counts.values()]\n",
    "    # Calculate Shannon Entropy\n",
    "    return entropy(probs, base=2)\n",
    "\n",
    "# 1. Simulate Files\n",
    "legitimate_header = b'\\x00\\x00\\x00\\x01\\x00\\x00\\x00\\x02' # Structured (Zeros)\n",
    "encrypted_payload = b'\\x8a\\xf1\\x23\\x99\\x1b\\xcd\\x5e\\x4a' # Random\n",
    "\n",
    "# 2. Compare Features\n",
    "print(f\"Legit Entropy:     {calc_entropy(legitimate_header):.2f}\") \n",
    "# Expected: Low (e.g., 1.0) due to repetition\n",
    "\n",
    "print(f\"Malware Entropy:   {calc_entropy(encrypted_payload):.2f}\") \n",
    "# Expected: High (e.g., 3.0) due to randomness"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Data Normalization and Scaling\n",
    "A critical, often overlooked step is **Normalization**. Features often have vastly different scales. For example, a \"URL Length\" might be 200, while \"Hour of Day\" is only 23. If we leave these numbers as-is, the model will assume the URL Length is mathematically more important. To fix this, we scale numerical features to a common range (typically 0 to 1).\n",
    "\n",
    "Listing below shows how failing to scale data can distort the numbers, and how `MinMaxScaler` fixes it."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.preprocessing import MinMaxScaler\n",
    "import pandas as pd\n",
    "\n",
    "# --- SCALING EXAMPLE ---\n",
    "\n",
    "# 1. Unscaled Data\n",
    "# 'Bytes' is huge (10,000), 'Hour' is tiny (23).\n",
    "df = pd.DataFrame({\n",
    "    'bytes': [100, 10000, 5000], \n",
    "    'hour':  [1, 23, 12]\n",
    "})\n",
    "\n",
    "print(\"Original (Model will ignore 'hour'):\")\n",
    "print(df)\n",
    "\n",
    "# 2. Apply Scaling\n",
    "# Squeezes everything between 0.0 and 1.0\n",
    "scaler = MinMaxScaler()\n",
    "df_scaled = pd.DataFrame(scaler.fit_transform(df), columns=df.columns)\n",
    "\n",
    "print(\"\\nScaled (Model treats features equally):\")\n",
    "print(df_scaled.round(2))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "The code below demonstrates the **Manual Feature Engineering** approach, showing how to create security features and properly scale them for a classical model."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import pandas as pd\n",
    "from sklearn.preprocessing import OneHotEncoder, MinMaxScaler\n",
    "\n",
    "# 1. Raw Data (Web Logs with Timestamps)\n",
    "df = pd.DataFrame({\n",
    "    'timestamp': ['2026-01-15 08:00:00', '2026-01-15 23:45:00', '2026-01-15 14:30:00'],\n",
    "    'method': ['GET', 'POST', 'GET'],\n",
    "    'url': ['/index.html', '/admin/login', '/api/data/' * 50] # Last one is a buffer overflow attempt\n",
    "})\n",
    "\n",
    "print(\"--- 1. Raw Data ---\")\n",
    "print(df)\n",
    "\n",
    "# --- FEATURE ENGINEERING ---\n",
    "\n",
    "# Feature 1: URL Length\n",
    "# Context: Extremely long URLs often indicate Buffer Overflows or SQL Injection strings.\n",
    "df['url_length'] = df['url'].apply(len)\n",
    "\n",
    "# Feature 2: Dangerous Keywords\n",
    "# Context: The presence of 'admin', 'root', or 'cmd' suggests targeted privilege escalation.\n",
    "df['has_admin_keyword'] = df['url'].apply(lambda x: 1 if 'admin' in x else 0)\n",
    "\n",
    "# Feature 3: Temporal Analysis (Request Hour)\n",
    "# Context: Anomalous activity often occurs outside business hours (e.g., 3 AM).\n",
    "df['timestamp'] = pd.to_datetime(df['timestamp'])\n",
    "df['request_hour'] = df['timestamp'].dt.hour\n",
    "\n",
    "# --- DATA PREPARATION (Encoding & Scaling) ---\n",
    "\n",
    "# 4. One-Hot Encoding for Categorical Data\n",
    "# We convert 'method' (GET, POST) into separate columns (method_GET, method_POST).\n",
    "# We do NOT use 0, 1, 2 for categories, as that implies rank (POST > GET).\n",
    "encoder = OneHotEncoder(sparse_output=False)\n",
    "method_encoded = encoder.fit_transform(df[['method']])\n",
    "method_df = pd.DataFrame(method_encoded, columns=encoder.get_feature_names_out(['method']))\n",
    "df = pd.concat([df, method_df], axis=1)\n",
    "\n",
    "# 5. Normalization (Scaling)\n",
    "# We scale 'url_length' and 'request_hour' to the 0-1 range.\n",
    "scaler = MinMaxScaler()\n",
    "cols_to_scale = ['url_length', 'request_hour']\n",
    "df[cols_to_scale] = scaler.fit_transform(df[cols_to_scale])\n",
    "\n",
    "print(\"\\n--- 2. Final Feature Vector (Ready for ML) ---\")\n",
    "# Drop the raw columns that the model cannot understand\n",
    "final_df = df.drop(['timestamp', 'method', 'url'], axis=1)\n",
    "print(final_df.round(2))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Representation Learning\n",
    "While feature engineering is powerful, it is also time-consuming and requires deep expert knowledge. **Representation learning** is the automated alternative used by Deep Learning.\n",
    "\n",
    "Instead of a human manually telling the model \"look for the string 'admin'\" or \"calculate entropy,\" a Deep Learning model (like a CNN or Transformer) ingests the raw data and *learns* the best features itself. It automatically discovers that certain byte patterns or text sequences are associated with attacks.\n",
    "\n",
    "* **Advantage:** Removes the need for manual feature crafting; can find complex patterns humans might miss.\n",
    "* **Disadvantage:** Requires much more data and computing power; the resulting model is a \"black box\" that is harder to interpret.\n",
    "\n",
    "> **Figure:** Manual Feature Engineering vs. Automated Representation Learning.\n",
    "\n",
    "Listing below demonstrates how a Neural Network automatically converts raw system API calls (integers) into dense numerical vectors. Unlike the \"Bag-of-Words\" example where we counted words, here the model starts with random numbers and *learns* the best representation during training."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import torch\n",
    "import torch.nn as nn\n",
    "\n",
    "# --- REPRESENTATION LEARNING EXAMPLE: API Call Embeddings ---\n",
    "\n",
    "# 1. Raw Data (Sequence of System API Calls)\n",
    "# Imagine we map calls to IDs: 0='NtOpenProcess', 1='NtWriteFile', 2='NtClose'\n",
    "# A malware trace might look like: Open -> Write -> Close\n",
    "raw_sequence = torch.tensor([[0, 1, 2]])\n",
    "\n",
    "print(f\"Raw Input (Human IDs): {raw_sequence}\")\n",
    "\n",
    "# 2. Define the Representation Layer (Embedding)\n",
    "# We ask the Neural Network: \"Learn a vector of size 3 for every API call.\"\n",
    "# num_embeddings=10 (vocabulary size), embedding_dim=3 (feature size)\n",
    "embedding_layer = nn.Embedding(num_embeddings=10, embedding_dim=3)\n",
    "\n",
    "# 3. Automatic Feature Extraction\n",
    "# The model converts the integers [0, 1, 2] into vectors automatically.\n",
    "# Crucially, these numbers are 'parameters' that improve as the model trains.\n",
    "learned_features = embedding_layer(raw_sequence)\n",
    "\n",
    "print(\"\\nLearned Representation (Model's Internal Features):\")\n",
    "# Each row represents one API call\n",
    "print(learned_features.detach().numpy())\n",
    "\n",
    "# Result:\n",
    "# [[-0.45,  1.21, -0.05]   <- The model's numerical view of 'NtOpenProcess'\n",
    "#  [ 0.88, -0.12,  0.33]   <- The model's numerical view of 'NtWriteFile'\n",
    "#  ... ]"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Evaluation Metrics Tailored for Security\n",
    "\n",
    "Measuring a model's performance correctly is critical. In cybersecurity, the cost of a mistake—missing a real attack or blocking a legitimate CEO—is extremely high. Furthermore, standard metrics often fail because security datasets are heavily imbalanced.\n",
    "\n",
    "## Beyond Simple Accuracy: The Base Rate Fallacy\n",
    "In a typical network, malicious packets are rare. Imagine a dataset with 99,900 benign packets and only 100 malicious ones.\n",
    "If a model simply predicts \"Benign\" for *everything*, it achieves 99.9% accuracy. However, it has failed 100% of its job because it missed every single attack. This is why we never rely on simple accuracy for security tasks.\n",
    "\n",
    "## Metric Type 1: Classification Metrics\n",
    "Classification is the most common task (e.g., Firewall: Block/Allow). To evaluate it, we first map the outcomes into a **Confusion Matrix**.\n",
    "\n",
    "From this matrix, we calculate specific metrics:\n",
    "\n",
    "* **Precision:** When the system alerts, how often is it real? (High precision reduces analyst fatigue).\n",
    "$$ Precision = \\frac{TP}{TP + FP} $$\n",
    "\n",
    "* **Recall (Sensitivity):** Out of all real attacks, how many did we catch? (High recall prevents breaches).\n",
    "$$ Recall = \\frac{TP}{TP + FN} $$\n",
    "\n",
    "* **F1-Score:** The harmonic mean of Precision and Recall. Best for comparing models on imbalanced data.\n",
    "$$ F1 = 2 \\times \\frac{Precision \\times Recall}{Precision + Recall} $$\n",
    "\n",
    "Listing below demonstrates evaluating an Intrusion Detection System (IDS)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.metrics import classification_report, confusion_matrix\n",
    "import numpy as np\n",
    "\n",
    "# --- CLASSIFICATION EVALUATION: IDS ---\n",
    "\n",
    "# 1. Ground Truth (0: Benign, 1: Attack)\n",
    "# In reality, attacks are rare (imbalanced).\n",
    "y_true = np.array([0, 0, 0, 0, 0, 0, 0, 1, 1, 1])\n",
    "\n",
    "# 2. Model Predictions\n",
    "# The model caught 2 attacks (TP), missed 1 (FN), and raised 1 false alarm (FP).\n",
    "y_pred = np.array([0, 0, 0, 0, 0, 0, 1, 1, 1, 0])\n",
    "\n",
    "# 3. Evaluation\n",
    "print(\"Confusion Matrix:\\n\", confusion_matrix(y_true, y_pred))\n",
    "print(\"\\nDetailed Report:\\n\", classification_report(y_true, y_pred, target_names=['Benign', 'Attack']))\n",
    "\n",
    "# Analysis:\n",
    "# Precision (Attack) = 2/3 (0.67) -> Of 3 alerts, 2 were real.\n",
    "# Recall (Attack)    = 2/3 (0.67) -> Of 3 real attacks, we caught 2."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Metric Type 2: Regression Metrics\n",
    "When predicting numerical values (e.g., predicting a Risk Score from 0 to 10), \"Accuracy\" does not apply. Instead, we measure the **Error** (distance between prediction and reality).\n",
    "\n",
    "* **Mean Absolute Error (MAE):** The average difference. E.g., \"Our predicted risk score is usually off by 1.5 points.\" Easy to interpret.\n",
    "* **Mean Squared Error (MSE):** Squares the errors before averaging. This heavily punishes large mistakes. In security, missing a risk score by 5 points is much worse than missing by 1 point, so MSE is often preferred.\n",
    "\n",
    "Listing below evaluates a model predicting Vulnerability Severity (CVSS)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.metrics import mean_absolute_error, mean_squared_error\n",
    "import numpy as np\n",
    "\n",
    "# --- REGRESSION EVALUATION: Risk Scoring ---\n",
    "\n",
    "# 1. Actual CVSS Scores (0.0 to 10.0)\n",
    "y_true = np.array([2.5, 5.0, 9.0, 4.0])\n",
    "\n",
    "# 2. Predicted Scores\n",
    "# The last prediction (4.0 vs 9.0) is a catastrophic error.\n",
    "y_pred = np.array([2.0, 5.2, 4.0, 4.1])\n",
    "\n",
    "# 3. Evaluation\n",
    "mae = mean_absolute_error(y_true, y_pred)\n",
    "mse = mean_squared_error(y_true, y_pred)\n",
    "\n",
    "print(f\"MAE: {mae:.2f}\") # Result: 1.45 (Average error is small)\n",
    "print(f\"MSE: {mse:.2f}\") # Result: 6.27 (High because of the one huge error)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Metric Type 3: Clustering Metrics\n",
    "Evaluating unsupervised learning is harder because we often lack \"correct\" labels.\n",
    "\n",
    "* **Silhouette Score:** Measures how distinct the clusters are. Ranges from -1 to 1.\n",
    "    * **+1:** Perfect clusters (Dense and well-separated).\n",
    "    * **0:** Overlapping clusters.\n",
    "    * **-1:** Data points assigned to the wrong cluster.\n",
    "\n",
    "Listing below evaluates how well a model grouped Malware Families."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.metrics import silhouette_score\n",
    "import numpy as np\n",
    "\n",
    "# --- CLUSTERING EVALUATION: Malware Grouping ---\n",
    "\n",
    "# 1. Feature Space (2D for simplicity: [File Size, Entropy])\n",
    "X_features = np.array([\n",
    "    [100, 0.5], [105, 0.6], [102, 0.55], # Group A (e.g., Adware)\n",
    "    [500, 7.5], [510, 7.6], [505, 7.4]   # Group B (e.g., Ransomware)\n",
    "])\n",
    "\n",
    "# 2. Cluster Labels assigned by model (e.g., K-Means)\n",
    "labels = np.array([0, 0, 0, 1, 1, 1])\n",
    "\n",
    "# 3. Evaluation (Internal Validity)\n",
    "# Do these groups make mathematical sense?\n",
    "score = silhouette_score(X_features, labels)\n",
    "\n",
    "print(f\"Silhouette Score: {score:.2f}\")\n",
    "# Result: ~0.98 (Very High). The groups are clearly separated."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## The Balancing Act: Threshold Tuning\n",
    "There is no \"perfect\" model. We must choose a trade-off based on the cost of errors. This choice is implemented by adjusting the **Decision Threshold**.\n",
    "\n",
    "Most models output a probability score (0.0 to 1.0). By default, we classify anything above 0.5 as \"Malicious.\" However, we can shift this line:\n",
    "* **Low Threshold (e.g., > 0.2):** The \"Paranoid\" setting. We classify items as malicious even if the model is only slightly suspicious.\n",
    "    * **Result:** High Recall (we miss very little), but Low Precision (many False Positives).\n",
    "    * **Use Case:** **Threat Hunting**. Analysts want to see everything suspicious, even if it creates more work.\n",
    "* **High Threshold (e.g., > 0.9):** The \"Conservative\" setting. We only act if the model is extremely confident.\n",
    "    * **Result:** High Precision (very few false alarms), but Lower Recall (we might miss subtle attacks).\n",
    "    * **Use Case:** **Automated Blocking (IPS)**. We cannot afford to block the CEO's email, so we only block when 99% sure.\n",
    "\n",
    "Listing below demonstrates how changing the threshold dramatically alters the model's behavior using the same set of predictions."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import numpy as np\n",
    "from sklearn.metrics import precision_score, recall_score, confusion_matrix\n",
    "\n",
    "# --- THRESHOLD TUNING EXAMPLE ---\n",
    "\n",
    "# 1. Ground Truth (10 Attacks, 10 Benign)\n",
    "y_true = np.array([0]*10 + [1]*10)\n",
    "\n",
    "# 2. Model Probabilities (0.0 to 1.0)\n",
    "# The model is unsure about several items (scores around 0.4 - 0.6)\n",
    "y_proba = np.array([\n",
    "    0.1, 0.2, 0.1, 0.1, 0.45, 0.55, 0.1, 0.1, 0.1, 0.1, # Benign samples\n",
    "    0.9, 0.8, 0.9, 0.95, 0.40, 0.60, 0.85, 0.9, 0.9, 0.9 # Malicious samples\n",
    "])\n",
    "\n",
    "def evaluate_threshold(threshold):\n",
    "    # Apply threshold: If proba > threshold, predict 1 (Attack)\n",
    "    y_pred = (y_proba > threshold).astype(int)\n",
    "    \n",
    "    prec = precision_score(y_true, y_pred)\n",
    "    rec = recall_score(y_true, y_pred)\n",
    "    tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()\n",
    "    \n",
    "    print(f\"--- Threshold: {threshold} ---\")\n",
    "    print(f\"Precision: {prec:.2f} | Recall: {rec:.2f}\")\n",
    "    print(f\"False Alarms (FP): {fp} | Missed Attacks (FN): {fn}\\n\")\n",
    "\n",
    "# Scenario A: Paranoid Mode (Threat Hunting)\n",
    "# We want to catch everything, even if we investigate false leads.\n",
    "evaluate_threshold(0.3)\n",
    "\n",
    "# Scenario B: Conservative Mode (Automated Blocking)\n",
    "# We strictly avoid blocking legitimate users (Zero False Positives desired).\n",
    "evaluate_threshold(0.8)\n",
    "\n",
    "# OUTPUT ANALYSIS:\n",
    "# Threshold 0.3: Recall is 0.90 (Caught 9/10), but we have 2 False Alarms.\n",
    "# Threshold 0.8: Precision is 1.00 (0 False Alarms), but Recall drops to 0.80."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Challenges in ML for Cybersecurity\n",
    "\n",
    "Applying machine learning in cybersecurity is not without its difficulties. Unlike standard academic datasets, security environments are adversarial, encrypted, and highly unstable.\n",
    "\n",
    "## Imbalanced Data\n",
    "Attack instances are usually far less common than normal events (often a 1:10,000 ratio). A model trained on this will become lazy, predicting \"Benign\" 100% of the time to achieve high accuracy while missing every attack.\n",
    "\n",
    "**Possible Solution:** We use techniques like **SMOTE** (Synthetic Minority Over-sampling Technique) to generate fake examples of attacks, forcing the model to pay attention to them."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from imblearn.over_sampling import SMOTE\n",
    "from collections import Counter\n",
    "\n",
    "# 1. Imbalanced Dataset (99 Benign, 1 Attack)\n",
    "X = [[1.0, 2.0]] * 99 + [[9.0, 9.0]] * 1\n",
    "y = [0] * 99 + [1] * 1\n",
    "\n",
    "print(f\"Original Counts: {Counter(y)}\")\n",
    "# Output: {0: 99, 1: 1}\n",
    "\n",
    "# 2. Apply SMOTE\n",
    "# It creates synthetic attacks by interpolating between existing ones\n",
    "smote = SMOTE(k_neighbors=1) # k=1 because we only have 1 sample!\n",
    "X_res, y_res = smote.fit_resample(X, y)\n",
    "\n",
    "print(f\"Resampled Counts: {Counter(y_res)}\")\n",
    "# Output: {0: 99, 1: 99} -> Perfectly balanced!"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## The Encryption Challenge (\"Going Dark\")\n",
    "Historically, security tools inspected the \"payload\" of packets (e.g., looking for the word \"malware\" in an HTML body). Today, over 90% of web traffic is encrypted (HTTPS/TLS 1.3), making payloads unreadable.\n",
    "\n",
    "**Possible Solution:** Machine Learning has shifted to **Encrypted Traffic Analysis (ETA)**. Instead of reading the content, models analyze the *metadata*:\n",
    "* Sequence of packet lengths.\n",
    "* Inter-arrival times (the pause between packets).\n",
    "* Handshake patterns (Cipher suites offered).\n",
    "\n",
    "Deep Learning models (like LSTMs) are particularly good at fingerprinting applications (e.g., distinguishing Netflix traffic from Data Exfiltration) purely based on these timing patterns.\n",
    "\n",
    "## Concept Drift\n",
    "**Concept drift** occurs when the behavior of attacks changes over time. An ML model is a snapshot of the past. If trained on malware from 2020, it will likely fail against ransomware from 2026.\n",
    "\n",
    "**Possible Solution:** Continuous Retraining pipelines are required. In security, a model is often considered \"stale\" after just a few weeks.\n",
    "\n",
    "> **Figure:** Concept Drift: As attackers change tactics, the old decision boundary becomes obsolete.\n",
    "\n",
    "## Adversarial Examples\n",
    "An **adversarial example** is an input carefully crafted to fool an AI. An attacker can add invisible \"noise\" to a malware file that changes its mathematical signature just enough to cross the decision boundary from \"Malware\" to \"Benign.\"\n",
    "\n",
    "> **Figure:** Adversarial Attacks: Adding invisible noise to fool the model.\n",
    "\n",
    "## Interpretability (Explainable AI - XAI)\n",
    "Deep Learning models are \"black boxes.\" If a firewall blocks the CEO's laptop saying \"99% Malicious,\" the security analyst needs to know *why* to verify the alert.\n",
    "\n",
    "**Possible Solution:** We use tools like *SHAP (SHapley Additive exPlanations)* to calculate which specific feature pushed the probability up."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import shap\n",
    "import xgboost\n",
    "\n",
    "# 1. Train a model (Black Box)\n",
    "# Note: Using X_train, y_train from previous sections\n",
    "model = xgboost.XGBClassifier().fit(X_train, y_train)\n",
    "\n",
    "# 2. Use SHAP to explain a single prediction\n",
    "explainer = shap.Explainer(model)\n",
    "shap_values = explainer(X_test)\n",
    "\n",
    "# 3. Visualize\n",
    "# This tells the analyst: \"Blocked because 'Packet_Size' > 1500\"\n",
    "shap.plots.waterfall(shap_values[0])"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Hands-on Project: Basic ML Workflow for Threat Classification\n",
    "\n",
    "This project will guide you through a complete, basic supervised machine learning workflow. We will build a simple Network Intrusion Detection System (NIDS) using a Random Forest classifier.\n",
    "\n",
    "> **Figure:** Basic ML Workflow for Threat Classification\n",
    "\n",
    "## Step 1: Data Loading\n",
    "First, we create our dataset. In a real-world scenario, you would load this from a CSV file (e.g., `pd.read_csv('network_logs.csv')`). Here, we create a small synthetic dataset representing network flows with features like Packet Size and SYN Flag counts."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import pandas as pd\n",
    "import numpy as np\n",
    "\n",
    "# 1. Create a dummy dataset mimicking network traffic\n",
    "data = {\n",
    "    # Feature A: Average size of packets in the flow (Bytes)\n",
    "    'packet_size': [64, 1500, 64, 200, 64, 1500, 500, 64, 1500, 64],\n",
    "    \n",
    "    # Feature B: Number of SYN flags (High counts often indicate scanning/DoS)\n",
    "    'syn_count': [0, 0, 1, 50, 0, 0, 20, 0, 0, 100],\n",
    "    \n",
    "    # Feature C: Duration of the connection (Seconds)\n",
    "    'duration': [0.1, 5.2, 0.1, 1.5, 0.2, 10.0, 0.5, 0.1, 8.5, 2.0],\n",
    "    \n",
    "    # Target Label: 0 = Benign, 1 = Attack\n",
    "    'label': [0, 0, 0, 1, 0, 0, 1, 0, 0, 1]\n",
    "}\n",
    "\n",
    "df_flows = pd.DataFrame(data)\n",
    "X = df_flows.drop('label', axis=1) # Features\n",
    "y = df_flows['label']             # Target labels\n",
    "\n",
    "print(\"--- Data Loaded ---\")\n",
    "print(X.head())"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Step 2: Data Splitting\n",
    "We split our data into a training set and a testing set. We use `stratify=y` to ensure that if 30% of our total data is malicious, then 30% of our training set and 30% of our test set will also be malicious. This is crucial for maintaining the \"signal\" in imbalanced security data."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.model_selection import train_test_split\n",
    "\n",
    "# 2. Split Data (70% Train, 30% Test)\n",
    "X_train, X_test, y_train, y_test = train_test_split(\n",
    "    X, y, test_size=0.3, random_state=42, stratify=y\n",
    ")\n",
    "\n",
    "print(f\"Training shape: {X_train.shape}\")\n",
    "print(f\"Testing shape:  {X_test.shape}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Step 3: Model Selection and Cross-Validation\n",
    "We choose a *Random Forest* classifier because it handles mixed numerical data well and resists overfitting. Before the final training, we use Cross-Validation to check if the model is stable or if it just got lucky with easy data points."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.ensemble import RandomForestClassifier\n",
    "from sklearn.model_selection import cross_val_score\n",
    "\n",
    "# 3. Initialize Model and Check Stability\n",
    "model = RandomForestClassifier(random_state=42)\n",
    "\n",
    "# cv=2 splits the training data into 2 parts to validate performance\n",
    "scores = cross_val_score(model, X_train, y_train, cv=2, scoring='accuracy')\n",
    "\n",
    "print(f\"Cross-Validation Scores: {scores}\")\n",
    "print(f\"Average Stability: {scores.mean():.2f}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Step 4: Hyperparameter Tuning\n",
    "We search for the best configuration using `GridSearchCV`. We try different numbers of trees (`n_estimators`) to see which yields the best F1-score."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.model_selection import GridSearchCV\n",
    "\n",
    "# 4. Hyperparameter Tuning\n",
    "# We test: 10 vs 50 trees, and varying depth\n",
    "param_grid = {'n_estimators': [10, 50], 'max_depth': [None, 5]}\n",
    "\n",
    "grid_search = GridSearchCV(\n",
    "    RandomForestClassifier(random_state=42), \n",
    "    param_grid, \n",
    "    cv=2, \n",
    "    scoring='f1' # Optimize for F1 because attacks are rare\n",
    ")\n",
    "grid_search.fit(X_train, y_train)\n",
    "\n",
    "best_model = grid_search.best_estimator_\n",
    "print(f\"Best Parameters: {grid_search.best_params_}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Step 5: Model Training & Evaluation\n",
    "Using the best settings found above, the model makes predictions on the **Testing Set** (data it has never seen). We use a classification report to see exactly how well it caught the attacks."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sklearn.metrics import classification_report\n",
    "\n",
    "# 5. Final Evaluation\n",
    "y_pred = best_model.predict(X_test)\n",
    "\n",
    "print(\"--- Final Test Report ---\")\n",
    "print(classification_report(y_test, y_pred, target_names=['Benign', 'Attack']))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Step 6: Feature Importance\n",
    "Finally, we ask the model: \"Which feature was most important?\" In cybersecurity, this explains *why* the model flagged an event. Here, we expect `syn_count` to be highly important for detecting attacks."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import matplotlib.pyplot as plt\n",
    "\n",
    "# 6. Visualize Explanations\n",
    "importances = best_model.feature_importances_\n",
    "indices = np.argsort(importances)[::-1]\n",
    "\n",
    "plt.figure(figsize=(8, 4))\n",
    "plt.title(\"What did the model learn?\")\n",
    "plt.bar(range(X.shape[1]), importances[indices], align=\"center\")\n",
    "plt.xticks(range(X.shape[1]), X.columns[indices], rotation=45)\n",
    "plt.ylabel(\"Importance Score\")\n",
    "plt.tight_layout()\n",
    "plt.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Complete Project Code\n",
    "The code below combines all the steps above into a single, executable script."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import pandas as pd\n",
    "import numpy as np\n",
    "import matplotlib.pyplot as plt\n",
    "from sklearn.model_selection import train_test_split, cross_val_score, GridSearchCV\n",
    "from sklearn.ensemble import RandomForestClassifier\n",
    "from sklearn.metrics import classification_report\n",
    "\n",
    "# --- 1. DATA LOADING ---\n",
    "data = {\n",
    "    'packet_size': [64, 1500, 64, 200, 64, 1500, 500, 64, 1500, 64],\n",
    "    'syn_count': [0, 0, 1, 50, 0, 0, 20, 0, 0, 100],\n",
    "    'duration': [0.1, 5.2, 0.1, 1.5, 0.2, 10.0, 0.5, 0.1, 8.5, 2.0],\n",
    "    'label': [0, 0, 0, 1, 0, 0, 1, 0, 0, 1]\n",
    "}\n",
    "df_flows = pd.DataFrame(data)\n",
    "X = df_flows.drop('label', axis=1)\n",
    "y = df_flows['label']\n",
    "\n",
    "print(\"--- Data Sample ---\")\n",
    "print(X.head())\n",
    "\n",
    "# --- 2. DATA SPLITTING ---\n",
    "X_train, X_test, y_train, y_test = train_test_split(\n",
    "    X, y, test_size=0.3, random_state=42, stratify=y\n",
    ")\n",
    "\n",
    "# --- 3 & 4. TUNING AND TRAINING ---\n",
    "# We combine cross-validation and grid search\n",
    "param_grid = {'n_estimators': [10, 50], 'max_depth': [None, 5]}\n",
    "grid_search = GridSearchCV(\n",
    "    RandomForestClassifier(random_state=42), \n",
    "    param_grid, \n",
    "    cv=2, \n",
    "    scoring='f1'\n",
    ")\n",
    "grid_search.fit(X_train, y_train)\n",
    "best_model = grid_search.best_estimator_\n",
    "\n",
    "print(f\"\\nBest Parameters: {grid_search.best_params_}\")\n",
    "\n",
    "# --- 5. EVALUATION ---\n",
    "y_pred = best_model.predict(X_test)\n",
    "print(\"\\n--- Final Test Report ---\")\n",
    "print(classification_report(y_test, y_pred, target_names=['Benign', 'Attack']))\n",
    "\n",
    "# --- 6. VISUALIZATION ---\n",
    "importances = best_model.feature_importances_\n",
    "indices = np.argsort(importances)[::-1]\n",
    "\n",
    "plt.figure(figsize=(8, 4))\n",
    "plt.title(\"Feature Importance: What drives the alerts?\")\n",
    "plt.bar(range(X.shape[1]), importances[indices], align=\"center\")\n",
    "plt.xticks(range(X.shape[1]), X.columns[indices], rotation=45)\n",
    "plt.ylabel(\"Importance\")\n",
    "plt.tight_layout()\n",
    "plt.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Conclusion\n",
    "This chapter has laid the essential groundwork for applying machine learning to cybersecurity. We have demystified the core concepts of AI, ML, and DL and explored the fundamental learning paradigms that form the basis of all intelligent systems. You learned about the key model families, from classical algorithms suited for structured data to deep learning architectures that excel with complex, raw data.\n",
    "\n",
    "Most importantly, we have focused on the practical realities of building these systems for security. You now understand the critical role of feature engineering, the necessity of choosing evaluation metrics like precision and recall over simple accuracy, and the significant challenges posed by imbalanced data and adversarial attacks. The hands-on project provided a concrete, step-by-step workflow that you can adapt for your own security data analysis tasks. With this foundation, you are now prepared to move on to more advanced applications of AI in building robust and intelligent defense systems."
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.8.5"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}
