{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "-zSRp65WHKjm"
      },
      "source": [
        "# 🧪 Lab: Python for Data Science & AI\n",
        "\n",
        "**Objective:** Master the essential stack for modern Data Science, from numerical computing to Large Language Models (LLMs).\n",
        "\n",
        "**Libraries Covered:**\n",
        "* **NumPy:** Numerical Math\n",
        "* **Pandas:** Tabular Data\n",
        "* **Seaborn & Plotly:** Visualization\n",
        "* **Scikit-Learn:** Classical ML\n",
        "* **PyTorch:** Deep Learning\n",
        "* **Hugging Face:** Working with LLMs"
      ],
      "id": "-zSRp65WHKjm"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "TFbd2EOrHKjo"
      },
      "source": [
        "## 1. Setup and Installation\n",
        "Run the cell below to ensure all required libraries are installed in your environment."
      ],
      "id": "TFbd2EOrHKjo"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "bxljOvSHHKjo"
      },
      "outputs": [],
      "source": [
        "%pip install numpy pandas seaborn plotly scikit-learn torch transformers langchain-community"
      ],
      "id": "bxljOvSHHKjo"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "JPPpZmjmHKjp"
      },
      "source": [
        "---"
      ],
      "id": "JPPpZmjmHKjp"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "Fj4hQousHKjp"
      },
      "source": [
        "## 2. Numerical Computing: NumPy\n",
        "**Use Case:** High-performance mathematical operations and matrix manipulation.\n",
        "\n",
        "### Example: Array Broadcasting\n",
        "NumPy allows operations on arrays of different shapes, which is crucial for vectorizing code without loops."
      ],
      "id": "Fj4hQousHKjp"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "9AjFNWgTHKjq"
      },
      "outputs": [],
      "source": [
        "import numpy as np\n",
        "\n",
        "# Create a matrix (3x3)\n",
        "matrix = np.array([[1, 2, 3],\n",
        "                   [4, 5, 6],\n",
        "                   [7, 8, 9]])\n",
        "\n",
        "# Create a vector (1x3)\n",
        "vector = np.array([1, 0, 1])\n",
        "\n",
        "# Broadcasting: Add vector to every row of the matrix\n",
        "result = matrix + vector\n",
        "\n",
        "print(\"Original Matrix:\\n\", matrix)\n",
        "print(\"\\nResult after broadcasting:\\n\", result)"
      ],
      "id": "9AjFNWgTHKjq"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "e06AUEo1HKjq"
      },
      "source": [
        "### 🧠 Exercise 1: Matrix Normalization\n",
        "**Task:**\n",
        "1. Create a random $5 \\times 5$ matrix using `np.random.rand`.\n",
        "2. Calculate the **mean** and **standard deviation** of the entire matrix.\n",
        "3. Normalize the matrix using the formula: $$Z = \\frac{X - \\mu}{\\sigma}$$"
      ],
      "id": "e06AUEo1HKjq"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "9aL7iR7_HKjr"
      },
      "outputs": [],
      "source": [
        "# TODO: Write your code here\n"
      ],
      "id": "9aL7iR7_HKjr"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "DeOVU2HrHKjr"
      },
      "source": [
        "---"
      ],
      "id": "DeOVU2HrHKjr"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "3RFPLyeHHKjs"
      },
      "source": [
        "## 3. Tabular Data: Pandas\n",
        "**Use Case:** Data manipulation, cleaning, and analysis (Excel for Python).\n",
        "\n",
        "### Example: GroupBy and Aggregation"
      ],
      "id": "3RFPLyeHHKjs"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "Q7VSDK4-HKjs"
      },
      "outputs": [],
      "source": [
        "import pandas as pd\n",
        "\n",
        "data = {\n",
        "    'Category': ['A', 'B', 'A', 'C', 'B', 'A'],\n",
        "    'Value': [10, 20, 15, 5, 25, 10],\n",
        "    'Status': ['Open', 'Closed', 'Open', 'Open', 'Closed', 'Closed']\n",
        "}\n",
        "\n",
        "df = pd.DataFrame(data)\n",
        "\n",
        "# Group by Category and calculate the average Value\n",
        "grouped_df = df.groupby('Category')['Value'].mean().reset_index()\n",
        "\n",
        "print(\"Average Value per Category:\")\n",
        "print(grouped_df)"
      ],
      "id": "Q7VSDK4-HKjs"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "7TtjnUBwHKjs"
      },
      "source": [
        "### 🧠 Exercise 2: Filtering and Aggregation\n",
        "**Task:**\n",
        "1. Filter the rows where `Status` is **'Open'**.\n",
        "2. Calculate the **sum** of the `Value` column for these filtered rows."
      ],
      "id": "7TtjnUBwHKjs"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "HRNMhFr-HKjt"
      },
      "outputs": [],
      "source": [
        "# TODO: Write your code here\n"
      ],
      "id": "HRNMhFr-HKjt"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "bVDVRgLGHKjt"
      },
      "source": [
        "---"
      ],
      "id": "bVDVRgLGHKjt"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "ZUaThPowHKjt"
      },
      "source": [
        "## 4. Visualization: Seaborn & Plotly\n",
        "**Use Case:** Statistical plots (`seaborn`) and Interactive dashboards (`plotly`).\n",
        "\n",
        "### Example A: Statistical Distribution (Seaborn)"
      ],
      "id": "ZUaThPowHKjt"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "NdtZqXHnHKjt"
      },
      "outputs": [],
      "source": [
        "import seaborn as sns\n",
        "import matplotlib.pyplot as plt\n",
        "\n",
        "# Generate dummy data\n",
        "data_dist = np.random.normal(loc=0, scale=1, size=1000)\n",
        "\n",
        "plt.figure(figsize=(8, 4))\n",
        "sns.histplot(data_dist, kde=True, color='teal')\n",
        "plt.title(\"Normal Distribution with Seaborn\")\n",
        "plt.show()"
      ],
      "id": "NdtZqXHnHKjt"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "O1IQ_-TiHKjt"
      },
      "source": [
        "### Example B: Interactive Plot (Plotly)\n",
        "Try hovering over the points in the chart below."
      ],
      "id": "O1IQ_-TiHKjt"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "7a7P_QsOHKjt"
      },
      "outputs": [],
      "source": [
        "import plotly.express as px\n",
        "\n",
        "# Using built-in Iris dataset\n",
        "df_iris = px.data.iris()\n",
        "\n",
        "fig = px.scatter(df_iris, x=\"sepal_width\", y=\"sepal_length\",\n",
        "                 color=\"species\", size=\"petal_length\",\n",
        "                 title=\"Interactive Iris Dataset Scatter Plot\")\n",
        "fig.show()"
      ],
      "id": "7a7P_QsOHKjt"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "9DL96hP4HKjt"
      },
      "source": [
        "### 🧠 Exercise 3: Correlation Heatmap\n",
        "**Task:**\n",
        "1. Calculate the correlation matrix of the `df_iris` dataset using `df.corr(numeric_only=True)`.\n",
        "2. Use `sns.heatmap` to visualize this correlation matrix."
      ],
      "id": "9DL96hP4HKjt"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "aFKsDyiAHKjt"
      },
      "outputs": [],
      "source": [
        "# TODO: Write your code here\n"
      ],
      "id": "aFKsDyiAHKjt"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "bVlJ_3V7HKju"
      },
      "source": [
        "---"
      ],
      "id": "bVlJ_3V7HKju"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "1ICvTOgVHKju"
      },
      "source": [
        "## 5. Classical Machine Learning: Scikit-Learn\n",
        "**Use Case:** Regression, Classification, Clustering.\n",
        "\n",
        "### Example: Random Forest Classification"
      ],
      "id": "1ICvTOgVHKju"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "TjeE_ynKHKju"
      },
      "outputs": [],
      "source": [
        "from sklearn.model_selection import train_test_split\n",
        "from sklearn.ensemble import RandomForestClassifier\n",
        "from sklearn.metrics import accuracy_score\n",
        "from sklearn.datasets import load_wine\n",
        "\n",
        "# Load Data\n",
        "wine_data = load_wine()\n",
        "X, y = wine_data.data, wine_data.target\n",
        "\n",
        "# Split Data (80% Train, 20% Test)\n",
        "X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n",
        "\n",
        "# Initialize and Train Model\n",
        "clf = RandomForestClassifier(n_estimators=100)\n",
        "clf.fit(X_train, y_train)\n",
        "\n",
        "# Predict and Evaluate\n",
        "preds = clf.predict(X_test)\n",
        "print(f\"Model Accuracy: {accuracy_score(y_test, preds):.2f}\")"
      ],
      "id": "TjeE_ynKHKju"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "xH4chItKHKju"
      },
      "source": [
        "### 🧠 Exercise 4: Dimensionality Reduction (PCA)\n",
        "**Task:**\n",
        "1. Import `PCA` from `sklearn.decomposition`.\n",
        "2. Initialize PCA with `n_components=2`.\n",
        "3. Fit and transform the variable `X` (Wine dataset) and print the shape of the result (Should be 178 rows, 2 columns)."
      ],
      "id": "xH4chItKHKju"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "J8nZh66lHKjv"
      },
      "outputs": [],
      "source": [
        "# TODO: Write your code here\n"
      ],
      "id": "J8nZh66lHKjv"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "Ph8Q6HHzHKjv"
      },
      "source": [
        "---"
      ],
      "id": "Ph8Q6HHzHKjv"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "WYRbTM14HKjv"
      },
      "source": [
        "## 6. Deep Learning: PyTorch\n",
        "**Use Case:** Neural Networks, Computer Vision, Complex AI systems.\n",
        "\n",
        "### Example: Defining a Neural Network"
      ],
      "id": "WYRbTM14HKjv"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "9svxUbSlHKjv"
      },
      "outputs": [],
      "source": [
        "import torch\n",
        "import torch.nn as nn\n",
        "\n",
        "# Define the Network\n",
        "class SimpleNet(nn.Module):\n",
        "    def __init__(self):\n",
        "        super(SimpleNet, self).__init__()\n",
        "        self.fc1 = nn.Linear(10, 5)  # Input layer (10 features) -> Hidden (5)\n",
        "        self.relu = nn.ReLU()        # Activation\n",
        "        self.fc2 = nn.Linear(5, 1)   # Hidden -> Output (1)\n",
        "\n",
        "    def forward(self, x):\n",
        "        out = self.fc1(x)\n",
        "        out = self.relu(out)\n",
        "        out = self.fc2(out)\n",
        "        return out\n",
        "\n",
        "# Initialize model\n",
        "model = SimpleNet()\n",
        "\n",
        "# Create a random input tensor (Batch size 1, 10 features)\n",
        "input_tensor = torch.randn(1, 10)\n",
        "\n",
        "# Forward pass\n",
        "output = model(input_tensor)\n",
        "print(\"Network Output:\", output.item())"
      ],
      "id": "9svxUbSlHKjv"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "sV7IkQgpHKjv"
      },
      "source": [
        "### 🧠 Exercise 5: Basic Tensor Math\n",
        "**Task:**\n",
        "1. Create two random tensors, `A` and `B`, both of size (3, 3).\n",
        "2. Perform matrix multiplication (`torch.matmul`).\n",
        "3. Print the result."
      ],
      "id": "sV7IkQgpHKjv"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "1nbaFzpHHKjv"
      },
      "outputs": [],
      "source": [
        "# TODO: Write your code here\n"
      ],
      "id": "1nbaFzpHHKjv"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "QvE69nYkHKjv"
      },
      "source": [
        "---"
      ],
      "id": "QvE69nYkHKjv"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "aqCCmOOIHKjv"
      },
      "source": [
        "## 7. Working with LLMs: Hugging Face\n",
        "**Use Case:** NLP, Sentiment Analysis, Chatbots.\n",
        "\n",
        "### Example: Sentiment Analysis Pipeline"
      ],
      "id": "aqCCmOOIHKjv"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "ymUQbYpJHKjv"
      },
      "outputs": [],
      "source": [
        "from transformers import pipeline\n",
        "\n",
        "# Load a specific sentiment-analysis pipeline\n",
        "classifier = pipeline(\"sentiment-analysis\", model=\"distilbert-base-uncased-finetuned-sst-2-english\")\n",
        "\n",
        "text = \"I absolutely love learning about data science, it makes me feel powerful!\"\n",
        "result = classifier(text)\n",
        "\n",
        "print(f\"Text: {text}\")\n",
        "print(f\"Sentiment: {result}\")"
      ],
      "id": "ymUQbYpJHKjv"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "joixYiGBHKjw"
      },
      "source": [
        "### 🧠 Exercise 6: Text Generation\n",
        "**Task:**\n",
        "1. Create a `pipeline` for `\"text-generation\"` using `model=\"gpt2\"`.\n",
        "2. Generate a continuation for the prompt: `\"The future of Artificial Intelligence is\"`."
      ],
      "id": "joixYiGBHKjw"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "UiuktCFEHKjw"
      },
      "outputs": [],
      "source": [
        "# TODO: Write your code here\n"
      ],
      "id": "UiuktCFEHKjw"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "s_KOTiTiHKjw"
      },
      "source": [
        "---"
      ],
      "id": "s_KOTiTiHKjw"
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.8.10"
    },
    "colab": {
      "provenance": []
    }
  },
  "nbformat": 4,
  "nbformat_minor": 5
}