In [None]:
!pip install --no-deps xformers "trl<0.9.0" peft accelerate bitsandbytes datasets
!pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"

In [None]:
import torch
import numpy as np
import pandas as pd
from datasets import load_dataset
from huggingface_hub import notebook_login
from transformers import TrainingArguments
from trl import SFTTrainer

In [None]:
from torch import __version__; from packaging.version import Version as V
xformers = "xformers==0.0.27" if V(__version__) < V("2.4.0") else "xformers"
!pip install --no-deps {xformers} trl peft accelerate bitsandbytes triton

In [None]:
from unsloth import FastLanguageModel

🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning.


In [None]:
# HuggingFace login -- training model/data & Pushing model
!pip install huggingface_hub
from huggingface_hub import notebook_login

notebook_login()

In [None]:
max_seq_length = None
dtype = None # None for auto detection. Float16 for Tesla T4, V100, Bfloat16 for Ampere+
load_in_4bit = True # Use 4bit quantization to reduce memory usage. Can be False.

In [None]:
# Loading the fine-tuned model and the tokenizer for inference
model, tokenizer = FastLanguageModel.from_pretrained(
        model_name = "LimYeri/CodeMind-Llama3-8B-unsloth_v4-one",
        max_seq_length = max_seq_length,
        dtype = dtype,
        load_in_4bit = load_in_4bit,
    )

==((====))==  Unsloth 2024.8: Fast Llama patching. Transformers = 4.44.2.
   \\   /|    GPU: NVIDIA A100-SXM4-40GB. Max memory: 39.564 GB. Platform = Linux.
O^O/ \_/ \    Pytorch: 2.4.0+cu121. CUDA = 8.0. CUDA Toolkit = 12.1.
\        /    Bfloat16 = TRUE. FA [Xformers = 0.0.27.post2. FA2 = False]
 "-____-"     Free Apache license: http://github.com/unslothai/unsloth
Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!


model.safetensors:   0%|          | 0.00/5.70G [00:00<?, ?B/s]

generation_config.json:   0%|          | 0.00/220 [00:00<?, ?B/s]

tokenizer_config.json:   0%|          | 0.00/51.1k [00:00<?, ?B/s]

tokenizer.json:   0%|          | 0.00/9.09M [00:00<?, ?B/s]

special_tokens_map.json:   0%|          | 0.00/345 [00:00<?, ?B/s]

adapter_model.safetensors:   0%|          | 0.00/671M [00:00<?, ?B/s]

Unsloth 2024.8 patched 32 layers with 32 QKV layers, 32 O layers and 32 MLP layers.


### Inference

In [None]:
from IPython.display import display, Markdown

In [None]:
# alpaca_prompt = Copied from above
FastLanguageModel.for_inference(model) # Enable native 2x faster inference
inputs = tokenizer(
[
    "<|start_header_id|>system<|end_header_id|> You are a coding instructor. Below is a coding test problem. Write a Python code to solve the given problem or provide a detailed explanation of the approach to solving it.<|eot_id|><|start_header_id|>user<|end_header_id|> I don't know how to solve LeetCode problem 268: Missing Number. Could you provide a python solution or explain the approach?<|eot_id|>"
], return_tensors = "pt").to("cuda")

outputs = model.generate(**inputs, max_new_tokens = 3000, use_cache = True)
answer = tokenizer.batch_decode(outputs, skip_special_tokens = True)
text = answer[0].split("assistant")[1].strip()
display(Markdown(text))

```python
class Solution:
    def missingNumber(self, nums: List[int]) -> int:
        return (len(nums)*(len(nums)+1)//2) - sum(nums)
```


The given code is a solution to the LeetCode problem 268, which asks to find the missing number in an array of integers. Let's break down the code and analyze it:

1. **Data Structures Used**:
   - The code uses a Python class `Solution` with a method `missingNumber`.
   - The input parameter `nums` is a list of integers.
   - The code does not use any additional data structures like dictionaries or sets.

2. **Code Explanation**:
   - The `missingNumber` method calculates the missing number in the given list of integers.
   - It calculates the sum of all numbers from 0 to the length of the input list using the formula `(len(nums)*(len(nums)+1)//2)`.
   - It then subtracts the sum of all elements in the input list from the sum of all numbers from 0 to the length of the list.
   - The result of this subtraction gives the missing number in the list.

3. **Time Complexity**:
   - Calculating the sum of all numbers from 0 to n using the formula `(n*(n+1)//2)` takes O(1) time.
   - Summing all elements in the input list takes O(n) time where n is the length of the list.
   - Therefore, the overall time complexity of the code is O(n).

4. **Space Complexity**:
   - The code uses only a constant amount of extra space for storing variables.
   - Hence, the space complexity of the code is O(1).

In conclusion, the given code efficiently solves the problem of finding the missing number in an array of integers with a time complexity of O(n) and a space complexity of O(1).

In [None]:
# alpaca_prompt = Copied from above
FastLanguageModel.for_inference(model) # Enable native 2x faster inference
inputs = tokenizer(
[
    "<|start_header_id|>system<|end_header_id|> You are a coding instructor. Below is a coding test problem. Write a Python code to solve the given problem or provide a detailed explanation of the approach to solving it.<|eot_id|><|start_header_id|>user<|end_header_id|> I don't know how to solve LeetCode problem 3: Longest Substring Without Repeating Characters. Could you provide a python solution or explain the approach?<|eot_id|>"
], return_tensors = "pt").to("cuda")

outputs = model.generate(**inputs, max_new_tokens = 3000, use_cache = True)
answer = tokenizer.batch_decode(outputs, skip_special_tokens = True)
text = answer[0].split("assistant")[1].strip()
display(Markdown(text))

```python
class Solution:
    def lengthOfLongestSubstring(self, s: str) -> int:
        maxLength = left = 0
        seen = {}
        
        for i in range(len(s)):
            l = left
            r = i
            
            if s[i] in seen and l <= seen[s[i]]:
                left = seen[s[i]] + 1
            else:
                maxLength = max(maxLength, r - l + 1)
                
            seen[s[i]] = i
        return maxLength
```


The given code is a solution to the LeetCode problem 3, which asks to find the length of the longest substring without repeating characters in a given string `s`.

Explanation of the code:
1. The `lengthOfLongestSubstring` function takes a string `s` as input and returns an integer, which is the length of the longest substring without repeating characters.
2. `maxLength`, `left`, and `seen` are initialized. `maxLength` stores the length of the longest substring found so far, `left` is the index of the left boundary of the current substring, and `seen` is a dictionary that stores the index of the last occurrence of each character in the string.
3. The code iterates through the characters of the input string `s` using a for loop.
4. For each character, it checks if the character is already seen and if it is within the current substring boundaries. If it is, it updates the left boundary to the next index after the last occurrence of the character.
5. It calculates the length of the current substring (`r - l + 1`) and updates `maxLength` if the current substring is longer.
6. It updates the index of the current character in the `seen` dictionary.
7. Finally, it returns the `maxLength` which represents the length of the longest substring without repeating characters.

Data structures used:
- Integer variables `maxLength`, `left`, and `r` are used to keep track of the length of the longest substring and the boundaries of the current substring.
- A dictionary `seen` is used to store the index of the last occurrence of each character in the string.

Time complexity analysis:
- The code iterates through each character of the input string once, so the time complexity is O(n), where n is the length of the input string `s`.

Space complexity analysis:
- The space complexity is O(min(n, m)), where n is the length of the input string `s` and m is the size of the character set (in this case, all possible characters in the input string). The space complexity is determined by the size of the `seen` dictionary, which can store at most m characters.

Overall, the code efficiently solves the problem by keeping track of the longest substring length and the indices of characters to avoid repeating characters.

In [None]:
# alpaca_prompt = Copied from above
FastLanguageModel.for_inference(model) # Enable native 2x faster inference
inputs = tokenizer(
[
    "<|start_header_id|>system<|end_header_id|> You are a coding instructor. Below is a coding test problem. Write a Python code to solve the given problem or provide a detailed explanation of the approach to solving it.<|eot_id|><|start_header_id|>user<|end_header_id|> I don't know how to solve LeetCode problem 121: Best Time to Buy and Sell Stock. Could you provide a python solution or explain the approach?<|eot_id|>"
], return_tensors = "pt").to("cuda")

outputs = model.generate(**inputs, max_new_tokens = 3000, use_cache = True)
answer = tokenizer.batch_decode(outputs, skip_special_tokens = True)
text = answer[0].split("assistant")[1].strip()
display(Markdown(text))

```python
class Solution:
    def maxProfit(self, prices: List[int]) -> int:
        maxProfit = 0
        curMin = prices[0]
        
        for i in range(1, len(prices)):
            maxProfit = max(maxProfit, prices[i] - curMin)
            curMin = min(curMin, prices[i])
        
        return maxProfit
```


The given code is a solution to the LeetCode problem 121, which aims to find the maximum profit that can be achieved by buying and selling a stock once. Let's break down the code and analyze it:

### Code Explanation:
1. The code defines a class `Solution` with a method `maxProfit` that takes a list of integers `prices` as input and returns an integer (maximum profit).
2. It initializes two variables `maxProfit` and `curMin` to 0 and the first price in the list, respectively.
3. It then iterates through the prices starting from the second element (index 1) using a for loop.
4. For each price, it calculates the maximum profit by taking the maximum of the current `maxProfit` and the difference between the current price and `curMin`.
5. It updates `curMin` to be the minimum of the current price and the previous `curMin`.
6. Finally, it returns the maximum profit calculated.

### Data Structures Used:
- The code uses two variables `maxProfit` and `curMin` to keep track of the maximum profit and the minimum price encountered so far.
- It also uses a list of integers `prices` to store the input stock prices.

### Time Complexity Analysis:
- The code iterates through the prices list once, so the time complexity is O(n), where n is the number of prices in the input list.

### Space Complexity Analysis:
- The code uses a constant amount of extra space for variables `maxProfit` and `curMin`, so the space complexity is O(1).

### Overall:
- The code efficiently solves the problem by iterating through the prices list and updating the maximum profit and minimum price encountered.
- It has a time complexity of O(n) and a space complexity of O(1), making it an optimal solution for finding the maximum profit that can be achieved by buying and selling a stock once.

In [None]:
# alpaca_prompt = Copied from above
FastLanguageModel.for_inference(model) # Enable native 2x faster inference
inputs = tokenizer(
[
    "<|start_header_id|>system<|end_header_id|> You are a coding instructor. Below is a coding test problem. Write a Python code to solve the given problem or provide a detailed explanation of the approach to solving it.<|eot_id|><|start_header_id|>user<|end_header_id|> I don't know how to solve LeetCode problem 200: Number of Islands. Could you provide a python solution or explain the approach?<|eot_id|>"
], return_tensors = "pt").to("cuda")

outputs = model.generate(**inputs, max_new_tokens = 3000, use_cache = True)
answer = tokenizer.batch_decode(outputs, skip_special_tokens = True)
text = answer[0].split("assistant")[1].strip()
display(Markdown(text))

```python
class Solution:
    def numIslands(self, grid: List[List[str]]) -> int:
        if not grid: return 0
        m, n = len(grid), len(grid[0])
        ans = 0
        for i in range(m):
            for j in range(n):
                if grid[i][j] == "1":
                    q = collections.deque([(i, j)])
                    grid[i][j] = "0"
                    while q:
                        x, y = q.popleft()
                        for dx, dy in (0, 1), (0, -1), (1, 0), (-1, 0):
                            xx, yy = x+dx, y+dy
                            if 0 <= xx < m and 0 <= yy < n and grid[xx][yy] == "1":
                                q.append((xx, yy))
                                grid[xx][yy] = "0"
                    ans += 1
        return ans
```


The given code is a Python solution to the LeetCode problem 200 - "Number of Islands". The problem asks us to find the number of islands in a 2D grid where an island is a group of connected '1's. We are required to count the number of islands in the given grid.

Let's break down the code and understand it in detail:

1. The `numIslands` function takes a 2D list `grid` as input and returns the number of islands as an integer.

2. The code first checks if the input grid is empty. If it is empty, it returns 0 as there are no islands.

3. It then initializes variables `m` and `n` to store the number of rows and columns in the grid respectively.

4. It initializes a variable `ans` to store the count of islands found.

5. The code then iterates through each cell in the grid using nested loops.

6. For each cell, if the cell contains '1', it starts a breadth-first search (BFS) from that cell to explore the connected '1's.

7. It uses a deque `q` to store the cells that need to be visited. It starts with the current cell and marks it as '0' to avoid revisiting it.

8. While the deque is not empty, it pops a cell from the front of the deque and explores its neighbors in all four directions - up, down, left, and right.

9. If a neighbor cell is within the grid boundaries and contains '1', it adds it to the deque, marks it as '0', and continues the BFS.

10. After exploring all connected '1's from the current cell, it increments the island count `ans`.

11. Finally, it returns the total count of islands found.

Data Structures Used:
- List: Used to represent the 2D grid of cells.
- Deque: Used to implement the BFS traversal for exploring connected '1's.

Time Complexity Analysis:
- Let'm' be the number of rows and 'n' be the number of columns in the grid.
- The code iterates through each cell in the grid once, resulting in O(m*n) iterations.
- In the worst case, each cell is added to the deque once and explored in all four directions, resulting in O(m*n) operations.
- Therefore, the overall time complexity of the code is O(m*n).

Space Complexity Analysis:
- The code uses a deque to store cells for BFS traversal, which can grow up to O(m*n) in the worst case.
- Additionally, the code uses constant space for other variables.
- Therefore, the overall space complexity of the code is O(m*n).

This code efficiently solves the problem by using BFS to find and count the number of islands in the given grid.

In [None]:
# alpaca_prompt = Copied from above
FastLanguageModel.for_inference(model) # Enable native 2x faster inference
inputs = tokenizer(
[
    "<|start_header_id|>system<|end_header_id|> You are a kind coding test teacher. Solve the given question and give the insight and approach.<|eot_id|><|start_header_id|>user<|end_header_id|> I don't know how to solve LeetCode problem 200: Number of Islands. Could you provide a python solution or explain the approach?<|eot_id|>"
], return_tensors = "pt").to("cuda")

outputs = model.generate(**inputs, max_new_tokens = 3000, use_cache = True)
answer = tokenizer.batch_decode(outputs, skip_special_tokens = True)
text = answer[0].split("assistant")[1].strip()
display(Markdown(text))

```python
class Solution:
    def numIslands(self, grid: List[List[str]]) -> int:
        m = len(grid)
        n = len(grid[0])
        
        def dfs(i, j):
            grid[i][j] = "0"
            
            if i-1 >= 0 and grid[i-1][j] == "1":
                dfs(i-1, j)
            
            if i+1 < m and grid[i+1][j] == "1":
                dfs(i+1, j)
            
            if j-1 >= 0 and grid[i][j-1] == "1":
                dfs(i, j-1)
            
            if j+1 < n and grid[i][j+1] == "1":
                dfs(i, j+1)
        
        res = 0
        for i in range(m):
            for j in range(n):
                if grid[i][j] == "1":
                    res += 1
                    dfs(i, j)
        
        return res
```


The provided code is a Python solution to the LeetCode problem 200 - "Number of Islands". The goal is to count the number of islands in a 2D grid where an island is a group of connected '1's.

### Code Explanation:
1. The code defines a class `Solution` with a method `numIslands` that takes a 2D list `grid` as input and returns the number of islands as an integer.
2. It initializes variables `m` and `n` to store the number of rows and columns in the grid, respectively.
3. It defines a helper function `dfs` (depth-first search) that recursively marks all connected '1's as '0's to avoid counting the same island multiple times.
4. The main logic iterates through each cell in the grid. If the cell contains '1', it increments the island count `res`, calls the `dfs` function to mark all connected '1's, and finally returns the total number of islands.

### Data Structures Used:
1. **List:** The input grid is represented as a list of lists in Python.
2. **Recursion:** Depth-first search (DFS) is used to traverse and mark connected '1's.

### Time Complexity Analysis:
- Let `m` be the number of rows and `n` be the number of columns in the grid.
- The code iterates through each cell in the grid once, resulting in a time complexity of O(m*n).
- The DFS function also visits each cell at most once, contributing to the overall linear time complexity.

### Space Complexity Analysis:
- The space complexity is determined by the recursive calls in the DFS function.
- In the worst-case scenario, if all cells are '1's and form separate islands, the maximum depth of recursion is O(m*n).
- Therefore, the space complexity is O(m*n) due to the recursive stack space.

### Overall:
- The code efficiently solves the problem by using DFS to traverse and count islands in the grid.
- It has a time complexity of O(m*n) and a space complexity of O(m*n).
- The solution is optimal in terms of time and space complexity for this specific problem.