Skip to content

Latest commit

 

History

33 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Learning Neural Networks from Andrej Karpathy Videos

Andrej Karpathy has a "Neural Networks: Zero to Hero" video series. I am replicating the code from scratch following his YouTube videos/lectures from the original repo. Original repository can be found at https://github.com/karpathy/nn-zero-to-hero/tree/master.


Setting up the environment and installing packages to run code in VSCode.

  • conda create -n nanogpt python==3.8
  • pip install torch==2.4.0 torchvision==0.19.0 torchaudio==2.4.0 --index-url https://download.pytorch.org/whl/cu124
  • pip install transformers==2.11.0 datasets==2.21.0
  • pip install tiktoken wandb tqdm

Popular videos


GPT folder: Let's build GPT: from scratch, in code, spelled out.

The lecture builds a Generatively Pretrained Transformer (GPT), following the paper "Attention is All You Need" and OpenAI's GPT-2 / GPT-3. GPT can be thought of as a next-word (token) predictor or a document (prompt) completer. The video explains attention mechanism, the core of the transformer architecture, in a intutive way.

  • Tokenizer: Character-level tokenizer
  • Dataset: Shakespeare text (input.txt)

Here is a simple way to understand attention mechanism:

Imagine you have a cat and, for fun, you keep a journal where you note all the places your cat sat during the day (unlikely, but bear with me). Now, you need to complete this sentence:

"The cat is sitting on the ____."

How would you figure out the answer, given options like "mat," "couch," "floor," "bed," or "chair"? Let's look at a few approaches:

  • You could refer to your journal and count how often each place appears after "sitting on." The place with the highest count, say "mat," would be your best guess. While simple, this method ignores the context of the current sentence i.e, are we talking about cat or dog.

  • Another approach might be to consider all the words in the sentence and average their "meanings", treating them equally without considering their importance or order. In simple terms, this means treating all the words equally without understanding which ones are more relevant. While this might give a rough guess, it fails to emphasize the context or relevance of specific words and their spatial arrangements in the sentence.

  • Now think about how you would naturally approach this. You would process the sentence by focusing on key elements—who we’re talking about ("cat") and what it’s doing ("sitting"). Not all words are equally important, and their positions in the sentence also matter. The attention mechanism captures this by calculating a weighted average of the information (values) from all words. The weights are determined based on how relevant each word (key) is to the word we’re focusing on (query). Words that are more important or closely related to the query get higher weights, while irrelevant ones get lower weights. This allows the model to focus on the right words and their relationships, enabling it to make accurate predictions based on context.


GPT-2 folder: Let's reproduce GPT-2.


About

a personal copy of karpathy's nn-zero-to-hero

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages