For a full description of the assignment, see the assignment handout at cs336_assignment1_basics.pdf
If you see any issues with the assignment handout or code, please feel free to raise a GitHub issue or open a pull request with a fix.
We manage our environments with uv to ensure reproducibility, portability, and ease of use.
Install uv here (recommended), or run pip install uv/brew install uv.
We recommend reading a bit about managing projects in uv here (you will not regret it!).
You can now run any code in the repo using
uv run <python_file_path>and the environment will be automatically solved and activated when necessary.
uv run pytestInitially, all tests should fail with NotImplementedErrors.
To connect your implementation to the tests, complete the
functions in ./tests/adapters.py.
Student implementations belong in cs336_basics/. The tests are kept unchanged;
tests/adapters.py is the small compatibility layer that imports and calls the
student code.
The tokenizer implementation currently connects these two assignment interfaces:
get_tokenizer(...)constructscs336_basics.tokenizer.BPETokenizer.run_train_bpe(...)callscs336_basics.tokenizer.train_bpe(...).
Run the completed tokenizer section with:
uv run pytest -q tests/test_tokenizer.py tests/test_train_bpe.pyThe same command runs in GitHub Actions for every pull request. As later sections are implemented, connect each new component through its adapter and add its test file to the workflow.
Download the TinyStories data and a subsample of OpenWebText
mkdir -p data
cd data
wget https://huggingface.co/datasets/roneneldan/TinyStories/resolve/main/TinyStoriesV2-GPT4-train.txt
wget https://huggingface.co/datasets/roneneldan/TinyStories/resolve/main/TinyStoriesV2-GPT4-valid.txt
wget https://huggingface.co/datasets/stanford-cs336/owt-sample/resolve/main/owt_train.txt.gz
gunzip owt_train.txt.gz
wget https://huggingface.co/datasets/stanford-cs336/owt-sample/resolve/main/owt_valid.txt.gz
gunzip owt_valid.txt.gz
cd ..