Skip to content
Discussion options

You must be logged in to vote

Thanks for the note. You are right. Looks like a bad typo in the figure. The next page is correct:

The output tensor has two rows corresponding to the two text samples. Each text sam-
ple consists of four tokens; each token is a 50,257-dimensional vector, which matches
the size of the tokenizer’s vocabulary.
The embedding has 50,257 dimensions because each of these dimensions refers to
a unique token in the vocabulary.

Replies: 1 comment

Comment options

You must be logged in to vote
0 replies
Answer selected by bowenxieoregon
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
2 participants