# **[spaCy 101: Everything you need to know](https://spacy.io/usage/spacy-101)**

## **Part-of-speech tags and dependencies**

After tokenization, spaCy can **parse** and **tag** a given `Doc`. This is where the trained pipeline and its statistical models come in, which enable spaCy to **make predictions of which tag or label most likely applies in this context**. A trained component includes binary data that is produced by showing a system enough examples for it to make predictions that generalize across the language – for example, a word following “the” in English is most likely a noun.

Linguistic annotations are available as Token attributes. Like many NLP libraries, spaCy **encodes all strings to hash values** to reduce memory usage and improve efficiency. So to get the readable string representation of an attribute, we need to add an underscore `_` to its name:

### **Example**

In [1]:
!pip install -U spacy --quiet

[?25l     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m0.0/6.6 MB[0m [31m?[0m eta [36m-:--:--[0m[2K     [91m━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m[90m╺[0m[90m━━━━━━━━━━━━[0m [32m4.5/6.6 MB[0m [31m135.3 MB/s[0m eta [36m0:00:01[0m[2K     [91m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m[91m╸[0m [32m6.6/6.6 MB[0m [31m144.3 MB/s[0m eta [36m0:00:01[0m[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m6.6/6.6 MB[0m [31m81.2 MB/s[0m eta [36m0:00:00[0m
[?25h

In [2]:
import spacy

nlp = spacy.load("en_core_web_sm")
doc = nlp("Apple is looking at buying U.K. startup for $1 billion")

for token in doc:
    print(token.text, token.pos_, token.tag_, token.dep_)

Apple PROPN NNP nsubj
is AUX VBZ aux
looking VERB VBG ROOT
at ADP IN prep
buying VERB VBG pcomp
U.K. PROPN NNP dobj
startup NOUN NN dep
for ADP IN prep
$ SYM $ quantmod
1 NUM CD compound
billion NUM CD pobj


In [3]:
# There is more information about a token to check
for token in doc:
    print(token.text, token.lemma_, token.pos_, token.tag_, token.dep_,
            token.shape_, token.is_alpha, token.is_stop)

Apple Apple PROPN NNP nsubj Xxxxx True False
is be AUX VBZ aux xx True True
looking look VERB VBG ROOT xxxx True False
at at ADP IN prep xx True True
buying buy VERB VBG pcomp xxxx True False
U.K. U.K. PROPN NNP dobj X.X. False False
startup startup NOUN NN dep xxxx True False
for for ADP IN prep xxx True True
$ $ SYM $ quantmod $ False False
1 1 NUM CD compound d False False
billion billion NUM CD pobj xxxx True False


### **Visualizing the dependency parse**

In [4]:
from spacy import displacy

displacy.render(doc, jupyter=True, style="dep")

### **Additional Resources**
- [spaCy | spaCy's NER model](https://spacy.io/universe/project/video-spacys-ner-model)
- [spaCy | EntityRecognizer](https://spacy.io/api/entityrecognizer)
- [spaCy | Trained Models & Pipelines](https://spacy.io/models/)
- [spaCy | Available trained pipelines for Portuguese](https://spacy.io/models/pt)
- [spaCy | Visualizers](https://spacy.io/usage/visualizers)

### **Example in Portuguese**

In [5]:
!python -m spacy download pt_core_news_md

Looking in indexes: https://pypi.org/simple, https://us-python.pkg.dev/colab-wheels/public/simple/
Collecting pt-core-news-md==3.5.0
  Downloading https://github.com/explosion/spacy-models/releases/download/pt_core_news_md-3.5.0/pt_core_news_md-3.5.0-py3-none-any.whl (42.4 MB)
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m42.4/42.4 MB[0m [31m17.3 MB/s[0m eta [36m0:00:00[0m
Installing collected packages: pt-core-news-md
Successfully installed pt-core-news-md-3.5.0
[38;5;2m✔ Download and installation successful[0m
You can now load the package via spacy.load('pt_core_news_md')


In [6]:
nlp = spacy.load('pt_core_news_md')
doc = nlp("Apple estuda comprar startup britânica por US$ 1 bilhão")

displacy.render(doc, jupyter=True, style="dep")