twcs2PersonaChat 🐦2🤖

This project allows to take the Twitter Customer Support and format it in the Persona Chat format. This is helpful to adapt the model described in this paper into an task oriented version.

Table of content

Getting Started
Usage
Contribute
- Contributors
Show you Support
License

🚀 Getting Started

0 - Download the project

Click here and extract the zip to your preferred dirctory.

1 - Install pipenv

The first step is to install pipenv. Go to the project directory and run: On mac: You can use homebrew:

brew install pipenv

or pip:

pip install pipenv

On Linux:

sudo apt install software-properties-common python-software-properties
sudo add-apt-repository ppa:pypa/ppa
sudo apt update
sudo apt install pipenv

2 - Install project requirements

On the project directory run:

pip install -r requirements.txt

3 - Run the project

To run the project:

python cli.py [module_name] [options]

👩‍💻 Usage

This project includes 3 modules: **getMetadata**, **preprocess**, and **personify**.

getMetadata

This module allows you to retrieve some metadata about the Twitter Customer Support to use it run:

python cli.py getMetadata

preprocess

This module allows you to preprocess the Twitter Customer Support. Here are the options you can use:

--emojis: Boolean, if True, removes all emojis from the dataset (default: True)
--emoticons: Boolean, if True, removes all emoticons from the dataset (default: True)
--urls: Boolean, if True, tags urls as '(URL)' from the dataset (default: True)
--html_tags: Boolean, if True, removes all html tags (default: True)
--acronyms: Boolean, if True, converts acronyms to their meaning. E.g.: SMH -> So much Hate (default: True)
--spelling: Boolean, if True, spellchecks the dataset (default: False)
--usernames: Boolean, if True, tags usernames (default: False)

To run:

python cli.py preprocess [options]

personify

This modules allows you to format the (preprocessed or not) dataset. The options are:

--brand: String, represents the name of a brand, only uses the interactions with a specific brand. If none, uses the whole dataset (default: None)
--limit: Integer, only uses a limited amount of conversations. If -1 uses the whole dataset (default: -1)

🤝 Contribute

If you have any ideas, just open an issue and tell us what you think!

If you'd like to contribute, please fork the repository and make changes as you'd like. Pull requests are warmly welcome.

Contributors

Show your support

⭐ Star us on GitHub — it helps!

License

This project is MIT licensed.

Name		Name	Last commit message	Last commit date
Latest commit History 12 Commits
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
acronyms.py		acronyms.py
chat_elements.py		chat_elements.py
cli.py		cli.py
metadataExtractor.py		metadataExtractor.py
personaChatCS.json		personaChatCS.json
personifier.py		personifier.py
preprocessed.csv		preprocessed.csv
preprocessor.py		preprocessor.py
requirements.txt		requirements.txt
sample.csv		sample.csv
utilities.py		utilities.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

twcs2PersonaChat 🐦2🤖

Table of content

🚀 Getting Started

0 - Download the project

1 - Install pipenv

2 - Install project requirements

3 - Run the project

👩‍💻 Usage

getMetadata

preprocess

personify

🤝 Contribute

Contributors

Show your support

License

About

Releases

Packages

Contributors 2

Languages

License

HLT-MAIA/twcs2PersonaChat

Folders and files

Latest commit

History

Repository files navigation

twcs2PersonaChat 🐦2🤖

Table of content

🚀 Getting Started

0 - Download the project

1 - Install pipenv

2 - Install project requirements

3 - Run the project

👩‍💻 Usage

getMetadata

preprocess

personify

🤝 Contribute

Contributors

Show your support

License

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Contributors 2

Languages

Packages