Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Folklore Archetype Classifier

final project for an annotation course (COSI 230 - Brandeis University)

Developed annotation guidelines with a group to have students in the class hand-annotate folklore to identify and classify flora and fauny in stories using brat. Using these annotations, we attempted to classify the archetypal roles flora and fauna play in these stories.


Table of Contents

  1. Dataset
  2. Models
  3. Usage
  4. Directory Structure

Dataset

Overview:

  • 479 annotations across 122 story text files.
  • This comprises our manually curated dataset of stories, sourced from Project Gutenberg (Okazaki, Seneca Myths and Folktales, etc) and cultural websites (e.g. www.native-languages.org), across the Cherokee, Filipino, Cherokee, Korean, Japanese, Seneca, and Maori cultures.
  • 66 characters were good, 40 evil, and 297 were neutral. Of these, 66 tags were also protagonists (though not all protagonistswere always good/evil), 35 were antagonists, and 302 were default. Importantly, almost all fauna were default and neutral with the exception of crops.

Annotation Process:

  • Two trained annotators dually annotated each set.
  • Flora and Fauna were pre-tagged in a script in the brat software, so annotators just needed to classify the role and alignments of the flora and fauna
  • Roles (protagonist, antagonist, default)
  • Alignments (good, neutral, evil)
  • Inter-annotator agreement was ensured through a partially automated system where I or another group member would handle any agreements the automated system flagged. We had moderate agreement according to Cohen's Kappa: (role = .66, alignment = .58).

Models

Statistical Models:

  • We tried the following models: logistic regression, multinomial and compliment naive bayes, and random forest

LLM Promping:

  • We used OpenAI’s GPT 3.5-turbo and set its temperature to 0 to ensure deterministic responses. We prompted the LLM with 50 tokens around the annotated spans as context, and asked it to clas- sify a certain animal’s role and alignment. Using this method, role acccuracy was 45%, alignment accuracy was 68%, and the F1 of the alignment predictions was 34% and the F1 of role predictions was 39%, performing decently above chance.
  • We simply used 1-shot learning for this task, and if it had gone beyond the span of this course I would have likely tried more advanced techniques, including few-shot and implementing RAG to increase the accuracy.

Usage

Installation

  1. Clone the repository
  2. For running the LLM prompting, add your api-key to the script, then run the notebook in the directory LLM-prompting
  3. For running the statistical models, run statistical_models.py in the directory statistical models

Directory Structure

project/
│
├── annotation/
│   ├── adjudication.ipynb/                                 # Script to flag disagreements and get statistics
│   ├── Changelog.docx/                                     # Log of annotation guideline updates
│   ├── Folklore_Team_Annotation_Guidelines.pdf/            # Guidelines
│   ├── gold_data/                                          # Includes brat annotation configuration and gold annotations
│   ├── Sample Annotations For Presentation.docx/           # Exampe annotations
│   ├── Simplified Archetypes.docx/                         # Explanation of each archetype
│
├── Annotation_Final_Report.pdf/                            # Conference style report on project
├── Final Presentation.pptx/                                # Final class presentation
|   
├── model tests/
│   ├── LLM-prompting/        # Model-specific files, run `model_v1.ipynb` or `stats.py` for use
│   ├── statistical-models/   # Model-specific files, run `statistical_models.py` for use
│
└── README.md                 # Project documentation

About

Hand-annotated dataset and models for classifying the archetypes for flora and fauna in traditional folktales.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages