Experiments with Neural Networks for Small and Large Scale Authorship Verification

Authors : Marjan Hosseinia and Arjun Mukherjee [pdf]

We propose two models for a special case of authorship verification problem. The task is to investigate whether the two documents of a given pair are written by the same author. We consider the authorship verification problem for both small and large scale datasets. The underlying small-scale problem has two main challenges: First, the authors of the documents are unknown to us because no previous writing samples are available. Second, the two documents are short (a few hundred to a few thousand words) and may differ considerably in the genre and/or topic. To solve it we propose transformation encoder to transform one document of the pair into the other. This document transformation generates a loss which is used as a recognizable feature to verify if the authors of the pair are identical. For the large scale problem where various authors are engaged and more examples are available with larger length, a parallel recurrent neural network is proposed. It compares the language models of the two documents. We evaluate our methods on various types of datasets including Authorship Identification datasets of PAN competition, Amazon reviews, and machine learning articles. Experiments show that both methods achieve stable and competitive performance compared to the baselines.

Prerequisits:

Python 2.7, Theano, Keras 2.0

To see the list of argumnets:

python main.py -h

To run Cross Validation:

python main.py

Datasets:

PAN Authorship Verification datasets PAN13,PAN14, PAN15
Amazon reviews
A subset of MLPA-400

Name		Name	Last commit message	Last commit date
Latest commit History 25 Commits
data		data
LICENSE		LICENSE
README.md		README.md
main.py		main.py
model.py		model.py
util.py		util.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Experiments with Neural Networks for Small and Large Scale Authorship Verification

Prerequisits:

To see the list of argumnets:

To run Cross Validation:

Datasets:

About

Releases

Packages

Languages

License

marjanhs/prnn

Folders and files

Latest commit

History

Repository files navigation

Experiments with Neural Networks for Small and Large Scale Authorship Verification

Prerequisits:

To see the list of argumnets:

To run Cross Validation:

Datasets:

About

Topics

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages