New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Special eval metrics and scripts #457
Comments
There are a few points to note for some datasets:
|
|
For TriviaQA, NQ, and WebQuestions, we are evaluating on the open domain variant of the task and should use the evaluation procedure shown here: https://github.com/google-research/google-research/tree/master/t5_closed_book_qa |
Posting a colab that calls the special eval scripts for the above datasets. https://colab.research.google.com/drive/1G2zxbvi96qxbOv6LNYvTTrcJxsBX_4Hr Edit:
|
Dataset splits for open-domain QA are a total mess, so writing this down here to track my own records:
|
Leo Gao mentions that
We aren't doing any kind of length normalization, so if we are underperforming on those tasks, we could consider it. |
For Drop, I noticed that the model would predict a number by its word instead of its numeric symbol. So I added that to the normalization process. Predictions from Before:
After:
According to GPT3 Paper for Zero-shot
So we are performing significantly better |
For the record, each example in coqa and quac is actually N examples, where N is the number of turns of the dialog. Our prompts for coqa and quac will only evaluate on one turn per example. We need to create new serialized versions of the datasets if we are going to evaluate on the full dataset. |
Got confirmation that this is indeed the case. Unfortunately, this split is not available in HF, apart from in the main nq dataset which is utterly colossal. The one in the nq_open dataset (and the one in kilt nq) is different. |
@craffel I've actually made prompts for coqa that tries to solve this.
But the downside is that it has to be a unique prompt for each number of turns. For coqa the maximum number of turns is 25 so there needs to be 25 unique prompts. The idea is to then collect the prediction to a json and run the official eval script. So far I've made around 15 unique prompts (just change the number). I can make a pull request if this approach makes sense. |
Thanks @lintangsutawika . @zaidalyafeai actually made HFDS variants of the tasks that include prior dialog turns as context. I think we can just use that as the base dataset. |
I wrote ten GPT-3 style record prompts, where the model has to rank all the possible choices of the query sentence with every possible entity filled in. They will only make sense for rank eval. I can try to run eval on them before we cache. Not sure if it will help but worth a try. #490 |
Datasets that are mostly done but we need to re-run eval and/or compute scores manually for all models:
Datasets where there is still work to be done:
Datasets I don't know the status of:
|
Since we have the string outputs of all tasks, in principal we should be able to run arbitrary metrics, especially for datasets require fancy metrics.@lintangsutawika has imported the official eval scripts for ReCoRD, SQuAD v2, Natural Questions, TriviaQA, and DROP.
Update: Even when using Lintang's eval scripts, all extractive QAs and closed-book (generative) QAs still have abnormally low numbers, namely:
Also, I think all eval of extractive QA from the training mixture also failed.
(Note that ARC is closed-book, but its performance is fine because it's multiple-choice. A great point in case that machine task categories care more about format way more than human skill/knowledge.)
Others with issues to keep an eye on:
The text was updated successfully, but these errors were encountered: