GitHub - eva-n27/BERT-for-Chinese-Question-Answering

BERT-for-Chinese-Question-Answering

本仓库的代码来源于PyTorch Pretrained Bert，仅做适配中文的QA任务的修改

主要修改的地方为read_squad_examples函数，由于SQuAD是英文的，因此源代码处理的方式是按照英文的方式，即此处。

另外，增加了训练中每隔save_checkpoints_steps次进行evaluate，并保存dev上效果最好的模型参数。

因此修改为：

1.先使用tokenizer先使用tokenizer.basic_tokenizer.tokenize对doc进行处理得到doc_tokens（代码161行）

2.对orig_answer_text使用tokenizer.basic_tokenizer.tokenize，然后再计算answer的start_position和end_position（代码172-191）

python3 run_squad.py \
  --do_train 
  --do_predict 
  --save_checkpoints_steps 3000 
  --train_batch_size 12 
  --num_train_epochs 5

python3 eval.py data/squad_dev.json output/predictions.json

欢迎各位大佬批评和指正，感谢

Name		Name	Last commit message	Last commit date
Latest commit History 7 Commits
LICENSE		LICENSE
README.md		README.md
__init__.py		__init__.py
convert_tf_checkpoint_to_pytorch.py		convert_tf_checkpoint_to_pytorch.py
eval.py		eval.py
modeling.py		modeling.py
optimization.py		optimization.py
requirements.txt		requirements.txt
run_squad.py		run_squad.py
sample_dev.json		sample_dev.json
sample_test.json		sample_test.json
sample_train.json		sample_train.json
tokenization.py		tokenization.py