-
Notifications
You must be signed in to change notification settings - Fork 3.4k
How to run the reverse_demical40 task? #3
Comments
Looks like the issue is that training is unstable and the loss hits nan. Probably need some different hyperoarameter settings. I'll investigate and get back to you but in the meantime, feel free to fiddle with the learning rate and other learning settings. |
You can override individual hparam settings by flag: --hparams='learning_rate=0.1,another_hparam=blah' |
Could you give an example of the reverse task using transformer? Here is my run.sh. The loss goes down to 0.00001 but its output is
|
I tried and I believe it's a decoding problem -- we use 1 to mean "end of sequence" in decoding, but the algorithmic generator only avoids 0s (padding). Will try to prepare a fix soon, thanks for reporting the problem! |
@RickyHan -- the most recent 1.0.4 version should include all needed corrections to make the above instructions work well. I tried and I find that transformer still has some problems with determining the end of inputs, as it's not marked in the algorithmic tasks. So it sometimes reverses a bit too much, but otherwise seems to work. I'm closing this, but could you please test and let me know if it works for you? And if it doesn't, please re-open. Thanks! |
Here is my run script:
Output:
The text was updated successfully, but these errors were encountered: