Retnet parameter dimension #57

allanj · 2023-08-14T05:05:50Z

I wonder why we need twice dimensions for $\mathbf{W}_V$

Yuxin-CV · 2023-08-14T06:12:48Z

Please note that the MSR block includes an additional swish gate compared to the MHSA block in the vanilla Transformer. If we do not double the dimension of v, the MSR block will have 5d^2 parameters, while the MHSA block in the vanilla Transformer only has 4d^2 parameters. Given this scenario, it becomes challenging to determine the width and depth of a retnet for fair comparison with a baseline vanilla Transformer of the same size. Therefore, the authors decide to double the value of W_v and halve the value of d_ffn to maintain the overall parameters of each retnet block equal to 12d^2.

Alternatively, another option is to keep W_v the same as W_k and set d_ffn to 3.5d. However, it is preferable to have a wider swish gate rather than a wider mlp as ffn. For more details, please refer to https://arxiv.org/abs/2202.10447. I believe it is even better to use MSR block only and set d_v = 3.33d.

allanj · 2023-08-14T06:14:49Z

Cool, pretty much makes sense to me. Thanks for the thorough explanationa.

allanj closed this as completed Aug 14, 2023

donglixp self-assigned this Aug 14, 2023

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Retnet parameter dimension #57

Retnet parameter dimension #57

allanj commented Aug 14, 2023

Yuxin-CV commented Aug 14, 2023 •

edited

allanj commented Aug 14, 2023

Retnet parameter dimension #57

Retnet parameter dimension #57

Comments

allanj commented Aug 14, 2023

Yuxin-CV commented Aug 14, 2023 • edited

allanj commented Aug 14, 2023

Yuxin-CV commented Aug 14, 2023 •

edited