deepspeed-chat: fix bf16 stage2 accuracy for bloom-560m #772

mosheisland · 2023-10-17T06:41:18Z

Bloom-560m model has high variance in its last LN layer weight. This causes accuracy issues in bf16 stage2 training. Therefore, reset the parameters of the last LN layer before training. This is a good practice in any case where we replace the classifier that follows the LN.

In addition, in case we are using only optimize lora, we need to force the training of the LN parameters that were reset.

Note that current fix uses plain initialization of final LN. A separate commit will provide support for zero3 initialization.

Change-Id: I323d8947907eb4a1cc0fa6354bdaf0cbbf33a68d

Bloom-560m model has high variance in its last LN layer weight. This causes accuracy issues in bf16 stage2 training. Therefore, reset the parameters of the last LN layer before training. This is a good practice in any case where we replace the classifier that follows the LN. In addition, in case we are using only optimize lora, we need to force the training of the LN parameters that were reset. Note that current fix uses plain initialization of final LN. A separate commit will provide support for zero3 initialization. Change-Id: I323d8947907eb4a1cc0fa6354bdaf0cbbf33a68d Signed-off-by: Moshe Island <misland@habana.ai>

lekurile

LGTM!

mosheisland requested review from jeffra, samyam, tjruwase, ShadenSmith, conglongli, awan-10, eltonzheng, minjiaz, RezaYazdaniAminabadi, duli2012, mrwyattii, yaozhewei, arashb and xiaoxiawu-microsoft as code owners October 17, 2023 06:41

tjruwase requested review from lekurile and removed request for arashb, ShadenSmith, jeffra, duli2012, samyam, conglongli, mrwyattii, yaozhewei, eltonzheng, minjiaz, RezaYazdaniAminabadi, xiaoxiawu-microsoft and tjruwase October 17, 2023 13:30

lekurile approved these changes Oct 17, 2023

View reviewed changes

tjruwase merged commit 185e25c into microsoft:master Oct 17, 2023
2 checks passed

mosheisland deleted the 9_fix_bloom_stage2_bf16_acc branch November 22, 2023 07:52

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

deepspeed-chat: fix bf16 stage2 accuracy for bloom-560m #772

deepspeed-chat: fix bf16 stage2 accuracy for bloom-560m #772

mosheisland commented Oct 17, 2023

lekurile left a comment

deepspeed-chat: fix bf16 stage2 accuracy for bloom-560m #772

deepspeed-chat: fix bf16 stage2 accuracy for bloom-560m #772

Conversation

mosheisland commented Oct 17, 2023

lekurile left a comment

Choose a reason for hiding this comment