Skip to content

Fix DeBERTa-v2 models failing to run at F16 - #3751

Open
tobocop2 wants to merge 1 commit into
huggingface:mainfrom
tobocop2:debertav2-f16-dtype
Open

tobocop2 wants to merge 1 commit into
huggingface:mainfrom
tobocop2:debertav2-f16-dtype

Conversation

@tobocop2

@tobocop2 tobocop2 commented Jul 18, 2026 •

Copy link
Copy Markdown

Problem

DebertaV2Model::forward fails for F16 models with dtype mismatch in add, lhs: F32, rhs: F16.

Fixes #3750.

Solution

Two spots in debertav2.rs build F32 scalar tensors and combine them with tensors of the model's dtype:

  • the score accumulator in disentangled_attention_bias
  • the mask fill values in XSoftmax::apply

Cast them to the working dtype — same pattern as scale in this file and the mask fill in bert.rs. The F32 path is unchanged.

Testing

New regression test runs a tiny DeBERTa-v2 forward at F32 and F16; the F16 case fails without this fix. Workspace tests, clippy, and fmt all pass.

@tobocop2
tobocop2 force-pushed the debertav2-f16-dtype branch 2 times, most recently from 57e2dcc to 0211b5f Compare July 18, 2026 02:33
@tobocop2 tobocop2 changed the title Fix F16 dtype mismatch in DeBERTa-v2 forward Fix DeBERTa-v2 models failing to run at F16 Jul 18, 2026
disentangled_attention_bias initialized its score accumulator as F32 and
XSoftmax::apply built its fill tensors as F32 regardless of the input
dtype, so forward failed with a dtype mismatch for F16 models. Cast the
scalars to the working dtype; the F32 path is unchanged.

Signed-off-by: tobocop2 <5562156+tobocop2@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DeBERTa-v2 models fail to run at F16

1 participant