confidence_eos_eot_inf does not correctly suppress EOS confidence
Description
I believe there is a bug in the implementation of confidence_eos_eot_inf in generate().
The current code is:
if confidence_eos_eot_inf:
logits_with_noise[:, :, 126081] = logits[:, :, 126348] = -torch.inf
At this point, x0 has already been sampled from logits_with_noise:
logits_with_noise = add_gumbel_noise(logits, temperature=temperature)
x0 = torch.argmax(logits_with_noise, dim=-1)
Later, when remasking == 'low_confidence', the confidence of the sampled token is computed from logits, not logits_with_noise:
p = F.softmax(logits, dim=-1)
x0_p = torch.squeeze(
torch.gather(p, dim=-1, index=torch.unsqueeze(x0, -1)), -1
)
Therefore, setting
logits_with_noise[:, :, 126081] = -torch.inf
after x0 has already been sampled has no effect on either token selection or the subsequent confidence calculation.
My understanding is that confidence_eos_eot_inf is intended to allow EOS/EoT tokens to be sampled normally, but assign them -inf confidence afterward so that they are selected only after higher-confidence tokens have been transferred.
If so, both EOS and EoT should be modified in logits, e.g.
if confidence_eos_eot_inf:
logits[:, :, 126081] = logits[:, :, 126348] = -torch.inf
This way, the sampled x0 remains unchanged, while the subsequent softmax(logits) gives EOS/EoT zero probability and therefore effectively -inf confidence for remasking purposes.
confidence_eos_eot_infdoes not correctly suppress EOS confidenceDescription
I believe there is a bug in the implementation of
confidence_eos_eot_infingenerate().The current code is:
At this point,
x0has already been sampled fromlogits_with_noise:Later, when
remasking == 'low_confidence', the confidence of the sampled token is computed fromlogits, notlogits_with_noise:Therefore, setting
after
x0has already been sampled has no effect on either token selection or the subsequent confidence calculation.My understanding is that
confidence_eos_eot_infis intended to allow EOS/EoT tokens to be sampled normally, but assign them-infconfidence afterward so that they are selected only after higher-confidence tokens have been transferred.If so, both EOS and EoT should be modified in
logits, e.g.This way, the sampled
x0remains unchanged, while the subsequentsoftmax(logits)gives EOS/EoT zero probability and therefore effectively-infconfidence for remasking purposes.