you are viewing a single comment's thread
view the rest of the comments
[–] 5 points 10 months ago (1 child)

I suppose answering "I don't know" to every prompt is at least more accurate than what we have now, but I don't think they'll want to risk that.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 10 months ago

    Of course. What the paper is suggesting is that during training and evaluation you should reward correct answers, punish wrong answers, and treat abstentions as somewhere in between. Current benchmarks punish abstentions and wrong answers equally, therefore models that guess instead of abstaining score higher on average.

  • source
  • parent