RLHF data may encode annotator stress
A new framework shows that human annotators' stress can systematically skew AI training data, subtly warping how models learn to behave.
A new framework shows that human annotators' stress can systematically skew AI training data, subtly warping how models learn to behave.