Tsinghua Proposes New RLHF Paradigm: Models Self-Align Without Human Labels, Achieves SOTA
Self-alignment papers usually move the label budget, not remove humans. Ask what the preference model still needs, and whether SOTA is on a public bench or a private suite.
Tsinghua Proposes New RLHF Paradigm: Models Self-Align Without Human Labels, Achieves SOTA