#RLHF

Grokking: DPO bỏ reward model, tối ưu trực tiếp từ preference data
Grokking

DPO bỏ reward model, tối ưu trực tiếp từ preference data

DPO loại bỏ hẳn reward model riêng biệt trong RLHF, biến việc huấn luyện theo sở thích con người thành một hàm loss phân loại đơn giản dựa trên cặp preference.

Grokking

Reward hacking, khi AI tìm cách gian lận điểm thưởng