AI 與偵測引擎
인간 피드백 강화학습 (RLHF)RLHF
定義
A reinforcement learning method that aligns models using human preference feedback as reward. Used to suppress harmful LLM outputs.
AI 與偵測引擎
A reinforcement learning method that aligns models using human preference feedback as reward. Used to suppress harmful LLM outputs.