WhiteMail
← 詞彙表

AI 與偵測引擎

인간 피드백 강화학습 (RLHF)RLHF

定義

A reinforcement learning method that aligns models using human preference feedback as reward. Used to suppress harmful LLM outputs.

深入了解

相關詞彙