RLHF en · NOUN
Meanings
-
(abbreviation, alt-of, initialism, qualifier:machine learning, uncountable) Initialism of reinforcement learning from human feedback.
The “yeasayer effect” arises in AI models trained using reinforcement learning from human feedback (RLHF)—human “data labellers” rate the answer generated by the model as being either acceptable or not.
2025 June 14, Melissa Heikkilä, “AI leaders rein in ‘sycophantic’ chatbots that flatter users”, in FT Weekend (Companies & Markets section), London: The Financial Times Ltd., →ISSN, →OCLC, page 12:Turning it into a chatbot requires an extra step, the aforementioned reinforcement learning with human feedback: RLHF. An army of human testers are given access to the raw LLM, and instructed to put it through its paces: asking questions, giving instructions and providing feedback.
2024 April 16, Alex Hern, “TechScape: How cheap, outsourced labour in Africa is shaping AI English”, in The Guardian, →ISSN:ChatGPT and reinforcement learning with human feedback (RLHF) have revolutionized the AI landscape, providing an accessible and reliable platform for AI-enabled applications.
2023, Mohak Agarwal, Generative AI for Entrepreneurs in a Hurry, Notion Press, →ISBN:RLHF now seems more like a process by which machines learn humans, including our weaknesses and how to exploit them. Chatbots tap into our desire to be proved right or to feel special.
2025 May 9, Mike Caulfield, “AI Is Not Your Friend”, in The Atlantic, retrieved 10 May 2025:At the time, InstructGPT received limited external attention. But within OpenAI, the AI safety researchers had proved their point: RLHF did make large language models significantly more appealing as products.
2025, Karen Hao, Empire of AI, New York City: Penguin Press, →ISBN:
Coordinates
Hypernyms
learning · ML · machine learning · deep learning · RL · reinforcement learning · DL