Skip to main content
Deliberate AcademyProfessional AI Education
AI Glossary

RLHF (Reinforcement Learning from Human Feedback)

What is RLHF (Reinforcement Learning from Human Feedback)?

A training technique where human raters evaluate model outputs and their preferences are used to reward-train the model to produce more helpful, accurate, and less harmful responses. Used to align ChatGPT, Claude, and similar models with human values.

Example in practice

When Anthropic human raters compared pairs of Claude responses and marked which was safer and more helpful, those preferences fed into RLHF training — shaping the model's future behavior.