# DPO: Direct Preference Optimization — Rafailov 2023, RLHF'in Cheaper Yeniden Doğuşu

> Source: https://sukruyusufkaya.com/learn/llm-muhendisligi/dpo-direct-preference-optimization-rafailov-2023
> Updated: 2026-08-11T06:32:42.395Z
> Category: LLM Mühendisliği
> Module: Modül 15: RLHF + DPO — Alignment & Preference Optimization
**TLDR:** DPO (Rafailov 2023): RLHF mathematical reformulation — no reward model, no RL. Direct preference loss. Llama-3 RLHF replacement. Math derivation, implementation simpler than PPO, comparable quality. Türkçe DPO pratik: $1K maliyetle 8B model alignment.

