# Speculative Decoding Production: Draft + Target Pairing + Accept Rate Ölçümü

> Source: https://sukruyusufkaya.com/learn/fine-tuning-cookbook/ftc-speculative-decoding-production
> Updated: 2026-08-22T07:43:38.045Z
> Category: Fine-Tuning Cookbook (Model-by-Model)
> Module: Part XV — Serving Engineering
**TLDR:** Speculative decoding (Leviathan et al. 2023, Chen et al. 2023) — küçük draft model 4-8 token'ı tahmin eder, target model bunu **doğrular**. Accept rate yüksekse 2-3x throughput. EAGLE-2 (Li et al. 2024), MEDUSA head training. RTX 4090'da Llama 3.1 8B target + Llama 3.2 1B draft: tok/s 175 → 290.

