# Code Eval: HumanEval + MBPP + BigCodeBench + LiveCodeBench + SWE-Bench-Lite

> Source: https://sukruyusufkaya.com/learn/fine-tuning-cookbook/ftc-code-eval-humaneval-mbpp-swe-bench
> Updated: 2026-08-19T12:36:35.228Z
> Category: Fine-Tuning Cookbook (Model-by-Model)
> Module: Part VIII — Code Models & Repo-Level FT
**TLDR:** Code LLM'in standart benchmark suite'i: HumanEval (164 Python problem), MBPP (974 Python), BigCodeBench (1140 calls 139 lib), LiveCodeBench (datas leak-resistant), SWE-Bench-Lite (300 real GitHub issue fix). Pass@1 vs pass@10 metric, code execution sandbox. RTX 4090'da bench koşma.

