# llama.cpp + Ollama: GGUF Serving + Modelfile + System Prompt Versioning

> Source: https://sukruyusufkaya.com/learn/fine-tuning-cookbook/ftc-llama-cpp-ollama-gguf-modelfile
> Updated: 2026-08-13T22:36:11.345Z
> Category: Fine-Tuning Cookbook (Model-by-Model)
> Module: Part XV — Serving Engineering
**TLDR:** llama.cpp + Ollama — CPU/Apple Silicon/edge için altın standart. GGUF format, Ollama'nın Modelfile sistemi (system prompt + tools versioning), Ollama API, OpenAI-uyumlu endpoint. RTX 4090'da Q4_K_M Llama 8B Ollama'da 95 tok/s (vLLM AWQ 175'in altında ama 'set up zero' faktörüyle production-ready).

