9 posts
The four core metrics for measuring RAG systems: faithfulness, answer relevancy, context precision and recall. Evaluation with RAGAS, thresholds and context trust.
Moving beyond brittle hand-written prompts: a practical guide to meta-prompting and metric-driven, programmatic prompt optimization with DSPy.
Building RAG is easy, proving it reliable is hard. Retrieval/generation metrics, reference-free evaluation with RAGAS, OpenTelemetry spans, and cost-per-successful-output.
In 2026 'vector as a feature' wins: PostgreSQL + pgvector suffices for most scenarios. Three architectural thresholds, evaluation traps, and selection criteria.
An agent at 90% in testing drops to 70% in production. A field guide to the reliability gap, pass^k, LLM-judge biases, and a production evaluation framework.
Measure your RAG system's real quality with four core metrics: faithfulness, answer relevance, context precision, and recall. Ragas, LLM-as-a-judge, and Turkish challenges.
With reasoning models, explicit chain-of-thought instructions often hurt. The shift to context engineering, outcome-oriented prompts and evaluation loops.
What is LLMOps? The discipline of running large language models in production: MLOps difference, prompt management, evaluation, LLM observability, cost optimization, guardrails, and a maturity model.
LLM-as-a-judge is the dominant automated evaluation method but biased; RAGAS human correlation is only 0.55. A reliable eval guide covering best practices, biases, in a Turkish/KVKK context.