3 posts
MMLU, HumanEval, SWE-bench Verified/Pro, ARC-AGI-2, GPQA Diamond, AIME, LiveCodeBench v6, Terminal-Bench 2.0, OSWorld, HLE, plus Turkish benchmarks (TR-MMLU, TUMLU) — what each one measures, the frontier thresholds, contamination and cherry-picking risks, and practical meaning for CTOs, investors, and engineers. 32+ references.
We benchmarked GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro on Turkish workloads end to end: TR-MMLU and TUMLU benchmark numbers, a 50-prompt real-world test across legal, finance, code, creative writing and Q&A, an A/B in a Turkish enterprise, TL-based cost analysis and a decision matrix for picking the right model for each Turkish task. 35+ references.
A 2026 snapshot of the Turkish open-source LLM ecosystem: Trendyol-LLM, Cosmos-Llama, KanarYa, Kumru AI, the TÜBİTAK BİLGEM domestic model, and the T3 AI Baykar defense model. Detailed decision guide covering MMLU-TR and TUMLU benchmarks, licensing, tokenization gap, VRAM requirements, self-hosting needs, and which model to pick for which use case.