3 posts
What is model serving? A guide to the serving layer: from managed APIs to self-hosted inference servers like vLLM, TGI and Ollama; batching, streaming, concurrency, throughput and latency trade-offs.
Head-to-head of Codeium's 2024 Windsurf Editor vs Cursor: Cascade agent architecture, Supercomplete, Riptide context engine, model access (Claude Opus 4 + GPT-5 + DeepSeek), pricing ($15 vs $20), enterprise + on-prem options, Turkish developer experience, KVKK + code leakage risk, and 10 scenario-based selection guide.
Detailed review of 8+ Chinese LLMs including Moonshot Kimi K2 (1T MoE), Zhipu GLM-4.5, 01.AI Yi-Large/Yi-Lightning, MiniMax abab, and Baichuan: architectures, benchmarks, pricing, open-weight vs API, Turkish fluency, KVKK + data residency legal-risk map, censorship behavior, and a 6-scenario usage guide.