For me, a system working isn't enough on its own. I want to understand why it works, where it fails, and whether the output is being measured correctly.
I answer that question through Turkish LLM development at Mad Cat Labs. On the MMX project I identified 12 metric inconsistencies in the evaluation pipeline and built a 200+ question benchmark from scratch. I designed a 200,000+ row synthetic training set with automated quality validation. I compared four Qwen variants from 2B to 27B and diagnosed why the 27B model underperformed — a Turkish tokenizer gap, not a model flaw.
Before this role I spent four years in digital analytics and performance marketing. That's where I learned the conditions and limits of reading data: which metric reflects a real change, which is noise, and when a number can be correct but the interpretation wrong. These questions translate directly into LLM evaluation work.
I share what I learn in Turkish. I write technical posts on Medium about LLM development because the gap in Turkish-language AI resources is real and I want to help close it.
I produce Turkish-language AI and tech content, translating complex topics into plain, accessible language.
AI and tech content. Everyday posts in Turkish.
Technical writing on LLM development, fine-tuning and data engineering.