Technical writing on LLM development, fine-tuning and data engineering, published on Medium.
OpenAI
Benchmark
July 2026
AI Can Now Solve a Third of Real Biology Questions: What GeneBench-Pro Shows
On OpenAI's GeneBench-Pro benchmark — built from 129 problems spanning
genomics, quantitative biology and translational medicine — the
strongest model, GPT-5.6 Sol, reached a 28.7% success rate at its
highest reasoning level, rising to 31.5% in Pro mode. I wrote about
what the results mean for AI's contribution to scientific research.
UN
AI Policy
July 2026
The UN Gathered in Geneva for AI Governance
I wrote about the first UN Global Dialogue on AI Governance, attended
by 169 countries, the four topics on the agenda, and what it means
for developers in Turkey.
Meta
Coding Models
July 2026
Meta Is Back in the Coding Race
Meta's new Muse Spark model has caught up to GPT-5.5 on SWE-Bench Pro,
but Claude Opus 4.8 is still ahead. I wrote about the current state of
the coding-model race, Meta's possible infrastructure-leasing plans,
and what it means for developers in Turkey.
LoRA
DPO
Qwen3.5
Why Did the Largest Model Lose? Fine-Tuning 4 Qwen3.5 Variants for Digital Marketing
In an experiment fine-tuning four Qwen3.5 variants from 2B to 27B on
the same dataset, the largest model didn't win. I wrote about how
hardware constraints and data quality shaped performance, and why the
9B model turned out to be the most balanced choice.