👋 Hi, I'm Tuba.

FROM DATA TO
MEASURABLE AI.

I research why Turkish LLMs fail — across the data, model and evaluation layers. I don't just train models; I question where data loses quality, under what conditions model behaviour changes, and whether the metrics in use actually reflect real performance. I diagnose the problem, compare technical options, build the case for my chosen approach, and turn the outcome into measurable systems.

Hi there! 👋
Portrait of Tuba Çelik, neobrutalist illustration
200K+ Training data rows (MMX project)
94% Format compliance — automated validation
200+ Question evaluation benchmark
2B–27B Model comparison range (Qwen)
AREAS OF EXPERTISE
01

AI-Powered
Design

I run end-to-end Turkish LLM development: synthetic data engineering, LoRA/DPO fine-tuning, benchmark design and on-prem deployment. I diagnose the problem, researching why a model fails across the data, architecture and evaluation layers.

02 — Areas of Expertise

WHAT I DO

LLM Development

I diagnose why a model fails: data quality, architecture choice, or evaluation system? End-to-end Turkish LLM workflows: synthetic data engineering, LoRA/DPO fine-tuning, metric-based evaluation and on-prem deployment. Active experience with Hugging Face, Unsloth, Qwen and Gemma families.

See projects →

Data and Analytics

I analyze where raw data breaks down. I build JSONL training sets and quality control pipelines with Python and Pandas; design dashboards with GA4 and Looker Studio that make current state visible. Analytics infrastructure is a prerequisite for meaningful model development.

See my projects →

Web and Automation

I automate repetitive workflows with n8n and build accessible, fast websites and tools with HTML/CSS/JS and Next.js. This portfolio site included — I prefer to build the tools that produce the output myself.

See examples →
Portfolio

SELECTED WORK

LLM evaluation pipeline redesign illustration
LLMOps Evaluation Benchmark In-house project

LLM Evaluation Pipeline Redesign

The scoring system was returning misleading results but the root cause was unclear. I audited the pipeline, identified 12 metric inconsistencies suppressing measured performance by roughly 15 points, and rebuilt it with standardized metrics across a 200+ question benchmark.

PROBLEM
Metrics were inconsistent; model looked underperforming but causes were hidden in tooling.
APPROACH
Full audit of scoring logic, metric normalization, new benchmark construction.
OUTCOME
Pipeline now produces reliable, comparable scores; evaluation is the foundation for all fine-tuning decisions.
Secure AI Chat data flow: PII detection, token vault and approval layers
NLP Security spaCy Token Vault Next.js Local MVP

Turkish PII-Protected Secure AI Chat

Local secure MVP · Synthetic test data

I questioned whether masking personal data is actually enough to protect it. Built a layered system: Turkish-specific PII detection with spaCy + tr_core_news_trf, reversible tokenization into an AES-256 encrypted Vault, human-in-the-loop approval workflow, and a same-origin proxy so sensitive data never reaches the LLM layer in plain text.

PROBLEM
Masking hides data but doesn't protect it if the mask can be reversed without control.
APPROACH
NER-based PII detection, encrypted token vault, approval workflow, same-origin proxy.
OUTCOME
No PII leakage detected in tested local scenarios. Architecture separates detection, storage and transmission layers.

No real LLM connected. Synthetic data only. Local deployment.

MMX Turkish Marketing LLM project illustration
LLM Fine-tuning Data Engineering In-house project

MMX: Turkish Marketing LLM

Turkish LLMs were scoring low in domain tests but the reason was unclear. I designed a 200,000+ row synthetic data pipeline across 12 domains with automated quality validation reaching 94% format compliance, ran LoRA/DPO fine-tuning experiments, and contributed to on-prem deployment.

PROBLEM
Weak domain scores with no clear cause — data gap, model limit, or evaluation flaw?
APPROACH
Synthetic pipeline, format QA, fine-tuning experiments, on-prem deployment.
OUTCOME
94% format compliance; clear quality baseline for each subsequent fine-tuning round.

ABOUT

At Mad Cat Labs I work on Turkish LLM development. On the MMX project we built a 200,000+ row training set from synthetic and real data, ran LoRA and DPO fine-tuning experiments across model families like Qwen and Gemma, designed evaluation pipelines, and took part in on-prem deployment.

I spent many years in digital analytics and performance marketing. That's where I learned how data gets read, transformed and turned into decisions. Now I'm carrying that foundation into AI: building automations with n8n, evaluating model outputs, and building tools that smooth out the development process.

I love sharing what I learn. Everyday content on Instagram (@tubac3l1k), longer technical writing on Medium. Both channels are in Turkish, because the gap in Turkish-language AI resources is huge and I want to help close it.

Portrait of Tuba Çelik, neobrutalist illustration

I CREATE CONTENT

I produce Turkish-language AI and tech content, translating complex topics into plain, accessible language.

@tubac3l1k

AI and tech content. Everyday posts in Turkish.

Medium Blog

Technical writing on LLM development, fine-tuning and data engineering.

SKILLS

🤖 AI and LLM
Hugging Face Unsloth LoRA QLoRA DPO PEFT Fine-tuning Ollama Open WebUI F1 / BLEU / Perplexity
📊 Data and Analytics
Python Pandas JSONL Synthetic Data SQL GA4 Looker Studio Power BI
🛠️ Infra and Automation
Docker Linux n8n Claude Code Cursor
🎨 Design and Web
HTML CSS JavaScript Neobrutalism Responsive Design Figma

EXPERIENCE

Data & AI Associate
Mad Cat Labs
February 2026 — Present
Delivering AI and analytics products end to end for 5 enterprise clients in a 5-person squad: problem definition, scoping, delivery and measurement. Ran product discovery for a national retail chain, defined 7 AI use-cases — all entered development — and put a working demo in the client's hands in 2 weeks. For a national logistics operator: demo in 3 days, then an image-analysis product covering 500 in-vehicle cameras continuously, replacing an entirely manual process. Led the model selection decision on a TÜBİTAK-supported marketing LLM: evaluated 4 Qwen 3.5 variants on a 500-question benchmark scored with LLM-as-judge; the 27B model did not win. Identified 12 metric inconsistencies in the evaluation pipeline and rebuilt it. Designed a 200,000+ record synthetic data pipeline; training on H200 GPUs.
Senior Product Analytics & Marketing Specialist
Link Dijital
February 2022 — February 2026
Analytics and performance products for 10 brands; progressed from specialist to senior over 4 years and took management responsibility in a 5-person team. Cut cost per install from TRY 100 to TRY 7 (93% in 2 weeks): ran 4 audience segments against 5 creative x 5 copy variants in parallel, tracked at ad-set level, shifted budget to the winner — attribution verified via Firebase and AppsFlyer. Mapped the customer journey end to end, built 6 behavioural segments with automated Mailchimp lifecycle journeys; contributed to 250 new customers/month on a single B2B brand. Built 10 automated Looker Studio reports used by 2 teams without going through me.

EDUCATION & CERTIFICATES

M.Sc. in Business Engineering (ongoing)
Yıldız Technical University · Industrial Engineering
Bachelor's in Business Administration
Beykoz University · 2017 — 2022
Data Scientist Bootcamp (ongoing)
Data Science School
Advanced Fine-Tuning for LLMs
edX · IBM
Machine Learning A-Z: AI, Python and R
Udemy
Deep Learning and NLP A-Z
Udemy
Google Analytics 360
Google

LET'S WORK TOGETHER

Have a project, or just want to say hi? Don't hesitate to write.