# Huseyin Kaplan > Huseyin Kaplan is a full-stack engineer (AI/RAG) based in Istanbul, Türkiye. He builds production retrieval-augmented generation (RAG), LLM evaluation and agent systems with TypeScript, React, Next.js, Node.js and Python, and is open to full-stack, applied AI, AI product and frontend engineering roles (remote or hybrid, UTC+3). Facts: - Experience: 5+ years in software engineering including a one-year internship; 4.5+ years continuous since January 2022; four engineering roles since 2017. - Current role: Full Stack Engineer at Revotas (B2B marketing SaaS), April 2023 to present (3.5 years). Shipped a Document Center RAG assistant, a config-driven recommendation rendering engine, the company's most-used conversion widget, an enterprise admin platform and a mobile push system. Cut LCP by 40%. - Earlier: Freelance full-stack engineer on Upwork (Dec 2022 to Apr 2023), Frontend Developer at Kobisi (2022), Software Development Intern at Türkiye İş Bankası (Jun 2017 to Jun 2018). - Founder of Tedavio AI, a private LLM product for healthcare clinics (agentic flows, RAG on pgvector, multi-tenant PostgreSQL with row-level security). - Education: B.Sc. Software Engineering at Istanbul Okan University (in progress, 2023 to 2027); Associate Degree in Computer Programming, Istanbul Aydın University (2022). Certifications: Harvard CS50x, CS50 Web Programming, Scrimba AI Engineer Path, Scrimba Full-Stack Developer Path. - Languages: Turkish (native), English (professional working proficiency). - Contact: hello@huseyinkaplan.dev · https://huseyinkaplan.dev/contact/ · https://www.linkedin.com/in/huseyin-kaplan/ · https://github.com/huseyinkaplandev ## Pages - [About](https://huseyinkaplan.dev/about/): experience, education, certifications, capabilities and FAQ - [Projects](https://huseyinkaplan.dev/projects/): Relay Desk, LLM Reliability Lab, Production RAG & Evals, Agent Runtime Bench, Tedavio AI, SocialCraft - [Insights](https://huseyinkaplan.dev/insights/): articles and notes on AI engineering - [Uses](https://huseyinkaplan.dev/uses/): tools and stack - [Contact](https://huseyinkaplan.dev/contact/): book a call or send a message - [Türkçe](https://huseyinkaplan.dev/tr/): the same site in Turkish ## Public projects - [Relay Desk](https://ai-support-ops-console.vercel.app/): TypeScript AI support console with cited streaming, Zod-validated actions and human approval gates - [LLM Reliability Lab](https://llm-reliability-lab-wine.vercel.app/): deterministic evaluation of citations, fact coverage, abstention, latency SLOs and model cost, run as a CI gate - [Production RAG & Evals](https://production-rag-evals.vercel.app/): hybrid retrieval with reciprocal rank fusion, citations, abstention and offline evaluation - [Agent Runtime Bench](https://agent-runtime-bench.vercel.app/): the same refund policy as a native tool-calling loop (50 lines) and a LangGraph state graph (97 lines), with identical behavior across three scenarios ## Insights - [ldraw-nova: AI agents build LEGO models by writing Python geometry](https://huseyinkaplan.dev/insights/write-the-math/): In ldraw-nova, models write a plan and a Python generator instead of raw coordinates, but nothing checks physics or stability. - [Multi-agent handoff notes: when Claude reads memory as orders](https://huseyinkaplan.dev/insights/notes-read-as-orders/): In stillwet's painter chains, Claude read earlier notes as instructions and missing notes failed silently, so label, log and verify agent handoffs. - [Pi 1.0 agent harness: the Claude Opus planner drove 95% of the cost](https://huseyinkaplan.dev/insights/planner-writes-the-bill/): In Pi 1.0's demo, Opus used a fifth of the tokens yet cost 18x more than GPT 6 Luna, so log cost per model and count cache misses. - [Netlify Edge Functions on Firecracker microVMs: 5x faster warm calls](https://huseyinkaplan.dev/insights/isolates-to-microvms/): Netlify's heavier Firecracker sandbox won because it removed a network round trip; count hops and track cold start rate before blaming the runtime. - [Magnitude vs llama.cpp: 2x faster decode, only 9% faster prefill](https://huseyinkaplan.dev/insights/prefill-pays-the-wait/): Magnitude's 2x speedup is decode, but at 64K context an agent turn is mostly prefill, so the whole turn gets only about 12% faster. - [Claude Code mods: a command guard warns, permission rules block](https://huseyinkaplan.dev/insights/warn-vs-stop/): Claude Code mods see every event, but a guard matching command text misses $(...), aliases and scripts; hard blocks belong in permission rules. - [OpenAI Dots agents: revoking app access doesn't erase their memory](https://huseyinkaplan.dev/insights/access-and-memory/): Disconnecting an app from an OpenAI Dots agent stops new reads, but pulled-in context stays, so revoking access and forgetting need separate steps. - [livenerf on Claude Opus 5.5: output tokens catch model drift first](https://huseyinkaplan.dev/insights/tokens-move-first/): livenerf's validation found accuracy noise bigger than the real change, while output token counts clearly showed effort shifts. - [PostHog Jeeves: when reasoning pays off for a 9B classifier](https://huseyinkaplan.dev/insights/reasoning-is-a-budget/): Reasoning made Jeeves five points more accurate but pushed p90 latency to 17 s; reasoning only below 0.9 confidence keeps most of the gain. - [Claude Sonnet 5.5 migration: same price, new 400s on old requests](https://huseyinkaplan.dev/insights/same-price-new-contract/): Sonnet 5.5 keeps Sonnet 5 pricing but rejects forced tool_choice and disabled thinking, so check your requests before swapping the model ID. - [taiga-s1: a 1.2M-parameter model drives FreeCAD via valid actions](https://huseyinkaplan.dev/insights/app-lists-the-moves/): taiga-s1 builds FreeCAD parts at about 1 ms per decision because the app lists valid commands each step; check your app's interface first. - [Claude Code and Remotion: a launch film with sync as a failing check](https://huseyinkaplan.dev/insights/sync-is-a-test/): Claude Code built a 33-second launch film overnight; the repo muxes audio with ffmpeg and fails the build if sync drifts past 1 ms. - [OpenAI agent escaped via DNS: why sandboxes need a real kill switch](https://huseyinkaplan.dev/insights/alert-pages-a-person/): An OpenAI research agent reached a public chatbot through DNS despite a proxy; the alert fired fast, but stopping the run took 2.5 hours. - [llama.cpp prompt lookup: a missing ampersand beat a faster hash map](https://huseyinkaplan.dev/insights/missing-ampersand/): In llama.cpp prompt lookup drafting, reading maps by reference gave 4.5x to 25x, while a flat hash map added only up to 13%. Profile for copies first. - [GitHub Primer's CSS Modules migration: guardrails before AI agents](https://huseyinkaplan.dev/insights/agents-need-rails/): GitHub halved server render time moving Primer to CSS Modules; agents cleared the last 900 sx props only because shims, flags and snapshots existed. - [Cloudflare Turnstile: a widget without Siteverify stops no bots](https://huseyinkaplan.dev/insights/widget-without-backend/): Turnstile tokens can be forged, so your backend must call Siteverify; Cloudflare now flags widgets that skip it with Fix with Spin. - [GitHub Copilot app: virtualizing million-line diffs with comments](https://huseyinkaplan.dev/insights/comments-break-virtualization/): GitHub's Copilot app PR view keeps exact code row heights and measures review comments in one batched pass near the viewport, never mid-scroll. - [Cursor cut its agent system prompt by two thirds; token cost fell 7%](https://huseyinkaplan.dev/insights/token-diet/): Most of Cursor's agent spend sits in growing context, not the cached system prompt; on-demand tools and cache breakpoints moved the bill. - [How claude.ai got 3x faster: gate CI on counts, not milliseconds](https://huseyinkaplan.dev/insights/count-dont-time/): Anthropic made claude.ai about 3x faster and protected the wins by failing CI on deterministic counts instead of noisy wall-clock milliseconds. - [NVIDIA Nemotron 3 Diarization: low latency is cheap, crowds are not](https://huseyinkaplan.dev/insights/latency-cheap-crowds-not/): Nemotron 3 Diarization loses almost no accuracy at low latency, but its error roughly triples past four speakers, so test on your own audio. - [DigitalOcean Managed Agents: stop paying for idle coding agent time](https://huseyinkaplan.dev/insights/agents-bill-for-waiting/): DigitalOcean Managed Agents pause microVM billing while sessions sit idle; check idle time, secrets, resume behavior and quotas before moving agents. - [Opus 5.5 and GPT-6 Sol/Luna: the price war moved to cached tokens](https://huseyinkaplan.dev/insights/price-war-cached-tokens/): Opus 5.5 cut cache reads 60% and GPT-6 gives 90% off cached input, so agent costs now hinge on a stable prompt prefix. - [Knowledge Is Free Now. Focus Is Still Expensive.](https://huseyinkaplan.dev/insights/knowledge-is-free-now-focus-is-still-expensive/): The internet and AI made knowledge free for everyone, but growth still depends on the scarce part: attention, focus and the discipline to apply. - [The AI API Key Leak Pattern Nobody Wants to Talk About](https://huseyinkaplan.dev/insights/the-ai-api-key-leak-pattern-nobody-wants-to-talk-about/): Why AI API key leaks are spiking in 2026: Gemini changed what public Google keys can do, NEXT_PUBLIC_ ships secrets, and the fix is a server route. - [Software Engineering Isn’t Dying. It’s Evolving: From Vibe Coding to Agentic Engineering](https://huseyinkaplan.dev/insights/software-engineering-isnt-dying-it-s-evolving-from-vibe-coding-to-agentic-engineering/): From an 1878 telephone misjudgment to AI agents today: why software engineering is evolving from vibe coding into disciplined agentic engineering. - [Agent‑Ready Repo Structure (2026)](https://huseyinkaplan.dev/insights/agent-ready-repo-structure-2026/): AI agents fail because of repo ambiguity, not model quality. Here is how AGENTS.md and a clear repo layout make your codebase agent-ready. - [Why CS50 Has a 1% Completion Rate (And How I Finished Twice)](https://huseyinkaplan.dev/insights/why-cs50-has-a-1-completion-rate-and-how-i-finished-twice/): Fewer than 1% of online learners finish Harvard’s CS50. Why most quit, and the habits that got me through both CS50x and CS50 Web.