Essays, breakdowns, and lessons from building production AI agents, modern full-stack applications, and automation systems.
AI coding tools have changed software development from typing code to directing code. This essay explores why tokenmaxxing is real, where it works, where it fails, and why engineering judgment still matters.
A detailed guide to evaluating LLM outputs using exact match, semantic checks, factuality, human review, and production-ready scoring pipelines.
Lessons from shipping multi-agent systems in production — architecture, tool-calling patterns, observability, and the failure modes that actually matter.