AIai-costModel routing: why choosing Haiku, Sonnet, or Opus matters more than your prompt80% of AI cost reduction comes from sending the right request to the right model — not from prompt engineering. A practical guide to model routing in production.Apr 23, 2026·7 min read
AIai-costPrompt caching: the 90% cost reduction nobody talks aboutAnthropic's ephemeral cache discount is mechanically simple but operationally hard. The placement pattern, the 1024-token threshold, and what can and cannot be cached.Apr 23, 2026·6 min read