The Chinese lab continues its cost-disruption strategy: V4 base model weights land on Hugging Face under a permissive license while hosted inference gets dramatically cheaper. Demo content.
The new Flash tier reads entire codebases and multi-hour video in a single request, and Google is pricing it aggressively against GPT-5-mini. Demo content.
Every major API now offers prompt caching, but most teams still pay full price. Here is how cache breakpoints, TTLs and prefix design actually work. Demo content.
A simplified three-tier structure — realtime, cached, batch — replaces the growing menu of per-model discounts. Most workloads get cheaper, some get sharply more expensive. Demo content.