Live

Google ships Gemini 3.5 Flash with 10M token context at commodity prices

The new Flash tier reads entire codebases and multi-hour video in a single request, and Google is pricing it aggressively against GPT-5-mini. Demo content.

MVMara Vidović
Published 6 Jul 2026, 14:20Updated 27 Jul 2026, 18:422 min read
Google ships Gemini 3.5 Flash with 10M token context at commodity prices

Google has released Gemini 3.5 Flash, pushing its fast-tier model to a 10 million token context window while keeping pricing at commodity levels.

Headline capabilities

  • 10M token context: whole monorepos, multi-hour video, or years of chat history in one request
  • Needle-in-haystack retrieval accuracy above 99% at full length, per Google's benchmarks
  • Streaming ingestion so long inputs process while still uploading

The strategic angle

Flash-tier models have become the workhorse of production AI. By making extreme context cheap, Google is betting that many RAG pipelines will be replaced by brute-force context stuffing.

"For a lot of teams, 10M tokens of cheap context is a simpler architecture than a vector database they have to keep in sync." — infrastructure engineer, quoted in launch coverage

Rivals are expected to answer within the quarter; leaked roadmaps suggest similar context expansion across the industry.

Source: Google DeepMind

Related stories

All →