Google ships Gemini 3.5 Flash with 10M token context at commodity prices
The new Flash tier reads entire codebases and multi-hour video in a single request, and Google is pricing it aggressively against GPT-5-mini. Demo content.
Google has released Gemini 3.5 Flash, pushing its fast-tier model to a 10 million token context window while keeping pricing at commodity levels.
Headline capabilities
- ▸10M token context: whole monorepos, multi-hour video, or years of chat history in one request
- ▸Needle-in-haystack retrieval accuracy above 99% at full length, per Google's benchmarks
- ▸Streaming ingestion so long inputs process while still uploading
The strategic angle
Flash-tier models have become the workhorse of production AI. By making extreme context cheap, Google is betting that many RAG pipelines will be replaced by brute-force context stuffing.
"For a lot of teams, 10M tokens of cheap context is a simpler architecture than a vector database they have to keep in sync." — infrastructure engineer, quoted in launch coverage
Rivals are expected to answer within the quarter; leaked roadmaps suggest similar context expansion across the industry.
Source: Google DeepMind


