codingBy HowDoIUseAI Team

Why Gemini 3.5 Pro's delay is rattling Google (and what's actually confirmed)

Gemini 3.5 Pro missed its June launch, Google lost key researchers, and Alphabet's stock took a hit. Here's what's real and what's rumor.

Sundar Pichai stood on stage at Google I/O in May 2026 and teased a model that sounded like it would put every competitor on notice. Then he asked the crowd for one more month. They groaned. And that groan turned out to be the opening note of one of the messiest product rollouts in Google's recent history — one that has since wiped tens of billions of dollars off Alphabet's market cap and cost the company some of its most decorated AI researchers.

If you've been trying to figure out what's actually true about Gemini 3.5 Pro versus what's just noise from aggregator channels repeating each other, this guide sorts it out. No badge-guessing, no vague "sources say" — just what's documented, what's plausible, and what you should genuinely ignore.

What actually happened with the Gemini 3.5 Pro launch?

The timeline itself isn't in dispute. Google unveiled Gemini 3.5 Pro in May at its annual developer conference, billing it as the company's most powerful model and promising a broad public rollout the very next month. That June deadline slipped by without any update from the company. At I/O, Sundar Pichai walked the audience through a model that sounded genuinely exciting: a 2-million-token context window, a new Deep Think reasoning mode, and frontier multimodal capabilities — and when pressed on the release date, Pichai said "give us until next month," and the live audience responded with audible groans.

The reason for the slip is the part worth paying attention to. In an effort to catch up with competitors' lead in the AI coding space, Google updated the model's training data late last month, but the test results fell short of expectations, forcing a delay in the release process. That's a coding-benchmark problem, not a vague "safety review" problem. The delay is not due to product roadmap adjustments, but rather Google's desire to continue improving the model's overall capabilities, particularly in AI coding, which is currently the most fiercely contested area.

Market reaction was swift and repeated. Different sessions produced different numbers — one report put a single-day drop at Alphabet shares closed down 4.4% on Thursday, erasing about $200 billion in market value, while a broader accounting across the whole saga puts the delays have caused a $425 billion drop in Alphabet's market valuation and led to talent retention challenges within Google's DeepMind division. Either way, this is a market treating a missed software deadline like a genuine crisis, not a rounding error.

Why did Google lose its top AI researchers right when this happened?

This is the part of the story that's fully documented, not speculative. In a single ten-day window in June 2026, Google DeepMind lost four researchers whose names carry serious weight in the field. Four senior DeepMind researchers departed in a single week — Noam Shazeer (Gemini co-lead) to OpenAI, Nobel laureate John Jumper to Anthropic, plus Jonas Adler and Alexander Pritzel also joining Anthropic.

Shazeer isn't a minor name to lose. Noam Shazeer — a Gemini co-lead and one of the eight original authors of the "Attention Is All You Need" transformer paper — left for OpenAI. Google had already paid an enormous premium to keep him once: Google previously paid more than $2 billion to acqui-hire Shazeer and part of his Character.ai team.

Then came Jumper. In his own words, posted publicly, "After nearly nine years, I have decided to leave Google DeepMind and join Anthropic." This wasn't a junior engineer chasing a bigger title. Jumper, who won a Nobel prize alongside Google's Demis Hassabis in 2024, is best known as the co-creator of AlphaFold, a breakthrough AI that has predicted over 200 million protein structures, cutting years off biological and medical research. Two of his AlphaFold collaborators, Jonas Adler and Alexander Pritzel, followed him to the same company within days.

Why does this matter beyond the headline? Frontier AI progress depends on a genuinely small pool of people who understand how to push these systems forward, and frontier AI capability is built by a relatively small number of people at the top of the field, and that population is now moving toward the competitors Google most needs to beat. Money is clearly part of the calculation too — pre-IPO equity at a company valued in the hundreds of billions offers a payday that even a senior Google salary with vested shares cannot easily match.

To Google's credit, this doesn't automatically mean the ship is sinking. "Falls short of internal goals" is the reporting's framing, not Google's — Google's public position is that it is taking the time required to ship the model correctly, a delay taken for quality is not the same as a failure to execute. That's a fair point worth holding onto while you evaluate the rest of the rumor mill.

Is the 2-million-token context window real or just a rumor?

This one sits closer to "credible but unconfirmed in production" than outright rumor. Google's own conference materials describe the number directly: Pro targets a 2-million-token context window, a "Deep Think" reasoning mode, and frontier multimodal understanding — the ability to work across text, images, and other formats. But announcing a target on stage isn't the same as shipping it to the public with a benchmark sheet attached.

Independent trackers have flagged the same gap. Neither specification has been confirmed in official Google documentation, and Google's official communications remain limited — at I/O 2026, the company confirmed that Gemini 3.5 Pro exists, is being used internally, and would follow the release of Gemini 3.5 Flash. As of one mid-July check, the public Gemini API lists gemini-3.5-flash and gemini-3.1-pro-preview as available model IDs — meaning Pro itself hadn't fully surfaced under its own model ID at that point. If the 2M window does land, it would be a genuine leap: the largest production context available from any major provider as of this writing, while Claude's current Opus 4.x models cap at 200K tokens.

What is Deep Think mode, and can you actually use it right now?

Deep Think isn't vaporware — it's a real, shipping feature, just not exclusive to the unreleased 3.5 Pro. Google already rolled out an earlier version broadly: Google rolled out Gemini 3 Deep Think mode to Google AI Ultra subscribers in the Gemini app, delivering a meaningful improvement in reasoning capabilities designed to tackle complex math, science and logic problems that challenge even the most advanced state-of-the-art models. On the benchmark side, it's genuinely strong: Gemini 3 Deep Think is industry leading on rigorous benchmarks like Humanity's Last Exam (41.0% without the use of tools) and ARC-AGI-2 (an unprecedented 45.1% with code execution), using advanced parallel reasoning to explore multiple hypotheses simultaneously.

You can try the current version today. According to Google's own support documentation, here's the process:

  1. Go to gemini.google.com on desktop
  2. Click the model name in the text box and select Pro
  3. Click the model name again and select Thinking Level → Deep Think
  4. Enter your prompt and click Submit

Google's help page notes the practical trade-off directly: Deep Think is an experimental capability, allowing you to try out Gemini's newest advanced reasoning — Google may discontinue or suspend Deep Think at any time without prior notice. For developers who want the technical documentation rather than the consumer app version, Google's Gemini API thinking docs and the DeepMind Deep Think model page both go deeper into how thinking budgets and reasoning levels actually work under the hood.

How does this compare to what OpenAI and Anthropic are doing?

This is where distribution versus capability becomes the real story. Google's rivals haven't been sitting still while Gemini 3.5 Pro stalls, and the competitive pressure compounds the researcher exodus. As one analysis put it plainly, Google's distribution advantage — Gemini on Search, Android, Workspace — is real and durable, but distribution is how you reach users at scale; it is not what determines whether your model is the strongest one.

That distribution point matters more than it sounds. Whatever ships, Gemini has a built-in installed base that no standalone lab can match, since it's baked into products already sitting on well over a billion devices. No amount of benchmark bragging from a competitor changes that math overnight — but it also doesn't fix a coding benchmark gap, and it doesn't bring back a Nobel laureate who already signed with Anthropic.

What should developers actually do while they wait?

Don't build your roadmap around a release date Google hasn't confirmed. If your workflow depends on frontier coding performance right now, evaluate what's actually shipped and benchmarked today rather than what's promised. Keep an eye on Google's official Gemini blog for genuine announcements instead of aggregator content, and check the Gemini API model list periodically to see which model IDs are actually live.

If you're testing Deep Think mode for research, science, or hard coding problems, start with a narrow, well-defined problem — it's not designed for quick back-and-forth chat, and Google says as much in its own docs.

The takeaway worth remembering

A model demo on a conference stage is a promise, not a product. The gap between what Google showed in May and what's verifiably running in production months later is the entire story here — and it's a useful reminder for anyone evaluating AI tools: benchmarks from a keynote slide aren't the same as benchmarks from an independent test. Track the model IDs, not the hype cycle, and you'll make better decisions than anyone chasing the next leaked spec sheet.