Skip to main content

Gemini 3.6 Flash Launches, But Users Aren't Impressed: Cheaper Tokens, Dumber AI?

If you've been following Gemini since last year, you've probably heard the joke: it's like watching your own brother slowly develop Alzheimer's. Normally, a new model launch would be the perfect chance to shut down such jokes. But after Google unveiled three new models—3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—the laughter only got louder. It's as if the company said, "We know you didn't believe in us, but we'll still disappoint you."

Let's start with the flagship: Gemini 3.6 Flash. It's a minor upgrade from the 3.5 Flash introduced at I/O in May, positioned as a workhorse model. The biggest selling point? Token savings. According to Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than its predecessor, and up to 65% less in complex coding scenarios. The price dropped too: $1.5 per million input tokens and $7.5 per million output tokens, down from $9. Faster and cheaper—sounds great, right?

Image

In benchmarks, code ability improved: DeepSWE rose from 37% to 49%, and the model now makes fewer unnecessary code changes. The MLE Bench for machine learning research jumped from 49.7% to 63.9%. Computer operations (OSWorld-Verified) went from 78.4% to 83%, and it's now a built-in tool in the Gemini API. A minor but welcome update: the knowledge cutoff date finally moved from January 2025 to March 2026. On the security front, Google says 3.6 Flash has upgraded Frontier Safety protection against chemical, biological, radiological, nuclear, and cyber threats, with better jailbreak resistance.

Image

Next up is 3.5 Flash-Lite, focused on speed and affordability. It's the fastest in the 3.5 series, generating 350 tokens per second. Pricing is ultra-low: $0.3 input and $2.5 output per million tokens. Quality is significantly better than the 3.1 Flash-Lite from March. It supports adjustable reasoning levels, so simple tasks run fast, while complex ones can think harder. Computer operations are also built-in.

Benchmark improvements are notable: Terminal-Bench2.1 jumped from 31% to 54%, long-context GDM-MRCR v2 from 60.1% to 72.2%, and real-task execution GDPval-AA v2 from 642 to 1140. Interestingly, this lower-tier model actually outperformed its higher-level sibling, 3Flash, in many agent and code tasks—like SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%). So tasks once handled by 3Flash now have a faster, stronger alternative.

Then there's 3.5 Flash Cyber, a specialized model for finding and fixing security vulnerabilities. Google's logic: AI finds bugs faster than humans can fix them, so why not use AI to patch them? It's fine-tuned from 3.5 Flash and paired with CodeMender, where multiple agents work together and summarize into a report. On the CyberGym benchmark, it reached top-tier levels while using fewer tokens. But it's not publicly available—Google, following Anthropic's lead, is keeping it locked down, offering limited access to trusted partners for internal testing. The goal is to give defenders time to patch before attackers exploit.

So where's the true flagship, 3.5 Pro? According to Google, it's still being tested with partners. But behind the scenes, things aren't smooth. Bloomberg reports that 3.5 Pro's code generation has never met internal expectations. Google updated training data in late June to fix code shortcomings, but results didn't improve. Rumors even say the original training plan was scrapped and started over. Meanwhile, Google AI spokesperson Logan Kilpatrick openly skipped 3.5 Pro in public statements, shifting focus to Gemini 4, claiming its pre-training is "the largest and most ambitious" and progress is exciting. Given the flagship delay, this feels like a distraction.

Looking at the Flash models alone, netizens aren't buying it. Third-party evaluator Artificial Analysis says the new models cut task time in half and improved token efficiency, with Flash-Lite's intelligence index up 11 points. But 3.6 Flash's intelligence level remained basically unchanged from 3.5 Flash. The price drop isn't enough to ignore the capability gap, nor is the capability rise enough to justify the cost. On X, users criticized that 3.6 Flash scored exactly the same as 3.5 Flash on Artificial Analysis, and worse than competitors like Meta Spark1.1, GLM-5.2, 5.6Luna, Sonnet5, Grok4.5, and 5.6Terra. One user called it "simply terrible."

User Angel pointed out that 3.6 Flash's usage cost is higher than GPT-5.6Sol medium, but its intelligence is lower—an awkward combination for a series built on value for money. Meanwhile, OpenAI's Tibo announced that due to reaching 10 million weekly active users, paid users of Codex and ChatGPT Work got a new quota reset. The contrast is stark.

Real-world testers added fuel to the fire. User Balder criticized the three models as worse than the previous 3.5 Flash, with problems including basic decoding errors, strange Chinese word choices, poor image and video quality, mismatch with instructions, severe limitations, frequent memory confusion, and incorrect tool calls. He even joked that Google might be freeing up computing power to sell to Anthropic—sarcastically calling it "what Pichai wants—Cloud First." X user Conor Dart tested an AI-generated game using Google's own Antigravity, resulting in worse wood textures and marble that looked almost similar but still unusable. He called it another failure, saying he'd rather wait for Gemini 3.5 or 3.6 Pro.

On one side, official accounts highlight efficiency, cheaper prices, and fewer tokens. On the other, most netizens give negative feedback, alongside bad news like the flagship delay, top talent leaving, and shrinking market value. It's hard not to wonder if Google messed up its technology path, focusing too much on cost-saving and forgetting about improving intelligence, relying on a few Flash models to stabilize the situation. Regardless, Gemini 4 pre-training has started. If this attempt also fails, the Alzheimer's joke could become a prophecy.

Key Points

  • Gemini 3.6 Flash offers 17% fewer tokens and lower prices, but intelligence scores remain flat.
  • 3.5 Flash-Lite outperforms its higher-tier sibling in many tasks, with significant benchmark gains.
  • 3.5 Flash Cyber is a specialized security model, but access is restricted.
  • Flagship 3.5 Pro is delayed due to code generation issues; Google shifts focus to Gemini 4.
  • User feedback is largely negative, citing regression in quality and value compared to competitors.