Up North AIUp North
Back to news

Gemini 3.7 Flash: Half the Price, Three Weeks Later

Open-Weight Models Are Closing the Gap Fast. Perplexity Doubles Down on Multi-Model Orchestration.

Share

Gemini 3.7 Flash: Half the Price, Three Weeks Later

Google didn't wait. Just three weeks after Gemini 3.6 Flash, it released 3.7 Flash on August 13 — its "most intelligent workhorse model," aimed squarely at coding, agents, software engineering, and multi-step tool orchestration [4]. Introductory pricing runs at $0.75/million input tokens and $3.75/million output tokens through the end of 2026 — literally half the previous rate [5][6].

This is a workhorse model, not a flagship — and that's the point. Google is racing to make Flash cheap enough that it becomes the default subagent everyone reaches for, embedded across Workspace, GitHub Copilot, and Gemini Enterprise. The three-week release cadence is itself the story: model iteration cycles are compressing to the point where "current" has a shelf life measured in weeks, not quarters [6].

Open-Weight Models Are Closing the Gap Fast

Mid-August saw a pileup of open releases — DeepSeek V4 Pro/Flash, Meta's Muse-Glimmer-30B, Qwen3.8 variants, NVIDIA Nemotron, Liquid AI LFM2.5 — all converging on the same targets: efficiency and agentic/coding performance [1][4]. This isn't a coincidence of timing; it's competitive pressure. Every closed-model release now gets an open-weight answer within days.

For builders, this matters more than any single benchmark win. When open models are credibly competitive on coding and orchestration tasks, the cost floor for running agentic systems drops for everyone — including the people building on closed APIs, because it forces pricing discipline across the board.

Perplexity Doubles Down on Multi-Model Orchestration

Perplexity's moves this week are a case study in where the puck is going: Grok 4.6 rolled out to Pro/Max users almost immediately, while the company's leadership publicly praised Gemini Flash models for subagent roles inside their "Computer" harness [1][4]. Nobody's betting on one model anymore — they're building routing layers that pick the right model for the right sub-task.

Engineer arranging diagrams from multiple laptops on a desk

That's the real product now: not the model, but the harness that decides which model handles which piece of a job.

What This Means For Your Business

Three releases, three days, one pattern: the market has moved past "which model is best" to "which model is best for this specific step in this specific pipeline." Grok 4.6 for long-running agentic reasoning, Gemini 3.7 Flash for cheap high-volume subagent work, open-weight models as a cost and leverage backstop — this is what post-code infrastructure actually looks like. Nobody's writing more code to make this work; they're writing better judgment calls about which system does what.

If you're still evaluating AI vendors by picking "the best model" and building around it, you're already behind. The companies pulling ahead are the ones treating model selection as a runtime decision, not a procurement decision — routing tasks dynamically, testing new releases within days of launch (as Perplexity clearly does), and treating price-per-task as seriously as capability. Three-week release cycles mean any static bet on a single vendor is stale before your integration is even finished.

The judgment work that matters now is architectural: which tasks need a frontier reasoner, which need a cheap fast subagent, and how you orchestrate between them without your engineering team becoming a full-time model-swapping shop. That's not a coding problem anymore. It's a design and governance problem.

Key takeaway: The model layer is becoming a commodity that changes weekly — the durable advantage is in how well you orchestrate across it, not which single model you picked.

See what we're exploring →

Sources

  1. https://x.ai/news/grok-4-6
  2. https://cursor.com/blog/grok-4-6
  3. https://x.ai/
  4. https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/
  5. https://venturebeat.com/technology/googles-gemini-3-7-flash-targets-coding-and-agents-with-a-50-introductory-price-cut
  6. https://arstechnica.com/ai/2026/08/google-announces-gemini-3-7-flash-just-three-weeks-after-previous-release/

Stay ahead of AI

No spam. Unsubscribe anytime.

Want to go deeper?

Reading the news is one thing. Exploring the frontier is another. See what we're building.