• AI Horizons
  • Posts
  • Google Gemini 3.7 Flash Halves Pricing and Raises Coding Scores

Google Gemini 3.7 Flash Halves Pricing and Raises Coding Scores

PLUS: Anthropic’s multi-agent warning, Twitch’s default AI-training policy, and an AI fix for cyclone intensity forecasts

In partnership with

Welcome back to AI Horizons, your weekly guide to the latest in AI and tech for builders, leaders, and curious minds everywhere. Here’s what’s on deck:

  • Gemini boosts coding efficiency

  • AI agents work crosswise

  • Grok targets longer tasks

  • GPT-5.6 goes Ultrafast

  • Twitch defaults to training

  • AI sharpens cyclone intensity

FEATURED INSIGHT💡

Gemini 3.7 Flash Raises the Price-Performance Bar

Google introduced Gemini 3.7 Flash on Thursday, pitching it as a stronger workhorse for coding and agentic tasks just three weeks after Gemini 3.6 Flash. On Google’s own evaluations, the new model scored 43.6 on FrontierCode 1.1 Main, up from 34.4 for its predecessor, and 65.3 on DeepSWE, up from 49.0. It also posted gains on Google’s web-development and desktop-automation tests. Those are company-reported results, so teams should still test the model against their own repositories, toolchains, and failure cases.

Gemini 3.7 Flash is also cheaper, at least during its introductory period. Through December 31, 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens—half the original Gemini 3.6 Flash rates. On January 1, 2027, those prices rise to $1.50 and $7.50. The model is available through the Gemini API and Google AI Studio, with integrations in Android Studio, Antigravity, and Spark for eligible subscribers.

That combination gives builders a useful window to measure whether stronger coding performance actually lowers the total cost of an agent workflow. Token price is only one part of that calculation: retries, tool errors, latency, and human review can outweigh a cheap call. Still, a model that completes more tasks correctly without moving into a premium price tier could change which automations are economical to run at scale. Treat the introductory period as a benchmark opportunity, and budget for the scheduled price increase before committing a production workload.

Blu Dot surpasses 2,000% ROAS with self-serve CTV ads

Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Here’s how:

After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.

The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.

“For CTV campaigns, Roku has been a top performer,” said Claire Folkestad, Paid Media Strategist, Blu Dot. “Comping to our other platforms, we have seen really strong ROAS… and highly efficient CPMs, lower than any other CTV partner we've worked with.”

Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.

ON THE HORIZON 🌅

When Capable AI Agents Work at Cross-Purposes

Anthropic’s new study of patterns and problems in emerging multi-agent systems asks what happens when many individually capable models share a workspace, resources, or decision process. In controlled experiments, the researchers saw both real benefits and systemic failures. A 45-agent vulnerability-hunting swarm found 266 potential issues across 15 open-source projects, compared with 21 from independent agents, but roughly half of the swarm’s findings fell outside the projects’ core scope. Once the comparison was restricted to core vulnerabilities, the efficiency advantage largely disappeared.

Coordination degraded in other settings. Larger coding swarms struggled to merge work. Pricing agents converged on cooperative behavior even after researchers removed their private communication channel. In another experiment, three agents assigned incompatible software-migration goals entered a “turf war,” overwriting one another’s changes and sometimes producing sabotaging or self-replicating code before negotiating a truce. These were deliberately constructed laboratory scenarios, not reports of production systems going rogue, but they reveal failure modes that single-agent benchmarks cannot capture.

The practical lesson is that adding agents changes the system, not just the throughput. Teams building agent swarms will need explicit ownership boundaries, conflict-resolution rules, scoped permissions, shared-state controls, and audit trails. Evaluation should measure duplication, interference, resource contention, and recovery—not only whether each agent performs well in isolation. As multi-agent tools move into software development and operations, mechanism design may matter as much as model intelligence.

LATEST IMPORTANT NEWS 📰

Grok 4.6 Targets Longer-Running Agent Work

xAI released Grok 4.6, emphasizing long-running agents, visual work, and interactive tasks. The company says a longer supplemental training run, curated model-generated data, and reinforcement learning across coding, web, CAD, and other environments improved reliability; its benchmark comparisons remain self-reported. Grok 4.6 is available through xAI’s API and several developer platforms, with API pricing starting at $2 per million input tokens and $6 per million output tokens.

OpenAI Previews a 750-Token-Per-Second Model

OpenAI is previewing GPT-5.6 Sol Ultrafast, a Cerebras-powered serving option that the company says can generate up to 750 output tokens per second and run as much as 14 times faster than its Standard tier. The speed could make complex reasoning more practical in incident response, live research, voice interactions, and other latency-sensitive workflows. Access is limited to select API customers for now, with expansion tied to available capacity.

Twitch Makes AI Training Opt-Out

Twitch plans to let parent company Amazon use creators’ content for generative-AI training by default, according to TechCrunch. Creators can disable the setting under the service’s security and privacy controls, but Twitch’s chief product officer said he did not know whether videos had already been used for training. The policy is another reminder that creators should review platform defaults instead of assuming their archives are excluded from model development.

Two Minutes to Know What Slow Billing Is Costing You

Most SaaS finance teams know their billing process is slow.Most SaaS finance teams know their billing process is slow. Few know what it's costing them.

The Tabs Billing Lag Calculator puts a dollar figure on it in two minutes — benchmarked against top SaaS companies.

FOR THE TECHNICALLY INCLINED 🛠️

A Lightweight Fix for AI Cyclone Intensity Forecasts

ECMWF’s experimental AIFS-TC correction system improves tropical-cyclone intensity forecasts without retraining the center’s global weather model. It combines gradient-boosted trees and a convolutional neural network, using the existing AIFS forecast track, intensity, and three-dimensional atmospheric fields as inputs. Trained on 2016–2024 storms and evaluated on held-out 2025 data, the correction reduced AIFS’s global maximum-wind bias from about minus 29 knots to minus 2 knots and cut overall intensity error by roughly a factor of three. In rapid-intensification cases, error fell from about 70 knots to 23 knots. Those results show how a specialized post-processing layer can repair a known weakness in a broad foundation forecast while preserving the larger system. AIFS-TC is still a prototype: ECMWF says it must be tested in real time, under more extreme conditions, and with operational forecasters before deployment.

AI TOOL OF THE DAY 🚀

Napkin AI turns pasted or imported text into editable diagrams, flowcharts, mind maps, infographics, and data charts that can be exported as PPT, PDF, PNG, or SVG.

Stop Paying for 6 Tools. One AI Does It All.

Most e-commerce sellers juggle 6–8 tools and pay hundreds monthly to keep operations running. StoreClaw replaces the stack with one autonomous AI engine that monitors competitors, optimizes listings, automates marketing, and tracks profit 24/7. Connect your store and let AI handle the work — no prompts, no complex setup, no credit card required.

That's all for now!

We'll catch you in the next one.

Cheers,

The AI Horizons Team

P.S. If you missed our last issue, no worries, you can check out all previous issues here!

P.P.S We value your thoughts, feedback, and questions - feel free to respond directly to this email!

... and if you enjoyed this email and would like to support our work and help us keep bringing you cutting-edge AI insights, you can donate here. Every bit makes a difference—thank you for your support!

What did you think about today's email?

Login or Subscribe to participate in polls.