LATEST NEWS
SELECTED FOR YOU

Google staff turn to Gemini 3.8 Flash as AI costs mount

ByAshish KumarAshish Kumar 3 mins read
Google staff test Gemini 3.8 Flash as cost pressure grips the AI market
  • Google employees are reportedly testing Gemini 3.8 Flash Preview internally, though Google has not officially confirmed the model or release date.
  • Google is accelerating Flash releases, with Gemini 3.6 Flash and 3.7 Flash arriving just weeks apart as it targets cheaper, high-volume AI workloads.
  • AI costs are under growing pressure: Gartner expects AI platform and model spending to reach $64.25B in 2026, while enterprises increasingly demand measurable returns.

Google employees have reportedly begun testing the Gemini 3.8 Flash Preview on its own coding platform, which is further confirmation that the leading AI companies are working on innovative and cost-effective models at a remarkable rate.

Speed is important for companies observing the rising cost of AI. According to Gartner, global expenditure on AI platforms and models is projected to reach $64.25 billion in 2026, representing an annual increase of 63.4% over 2025, as companies are becoming increasingly picky about the value of their expenditures.

Business Insider learned from images that the “Gemini 3.8 Flash Preview” model was posted on Google’s internal Jetski platform. An employee has reported that it seems to be better than version 3.7 Flash, but warned that this might not be the correct assessment. Google refused to make any comments. Hence, the version number and eventual public release remain unknown.

A near-monthly cadence aimed at the low-cost tier

Google has been on the go. Gemini 3.6 Flash was introduced on July 21, and 3.7 Flash became available on August 13, all taking a little over three weeks.

That’s done on purpose. During Alphabet’s second-quarter earnings call, CEO Sundar Pichai stated that Google expected to release models “almost at a monthly cadence” while working on Gemini 4.

Flash also represents the area where Google is having its greatest pricing effort. Google calls the series a “workhorse” for coding and agents – applications where repeat usages of the language model leads to significant token costs.

Gemini 3.7 Flash launched at an introductory $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026. That is exactly half the $1.50 input and $7.50 output launch pricing of Gemini 3.6 Flash.

Racing on price because the frontier is out of reach

Google’s approach is evident in the way its public product portfolio has gaps at the very top. On May 19, the company announced that Gemini 3.5 Pro was being used inside the company and would come out in June. However, by July 21, the model was still being tested with partners and Google announced that it would be rolled out to the public once it was ready.

For now, that leaves Google competing aggressively on price and release speed while its flagship remains pending.

Artificial Analysis gives Gemini 3.7 Flash at high reasoning effort an Intelligence Index score of 56, compared with 61 for GPT-5.6 Sol at maximum effort and 62 for Claude Fable 5.

OpenAI’s Luna tier remains considerably cheaper than Google’s Flash offer at $0.20 per million input tokens and $1.20 per million output tokens. More broadly, the comparison shows how wide AI pricing has become, from cents to tens of dollars per million tokens depending on capability.

Why buyers are the ones setting the terms

The market backdrop helps explain Google’s approach. Gartner expects AI model and platform spending to surge this year, but analyst Arunasree Cheparthi said enterprise budgets are facing “greater scrutiny,” with more attention on efficiency, cost control and measurable outcomes.

Ramp’s spending data shows how that pressure can affect even top-performing models. In its first month, Claude Fable 5 accounted for just 6% of the tokens businesses bought from Anthropic and 11.4% of Anthropic model spending, despite launching at the top of Artificial Analysis’ intelligence ranking. Anthropic charges $10 per million input tokens and $50 per million output tokens for the model.

Ramp economist Ara Kharazian said the pattern suggests there is an upper limit to what businesses will pay for raw capability. That favors models that are cheap, fast and good enough for the job — precisely the part of the market Google is trying to capture.

The announcement of a public Gemini 3.8 Flash release is still uncertain. Currently, the information regarding its demonstration comes from the claims made by Business Insider, meaning that there has not yet been an official recognition of this fact by Google, which is of utmost importance given the frequently changing names of models and release dates in the market.

 

The smartest crypto minds already read our newsletter. Want in? Join them.

FAQs

What is Gemini 3.8 Flash?

It is an unreleased Google AI model that company staff are testing in preview on Google's internal coding platform Jetski, according to Business Insider. One employee said it felt better than 3.7 Flash, but Google declined to comment and a public launch is not guaranteed.

Why is Google focusing on Flash models instead of a flagship?

Google's promised Gemini 3.5 Pro flagship, pledged to developers in May, never shipped, leaving the company without a frontier model as Anthropic and OpenAI lead, per Cryptopolitan. Google has instead doubled down on its cheaper, faster "workhorse" Flash models for coding and agents.

How much does Gemini 3.7 Flash cost?

Google priced 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, half the launch cost of 3.6 Flash, according to Google's announcement. OpenAI's GPT-5.6 Luna tier is cheaper still at $0.20 and $1.20.

Share this article

Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Ashish Kumar

Ashish Kumar

Ashish Kumar is a crypto and financial journalist with eight years of newsroom experience. He covers what’s happening with crypto markets, regulation, DeFi, and exchange ecosystems. He has worked with Coingape, Todayq, and Newsroompost. Ashish holds a PGDP in English Journalism from the IIMC. He has also interviewed industry figures including Arthur Hayes, Yat Siu, Austin Federa, and more.

MORE … NEWS