Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

Google Makes Voice Model Gemini 3.8 Live and Its Reasoning Variant Generally Available in the API - Same Rates for Both

Image: Google official blog

Google Makes Voice Model Gemini 3.8 Live and Its Reasoning Variant Generally Available in the API - Same Rates for Both

On September 15, 2026, Google announced the voice dialogue model Gemini 3.8 Live and the reasoning-focused 3.8 Live Extended Thinking, and made both generally available in the Gemini API. The pricing page lists the two at the same rates, and the models can run tools in the background while the conversation continues.

On September 15, 2026, Google announced two models for spoken conversation: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, the latter with stronger reasoning 1. In the Gemini API release notes, both are listed as generally available (GA) as of the same date 2. Both are audio-to-audio models, taking speech in and returning speech, and are used through the Live API.

Google calls the pair its “most advanced live dialogue models yet” and presents them to developers and enterprises as building blocks for voice agents that can hold up in production 1.

How the two models divide the work

3.8 Live is the one aimed at scale and low cost, pairing smooth dialogue with an understanding of visual input. 3.8 Live Extended Thinking targets complex tasks and adds more multi-step reasoning 1. In the release notes, Google recommends 3.8 Live as the default for most low-latency voice agents, and Extended Thinking for cases that need deeper reasoning running behind the conversation 2. The model IDs are gemini-3.8-live and gemini-3.8-live-extended-thinking.

3.8 Live works with visual input almost as it arrives, and when the speaker changes language, it recognizes which of its 97 supported languages is in use and switches on its own without ending the conversation. For anyone building a voice agent, the notable part is that it can call tools and APIs in the background in parallel with the conversation: it can respond as soon as a request comes in and keep talking until the work finishes 1. The release notes also list asynchronous function calling by default among 3.8 Live’s features 2.

Extended Thinking reasons and speaks at the same time. Google describes a design that keeps the conversation from breaking up: the model first gives a short acknowledgment such as “Let me check that,” then narrates how a multi-step background task is progressing 1.

According to the model card, both models are based on Gemini 3 Pro. Inputs are audio, images, video and text, with a context window of up to 128K tokens, and outputs are audio and text, up to 64K tokens. The knowledge cutoff is January 2025, and the known limitations listed include hallucinations as well as occasional slowness or timeouts 3.

All benchmark figures come from Google. 3.8 Live Extended Thinking is reported to rank first overall on Artificial Analysis’ Speech to Speech Quality Index with 82.6, to score 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking for agentic task completion, and 97.7% on Big Bench Audio. 3.8 Live is said to place second in the Speech Agent Arena 1.

Where to use them and what they cost

The rollout began the same day. For developers, both models are available in the Gemini API and Google AI Studio. For enterprises, they are in private preview in Gemini Enterprise, with Gemini Enterprise for Customer Experience listed as coming soon 1.

The consumer entry points differ between the two. 3.8 Live goes into Search Live in Google Search. Extended Thinking goes into Gemini Live in the Gemini app, into Docs in Google Workspace for Google AI Pro and Ultra subscribers, and into Gmail and Keep for all Google AI subscribers. Availability for Workspace business customers is listed as coming soon 1.

On the Gemini API pricing page, 3.8 Live, 3.8 Live Extended Thinking and 3.1 Flash Live Preview share a single pricing table, with no per-model differences shown 4. Paid-tier rates (per 1 million tokens) are as follows.

ItemPaid-tier rate
Input (text)$0.75
Input (audio)$3.00, or $0.005 per minute
Input (image/video)$1.00, or $0.002 per minute
Output (text, including thinking tokens)$4.50
Output (audio, including thinking tokens)$12.00, or $0.018 per minute

The free tier costs nothing for input or output, but free-tier data is used to improve Google’s products, while paid-tier data is not. Grounding with Google Search is free up to 5,000 requests a month, shared across all Gemini 3.x models, and then costs $14 per 1,000 requests 4.

Output rates are stated as including thinking tokens. Extended Thinking, the stronger reasoner, can be chosen at the same rate, but if its reasoning produces more output tokens, the bill could end up higher than with 3.8 Live. Deciding which one to use by default will likely come down to measuring actual token consumption on your own conversation scenarios.

How it differs from OpenAI’s GPT-Live 1

OpenAI’s voice model GPT-Live 1, which the company first announced for ChatGPT Voice in July, is also offered through the API. It is a full-duplex model that listens and speaks at once and can hand reasoning and tool use off to a backend agent. Voice sessions are billed per second at $0.05 per minute, and backend model use and tool use are billed separately 5.

Gemini 3.8 Live is built so that the model itself runs tool calls in the background while it talks, and its pricing is expressed as per-token rates plus per-minute rates for audio and video 14. Because the billing units and what each rate covers are different, the two companies’ prices cannot be compared on per-minute figures alone.

On Google’s side, Gemini 3.8 Flash and Flash Cyber were released on September 2, and less than two weeks later a 3.8 model for voice dialogue has joined them. 3.8 Live is a multimodal dialogue model that handles audio, video and text together, and Google notes that developer platforms including Agora, LangChain, LiveKit, Pipecat and Vercel let developers build voice interfaces on the Gemini Live API 1.

Google also says that all audio generated by its AI products carries a SynthID watermark 1.

Sources

  1. Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking - Google official blog (September 15, 2026)
  2. Gemini API release notes - Official Gemini API release notes (September 15, 2026 entry)
  3. Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card - Official Google DeepMind model card
  4. Gemini Developer API pricing - Official Gemini API pricing page
  5. GPT-Live 1 - OpenAI official model page (for comparison)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →