Article

Voice search is back - and this time there are no ten links to fight over, just one spoken answer

On 15 September 2026 Google launched Gemini 3.8 Live and brought conversational voice search to Search Live: multimodal, able to solve tasks step by step and to switch between 97 languages mid-conversation. This isn't 2016 voice reading a featured snippet - it's a single, spoken answer with no second option. Classic SEO of rankings and clicks has no answer for it; what answers is being the citable, self-contained source the model picks, in every language your audience asks in.

From voice that read a snippet to voice that solves
2016 voice read a snippet; 2026 voice talks, sees and solves - and returns one answer, not a list.

Voice search failed its first promise a decade ago. It's back, and this time it's different: it doesn't return a list of links to fight over, it returns one spoken answer, with no second option. Classic SEO - rankings, CTR, ten positions - has nothing to say about a surface where there are no positions. And most sites still write for the screen.

What changed on 15 September

Google launched Gemini 3.8 Live and an Extended Thinking version, its most advanced live dialogue models, and wired them into Search Live. In practice: they process visual input in near real-time (the camera is part of the question), detect and switch between 97 languages mid-conversation, run tools and API calls in the background while still talking, and the Extended Thinking version reasons and speaks at the same time, using natural cues like "let me check that" so the dialogue doesn't break. They're in Search Live, the Gemini app, Workspace and the API. Google describes Search Live as step-by-step, real-time troubleshooting help.

Why it's different this time

Because in 2016 "voice search" was Assistant or Siri reading aloud a featured snippet that already existed in text search. You optimised the snippet and got the voice answer for free. Not anymore: the answer is generated, conversational, sees what you point the camera at, solves tasks and changes language with you. It stopped being a reading format and became a surface of its own.

What didn't change, and it's what matters: it still returns one answer, not a list. In text search there are still ten results and a second page to hide on. In a spoken answer there is no second place - either you're the source the model uses, or you don't exist for that user in that moment. The "position zero" of featured snippets was a prominent spot; this is the only spot.

What it breaks in your SEO

It breaks two things at once. First, optimisation: rankings, click-oriented meta titles, CTR, headline tests - none of it applies to a sentence spoken aloud. Second, measurement: no click, no classic referral session, and your AI report in Search Console and GA4 still have no honest line for "I was cited in a spoken answer".

A rigour caveat, because I won't fake what I don't know: as of now there's no reliable public data on what share of search goes through Search Live or voice, nor on how Search Live picks and cites sources. It's early. What's known is the mechanism (a generated, multimodal, multilingual answer); the size of the channel, not yet. Anyone selling you a "voice SEO strategy" with percentages is selling smoke.

How to be the spoken answer

The good news is that preparing isn't a new channel to build from scratch - it's the same GEO work I already argue for, tightened for the spoken format. It's my CITAR method applied to voice, and it comes down to four things:

  1. Citable, self-contained information: 40-60 word answers that make sense read aloud, without depending on what's around them on the page. If the sentence only makes sense with the image next to it or the previous paragraph, it's no good for voice.
  2. A clear entity: the model only cites with confidence those it can resolve as an unambiguous entity. It's the same entity authority work - without it, voice skips you.
  3. Content that solves step by step: Search Live is live troubleshooting. Content that answers "how do I do X" with verifiable steps is what a spoken answer can reuse; a brand text praising itself is not.
  4. Real multilingual coverage: the model switches between 97 languages mid-conversation. If you only have strong content in one language, you vanish the moment the user switches. This is where international SEO stops being an extra and starts deciding whether you're cited.

None of this is voice-exclusive - it's what already makes a site citable by AI in text. Voice is just less forgiving: with no click to recover from a bad first impression, self-containment and clarity stop being best practices and become the condition of entry.

What we don't know yet, and what that should make you do

We don't know how Search Live selects and cites sources, nor how much real traffic it moves - and Google hasn't detailed it. This reading holds until that's public; when it is, I'll update the article. Until then, the sane move isn't to launch a "voice campaign", it's to make sure your best content is citable, self-contained and exists in the languages your audience asks in. Anyone who keeps optimising only for the screen won't lose rankings - they'll lose the surface where the decision becomes spoken and has no second option, and won't even see it in the report.

Sources

Frequently asked questions

Is it worth investing in voice SEO already?

It's early and there's no reliable public traffic-share data. But the preparation (citable, self-contained content, a clear entity, multilingual) is the same GEO work you should already be doing - so it's not an extra cost, it's the same work tightened for the spoken format. GEO guide →

Do I need special schema for voice search?

There's no magic "voice schema". Schema.org's Speakable exists, but it has always had limited, experimental availability. What helps is the usual: correct structured data and self-contained content a spoken answer can reuse. Structured data →

Does multilingual really matter for voice?

Yes, more than in text search. The model switches between 97 languages mid-conversation; if you only have strong content in one language, you drop out the moment the user switches. That's where international SEO starts deciding the citation. International SEO →