Gemini 3.8 TTS reads your instructions aloud? Move delivery into metadata

Migrate to Gemini 3.8 Flash TTS with literal transcript text, speech_metadata and the Interactions API. Check audio output without mixing older request formats.

A stencil sheet drawn in black ink on a pale card, a sage background and the title Gemini 3.8 TTS.

For Gemini 3.8 Flash TTS, put the words to be spoken in the transcript and sustained delivery instructions in speech metadata. If you paste “speak calmly” into text that the model treats as a literal transcript, that instruction can become part of the speech you asked it to produce. The Gemini 3.8 Flash TTS migration guidance explicitly separates transcript text from delivery metadata.

Google announced the new TTS models in its September 22, 2026 API release notes. This article gives a documentation-aligned migration example for the Interactions REST API. The request structure and decoding logic can be checked locally; this article does not claim a live listening test, improved voice quality or current Ofox support for this endpoint.

Separate what is said from how it is said

InformationWhere it belongs in this example
Words the listener should hearThe text content’s text field
Sustained delivery, such as calm and clearA speech_metadata annotation’s style
Selected voicegeneration_config.speech_config
Requested audio outputresponse_format

Keep that separation even when the instruction is short. It makes the request easier to inspect and avoids confusing an instruction with the script itself. A transcript that includes a character saying “speak calmly” is different: those words belong in the transcript because the listener is supposed to hear them.

The model documentation also describes point-in-time vocal events. Do not assume that every older prompt tag or every arbitrary annotation is supported. Use the current reference for the model and API family you selected.

Use one API family from request to response

The following example uses the Interactions API. Do not mix its input array and annotations with a GenerateContent request body or with fields from an older SDK example. The official speech-generation guide is the source for this request family.

Save this as request.json:

{
  "model": "gemini-3.8-flash-tts",
  "input": [{
    "type": "user_input",
    "content": [{
      "type": "text",
      "text": "The next train leaves at noon.",
      "annotations": [{
        "type": "speech_metadata",
        "style": "calm and clear"
      }]
    }]
  }],
  "response_format": {"type": "audio"},
  "generation_config": {
    "speech_config": [{"voice": "Kore"}]
  }
}

For an authorized direct Google account, the request is:

curl --fail-with-body --silent --show-error \
  'https://generativelanguage.googleapis.com/v1beta/interactions' \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H 'Content-Type: application/json' \
  --data-binary @request.json > response.json

Supply your own key securely through the environment. Running this command makes a provider API request and can incur usage charges. A successful local JSON parse does not prove that your account has model access or that a provider has accepted the request.

This is a single-speaker example. The official multi-speaker configuration has a different structure; do not turn the speech_config array above into an improvised dialogue schema. Each dialogue turn must use a speaker that matches the configured speakers.

Decode audio only after checking the response

An HTTP error saved into response.json is not audio. Inspect the HTTP result and the response structure before decoding base64 data. For Interactions REST responses, inspect model_output steps and their content blocks. Do not assume that an SDK convenience property such as output_audio exists in the raw JSON.

A defensive decoder should:

  1. Reject an error response instead of writing it as an audio file.
  2. Select audio content from model_output steps, ignoring text and tool content.
  3. Verify the returned MIME type before deciding on a file extension.
  4. Decode the base64 payload and preserve its native container.

The migration guide says unary output defaults to WAV. Do not automatically add a WAV header copied from an older raw-PCM example: a second header can corrupt an already valid WAV file. If you intentionally request another format, use its actual MIME type and container rather than renaming its bytes.

Make the first migration test small

Begin with a single speaker, a short script and one delivery instruction. Listen for omitted words, extra instruction text, pronunciation and unexpected changes of voice. Save the request, model ID, timestamp and output file together. These are proposed acceptance checks, not results we obtained for this article.

Only then extend the script or add speakers. Change one variable at a time: moving an instruction into metadata while also changing voices, splitting the script and switching API families makes a failure difficult to diagnose.

For a voiceover workflow, compare the generated audio against the exact text before editing it into a video. The faceless video workflow guide covers the larger production pipeline. TTS migration is one step in that pipeline; it does not guarantee timing, pronunciation or consent for a particular voice.

What to verify before routing through another provider

An OpenAI-compatible text endpoint does not imply support for Google’s Interactions API or its speech metadata fields. Check the provider’s specific audio route, supported model ID and output format. This article’s direct Google example is not an Ofox endpoint recipe.

Keep model selection separate from API-format migration. Gemini 3.8 Flash-Lite TTS is a related offering, but support and behavior must be checked against its own reference before changing the model string. Our multimodal API overview provides broader context without replacing the current provider documentation.

Frequently Asked Questions

Why might the model read a delivery instruction aloud?
The new model treats input text as a literal transcript. Put sustained delivery guidance in the documented speech metadata location rather than inserting it into the script.
Can I copy this JSON into GenerateContent?
No. This example uses Interactions input and annotation fields. GenerateContent uses a different request structure; use that API's documented example end to end.
Is the output always raw PCM?
No. The migration guidance says unary output defaults to WAV. Inspect the actual response format and do not add a second WAV header.