Skip to content

feat(ai): support speech metadata on text parts - #16749

Draft
andrewheard wants to merge 9 commits into
mainfrom
ah/ai-speech-metadata
Draft

andrewheard wants to merge 9 commits into
mainfrom
ah/ai-speech-metadata

Conversation

@andrewheard

@andrewheard andrewheard commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Adds support for turn-level speech metadata (SpeechMetadata) on TextPart for Gemini text-to-speech (TTS) models in the FirebaseAILogic Swift SDK, enabling multi-speaker dialogue routing and sustained delivery styling. See the speech generation guide for more details.

Developer-Facing Changes

  • SpeechMetadata [Public Preview]: Added the SpeechMetadata struct to configure turn-level speech synthesis options for individual dialogue turns. Supports assigning turns to specific speakers (speaker) during multi-speaker generation and specifying sustained vocal delivery guidance (style), such as emotion, tempo, or tone.
  • TextPart Integration: Updated TextPart to optionally store and expose speechMetadata: SpeechMetadata?, with an updated public initializer init(_:speechMetadata:).
  • Multi-Speaker Dialogue Routing: Enables structured multi-turn conversation synthesis with SpeechConfig(multiSpeakerVoiceConfig:). Each TextPart in a multi-speaker request can be attributed to a configured voice profile using SpeechMetadata(speaker:).
  • Sustained Delivery Styling: Decouples sustained tone and emotional delivery instructions from transcript text. In Gemini 3.8 TTS models, TextPart.text is treated strictly as a verbatim transcript to speak aloud; passing styling instructions via SpeechMetadata.style prevents the model from reading directorial guidance aloud as speech.
  • Comprehensive DocC Documentation: Authored extensive documentation and usage guides across SpeechMetadata, TextPart, SpeechConfig, MultiSpeakerVoiceConfig, and SpeakerVoiceConfig, linking out to canonical Firebase guides and detailing the operational boundary between sustained turn-level style and point-in-time inline audio tags (such as <laugh>, <sigh>, and <short pause>).

Usage Examples

The following examples demonstrate how to use each feature:

1. Single-Speaker Synthesis with Sustained Style Instructions
import FirebaseAILogic

let ai = FirebaseAI.firebaseAI()
let model = ai.generativeModel(
  modelName: "gemini-3.8-flash-tts",
  generationConfig: GenerationConfig(
    responseModalities: [.audio],
    speechConfig: SpeechConfig(voiceName: "Kore")
  )
)

// Style instructions guide delivery (tone, pacing, emotion) without being read aloud.
let metadata = SpeechMetadata(style: "Calm, reassuring, and soothing whisper.")
let prompt = TextPart(
  "Take a deep breath and close your eyes as we begin our guided meditation.",
  speechMetadata: metadata
)

let response = try await model.generateContent(prompt)

if let candidate = response.candidates.first,
   let audioPart = candidate.content.parts.first as? InlineDataPart {
  // Process synthesized audio (audio/wav)
  print("Received \(audioPart.data.count) bytes of audio data.")
}
2. Multi-Speaker Dialogue Turn Routing
import FirebaseAILogic

// 1. Configure the distinct speaker voices in SpeechConfig
let speechConfig = SpeechConfig(
  multiSpeakerVoiceConfig: MultiSpeakerVoiceConfig(
    speakerVoiceConfigs: [
      SpeakerVoiceConfig(speaker: "Host", voiceName: "Puck"),
      SpeakerVoiceConfig(speaker: "Guest", voiceName: "Aoede"),
    ]
  )
)

let ai = FirebaseAI.firebaseAI()
let model = ai.generativeModel(
  modelName: "gemini-3.8-flash-tts",
  generationConfig: GenerationConfig(
    responseModalities: [.audio],
    speechConfig: speechConfig
  )
)

// 2. Assign each dialogue turn to a configured speaker using SpeechMetadata
let conversationTurns = [
  ModelContent(role: "user", parts: [
    TextPart(
      "Welcome back to Science Today! Today we're exploring deep space propulsion.",
      speechMetadata: SpeechMetadata(speaker: "Host", style: "Enthusiastic and welcoming.")
    ),
    TextPart(
      "Thanks for having me! It's an exciting time for ion drive technology.",
      speechMetadata: SpeechMetadata(speaker: "Guest", style: "Knowledgeable and upbeat.")
    ),
    TextPart(
      "What makes this new breakthrough so significant?",
      speechMetadata: SpeechMetadata(speaker: "Host", style: "Curious and engaging.")
    ),
  ])
]

let response = try await model.generateContent(conversationTurns)
3. Streaming Speech Synthesis with Turn Metadata
import FirebaseAILogic

let ai = FirebaseAI.firebaseAI()
let model = ai.generativeModel(
  modelName: "gemini-3.8-flash-tts",
  generationConfig: GenerationConfig(
    responseModalities: [.audio],
    speechConfig: SpeechConfig(voiceName: "Charon")
  )
)

let prompt = TextPart(
  "Breaking news: Astronomers have detected atmospheric water vapor on a nearby exoplanet.",
  speechMetadata: SpeechMetadata(style: "Formal news anchor, clear and authoritative.")
)

let stream = try model.generateContentStream(prompt)

for try await chunk in stream {
  if let candidate = chunk.candidates.first,
     let audioPart = candidate.content.parts.first as? InlineDataPart {
    // Process streaming audio chunks (audio/l16; rate=24000; channels=1)
    print("Received audio chunk: \(audioPart.data.count) bytes")
  }
}
4. Inline Vocal Tags vs. Sustained Style Guidance
import FirebaseAILogic

let ai = FirebaseAI.firebaseAI()
let model = ai.generativeModel(
  modelName: "gemini-3.8-flash-tts",
  generationConfig: GenerationConfig(
    responseModalities: [.audio],
    speechConfig: SpeechConfig(voiceName: "Fenrir")
  )
)

// Sustained delivery (emotion/mood) is set in SpeechMetadata.
// Momentary point-in-time vocal sounds and pauses are placed directly in transcript text.
let metadata = SpeechMetadata(style: "Amused, playful, and lighthearted storytelling.")
let prompt = TextPart(
  "I couldn't believe it when the cat jumped straight into the box. <laugh> It didn't even fit! </laugh> <short pause> Absolutely ridiculous.",
  speechMetadata: metadata
)

let response = try await model.generateContent(prompt)

Technical Details

This pull request introduces turn-level speech metadata to support the Gemini 3.8 Flash TTS and Flash-Lite TTS model families while keeping the public API clean, type-safe, and decoupled from backend protocol models.

Key Architectural Points:

  • Clean Abstraction on TextPart: While other content parts (such as inline binary data or function calls) do not require speech metadata, TextPart represents speech transcripts and now encapsulates optional turn-level metadata via its speechMetadata property.
  • Encapsulation of Metadata Properties: SpeechMetadata.speaker and SpeechMetadata.style are stored internally and initialized via init(speaker:style:). This enforces clean immutability while allowing future extension without breaking changes.
  • Decoupled Internal Wire Representation: The internal transport type InternalPart maps speechMetadata to and from JSON payloads ({"speaker": "...", "style": "..."}) conforming to the Gemini API specifications, keeping wire serialization logic completely isolated from public API types.
  • Streaming History Aggregation: Updated History.swift so that adjacent text chunks produced during streaming responses are merged intelligently based on their speechMetadata values without concatenating chunks across differing speakers or styles.
Deep Dive: Type Architecture & Encapsulation (SpeechMetadata & TextPart)
  • Defined SpeechMetadata as a public struct conforming to Equatable, Hashable, and Sendable.
  • Marked the struct with **[Public Preview]** to signal its preview status alongside Gemini 3.8 TTS capabilities.
  • Stored speaker and style properties internally to maintain tight encapsulation while providing a comprehensive public initializer with default nil arguments.
  • Extended TextPart with public let speechMetadata: SpeechMetadata? and provided an overloaded initializer public init(_ text: String, speechMetadata: SpeechMetadata? = nil). Existing call sites passing only a string literal or single argument compile without modification.
Deep Dive: Wire Protocol & Serialization (InternalPart)
  • Added speechMetadata: SpeechMetadata? to InternalPart and added speechMetadata to CodingKeys.
  • Wire serialization maps directly to the backend schema expected by Vertex AI and Google AI Gemini endpoints:
    {
      "text": "Hello world",
      "speechMetadata": {
        "speaker": "Speaker1",
        "style": "calm and welcoming"
      }
    }
  • Nil fields are omitted during encoding to maintain compact payloads.
  • InternalPart conversions in ModelContent.swift convert between public TextPart and InternalPart seamlessly for both inbound responses and outbound requests.
Deep Dive: Streaming History Aggregation (History.swift)
  • In streaming content generation (generateContentStream), the SDK aggregates incremental response chunks into unified turns in History.
  • When aggregating consecutive parts in aggregatedChunks:
    • If two consecutive chunks are both TextParts sharing the exact same speechMetadata and thought classification, their text contents are concatenated.
    • If consecutive chunks have differing speechMetadata (such as a speaker transition in multi-speaker responses or distinct style directives), the chunks are preserved as distinct parts rather than improperly merged.
Testing Strategy
  • Unit Tests (FirebaseAILogicUnit):
    • PartTests: Verified TextPart initialization, equatability, and speechMetadata property propagation.
    • InternalPartTests: Tested JSON encoding and decoding of InternalPart with full, partial, and empty SpeechMetadata structures, ensuring round-trip fidelity and correct omission of empty keys.
    • ChatTests: Verified streaming chunk combination logic in History when aggregating chunks with matching, differing, and absent speech metadata.
    • All unit tests passed cleanly (swift test --test-product FirebaseAILogicUnit exits with 0 failures).
  • Integration Tests (TextToSpeechTests):
    • Replaced legacy gemini-3.1-flash-tts-preview references with gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.
    • Removed legacy TTS integration tests from GenerateContentIntegrationTests.swift and introduced a dedicated TextToSpeechTests integration test suite.
    • Tested single-speaker unary and streaming generation with SpeechMetadata(style:).
    • Tested single-speaker unary and streaming generation without metadata across languages (en-GB, fr-CA).
    • Tested multi-speaker unary and streaming dialogue generation with SpeechMetadata(speaker:style:) across multiple configured voices.
    • Tested error handling when supplying an invalid number of speakers in MultiSpeakerVoiceConfig.
    • Confirmed that multiSpeaker tests targeting InstanceConfig.enterprise_v1beta_global run and validate audio output payloads (audio/wav for unary, audio/l16 for streaming).

Add `SpeechMetadata` to `TextPart` to support turn- and part-level
speech configurations (such as speaker and style) in Gemini TTS models.

- Plumb `speechMetadata` through `InternalPart` and `ModelContent` with
  wire-format Codable serialization using the `"speechMetadata"` key.
- Update `History.aggregatedChunks` to avoid combining streaming text
  chunks that carry speech metadata.
- Add unit tests for encoding and decoding `TextPart` and `InternalPart`
  with full and partial speech metadata.
@gemini-code-assist

Copy link
Copy Markdown
Contributor
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

@danger-firebase-ios

Copy link
Copy Markdown
1 Warning
⚠️ New public headers were added, did you remember to add them to the umbrella header?

Generated by 🚫 Danger

Add detailed DocC documentation for SpeechMetadata, TextPart additions,
and existing speech configuration types for Gemini text-to-speech (TTS).

Key documentation updates:
- Document SpeechMetadata and TextPart.speechMetadata, explaining the
  verbatim transcript nature of input text and turn-level control.
- Detail the distinction between sustained turn-level delivery style
  and momentary inline audio tags (such as <laugh> and <short pause>).
- Call out that speaker is mandatory on every TextPart when generating
  multi-speaker speech.
- Update SpeechConfig, MultiSpeakerVoiceConfig, and SpeakerVoiceConfig
  with canonical Firebase documentation links and Public Preview
  callouts.
@andrewheard andrewheard linked an issue Oct 1, 2026 that may be closed by this pull request
Add integration tests for Gemini 3.8 Flash TTS and Flash-Lite TTS models
in the Firebase AI test app.

- Update `ModelNames` constants to replace
  `gemini-3.1-flash-tts-preview` with `gemini-3.8-flash-tts` and
  `gemini-3.8-flash-lite-tts`.
- Remove legacy 3.1 TTS integration tests from
  `GenerateContentIntegrationTests`.
- Add a dedicated `TextToSpeechTests` suite covering unary and streaming
  synthesis for single-speaker and multi-speaker configurations, with
  and without turn-level `SpeechMetadata`.
- Verify error handling for invalid multi-speaker voice counts.
Document support for turn-level speech metadata (`SpeechMetadata`) on
`TextPart` for Gemini text-to-speech (TTS) models in Firebase AI Logic.
Address readability and safety feedback from PR review.

- Conform `SpeechMetadata` to `Hashable`.
- Restore `TextPart.init(_:)` to maintain source compatibility for
  unapplied function references (such as `strings.map(TextPart.init)`).
- Remove default `nil` arguments for `isThought` and `thoughtSignature`
  on internal `InternalPart` and `TextPart` initializers to prevent
  silent omissions at conversion call sites.
- Add unit test verifying adjacent `speechMetadata` chunks remain
  discrete and unmerged during `History` aggregation.
- Tighten streaming assertions in `TextToSpeechTests` to verify
  non-audio chunks finish cleanly, remove leftover debug print, and
  modernize error handling test to use `#require(throws:)`.
# Conflicts:
#	FirebaseAI/CHANGELOG.md
#	FirebaseAI/Tests/Unit/ChatTests.swift
#	FirebaseAI/Tests/Unit/Types/InternalPartTests.swift
Run multi-speaker text-to-speech tests against default instance
configurations following backend rollout completion.

- Update `multiSpeaker_withTurnMetadata` and
  `multiSpeakerStream_withTurnMetadata` in `TextToSpeechTests` to use
  `InstanceConfig.defaultConfigs` instead of restricting to
  `enterprise_v1beta_global`.
- Remove temporary TODO comments for backend rollout (#16775).
- Wrap doc comment line in `TextPart.init(_:)`.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[FR]: Add SpeechMetadata support in Firebase AI

1 participant