feat(ai): support speech metadata on text parts - #16749
andrewheard wants to merge 9 commits into
Conversation
Add `SpeechMetadata` to `TextPart` to support turn- and part-level speech configurations (such as speaker and style) in Gemini TTS models. - Plumb `speechMetadata` through `InternalPart` and `ModelContent` with wire-format Codable serialization using the `"speechMetadata"` key. - Update `History.aggregatedChunks` to avoid combining streaming text chunks that carry speech metadata. - Add unit tests for encoding and decoding `TextPart` and `InternalPart` with full and partial speech metadata.
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. |
Generated by 🚫 Danger |
Add detailed DocC documentation for SpeechMetadata, TextPart additions, and existing speech configuration types for Gemini text-to-speech (TTS). Key documentation updates: - Document SpeechMetadata and TextPart.speechMetadata, explaining the verbatim transcript nature of input text and turn-level control. - Detail the distinction between sustained turn-level delivery style and momentary inline audio tags (such as <laugh> and <short pause>). - Call out that speaker is mandatory on every TextPart when generating multi-speaker speech. - Update SpeechConfig, MultiSpeakerVoiceConfig, and SpeakerVoiceConfig with canonical Firebase documentation links and Public Preview callouts.
Add integration tests for Gemini 3.8 Flash TTS and Flash-Lite TTS models in the Firebase AI test app. - Update `ModelNames` constants to replace `gemini-3.1-flash-tts-preview` with `gemini-3.8-flash-tts` and `gemini-3.8-flash-lite-tts`. - Remove legacy 3.1 TTS integration tests from `GenerateContentIntegrationTests`. - Add a dedicated `TextToSpeechTests` suite covering unary and streaming synthesis for single-speaker and multi-speaker configurations, with and without turn-level `SpeechMetadata`. - Verify error handling for invalid multi-speaker voice counts.
Document support for turn-level speech metadata (`SpeechMetadata`) on `TextPart` for Gemini text-to-speech (TTS) models in Firebase AI Logic.
Address readability and safety feedback from PR review. - Conform `SpeechMetadata` to `Hashable`. - Restore `TextPart.init(_:)` to maintain source compatibility for unapplied function references (such as `strings.map(TextPart.init)`). - Remove default `nil` arguments for `isThought` and `thoughtSignature` on internal `InternalPart` and `TextPart` initializers to prevent silent omissions at conversion call sites. - Add unit test verifying adjacent `speechMetadata` chunks remain discrete and unmerged during `History` aggregation. - Tighten streaming assertions in `TextToSpeechTests` to verify non-audio chunks finish cleanly, remove leftover debug print, and modernize error handling test to use `#require(throws:)`.
# Conflicts: # FirebaseAI/CHANGELOG.md # FirebaseAI/Tests/Unit/ChatTests.swift # FirebaseAI/Tests/Unit/Types/InternalPartTests.swift
Run multi-speaker text-to-speech tests against default instance configurations following backend rollout completion. - Update `multiSpeaker_withTurnMetadata` and `multiSpeakerStream_withTurnMetadata` in `TextToSpeechTests` to use `InstanceConfig.defaultConfigs` instead of restricting to `enterprise_v1beta_global`. - Remove temporary TODO comments for backend rollout (#16775). - Wrap doc comment line in `TextPart.init(_:)`.
Adds support for turn-level speech metadata (
SpeechMetadata) onTextPartfor Gemini text-to-speech (TTS) models in theFirebaseAILogicSwift SDK, enabling multi-speaker dialogue routing and sustained delivery styling. See the speech generation guide for more details.Developer-Facing Changes
SpeechMetadata[Public Preview]: Added theSpeechMetadatastruct to configure turn-level speech synthesis options for individual dialogue turns. Supports assigning turns to specific speakers (speaker) during multi-speaker generation and specifying sustained vocal delivery guidance (style), such as emotion, tempo, or tone.TextPartIntegration: UpdatedTextPartto optionally store and exposespeechMetadata: SpeechMetadata?, with an updated public initializerinit(_:speechMetadata:).SpeechConfig(multiSpeakerVoiceConfig:). EachTextPartin a multi-speaker request can be attributed to a configured voice profile usingSpeechMetadata(speaker:).TextPart.textis treated strictly as a verbatim transcript to speak aloud; passing styling instructions viaSpeechMetadata.styleprevents the model from reading directorial guidance aloud as speech.SpeechMetadata,TextPart,SpeechConfig,MultiSpeakerVoiceConfig, andSpeakerVoiceConfig, linking out to canonical Firebase guides and detailing the operational boundary between sustained turn-levelstyleand point-in-time inline audio tags (such as<laugh>,<sigh>, and<short pause>).Usage Examples
The following examples demonstrate how to use each feature:
1. Single-Speaker Synthesis with Sustained Style Instructions
2. Multi-Speaker Dialogue Turn Routing
3. Streaming Speech Synthesis with Turn Metadata
4. Inline Vocal Tags vs. Sustained Style Guidance
Technical Details
This pull request introduces turn-level speech metadata to support the Gemini 3.8 Flash TTS and Flash-Lite TTS model families while keeping the public API clean, type-safe, and decoupled from backend protocol models.
Key Architectural Points:
TextPart: While other content parts (such as inline binary data or function calls) do not require speech metadata,TextPartrepresents speech transcripts and now encapsulates optional turn-level metadata via itsspeechMetadataproperty.SpeechMetadata.speakerandSpeechMetadata.styleare stored internally and initialized viainit(speaker:style:). This enforces clean immutability while allowing future extension without breaking changes.InternalPartmapsspeechMetadatato and from JSON payloads ({"speaker": "...", "style": "..."}) conforming to the Gemini API specifications, keeping wire serialization logic completely isolated from public API types.History.swiftso that adjacent text chunks produced during streaming responses are merged intelligently based on theirspeechMetadatavalues without concatenating chunks across differing speakers or styles.Deep Dive: Type Architecture & Encapsulation (
SpeechMetadata&TextPart)SpeechMetadataas a public struct conforming toEquatable,Hashable, andSendable.**[Public Preview]**to signal its preview status alongside Gemini 3.8 TTS capabilities.speakerandstyleproperties internally to maintain tight encapsulation while providing a comprehensive public initializer with defaultnilarguments.TextPartwithpublic let speechMetadata: SpeechMetadata?and provided an overloaded initializerpublic init(_ text: String, speechMetadata: SpeechMetadata? = nil). Existing call sites passing only a string literal or single argument compile without modification.Deep Dive: Wire Protocol & Serialization (
InternalPart)speechMetadata: SpeechMetadata?toInternalPartand addedspeechMetadatatoCodingKeys.{ "text": "Hello world", "speechMetadata": { "speaker": "Speaker1", "style": "calm and welcoming" } }InternalPartconversions inModelContent.swiftconvert between publicTextPartandInternalPartseamlessly for both inbound responses and outbound requests.Deep Dive: Streaming History Aggregation (
History.swift)generateContentStream), the SDK aggregates incremental response chunks into unified turns inHistory.aggregatedChunks:TextParts sharing the exact samespeechMetadataand thought classification, their text contents are concatenated.speechMetadata(such as a speaker transition in multi-speaker responses or distinct style directives), the chunks are preserved as distinct parts rather than improperly merged.Testing Strategy
FirebaseAILogicUnit):PartTests: VerifiedTextPartinitialization, equatability, andspeechMetadataproperty propagation.InternalPartTests: Tested JSON encoding and decoding ofInternalPartwith full, partial, and emptySpeechMetadatastructures, ensuring round-trip fidelity and correct omission of empty keys.ChatTests: Verified streaming chunk combination logic inHistorywhen aggregating chunks with matching, differing, and absent speech metadata.swift test --test-product FirebaseAILogicUnitexits with 0 failures).TextToSpeechTests):gemini-3.1-flash-tts-previewreferences withgemini-3.8-flash-ttsandgemini-3.8-flash-lite-tts.GenerateContentIntegrationTests.swiftand introduced a dedicatedTextToSpeechTestsintegration test suite.SpeechMetadata(style:).en-GB,fr-CA).SpeechMetadata(speaker:style:)across multiple configured voices.MultiSpeakerVoiceConfig.multiSpeakertests targetingInstanceConfig.enterprise_v1beta_globalrun and validate audio output payloads (audio/wavfor unary,audio/l16for streaming).