diff --git a/CONTEXT.md b/CONTEXT.md index 6a5eabc..8741eb1 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -59,6 +59,10 @@ _Avoid_: Bundled model, Talkify model, Whisper model An opt-in pass that rewrites finished **Direct Dictation** text through the **On-device model** before insertion, using the session's **Shaping Prompt**. _Avoid_: AI cleanup, post-processing, autocorrect +**Spelling replacement**: +One user-authored from→to pair, applied as a whole-word swap after recognition and before **Prompt Shaping**. The left side may be a phrase. The list is empty by default. +_Avoid_: Autocorrect, vocabulary, dictionary, lexicon + **Shaping Prompt**: One named entry in the user's editable library, holding the wording that tells the **On-device model** what to do with the words. The framing that keeps it rewriting rather than answering is not part of it and is never editable. _Avoid_: System prompt, preset, template @@ -150,7 +154,7 @@ _Avoid_: Transcript history, cloud analytics - Settings changes apply to the live app and persist immediately; Settings has no Save step - The Appearance preview uses the same preferences and HUD surface as Direct Dictation - Settings changes update the Appearance preview immediately, while an active Direct Dictation session keeps the choices captured at session start -- Dictation session settings include the voice visual, waveform style, glow palette, glow center, reveal style, long-draft behavior, HUD size, sound set, sound enabled state, sound volume, insertion destination, the transcription history choice, the shaping choice, and whether the session lowers other audio +- Dictation session settings include the voice visual, waveform style, glow palette, glow center, reveal style, long-draft behavior, HUD size, sound set, sound enabled state, sound volume, insertion destination, the transcription history choice, the shaping choice, the **Spelling replacement** list, and whether the session lowers other audio - **Direct Dictation** can lower other audio while it listens and put it back when the session ends, cancels or fails; it is off by default because the control it moves is system-wide - macOS has no per-application ducking, so lowering other audio moves the default output device's own volume and quiets Talkify's session sounds along with everything else - A lowered volume is restored only while it is still the value Talkify set: a volume the user changed mid-session is theirs, the same rule the clipboard restore follows @@ -199,7 +203,15 @@ _Avoid_: Transcript history, cloud analytics - Settings preserves its selected section and window frame while the app runs, and opens on Appearance after a fresh launch - The voice-reactive visual must make silence and a dead microphone look different - With Reduce Motion enabled, the HUD replaces the animated visual with a quiet level meter and skips expand/collapse animation -- A session inserts raw finalized text unless the user turned **Prompt Shaping** on: no filler-word list, no autocorrect, and no rewriting anybody did not ask for +- A session inserts raw finalized text unless the user turned **Prompt Shaping** on or added a **Spelling replacement**: no filler-word list, no autocorrect, and no rewriting anybody did not ask for +- A **Spelling replacement** is a whole-word swap the user typed: the misspelling Apple Speech produced, and the spelling that should land instead +- **Spelling replacement** has its own Settings section; an empty list is exactly today's behavior +- Matching ignores case, surrounding spaces are trimmed, and a possessive keeps its 's, so Calman, calman, and Calman's all become Kalman / Kalman's +- The misspelling can be several words or a hyphenated one, because splitting a name is how Apple Speech usually gets it wrong: "ex code" and "e-mail" are as common as a single wrong token +- Every pair is measured against what was said rather than against another pair's output, and where two pairs both fit, the longer misspelling wins, so "git hub" becomes GitHub rather than Git hub +- A pair with a blank from or to is skipped, so a row still being filled in does not rewrite anything +- **Spelling replacement** runs after recognition and before **Prompt Shaping**, on the live draft and on the inserted text, and on a finished **Drop Transcription** +- Prompt Shaping sees the replaced spelling, and is skipped if that spelling is empty - **Prompt Shaping** is the text cleanup the roadmap deferred, shipped as an explicit beta and off by default, so the default session still inserts exactly what was said - While **Prompt Shaping** is on, the session's **Shaping Prompt** rewrites finished text through the **On-device model** between recognition and insertion, and nothing leaves the Mac: the model is Apple's, it runs here, and Talkify makes no request on its behalf - Any prompt shaping unavailability, failure, or slow answer inserts the raw words unchanged @@ -208,7 +220,7 @@ _Avoid_: Transcript history, cloud analytics - Deleting the selected shaping prompt, or restoring the defaults over it, moves the selection to the first prompt in the library - A selected shaping prompt that still resolves to nothing inserts the raw words unchanged - A shaping prompt can be run against a sample sentence in Settings, through the same service a session uses, so a prompt is written by reading what it does rather than by guessing -- Transcription history keeps the words as spoken; a translated session's second line is what actually landed, so it is written after shaping and before insertion +- Transcription history keeps the words as recognized, before **Spelling replacement** or **Prompt Shaping**; a translated session's second line is what actually landed, so it is written after shaping and before insertion - Shaping runs before translation: a prompt carries its own language and its one-shot example in that language, and a translator handed cleaned-up words has less to get wrong - The shaping choice is captured in the Dictation session settings snapshot, along with the whole prompt library while shaping is on - A session that will shape keeps the HUD up through the shaping phase, naming the prompt in the same caption the pick rode in, and the HUD leaves early when the answer lands diff --git a/README.md b/README.md index 617f5b1..ffd72f3 100644 --- a/README.md +++ b/README.md @@ -85,6 +85,7 @@ xcodebuild test -project Talkify.xcodeproj -scheme Talkify -destination 'platfor | Transcribe a file | Drag audio or video at the notch and drop it | | Cancel mid-session | **Esc** | | Shape what you dictate | Turn it on in **Settings → Prompt Shaping** | +| Fix a misspelled name | Add a pair in **Settings → Spelling replacements** | | Pick the shaping prompt mid-session | **←** / **→** while dictating, with shaping on | | Read selected text aloud | **⌥ ⎋** (toggles; also in the menu) | | Read it aloud translated | Same key, with **Translate before speaking** on | diff --git a/Talkify/App/AppSettings.swift b/Talkify/App/AppSettings.swift index 4139e61..c72b887 100644 --- a/Talkify/App/AppSettings.swift +++ b/Talkify/App/AppSettings.swift @@ -42,6 +42,7 @@ final class AppSettings { static let promptShapingEnabled = "dictationPromptShapingEnabled" static let promptShapingPrompt = "dictationPromptShapingPrompt" static let shapingPrompts = "dictationShapingPrompts" + static let spellingReplacements = "dictationSpellingReplacements" } @ObservationIgnored @@ -137,6 +138,16 @@ final class AppSettings { shapingPrompts = ShapingPrompt.defaults } + /// User-authored whole-word swaps applied after recognition. Empty is + /// today's behavior: the session inserts what the Speech Model produced. + var spellingReplacements: [SpellingReplacement] { + didSet { + if let data = try? JSONEncoder().encode(spellingReplacements) { + defaults.set(data, forKey: Keys.spellingReplacements) + } + } + } + var voiceVisual: HUDVoiceVisualStyle { didSet { defaults.set(voiceVisual.rawValue, forKey: Keys.voiceVisual) } } @@ -310,6 +321,7 @@ final class AppSettings { promptShapingPromptID = defaults.string(forKey: Keys.promptShapingPrompt) ?? ShapingPrompt.defaults[0].id shapingPrompts = Self.storedShapingPrompts(in: defaults) ?? ShapingPrompt.defaults + spellingReplacements = Self.storedSpellingReplacements(in: defaults) voiceVisual = Self.stored(in: defaults, key: Keys.voiceVisual) ?? .waveform waveformStyle = Self.stored(in: defaults, key: Keys.waveformStyle) ?? .chartLine revealStyle = Self.stored(in: defaults, key: Keys.revealStyle) ?? .slide @@ -378,6 +390,17 @@ final class AppSettings { return prompts } + /// A missing or unreadable value is an empty list, which is the same as + /// never having had the feature: nothing is rewritten. + private static func storedSpellingReplacements( + in defaults: UserDefaults + ) -> [SpellingReplacement] { + guard let data = defaults.data(forKey: Keys.spellingReplacements), + let pairs = try? JSONDecoder().decode([SpellingReplacement].self, from: data) + else { return [] } + return pairs + } + private static func store(_ binding: KeyBinding, in defaults: UserDefaults, key: String) { if let data = try? JSONEncoder().encode(binding) { defaults.set(data, forKey: key) @@ -461,6 +484,9 @@ struct DictationSessionSettings: Equatable { /// the arrow keys can cycle the session's pick without reading a library /// that may change mid-session. let shapingLibrary: [ShapingPrompt] + /// Whole-word swaps captured with everything else, so editing the list + /// mid-session cannot rewrite a phrase already on its way to insertion. + let spellingReplacements: [SpellingReplacement] let voiceVisual: HUDVoiceVisualStyle let waveformStyle: HUDWaveformStyle let revealStyle: HUDRevealStyle @@ -494,6 +520,7 @@ struct DictationSessionSettings: Equatable { ? settings.shapingPrompts.prompt(for: settings.promptShapingPromptID) : nil shapingLibrary = settings.promptShapingEnabled ? settings.shapingPrompts : [] + spellingReplacements = settings.spellingReplacements voiceVisual = settings.voiceVisual waveformStyle = settings.waveformStyle revealStyle = settings.revealStyle diff --git a/Talkify/Dictation/DirectDictationController.swift b/Talkify/Dictation/DirectDictationController.swift index ecc0322..65f8e6f 100644 --- a/Talkify/Dictation/DirectDictationController.swift +++ b/Talkify/Dictation/DirectDictationController.swift @@ -676,8 +676,18 @@ final class DirectDictationController { let hasVisibleText = !displayText .trimmingCharacters(in: .whitespacesAndNewlines) .isEmpty - pendingLiveText = update.finalizedText - pendingVolatileText = update.volatileText + // Keep the committed/volatile split the HUD uses, and rewrite each + // half so a pair like Hetty→Hedy shows in the live draft, not only + // in the paste. Insertion still applies to the full spoken string. + let replacements = currentSessionSettings?.spellingReplacements ?? [] + pendingLiveText = SpellingReplacements.apply( + update.finalizedText, + using: replacements + ) + pendingVolatileText = SpellingReplacements.apply( + update.volatileText, + using: replacements + ) send(.updateReceived(hasVisibleText: hasVisibleText)) pendingLiveText = nil pendingVolatileText = "" @@ -737,27 +747,38 @@ final class DirectDictationController { defer { finishTask = nil } do { let spoken = try await dependencies.finishRecognition() + // Delivery follows the session snapshot, so a Settings change + // mid-session applies to the next session (ADR-0004). + let session = currentSessionSettings ?? settings.sessionSettings + // Held separately from `text`, which shaping and translation go on to + // overwrite: this is the spelling the user asked for, and it is what + // a failed translation has to fall back to. + let replaced = SpellingReplacements.apply( + spoken, + using: session.spellingReplacements + ) + var text = replaced + // Shape the words that will be inserted, not the recognizer's + // misspelling: an empty replacement must not send a blank + // transcript to the On-device model, which can answer with + // words nobody spoke. + let willShape = chosenPrompt != nil + && !text.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty // A session about to shape keeps the HUD up saying so; every other // session dismisses here exactly as before. - let willShape = chosenPrompt != nil && !spoken.isEmpty if let chosenPrompt, willShape { dependencies.showShaping(chosenPrompt.name) } else { dependencies.hideHUD() } - // Delivery follows the session snapshot, so a Settings change - // mid-session applies to the next session (ADR-0004). - let session = currentSessionSettings ?? settings.sessionSettings - // The one place between recognition and insertion where the words may - // change. A transform that fails delivers nothing: the trigger - // promised a translation, and pasting the untranslated words instead - // lands the wrong language in someone else's document. - var text = spoken - // Shaping first, then translation. A prompt is written in one language, - // with a one-shot example in it, so handing it a translation of the - // words asks it to work in a language it was not written for; and a - // translator given cleaned-up words has less to get wrong. + // Replacements, then shaping, then translation. The list is a + // spelling fix for what the Speech Model produced, so shaping and + // translation both see the name the user wrote. A prompt is written + // in one language, with a one-shot example in it, so handing it a + // translation of the words asks it to work in a language it was not + // written for; and a translator given cleaned-up words has less to + // get wrong. if let chosenPrompt, willShape { text = await dependencies.shapeText(text, chosenPrompt) } @@ -768,12 +789,13 @@ final class DirectDictationController { do { text = try await translation.translate(text, with: pair) } catch { - // The rescue is the words as spoken, not as shaped: shaping is a - // convenience and the raw words are what must survive. + // The rescue drops shaping, which is a convenience, and keeps the + // replacements, which are not: the user typed that spelling and it + // is the one they want wherever the words land. await recordHistory(spoken: spoken, delivered: nil, session: session) // Clipboard-only whatever the session's destination: nothing is // pasted, and the words survive where the user can reach them. - let rescue = await dependencies.insertText(spoken, nil, .clipboardOnly) + let rescue = await dependencies.insertText(replaced, nil, .clipboardOnly) // Ended, not failed: a failure action would drive the machine to // cancelling and cancel a session that has already finished. send(.sessionEnded) diff --git a/Talkify/Dictation/SpellingReplacement.swift b/Talkify/Dictation/SpellingReplacement.swift new file mode 100644 index 0000000..9ea29c5 --- /dev/null +++ b/Talkify/Dictation/SpellingReplacement.swift @@ -0,0 +1,118 @@ +import Foundation + +/// One user-authored from→to pair, applied as a whole-word swap after +/// recognition. The Speech Model cannot be taught a word it does not +/// know; this is the list that fixes what it keeps misspelling, and +/// only the words the user typed. +/// +/// Whole-word on purpose: a substring swap has no edge, and a list +/// holding "mm" would eat it inside "comment". An incomplete pair — +/// blank `from` or blank `to` — is skipped so a row being typed does +/// not rewrite anything yet. +struct SpellingReplacement: Codable, Equatable, Identifiable, Sendable { + var id: String + var from: String + var to: String +} + +enum SpellingReplacements { + /// Rewrites `text` in one left-to-right pass. Each position offers every + /// pair the same chance and the longest `from` wins, so a pair for + /// "git hub" beats one for "git" wherever both fit; equal lengths go + /// to list order. One pass on purpose: replacing into the running + /// result would let a later pair rewrite an earlier pair's output, and + /// the Settings list has no reordering for the user to fix it with. + /// + /// Matching is case-insensitive; the replacement is the trimmed `to` + /// string, so "calman" and "Calman" both become "Kalman" if that is what + /// the user wrote. A `from` may hold spaces or hyphens — "ex code" and + /// "e-mail" are what the recognizer produces as often as a single token + /// is. + static func apply(_ text: String, using pairs: [SpellingReplacement]) -> String { + let active = pairs.compactMap { pair -> Pair? in + let from = pair.from.trimmingCharacters(in: .whitespacesAndNewlines) + let to = pair.to.trimmingCharacters(in: .whitespacesAndNewlines) + guard !from.isEmpty, !to.isEmpty else { return nil } + return Pair(from: from, to: to, length: from.count) + } + guard !active.isEmpty, !text.isEmpty else { return text } + + var result = "" + result.reserveCapacity(text.count) + var index = text.startIndex + var previous: Character? + while index < text.endIndex { + // A match can only open where a word does, which is also what keeps + // "mm" out of the middle of "comment". + if !isWordCharacter(previous), + let match = longestMatch(in: text, at: index, using: active) { + result += match.to + // The boundary the next position sees is the text's, not the + // replacement's: what was matched stays matched. + previous = text[text.index(before: match.end)] + index = match.end + } else { + previous = text[index] + result.append(text[index]) + index = text.index(after: index) + } + } + return result + } + + private struct Pair { + let from: String + let to: String + let length: Int + } + + private struct Match { + let to: String + let end: String.Index + } + + private static func longestMatch( + in text: String, + at index: String.Index, + using pairs: [Pair] + ) -> Match? { + var best: Match? + var bestLength = 0 + for pair in pairs where pair.length > bestLength { + guard + let end = text.index(index, offsetBy: pair.length, limitedBy: text.endIndex), + text.compare(pair.from, options: .caseInsensitive, range: index.. = ["'", "\u{2019}"] + + /// Letters, digits, and the apostrophe, so "don't" is one word and a + /// pair for "don" cannot cut it in half. + private static func isWordCharacter(_ character: Character?) -> Bool { + guard let character else { return false } + return character.isLetter || character.isNumber || apostrophes.contains(character) + } + + private static func endsWord(_ text: String, at end: String.Index) -> Bool { + guard end < text.endIndex, isWordCharacter(text[end]) else { return true } + // A possessive stays attached to the name, because English uses that + // form for a product and the recognizer misspells it the same way. The + // 's is left where it is, so it rides the new spelling untouched. + guard apostrophes.contains(text[end]) else { return false } + let afterMark = text.index(after: end) + guard afterMark < text.endIndex, text[afterMark].lowercased() == "s" else { + return false + } + let afterS = text.index(after: afterMark) + return afterS == text.endIndex || !isWordCharacter(text[afterS]) + } +} diff --git a/Talkify/DropTranscription/DropTranscriptionController.swift b/Talkify/DropTranscription/DropTranscriptionController.swift index 0ded059..ce5d83d 100644 --- a/Talkify/DropTranscription/DropTranscriptionController.swift +++ b/Talkify/DropTranscription/DropTranscriptionController.swift @@ -187,6 +187,7 @@ final class DropTranscriptionController { // earned; a second file takes the shape, so the first goes to disk. offer.commit() let localeIdentifier = settings.localeIdentifierForDrop(languageIndex: languageIndex) + let replacements = settings.spellingReplacements // Nothing in the shape while a job runs. The drop target retracts and the // menu bar ghost carries progress from there. @@ -210,17 +211,26 @@ final class DropTranscriptionController { ) try Task.checkCancellation() + // Captured at job start so editing the list while a file runs cannot + // rewrite a transcript that was already recognized under the old one. + // Rebuilt rather than counted here, so the card keeps counting words + // the one way the service does. + let replaced = FileTranscriptionService.Transcript( + text: SpellingReplacements.apply(transcript.text, using: replacements), + duration: transcript.duration + ) + // Staged, not saved: the card is where the user picks a destination, // and the Save-to location is what happens if they pick none // (CONTEXT.md). The file is real from here on either way. - let staged = try StagedTranscript.stage(text: transcript.text, source: url) + let staged = try StagedTranscript.stage(text: replaced.text, source: url) offer.present( staged, as: DropHUDContent.Transcript( url: staged.url, - text: transcript.text, - wordCount: transcript.wordCount, - duration: transcript.duration + text: replaced.text, + wordCount: replaced.wordCount, + duration: replaced.duration ) ) } catch is CancellationError { diff --git a/Talkify/Settings/Sections/ReplacementsSettingsView.swift b/Talkify/Settings/Sections/ReplacementsSettingsView.swift new file mode 100644 index 0000000..4430eb1 --- /dev/null +++ b/Talkify/Settings/Sections/ReplacementsSettingsView.swift @@ -0,0 +1,75 @@ +import SwiftUI + +/// The Spelling replacements section: whole-word swaps after recognition. +/// +/// Apple Speech cannot be taught a word it does not know. These pairs fix +/// the spelling it produces, and only when the user wrote both sides. +struct ReplacementsSettingsView: View { + @Bindable var settings: AppSettings + + @Environment(\.colorSchemeContrast) private var contrast + + var body: some View { + VStack(alignment: .leading, spacing: 18) { + SettingsCard(title: "Words") { + ForEach($settings.spellingReplacements) { $pair in + replacementRow($pair) + } + + SettingsRow( + title: "Add a replacement", + description: settings.spellingReplacements.isEmpty + ? "Nothing is rewritten until you add a pair." + : "\(settings.spellingReplacements.count) pair" + + (settings.spellingReplacements.count == 1 ? "" : "s") + ) { + Button("Add") { add() } + .buttonStyle(SettingsButtonStyle()) + } + } + + Text( + "After recognition, each complete pair swaps that whole word or " + + "phrase for the spelling you typed. Matching ignores case, a " + + "possessive keeps its 's, and nothing leaves this Mac." + ) + .font(.caption) + .foregroundStyle(.white.opacity(contrast == .increased ? 0.7 : 0.45)) + .fixedSize(horizontal: false, vertical: true) + .padding(.horizontal, 6) + } + } + + private func replacementRow(_ pair: Binding) -> some View { + HStack(spacing: 10) { + TextField("Misspelled word or phrase", text: pair.from) + .textFieldStyle(.roundedBorder) + .frame(minWidth: 110) + Image(systemName: "arrow.right") + .font(.system(size: 11, weight: .semibold)) + .foregroundStyle(.white.opacity(contrast == .increased ? 0.7 : 0.35)) + TextField("Correct", text: pair.to) + .textFieldStyle(.roundedBorder) + .frame(minWidth: 110) + Spacer(minLength: 8) + Button("Remove") { remove(pair.wrappedValue.id) } + .buttonStyle(SettingsButtonStyle()) + } + .padding(.vertical, 13) + .overlay(alignment: .bottom) { + Rectangle() + .fill(.white.opacity(contrast == .increased ? 0.16 : 0.07)) + .frame(height: 1) + } + } + + private func add() { + settings.spellingReplacements.append( + SpellingReplacement(id: UUID().uuidString, from: "", to: "") + ) + } + + private func remove(_ id: String) { + settings.spellingReplacements.removeAll { $0.id == id } + } +} diff --git a/Talkify/Settings/SettingsSections.swift b/Talkify/Settings/SettingsSections.swift index 8df9ed7..f3326b3 100644 --- a/Talkify/Settings/SettingsSections.swift +++ b/Talkify/Settings/SettingsSections.swift @@ -18,6 +18,7 @@ enum SettingsSection: String, CaseIterable, Identifiable { case sounds case dictation case promptShaping + case replacements case dropTranscription case readAloud case language @@ -35,6 +36,7 @@ enum SettingsSection: String, CaseIterable, Identifiable { case .sounds: "Sounds" case .dictation: "Dictation" case .promptShaping: "Prompt Shaping" + case .replacements: "Spelling replacements" case .dropTranscription: "Drop Transcription" case .readAloud: "Read Aloud" case .language: "Language" @@ -51,6 +53,7 @@ enum SettingsSection: String, CaseIterable, Identifiable { case .sounds: "Choose and preview the session sounds" case .dictation: "Choose where finished dictation text goes" case .promptShaping: "Rewrite what you dictate, on device" + case .replacements: "Fix words Apple Speech keeps misspelling" case .dropTranscription: "Transcribe audio and video files" case .readAloud: "Choose the voice that reads selected text" case .language: "Choose your languages" @@ -67,6 +70,7 @@ enum SettingsSection: String, CaseIterable, Identifiable { case .sounds: "waveform" case .dictation: "text.cursor" case .promptShaping: "wand.and.sparkles" + case .replacements: "arrow.left.arrow.right" case .dropTranscription: "square.and.arrow.down" case .readAloud: "speaker.wave.2" case .language: "globe" diff --git a/Talkify/Settings/SettingsView.swift b/Talkify/Settings/SettingsView.swift index 0837fe8..00996db 100644 --- a/Talkify/Settings/SettingsView.swift +++ b/Talkify/Settings/SettingsView.swift @@ -257,6 +257,8 @@ private struct SettingsContent: View { DictationSettingsView(settings: settings) case .promptShaping: PromptShapingSettingsView(settings: settings) + case .replacements: + ReplacementsSettingsView(settings: settings) case .dropTranscription: DropTranscriptionSettingsView(settings: settings) case .readAloud: diff --git a/TalkifyTests/AppSettingsTests.swift b/TalkifyTests/AppSettingsTests.swift index 52d636c..eab6dff 100644 --- a/TalkifyTests/AppSettingsTests.swift +++ b/TalkifyTests/AppSettingsTests.swift @@ -25,6 +25,7 @@ struct AppSettingsTests { #expect(settings.readAloudVoiceID.isEmpty) #expect(settings.dictationTriggerBinding == .fnTrigger) #expect(settings.readAloudBinding == .optionEscape) + #expect(settings.spellingReplacements.isEmpty) } @Test func everyPreferenceRoundTrips() { @@ -466,6 +467,39 @@ struct AppSettingsTests { #expect(settings.promptShapingPromptID == "mine") } + @Test func spellingReplacementsPersistAndReadBack() { + let defaults = freshDefaults() + let settings = AppSettings(defaults: defaults) + settings.spellingReplacements = [ + SpellingReplacement(id: "1", from: "ex code", to: "Xcode"), + ] + + let reloaded = AppSettings(defaults: defaults) + #expect(reloaded.spellingReplacements == settings.spellingReplacements) + #expect(defaults.data(forKey: "dictationSpellingReplacements") != nil) + } + + @Test func undecodableStoredSpellingReplacementsBecomeAnEmptyList() { + let defaults = freshDefaults() + defaults.set(Data("not json".utf8), forKey: "dictationSpellingReplacements") + + let settings = AppSettings(defaults: defaults) + #expect(settings.spellingReplacements.isEmpty) + } + + @Test func sessionSnapshotCapturesSpellingReplacements() { + let settings = AppSettings(defaults: freshDefaults()) + settings.spellingReplacements = [ + SpellingReplacement(id: "1", from: "ex code", to: "Xcode"), + ] + let snapshot = settings.sessionSettings + settings.spellingReplacements = [] + #expect(snapshot.spellingReplacements == [ + SpellingReplacement(id: "1", from: "ex code", to: "Xcode"), + ]) + #expect(settings.sessionSettings.spellingReplacements.isEmpty) + } + @Test func promptShapingRoundTripsUnderItsKeys() { let defaults = freshDefaults() let settings = AppSettings(defaults: defaults) diff --git a/TalkifyTests/DirectDictationControllerTests.swift b/TalkifyTests/DirectDictationControllerTests.swift index 5f6d56e..d96f0b9 100644 --- a/TalkifyTests/DirectDictationControllerTests.swift +++ b/TalkifyTests/DirectDictationControllerTests.swift @@ -1098,6 +1098,40 @@ struct DirectDictationControllerTests { controller.stop() } + /// The rescue drops shaping but keeps the replacements: the user typed that + /// spelling, so it is the one the clipboard has to end up holding. + @Test func aFailedTranslationCopiesTheReplacedSpelling() async { + let recorder = Recorder() + let prewarmed = OSAllocatedUnfairLock(initialState: false) + let settings = AppSettings(defaults: freshDefaults()) + settings.translationTargetIdentifier = "es" + settings.spellingReplacements = [ + SpellingReplacement(id: "1", from: "Calman", to: "Kalman") + ] + let controller = makeController( + settings: settings, + dependencies: makeDependencies( + recorder: recorder, + prewarmed: prewarmed, + finishRecognition: { "Calman shipped" }, + translateBody: { _ in throw TranslationFailure.timedOut } + ) + ) + await prepareWithTranslation(controller, prewarmed: prewarmed) + + controller.handle(.triggerPressed(.translate)) + controller.handle(.triggerReleased(.translate)) + await waitUntil("Session never latched") { + controller.sessionStateForTesting == .recording(.latched) + } + controller.handle(.triggerPressed(.translate)) + await waitUntil("Rescue never happened") { !recorder.insertedTexts.isEmpty } + + #expect(recorder.insertedTexts == ["Kalman shipped"]) + #expect(recorder.insertedDestinations == [.clipboardOnly]) + controller.stop() + } + /// The trap: routing the failure through fail() would drive the machine to /// cancelling and cancel a session that has already finished. @Test func aFailedTranslationEndsTheSessionRatherThanCancellingIt() async { @@ -1340,6 +1374,149 @@ struct DirectDictationControllerTests { controller.stop() } + /// A spelling replacement rewrites the recognized word before insertion, + /// and history still keeps what the Speech Model produced. + @Test func finishAppliesSpellingReplacementsBeforeInsertion() async { + let recorder = Recorder() + let prewarmed = OSAllocatedUnfairLock(initialState: false) + let historyEntries = OSAllocatedUnfairLock<[HistoryEntry]>( + initialState: [] + ) + let settings = AppSettings(defaults: freshDefaults()) + settings.dictationHistoryEnabled = true + settings.spellingReplacements = [ + SpellingReplacement(id: "1", from: "Calman", to: "Kalman"), + ] + let controller = makeController( + settings: settings, + dependencies: makeDependencies( + recorder: recorder, + prewarmed: prewarmed, + finishRecognition: { "ship Calman friday" }, + historyEntries: historyEntries + ) + ) + await prepare(controller, prewarmed: prewarmed) + + controller.toggleFromMenu() + await waitUntil("Session never reached recording") { + controller.sessionStateForTesting == .recording(.latched) + } + controller.toggleFromMenu() + await waitUntil("Finish never delivered") { !recorder.insertedTexts.isEmpty } + + #expect(recorder.insertedTexts == ["ship Kalman friday"]) + #expect(historyEntries.withLock { $0 }.map(\.text) == ["ship Calman friday"]) + controller.stop() + } + + /// Shaping sees the replaced spelling, not the misspelling the Speech + /// Model produced. + @Test func spellingReplacementsRunBeforeShaping() async { + let recorder = Recorder() + let prewarmed = OSAllocatedUnfairLock(initialState: false) + let shapedInput = OSAllocatedUnfairLock(initialState: nil) + let settings = AppSettings(defaults: freshDefaults()) + settings.promptShapingEnabled = true + settings.promptShapingPromptID = "tighten-grammar" + settings.spellingReplacements = [ + SpellingReplacement(id: "1", from: "Calman", to: "Kalman"), + ] + let controller = makeController( + settings: settings, + dependencies: makeDependencies( + recorder: recorder, + prewarmed: prewarmed, + finishRecognition: { "ship calman friday" }, + shapeText: { text, _ in + shapedInput.withLock { $0 = text } + return "shaped" + } + ) + ) + await prepare(controller, prewarmed: prewarmed) + + controller.toggleFromMenu() + await waitUntil("Session never reached recording") { + controller.sessionStateForTesting == .recording(.latched) + } + controller.toggleFromMenu() + await waitUntil("Finish never delivered") { !recorder.insertedTexts.isEmpty } + + #expect(shapedInput.withLock { $0 } == "ship Kalman friday") + #expect(recorder.insertedTexts == ["shaped"]) + controller.stop() + } + + /// Editing the list mid-session must not rewrite the session already + /// under way (ADR-0004). + @Test func spellingReplacementsFollowTheSessionSnapshot() async { + let recorder = Recorder() + let prewarmed = OSAllocatedUnfairLock(initialState: false) + let settings = AppSettings(defaults: freshDefaults()) + settings.spellingReplacements = [ + SpellingReplacement(id: "1", from: "Calman", to: "Kalman"), + ] + let controller = makeController( + settings: settings, + dependencies: makeDependencies( + recorder: recorder, + prewarmed: prewarmed, + finishRecognition: { "Calman" } + ) + ) + await prepare(controller, prewarmed: prewarmed) + + controller.toggleFromMenu() + await waitUntil("Session never reached recording") { + controller.sessionStateForTesting == .recording(.latched) + } + settings.spellingReplacements = [] + controller.toggleFromMenu() + await waitUntil("Finish never delivered") { !recorder.insertedTexts.isEmpty } + + #expect(recorder.insertedTexts == ["Kalman"]) + controller.stop() + } + + /// A row with no replacement yet is incomplete, so it must not delete + /// the word or hand an empty transcript to shaping. + @Test func anIncompleteReplacementLeavesTheRecognizedWordAlone() async { + let recorder = Recorder() + let prewarmed = OSAllocatedUnfairLock(initialState: false) + let shapedInput = OSAllocatedUnfairLock(initialState: nil) + let settings = AppSettings(defaults: freshDefaults()) + settings.promptShapingEnabled = true + settings.promptShapingPromptID = "tighten-grammar" + settings.spellingReplacements = [ + SpellingReplacement(id: "1", from: "Calman", to: ""), + ] + let controller = makeController( + settings: settings, + dependencies: makeDependencies( + recorder: recorder, + prewarmed: prewarmed, + finishRecognition: { "Calman" }, + shapeText: { text, _ in + shapedInput.withLock { $0 = text } + return "shaped" + } + ) + ) + await prepare(controller, prewarmed: prewarmed) + + controller.toggleFromMenu() + await waitUntil("Session never reached recording") { + controller.sessionStateForTesting == .recording(.latched) + } + controller.toggleFromMenu() + await waitUntil("Finish never delivered") { !recorder.insertedTexts.isEmpty } + + #expect(shapedInput.withLock { $0 } == "Calman") + #expect(recorder.insertedTexts == ["shaped"]) + controller.stop() + } + /// Shaping off — the default — never touches the text. @Test func finishInsertsRawTextWhileShapingIsOff() async { let recorder = Recorder() diff --git a/TalkifyTests/SettingsWindowTests.swift b/TalkifyTests/SettingsWindowTests.swift index 7cdf290..3945192 100644 --- a/TalkifyTests/SettingsWindowTests.swift +++ b/TalkifyTests/SettingsWindowTests.swift @@ -30,7 +30,7 @@ struct SettingsWindowTests { @Test func settingsSectionsStayFocusedOnImplementedFeatures() { let expected: [SettingsSection] = [ - .general, .appearance, .sounds, .dictation, .promptShaping, + .general, .appearance, .sounds, .dictation, .promptShaping, .replacements, .dropTranscription, .readAloud, .language, .shortcuts, .updates, .insights, ] #expect(SettingsSection.allCases == expected) diff --git a/TalkifyTests/SpellingReplacementTests.swift b/TalkifyTests/SpellingReplacementTests.swift new file mode 100644 index 0000000..19a40da --- /dev/null +++ b/TalkifyTests/SpellingReplacementTests.swift @@ -0,0 +1,95 @@ +import Testing +@testable import Talkify + +struct SpellingReplacementTests { + @Test func emptyListLeavesTheTextAlone() { + #expect(SpellingReplacements.apply("Calman shipped", using: []) == "Calman shipped") + } + + @Test func swapsAWholeWordRegardlessOfCase() { + let pairs = [SpellingReplacement(id: "1", from: "Calman", to: "Kalman")] + #expect(SpellingReplacements.apply("Calman shipped", using: pairs) == "Kalman shipped") + #expect(SpellingReplacements.apply("calman shipped", using: pairs) == "Kalman shipped") + #expect(SpellingReplacements.apply("CALMAN shipped", using: pairs) == "Kalman shipped") + } + + @Test func replacesEveryOccurrence() { + let pairs = [SpellingReplacement(id: "1", from: "Calman", to: "Kalman")] + #expect( + SpellingReplacements.apply("Calman met Calman", using: pairs) == "Kalman met Kalman" + ) + } + + @Test func doesNotReplaceAWordThatOnlyContainsTheFrom() { + let pairs = [SpellingReplacement(id: "1", from: "mm", to: "millimetres")] + #expect( + SpellingReplacements.apply("leave a comment", using: pairs) == "leave a comment" + ) + } + + @Test func keepsSurroundingPunctuation() { + let pairs = [SpellingReplacement(id: "1", from: "Calman", to: "Kalman")] + #expect(SpellingReplacements.apply("ship Calman.", using: pairs) == "ship Kalman.") + #expect(SpellingReplacements.apply("Calman, then", using: pairs) == "Kalman, then") + } + + @Test func skipsAPairWhoseFromIsBlank() { + let pairs = [SpellingReplacement(id: "1", from: " ", to: "Kalman")] + #expect(SpellingReplacements.apply("Calman", using: pairs) == "Calman") + } + + @Test func skipsAPairWhoseToIsBlank() { + let pairs = [SpellingReplacement(id: "1", from: "Calman", to: " ")] + #expect(SpellingReplacements.apply("Calman shipped", using: pairs) == "Calman shipped") + } + + @Test func replacesAPossessiveAndKeepsTheSuffix() { + let pairs = [SpellingReplacement(id: "1", from: "Calman", to: "Kalman")] + #expect(SpellingReplacements.apply("Calman's launch", using: pairs) == "Kalman's launch") + #expect( + SpellingReplacements.apply("Calman\u{2019}s launch", using: pairs) + == "Kalman\u{2019}s launch" + ) + } + + @Test func aPairNeverRewritesAnotherPairsOutput() { + let pairs = [ + SpellingReplacement(id: "1", from: "colour", to: "color"), + SpellingReplacement(id: "2", from: "color", to: "Color"), + ] + #expect(SpellingReplacements.apply("colour", using: pairs) == "color") + #expect(SpellingReplacements.apply("color", using: pairs) == "Color") + } + + @Test func swapsAPhraseTheRecognizerSplitUp() { + let pairs = [SpellingReplacement(id: "1", from: "ex code", to: "Xcode")] + #expect(SpellingReplacements.apply("open ex code", using: pairs) == "open Xcode") + #expect(SpellingReplacements.apply("Ex Code's build", using: pairs) == "Xcode's build") + } + + @Test func swapsAHyphenatedWord() { + let pairs = [SpellingReplacement(id: "1", from: "e-mail", to: "email")] + #expect(SpellingReplacements.apply("send an e-mail", using: pairs) == "send an email") + #expect(SpellingReplacements.apply("e-mailing you", using: pairs) == "e-mailing you") + } + + @Test func theLongestMatchingPairWins() { + let pairs = [ + SpellingReplacement(id: "1", from: "git", to: "Git"), + SpellingReplacement(id: "2", from: "git hub", to: "GitHub"), + ] + #expect(SpellingReplacements.apply("git hub repo", using: pairs) == "GitHub repo") + #expect(SpellingReplacements.apply("git clone", using: pairs) == "Git clone") + } + + @Test func leavesAContractionWhole() { + let pairs = [SpellingReplacement(id: "1", from: "don", to: "Don")] + #expect(SpellingReplacements.apply("don't stop", using: pairs) == "don't stop") + #expect(SpellingReplacements.apply("don stopped", using: pairs) == "Don stopped") + } + + @Test func trimsTheStoredEndsBeforeMatching() { + let pairs = [SpellingReplacement(id: "1", from: " Calman ", to: " Kalman ")] + #expect(SpellingReplacements.apply("Calman", using: pairs) == "Kalman") + } +} diff --git a/docs/adr/0009-spelling-replacements.md b/docs/adr/0009-spelling-replacements.md new file mode 100644 index 0000000..f09a963 --- /dev/null +++ b/docs/adr/0009-spelling-replacements.md @@ -0,0 +1,64 @@ +# User-authored spelling replacements + +Direct Dictation gains a **Spelling replacement** list: whole-word from→to +pairs the user types in Settings. After recognition, and before Prompt +Shaping or translation, each pair is applied to the live draft and to the +text that will be inserted. Drop Transcription uses the same list on a +finished file job. + +This is the feature #19 tried to ship as recognition biasing. Apple's +`AnalysisContext.contextualStrings` are on the macOS 26 stack, but +`SpeechTranscriber` does not take them into account — measured in that +issue, and confirmed by Apple. A list of terms handed to the analyzer +does not change "Hetty" into "Hedy". The only thing that does, today, is +rewriting the text after the Speech Model has produced it. + +#19 closed find-and-replace because a substring swap has no edge: a list +holding "mm" would eat it inside "comment". Whole-word matching is that +edge. A match has to start where a word starts and end where one ends, +counting letters, digits and the apostrophe as word characters. It does +not stop a user who adds the word "mm" from rewriting "3 mm"; that is the +pair they typed. It does stop a pair from rewriting a longer token that +only contains the letters. + +The edge is the boundary, not the token, so the left side may hold spaces +or hyphens. That matters more than it sounds: the recognizer's usual +failure on a name is to split it, and a rule that only ever matched one +token would have no answer for "ex code" or "e-mail" — the forms the user +actually has to correct. A boundary rule covers both without a second +mechanism. + +Every pair is measured against what was said, in one left-to-right pass, +and the longest matching left side wins. Feeding each pair the previous +one's output would let "colour→color" and a later "color→Color" turn +every "colour" into "Color", and the Settings list has no reordering for +the user to fix it with. Longest-match is what keeps a pair for "git" +from turning "git hub" into "Git hub". + +Case-insensitive on the misspelling, trimmed on both sides, possessive +suffix preserved, because English uses that form for a name and the +recognizer misspells it the same way; the 's is left in place and rides +the new spelling. An incomplete pair is skipped, so a row being typed +cannot delete a word or send an empty transcript to Prompt Shaping. + +The list is snapshotted with the rest of Dictation session settings +(ADR-0004). An empty list is the previous behavior. History keeps the +words as recognized, so a replacement is visible as a difference between +the history file and what was inserted, the same way shaping is. + +This argues with CONTEXT.md's "no autocorrect" rule the same way Prompt +Shaping does: the default session is unchanged, and the rewrite is only +the pair the user wrote. + +## Consequences + +- This is a spelling fix, not a Speech Model. It swaps what the + recognizer produced for what the user wrote, and only where the user + wrote the pair; it cannot generalize to a mangling they have not seen + yet. +- A failed translation rescues the replaced spelling, not the recognized + one. Shaping is dropped there because it is a convenience; a + replacement is not. +- Custom pronunciations (`SFCustomLanguageModelData`) stay out of scope. + They need X-SAMPA and attach to `DictationTranscriber`, which Talkify + does not use.