Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 15 additions & 3 deletions CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,10 @@ _Avoid_: Bundled model, Talkify model, Whisper model
An opt-in pass that rewrites finished **Direct Dictation** text through the **On-device model** before insertion, using the session's **Shaping Prompt**.
_Avoid_: AI cleanup, post-processing, autocorrect

**Spelling replacement**:
One user-authored from→to pair, applied as a whole-word swap after recognition and before **Prompt Shaping**. The left side may be a phrase. The list is empty by default.
_Avoid_: Autocorrect, vocabulary, dictionary, lexicon

**Shaping Prompt**:
One named entry in the user's editable library, holding the wording that tells the **On-device model** what to do with the words. The framing that keeps it rewriting rather than answering is not part of it and is never editable.
_Avoid_: System prompt, preset, template
Expand Down Expand Up @@ -150,7 +154,7 @@ _Avoid_: Transcript history, cloud analytics
- Settings changes apply to the live app and persist immediately; Settings has no Save step
- The Appearance preview uses the same preferences and HUD surface as Direct Dictation
- Settings changes update the Appearance preview immediately, while an active Direct Dictation session keeps the choices captured at session start
- Dictation session settings include the voice visual, waveform style, glow palette, glow center, reveal style, long-draft behavior, HUD size, sound set, sound enabled state, sound volume, insertion destination, the transcription history choice, the shaping choice, and whether the session lowers other audio
- Dictation session settings include the voice visual, waveform style, glow palette, glow center, reveal style, long-draft behavior, HUD size, sound set, sound enabled state, sound volume, insertion destination, the transcription history choice, the shaping choice, the **Spelling replacement** list, and whether the session lowers other audio
- **Direct Dictation** can lower other audio while it listens and put it back when the session ends, cancels or fails; it is off by default because the control it moves is system-wide
- macOS has no per-application ducking, so lowering other audio moves the default output device's own volume and quiets Talkify's session sounds along with everything else
- A lowered volume is restored only while it is still the value Talkify set: a volume the user changed mid-session is theirs, the same rule the clipboard restore follows
Expand Down Expand Up @@ -199,7 +203,15 @@ _Avoid_: Transcript history, cloud analytics
- Settings preserves its selected section and window frame while the app runs, and opens on Appearance after a fresh launch
- The voice-reactive visual must make silence and a dead microphone look different
- With Reduce Motion enabled, the HUD replaces the animated visual with a quiet level meter and skips expand/collapse animation
- A session inserts raw finalized text unless the user turned **Prompt Shaping** on: no filler-word list, no autocorrect, and no rewriting anybody did not ask for
- A session inserts raw finalized text unless the user turned **Prompt Shaping** on or added a **Spelling replacement**: no filler-word list, no autocorrect, and no rewriting anybody did not ask for
- A **Spelling replacement** is a whole-word swap the user typed: the misspelling Apple Speech produced, and the spelling that should land instead
- **Spelling replacement** has its own Settings section; an empty list is exactly today's behavior
- Matching ignores case, surrounding spaces are trimmed, and a possessive keeps its 's, so Calman, calman, and Calman's all become Kalman / Kalman's
- The misspelling can be several words or a hyphenated one, because splitting a name is how Apple Speech usually gets it wrong: "ex code" and "e-mail" are as common as a single wrong token
- Every pair is measured against what was said rather than against another pair's output, and where two pairs both fit, the longer misspelling wins, so "git hub" becomes GitHub rather than Git hub
- A pair with a blank from or to is skipped, so a row still being filled in does not rewrite anything
- **Spelling replacement** runs after recognition and before **Prompt Shaping**, on the live draft and on the inserted text, and on a finished **Drop Transcription**
- Prompt Shaping sees the replaced spelling, and is skipped if that spelling is empty
- **Prompt Shaping** is the text cleanup the roadmap deferred, shipped as an explicit beta and off by default, so the default session still inserts exactly what was said
- While **Prompt Shaping** is on, the session's **Shaping Prompt** rewrites finished text through the **On-device model** between recognition and insertion, and nothing leaves the Mac: the model is Apple's, it runs here, and Talkify makes no request on its behalf
- Any prompt shaping unavailability, failure, or slow answer inserts the raw words unchanged
Expand All @@ -208,7 +220,7 @@ _Avoid_: Transcript history, cloud analytics
- Deleting the selected shaping prompt, or restoring the defaults over it, moves the selection to the first prompt in the library
- A selected shaping prompt that still resolves to nothing inserts the raw words unchanged
- A shaping prompt can be run against a sample sentence in Settings, through the same service a session uses, so a prompt is written by reading what it does rather than by guessing
- Transcription history keeps the words as spoken; a translated session's second line is what actually landed, so it is written after shaping and before insertion
- Transcription history keeps the words as recognized, before **Spelling replacement** or **Prompt Shaping**; a translated session's second line is what actually landed, so it is written after shaping and before insertion
- Shaping runs before translation: a prompt carries its own language and its one-shot example in that language, and a translator handed cleaned-up words has less to get wrong
- The shaping choice is captured in the Dictation session settings snapshot, along with the whole prompt library while shaping is on
- A session that will shape keeps the HUD up through the shaping phase, naming the prompt in the same caption the pick rode in, and the HUD leaves early when the answer lands
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,7 @@ xcodebuild test -project Talkify.xcodeproj -scheme Talkify -destination 'platfor
| Transcribe a file | Drag audio or video at the notch and drop it |
| Cancel mid-session | **Esc** |
| Shape what you dictate | Turn it on in **Settings → Prompt Shaping** |
| Fix a misspelled name | Add a pair in **Settings → Spelling replacements** |
| Pick the shaping prompt mid-session | **←** / **→** while dictating, with shaping on |
| Read selected text aloud | **⌥ ⎋** (toggles; also in the menu) |
| Read it aloud translated | Same key, with **Translate before speaking** on |
Expand Down
27 changes: 27 additions & 0 deletions Talkify/App/AppSettings.swift
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,7 @@ final class AppSettings {
static let promptShapingEnabled = "dictationPromptShapingEnabled"
static let promptShapingPrompt = "dictationPromptShapingPrompt"
static let shapingPrompts = "dictationShapingPrompts"
static let spellingReplacements = "dictationSpellingReplacements"
}

@ObservationIgnored
Expand Down Expand Up @@ -137,6 +138,16 @@ final class AppSettings {
shapingPrompts = ShapingPrompt.defaults
}

/// User-authored whole-word swaps applied after recognition. Empty is
/// today's behavior: the session inserts what the Speech Model produced.
var spellingReplacements: [SpellingReplacement] {
didSet {
if let data = try? JSONEncoder().encode(spellingReplacements) {
defaults.set(data, forKey: Keys.spellingReplacements)
}
}
}

var voiceVisual: HUDVoiceVisualStyle {
didSet { defaults.set(voiceVisual.rawValue, forKey: Keys.voiceVisual) }
}
Expand Down Expand Up @@ -310,6 +321,7 @@ final class AppSettings {
promptShapingPromptID = defaults.string(forKey: Keys.promptShapingPrompt)
?? ShapingPrompt.defaults[0].id
shapingPrompts = Self.storedShapingPrompts(in: defaults) ?? ShapingPrompt.defaults
spellingReplacements = Self.storedSpellingReplacements(in: defaults)
voiceVisual = Self.stored(in: defaults, key: Keys.voiceVisual) ?? .waveform
waveformStyle = Self.stored(in: defaults, key: Keys.waveformStyle) ?? .chartLine
revealStyle = Self.stored(in: defaults, key: Keys.revealStyle) ?? .slide
Expand Down Expand Up @@ -378,6 +390,17 @@ final class AppSettings {
return prompts
}

/// A missing or unreadable value is an empty list, which is the same as
/// never having had the feature: nothing is rewritten.
private static func storedSpellingReplacements(
in defaults: UserDefaults
) -> [SpellingReplacement] {
guard let data = defaults.data(forKey: Keys.spellingReplacements),
let pairs = try? JSONDecoder().decode([SpellingReplacement].self, from: data)
else { return [] }
return pairs
}

private static func store(_ binding: KeyBinding, in defaults: UserDefaults, key: String) {
if let data = try? JSONEncoder().encode(binding) {
defaults.set(data, forKey: key)
Expand Down Expand Up @@ -461,6 +484,9 @@ struct DictationSessionSettings: Equatable {
/// the arrow keys can cycle the session's pick without reading a library
/// that may change mid-session.
let shapingLibrary: [ShapingPrompt]
/// Whole-word swaps captured with everything else, so editing the list
/// mid-session cannot rewrite a phrase already on its way to insertion.
let spellingReplacements: [SpellingReplacement]
let voiceVisual: HUDVoiceVisualStyle
let waveformStyle: HUDWaveformStyle
let revealStyle: HUDRevealStyle
Expand Down Expand Up @@ -494,6 +520,7 @@ struct DictationSessionSettings: Equatable {
? settings.shapingPrompts.prompt(for: settings.promptShapingPromptID)
: nil
shapingLibrary = settings.promptShapingEnabled ? settings.shapingPrompts : []
spellingReplacements = settings.spellingReplacements
voiceVisual = settings.voiceVisual
waveformStyle = settings.waveformStyle
revealStyle = settings.revealStyle
Expand Down
58 changes: 40 additions & 18 deletions Talkify/Dictation/DirectDictationController.swift
Original file line number Diff line number Diff line change
Expand Up @@ -676,8 +676,18 @@ final class DirectDictationController {
let hasVisibleText = !displayText
.trimmingCharacters(in: .whitespacesAndNewlines)
.isEmpty
pendingLiveText = update.finalizedText
pendingVolatileText = update.volatileText
// Keep the committed/volatile split the HUD uses, and rewrite each
// half so a pair like Hetty→Hedy shows in the live draft, not only
// in the paste. Insertion still applies to the full spoken string.
let replacements = currentSessionSettings?.spellingReplacements ?? []
pendingLiveText = SpellingReplacements.apply(
update.finalizedText,
using: replacements
)
pendingVolatileText = SpellingReplacements.apply(
update.volatileText,
using: replacements
)
send(.updateReceived(hasVisibleText: hasVisibleText))
pendingLiveText = nil
pendingVolatileText = ""
Expand Down Expand Up @@ -737,27 +747,38 @@ final class DirectDictationController {
defer { finishTask = nil }
do {
let spoken = try await dependencies.finishRecognition()
// Delivery follows the session snapshot, so a Settings change
// mid-session applies to the next session (ADR-0004).
let session = currentSessionSettings ?? settings.sessionSettings
// Held separately from `text`, which shaping and translation go on to
// overwrite: this is the spelling the user asked for, and it is what
// a failed translation has to fall back to.
let replaced = SpellingReplacements.apply(
spoken,
using: session.spellingReplacements
)
var text = replaced
// Shape the words that will be inserted, not the recognizer's
// misspelling: an empty replacement must not send a blank
// transcript to the On-device model, which can answer with
// words nobody spoke.
let willShape = chosenPrompt != nil
&& !text.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
// A session about to shape keeps the HUD up saying so; every other
// session dismisses here exactly as before.
let willShape = chosenPrompt != nil && !spoken.isEmpty
if let chosenPrompt, willShape {
dependencies.showShaping(chosenPrompt.name)
} else {
dependencies.hideHUD()
}
// Delivery follows the session snapshot, so a Settings change
// mid-session applies to the next session (ADR-0004).
let session = currentSessionSettings ?? settings.sessionSettings

// The one place between recognition and insertion where the words may
// change. A transform that fails delivers nothing: the trigger
// promised a translation, and pasting the untranslated words instead
// lands the wrong language in someone else's document.
var text = spoken
// Shaping first, then translation. A prompt is written in one language,
// with a one-shot example in it, so handing it a translation of the
// words asks it to work in a language it was not written for; and a
// translator given cleaned-up words has less to get wrong.
// Replacements, then shaping, then translation. The list is a
// spelling fix for what the Speech Model produced, so shaping and
// translation both see the name the user wrote. A prompt is written
// in one language, with a one-shot example in it, so handing it a
// translation of the words asks it to work in a language it was not
// written for; and a translator given cleaned-up words has less to
// get wrong.
if let chosenPrompt, willShape {
text = await dependencies.shapeText(text, chosenPrompt)
}
Expand All @@ -768,12 +789,13 @@ final class DirectDictationController {
do {
text = try await translation.translate(text, with: pair)
} catch {
// The rescue is the words as spoken, not as shaped: shaping is a
// convenience and the raw words are what must survive.
// The rescue drops shaping, which is a convenience, and keeps the
// replacements, which are not: the user typed that spelling and it
// is the one they want wherever the words land.
await recordHistory(spoken: spoken, delivered: nil, session: session)
// Clipboard-only whatever the session's destination: nothing is
// pasted, and the words survive where the user can reach them.
let rescue = await dependencies.insertText(spoken, nil, .clipboardOnly)
let rescue = await dependencies.insertText(replaced, nil, .clipboardOnly)
// Ended, not failed: a failure action would drive the machine to
// cancelling and cancel a session that has already finished.
send(.sessionEnded)
Expand Down
118 changes: 118 additions & 0 deletions Talkify/Dictation/SpellingReplacement.swift
Original file line number Diff line number Diff line change
@@ -0,0 +1,118 @@
import Foundation

/// One user-authored from→to pair, applied as a whole-word swap after
/// recognition. The Speech Model cannot be taught a word it does not
/// know; this is the list that fixes what it keeps misspelling, and
/// only the words the user typed.
///
/// Whole-word on purpose: a substring swap has no edge, and a list
/// holding "mm" would eat it inside "comment". An incomplete pair —
/// blank `from` or blank `to` — is skipped so a row being typed does
/// not rewrite anything yet.
struct SpellingReplacement: Codable, Equatable, Identifiable, Sendable {
var id: String
var from: String
var to: String
}

enum SpellingReplacements {
/// Rewrites `text` in one left-to-right pass. Each position offers every
/// pair the same chance and the longest `from` wins, so a pair for
/// "git hub" beats one for "git" wherever both fit; equal lengths go
/// to list order. One pass on purpose: replacing into the running
/// result would let a later pair rewrite an earlier pair's output, and
/// the Settings list has no reordering for the user to fix it with.
///
/// Matching is case-insensitive; the replacement is the trimmed `to`
/// string, so "calman" and "Calman" both become "Kalman" if that is what
/// the user wrote. A `from` may hold spaces or hyphens — "ex code" and
/// "e-mail" are what the recognizer produces as often as a single token
/// is.
static func apply(_ text: String, using pairs: [SpellingReplacement]) -> String {
let active = pairs.compactMap { pair -> Pair? in
let from = pair.from.trimmingCharacters(in: .whitespacesAndNewlines)
let to = pair.to.trimmingCharacters(in: .whitespacesAndNewlines)
guard !from.isEmpty, !to.isEmpty else { return nil }
return Pair(from: from, to: to, length: from.count)
}
guard !active.isEmpty, !text.isEmpty else { return text }

var result = ""
result.reserveCapacity(text.count)
var index = text.startIndex
var previous: Character?
while index < text.endIndex {
// A match can only open where a word does, which is also what keeps
// "mm" out of the middle of "comment".
if !isWordCharacter(previous),
let match = longestMatch(in: text, at: index, using: active) {
result += match.to
// The boundary the next position sees is the text's, not the
// replacement's: what was matched stays matched.
previous = text[text.index(before: match.end)]
index = match.end
} else {
previous = text[index]
result.append(text[index])
index = text.index(after: index)
}
}
return result
}

private struct Pair {
let from: String
let to: String
let length: Int
}

private struct Match {
let to: String
let end: String.Index
}

private static func longestMatch(
in text: String,
at index: String.Index,
using pairs: [Pair]
) -> Match? {
var best: Match?
var bestLength = 0
for pair in pairs where pair.length > bestLength {
guard
let end = text.index(index, offsetBy: pair.length, limitedBy: text.endIndex),
text.compare(pair.from, options: .caseInsensitive, range: index..<end)
== .orderedSame,
endsWord(text, at: end)
else { continue }
best = Match(to: pair.to, end: end)
bestLength = pair.length
}
return best
}

/// Straight and typographic apostrophes, the two that hold a word
/// together in English.
private static let apostrophes: Set<Character> = ["'", "\u{2019}"]

/// Letters, digits, and the apostrophe, so "don't" is one word and a
/// pair for "don" cannot cut it in half.
private static func isWordCharacter(_ character: Character?) -> Bool {
guard let character else { return false }
return character.isLetter || character.isNumber || apostrophes.contains(character)
}

private static func endsWord(_ text: String, at end: String.Index) -> Bool {
guard end < text.endIndex, isWordCharacter(text[end]) else { return true }
// A possessive stays attached to the name, because English uses that
// form for a product and the recognizer misspells it the same way. The
// 's is left where it is, so it rides the new spelling untouched.
guard apostrophes.contains(text[end]) else { return false }
let afterMark = text.index(after: end)
guard afterMark < text.endIndex, text[afterMark].lowercased() == "s" else {
return false
}
let afterS = text.index(after: afterMark)
return afterS == text.endIndex || !isWordCharacter(text[afterS])
}
}
Loading