Skip to content

Short-reply model: sub-1B Qwen3-0.6B fine-tune, Core ML host, hotkey demo - #23

Merged
Alex-Wengg merged 2 commits into
mainfrom
feat/short-reply-demo
Oct 2, 2026
Merged

Alex-Wengg merged 2 commits into
mainfrom
feat/short-reply-demo

Conversation

@Alex-Wengg

@Alex-Wengg Alex-Wengg commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Swift host and menu-bar demo for the short-reply model: a sub-1B Qwen3-0.6B fine-tune that drafts one short reply to a social post, on device. Package on HF: FluidInference/short-reply-0.6b-coreml — 724 MB, prefill on the Neural Engine (1,920/1,921 ops), decode on the GPU, ~190 ms per reply from Swift, 89/100 replies identical to the PyTorch checkpoint. Training and conversion live outside this repo.

  • ShortReplyManager (macOS 15): loads the multifunction package, tokenizes the Qwen3 chat prompt (QwenBPETokenizer gains decode), left-pads to the fixed prefill length, writes the prefill K/V into the decoder MLState, greedy decode with an echo guard; seeded sampling for "regenerate".
  • ShortReplyDemo: menu-bar app. With a reply box focused (X in Chrome/Safari, Slack…), 9 drafts and pastes the reply into it, 0 regenerates (replaces the box). Reads the post above the box via Accessibility; demo.sh --x opens x.com plus a macmon + log terminal. mock-feed/ is a local fictional feed for recordings.
  • ShortReplyCheck: replays benchmark posts for parity/latency.
  • README: "Short replies" section.

Notes for review: drafts are for a person to review; nothing is posted automatically. ShortReplyTests skip below macOS 15; the tokenizer-parity test needs SHORT_REPLY_MODEL_DIR. Long posts are cut to the 160-token prefill budget; a 256-token prefill function is the follow-up if needed.

🤖 Generated with Claude Code

ShortReplyManager (macOS 15) runs the FluidInference/short-reply-0.6b-coreml
package: Qwen3 chat prompt via QwenBPETokenizer (which gains decode),
left-padded to the fixed prefill length, prefill K/V written into the decoder
MLState, greedy decode with an echo guard, seeded sampling for regenerate.
ShortReplyDemo is a menu-bar app: with a reply box focused, 9 drafts and
pastes the reply, 0 regenerates; the post above the box is read through
Accessibility (Chrome and Safari); demo.sh --x opens x.com plus a macmon +
log terminal; mock-feed/ is a local fictional feed for recordings.
ShortReplyCheck replays benchmark posts for parity and latency. README gains
a Short replies section.
@Alex-Wengg
Alex-Wengg force-pushed the feat/short-reply-demo branch from ae401cf to 8c14be4 Compare October 2, 2026 01:55
Package.swift / README: keep main's Intern-Decision target and section (the
branch rebuild had taken older copies). Host: compile the package once via
KevManager.compiled; echo guard on greedy replies (seeded resample, as
reply.py); long-post trimming re-checks the tokenized prompt and flags the
character cut; empty replies are an error; prefill K/V copied by the arrays'
strides with layout checks; NaN-safe sampling; regenerate returns the last
attempt instead of re-decoding a known duplicate. Demo: regenerate keys off
the raw post text (normalized text no longer resets sampling), ignores keys
while a draft is in flight, Quit keeps its responder-chain target, bare 9/0
keys can be switched off (⌃⌥R stays), Chromium-only AXEnhancedUserInterface,
post lookup visits each element once. README fence fixed; tokenizer decode
reuses the special-id set.
@Alex-Wengg
Alex-Wengg merged commit 415ae5d into main Oct 2, 2026
1 check passed
@Alex-Wengg
Alex-Wengg deleted the feat/short-reply-demo branch October 2, 2026 02:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant