Short-reply model: sub-1B Qwen3-0.6B fine-tune, Core ML host, hotkey demo - #23
Merged
Merged
Conversation
ShortReplyManager (macOS 15) runs the FluidInference/short-reply-0.6b-coreml package: Qwen3 chat prompt via QwenBPETokenizer (which gains decode), left-padded to the fixed prefill length, prefill K/V written into the decoder MLState, greedy decode with an echo guard, seeded sampling for regenerate. ShortReplyDemo is a menu-bar app: with a reply box focused, 9 drafts and pastes the reply, 0 regenerates; the post above the box is read through Accessibility (Chrome and Safari); demo.sh --x opens x.com plus a macmon + log terminal; mock-feed/ is a local fictional feed for recordings. ShortReplyCheck replays benchmark posts for parity and latency. README gains a Short replies section.
Alex-Wengg
force-pushed
the
feat/short-reply-demo
branch
from
October 2, 2026 01:55
ae401cf to
8c14be4
Compare
Package.swift / README: keep main's Intern-Decision target and section (the branch rebuild had taken older copies). Host: compile the package once via KevManager.compiled; echo guard on greedy replies (seeded resample, as reply.py); long-post trimming re-checks the tokenized prompt and flags the character cut; empty replies are an error; prefill K/V copied by the arrays' strides with layout checks; NaN-safe sampling; regenerate returns the last attempt instead of re-decoding a known duplicate. Demo: regenerate keys off the raw post text (normalized text no longer resets sampling), ignores keys while a draft is in flight, Quit keeps its responder-chain target, bare 9/0 keys can be switched off (⌃⌥R stays), Chromium-only AXEnhancedUserInterface, post lookup visits each element once. README fence fixed; tokenizer decode reuses the special-id set.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Swift host and menu-bar demo for the short-reply model: a sub-1B Qwen3-0.6B fine-tune that drafts one short reply to a social post, on device. Package on HF: FluidInference/short-reply-0.6b-coreml — 724 MB, prefill on the Neural Engine (1,920/1,921 ops), decode on the GPU, ~190 ms per reply from Swift, 89/100 replies identical to the PyTorch checkpoint. Training and conversion live outside this repo.
ShortReplyManager(macOS 15): loads the multifunction package, tokenizes the Qwen3 chat prompt (QwenBPETokenizergainsdecode), left-pads to the fixed prefill length, writes the prefill K/V into the decoderMLState, greedy decode with an echo guard; seeded sampling for "regenerate".ShortReplyDemo: menu-bar app. With a reply box focused (X in Chrome/Safari, Slack…), 9 drafts and pastes the reply into it, 0 regenerates (replaces the box). Reads the post above the box via Accessibility;demo.sh --xopens x.com plus a macmon + log terminal.mock-feed/is a local fictional feed for recordings.ShortReplyCheck: replays benchmark posts for parity/latency.Notes for review: drafts are for a person to review; nothing is posted automatically.
ShortReplyTestsskip below macOS 15; the tokenizer-parity test needsSHORT_REPLY_MODEL_DIR. Long posts are cut to the 160-token prefill budget; a 256-token prefill function is the follow-up if needed.🤖 Generated with Claude Code