Skip to content

Intern-Decision-0.8B on Core ML: runtime, model store, Showdown fine-tune snapshot - #21

Merged
Alex-Wengg merged 4 commits into
mainfrom
feat/intern-decision
Sep 30, 2026
Merged

Alex-Wengg merged 4 commits into
mainfrom
feat/intern-decision

Conversation

@Alex-Wengg

@Alex-Wengg Alex-Wengg commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Runtime for FluidInference/intern-decision-0.8b-coreml (text path of internlm/Intern-Decision-0.8B, Apache-2.0) and its Pokémon Showdown fine-tune FluidInference/intern-decision-0.8b-showdown-coreml.

  • InternDecisionManager: one Core ML call per request (state + up to 16 typed questions), restricted softmax at each <decision> marker with the checkpoint's temperature. Prompt is byte-identical to the checkpoint's own compiler: OrderedJSON reproduces Python's json.dumps / str / repr (key order, escapes, float repr).
  • InternDecisionModelStore.ensure(.stock | .showdown): pinned, checksummed snapshots.
  • InternDecisionCheck parity/bench CLI; InternDecisionPromptTests against fixtures from the reference compiler.
  • Tools/showdown/: optional Python harness (poke-env player, deciders, teacher collection, LoRA distillation, export, demo) that produced the Showdown snapshot. The Swift package does not depend on it.

Checks: 36 records / 60 fields vs the checkpoint's fp32 engine, 0 token mismatches, 0 top-answer changes, max |Δp| 0.0044. Model-card request (319 tokens, 3 fields) 61 ms p50 on an M5 Pro GPU vs 150 ms for the checkpoint's PyTorch path on MPS.

Review round applied (surrogate-pair trap, Python str() for container descriptions, control-token rejection, RoPE cache, single load per bucket, stale .mlmodelc cleanup, CLI guards, JSONValue → OrderedJSON). Deferred to a follow-up: sharing the snapshot download loop and row-input builders with the Kev store/manager.

🤖 Generated with Claude Code

Alex-Wengg and others added 4 commits September 30, 2026 00:51
…checks

Runtime for FluidInference/intern-decision-0.8b-coreml (text path of
internlm/Intern-Decision-0.8B). A request is one prompt with every typed
question; the Core ML package returns the answer-symbol logits before each
<decision> marker and the host applies the restricted softmax and the
checkpoint's temperature.

The prompt must match the checkpoint's compile_row and chat template byte
for byte, including Python's json.dumps rendering of the state, so
OrderedJSON adds a JSON value that keeps key order with a json.dumps-exact
dumper (indent, escapes, float repr) and an order-keeping parser.
InternDecisionQuestion mirrors inference._options (choice, noul defaults,
score list and keyed forms) and validation mirrors validate_request.

Checks against the checkpoint's fp32 engine: 12 prompt fixtures from its
compiler (test), and 36 records / 60 fields from its published suites with
0 token mismatches, 0 top-answer changes, max |dp| 0.0044 (InternDecisionCheck
parity). The model-card request shape (319 tokens, three fields) takes
61 ms on an M5 Pro GPU in the 320-token bucket.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
InternDecisionModelStore.ensure(.showdown) downloads
FluidInference/intern-decision-0.8b-showdown-coreml (Intern-Decision-0.8B
distilled from the 4B on Pokémon Showdown battles; 512/640/1024-token
buckets) with the same checksummed layout as the stock snapshot, so
InternDecisionManager loads either unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Review fixes:
- OrderedJSON (renamed from JSONValue to avoid clashing with clients'
  own JSONValue types): malformed surrogate pairs now throw instead of
  trapping; pythonStr/pythonRepr reproduce Python str()/repr() so array
  and object option descriptions render as the reference compiler does.
- InternDecisionManager: control tokens in caller data are rejected
  (documented deviation from the reference, which only checks the
  decision marker); RoPE tables are built once per bucket; concurrent
  first loads of a bucket share one MLModel.load task.
- InternDecisionModelStore: a replaced .mlpackage drops the stale
  .mlmodelc compiled from its predecessor, so a revision bump cannot keep
  running old weights.
- InternDecisionCheck: no force unwrap on malformed records; bench
  refuses iterations < 1.
Deferred: sharing the snapshot download loop and the row-input builders
with KevModelStore/KevManager (five near-identical copies in the module).

Tools/showdown: the poke-env harness, deciders, teacher collection, LoRA
distillation, export, demo and footprint scripts that produced
FluidInference/intern-decision-0.8b-showdown-coreml, with the Core ML
export helpers vendored so the folder runs on its own. The Swift package
does not depend on it.

Swift parity after the changes: 36 records / 60 fields, 0 token
mismatches, 0 flips, max |dp| 0.0044; bench unchanged at 90 ms.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Alex-Wengg
Alex-Wengg merged commit 75d3645 into main Sep 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant