Intern-Decision-0.8B on Core ML: runtime, model store, Showdown fine-tune snapshot - #21
Merged
Merged
Conversation
…checks Runtime for FluidInference/intern-decision-0.8b-coreml (text path of internlm/Intern-Decision-0.8B). A request is one prompt with every typed question; the Core ML package returns the answer-symbol logits before each <decision> marker and the host applies the restricted softmax and the checkpoint's temperature. The prompt must match the checkpoint's compile_row and chat template byte for byte, including Python's json.dumps rendering of the state, so OrderedJSON adds a JSON value that keeps key order with a json.dumps-exact dumper (indent, escapes, float repr) and an order-keeping parser. InternDecisionQuestion mirrors inference._options (choice, noul defaults, score list and keyed forms) and validation mirrors validate_request. Checks against the checkpoint's fp32 engine: 12 prompt fixtures from its compiler (test), and 36 records / 60 fields from its published suites with 0 token mismatches, 0 top-answer changes, max |dp| 0.0044 (InternDecisionCheck parity). The model-card request shape (319 tokens, three fields) takes 61 ms on an M5 Pro GPU in the 320-token bucket. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
InternDecisionModelStore.ensure(.showdown) downloads FluidInference/intern-decision-0.8b-showdown-coreml (Intern-Decision-0.8B distilled from the 4B on Pokémon Showdown battles; 512/640/1024-token buckets) with the same checksummed layout as the stock snapshot, so InternDecisionManager loads either unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Review fixes: - OrderedJSON (renamed from JSONValue to avoid clashing with clients' own JSONValue types): malformed surrogate pairs now throw instead of trapping; pythonStr/pythonRepr reproduce Python str()/repr() so array and object option descriptions render as the reference compiler does. - InternDecisionManager: control tokens in caller data are rejected (documented deviation from the reference, which only checks the decision marker); RoPE tables are built once per bucket; concurrent first loads of a bucket share one MLModel.load task. - InternDecisionModelStore: a replaced .mlpackage drops the stale .mlmodelc compiled from its predecessor, so a revision bump cannot keep running old weights. - InternDecisionCheck: no force unwrap on malformed records; bench refuses iterations < 1. Deferred: sharing the snapshot download loop and the row-input builders with KevModelStore/KevManager (five near-identical copies in the module). Tools/showdown: the poke-env harness, deciders, teacher collection, LoRA distillation, export, demo and footprint scripts that produced FluidInference/intern-decision-0.8b-showdown-coreml, with the Core ML export helpers vendored so the folder runs on its own. The Swift package does not depend on it. Swift parity after the changes: 36 records / 60 fields, 0 token mismatches, 0 flips, max |dp| 0.0044; bench unchanged at 90 ms. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Runtime for FluidInference/intern-decision-0.8b-coreml (text path of internlm/Intern-Decision-0.8B, Apache-2.0) and its Pokémon Showdown fine-tune FluidInference/intern-decision-0.8b-showdown-coreml.
InternDecisionManager: one Core ML call per request (state + up to 16 typed questions), restricted softmax at each<decision>marker with the checkpoint's temperature. Prompt is byte-identical to the checkpoint's own compiler:OrderedJSONreproduces Python'sjson.dumps/str/repr(key order, escapes, float repr).InternDecisionModelStore.ensure(.stock | .showdown): pinned, checksummed snapshots.InternDecisionCheckparity/bench CLI;InternDecisionPromptTestsagainst fixtures from the reference compiler.Tools/showdown/: optional Python harness (poke-env player, deciders, teacher collection, LoRA distillation, export, demo) that produced the Showdown snapshot. The Swift package does not depend on it.Checks: 36 records / 60 fields vs the checkpoint's fp32 engine, 0 token mismatches, 0 top-answer changes, max |Δp| 0.0044. Model-card request (319 tokens, 3 fields) 61 ms p50 on an M5 Pro GPU vs 150 ms for the checkpoint's PyTorch path on MPS.
Review round applied (surrogate-pair trap, Python
str()for container descriptions, control-token rejection, RoPE cache, single load per bucket, stale.mlmodelccleanup, CLI guards,JSONValue→OrderedJSON). Deferred to a follow-up: sharing the snapshot download loop and row-input builders with the Kev store/manager.🤖 Generated with Claude Code