Skip to content

docs: demo ideas, game trial results, and models to convert - #17

Open
Alex-Wengg wants to merge 8 commits into
mainfrom
docs/demo-ideas
Open

Alex-Wengg wants to merge 8 commits into
mainfrom
docs/demo-ideas

Conversation

@Alex-Wengg

@Alex-Wengg Alex-Wengg commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Replaces #2. Models.md from that PR is superseded by DecisionModelSupport.md; the game findings are kept here in a shorter form.

  • Documentation/DemoIdeas.md:
    • What the trials showed: stock vs game-trained results for Tetris, Snake, Connect Four, Flappy/runner, Wordle, plus two cautions (linear scorer ties the 706K model on Tetris; 10-seed runs are noisy).
    • Demos to post: ranked by clip value (train-a-game-brain, fine-tuned Lex Snake, Sort Anything, model race, real-app forms, describe-a-decision, merge puzzles, Codenames, Minesweeper calibration). Real-time games stay ruled out.
    • Different decision shapes: Wordle/Connections, Minesweeper, chess mate-in-1–2, customer-support desk, model router, each with what it tests beyond board geometry.
    • Next up: record Connect Four, Lex Snake, and the Safari form fill before building anything new; Sort Anything live-edit follow-up; voice → decision → Accessibility action combo.
    • Doom (tried) and Minecraft: ViZDoom results (stock GLiClass 0 kills on bare labels; SauerkrautLM-Doom 1.3M on Core ML 20.54 kills, matches PyTorch on 100/100 seeds, 17–38× faster), demo built on local branch; VPT/STEVE-1 as the Minecraft Core ML candidate.
    • Models to convert or use: laya-browser, Kev 0.8B/4B, KaLM-Jev, Verdict 2.0, QwenJev, Simple Jev, ProgramAsWeights, GLiNER/GLiClass siblings, plus a conversion order (laya-browser first).
  • README.md: links the new doc.

Docs only. Connect Four / Tetris trained models and the Snake demo app are on local branches, marked as such.

🤖 Generated with Claude Code

Replaces the closed #2. Its model inventory is superseded by
DecisionModelSupport.md; the game-trial findings are kept in a shorter
form and updated with the game-trained results (fine-tuned Lex on Snake,
tiny and fine-tuned GLiClass on Connect Four and Tetris). Adds a ranked
list of demos to post and open Jev-style models to convert.
Wordle/Connections, Minesweeper, chess mate-in-1-2, customer-support desk,
and a model router, each tagged with what it tests beyond board geometry.
Next-up list puts the Connect Four, Lex Snake, and Safari form clips before
any new build, with a Sort Anything follow-up that edits categories mid-run.
Adds the Parakeet EOU -> decision model -> Accessibility action combo and a
conversion priority for the models table.
defend_the_center, text state, three actions, 35 Hz engine with held
actions, and a hand-coded baseline on the same seed. Minecraft skipped:
the dragon run is planner + scripted motors.
Replaces the Doom plan with measured results: stock GLiClass (0 kills on
bare labels, 11.98 with consequence labels) vs SauerkrautLM-Doom 1.3M on
Core ML (20.54 kills, fp32 kill-for-kill with PyTorch on 100/100 seeds,
17-38x faster than 1-thread PyTorch). Notes the depth-only input and the
built demo. Minecraft: VPT/STEVE-1 is the one Core ML candidate; OpenHA
too large, Jev dragon runs are harness-driven.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant