Your tailor for local AI.
Start with a model. Find the configuration that fits your machine, your work, and your priorities. fitr measures the fit, tests the output, and keeps the evidence behind each choice.
Model names and public leaderboards do not answer the local questions. Will this artifact fit at the context you need? Is the runtime really using the accelerator? Are tool calls reaching the tool channel? Is the faster quant still correct? Does a candidate complete a bounded workflow when another system independently verifies the result?
fitr measures those questions against the model bytes, runtime, context, placement, and device that produced the evidence. Missing evidence stays missing. An unmeasured model stays a candidate, not a recommendation.
Each role has its own requirements and preferences. A coding agent, daily driver and classifier need different fittings. Quality requirements come first; speed and resource preferences help choose among models that meet them. See the personal fitting direction for the proposed Fit and Extended fit scopes, usable context, compaction and harness scorecards.
The wide Board keeps the comparable configurations, selected evidence, exact measurements, unresolved requirements, and one next action on one screen. The same facts remain available as compact terminal output, JSON, and HTML.
Ordinary measurements need a running supported backend with at least one model. Auto mode can own an installed Windows Ollama runtime. See backend requirements and identity limits.
macOS and Linux:
curl -fsSL https://raw.githubusercontent.com/blisspixel/fitr/main/install.sh | shWindows PowerShell:
irm https://raw.githubusercontent.com/blisspixel/fitr/main/install.ps1 | iexfitr is one static binary. It needs no Python environment and sends no
telemetry. Installers and fitr update verify the selected release asset
against its published SHA-256 checksum. Pinning a version, relocating the
binary, updating, and building from source are covered in the
install guide.
fitr # inventory, evidence state, one next action
fitr qwen3:30b # fit advice for one model
fitr run qwen3:30b --ctx 16384 --capacity-budget-gb 22
# behavior, performance, and safe-budget fit
fitr apply qwen3:30b # print, but do not execute, persistence steps
fitr board # compare compatible measured evidence
fitr decide qwen3:30b --spec local-coding.jsonEvery inventory row ends with one useful next command. A decision spec then applies your workload requirements to sealed evidence without changing the original measurement. Requirements may cover behavior, capability support, effective context, a specific latency state, and exact-context resident memory. The answer is eligible, ineligible, or unresolved, with exit code 4 reserved for required evidence that has not yet been established.
Read usage for all commands and flags, or decision specifications for the strict schema and requirement semantics.
A post, video, podcast, model card, or selected email excerpt can become an idea to investigate for coding, a daily driver, classification, or another role. Capture the claim and intended harness without turning it into evidence:
fitr discover add https://huggingface.co/Qwen/Qwen3.8-27B --role coding --model qwen3.8:27b --harness Pi
fitr discover list
fitr discover plan --role codingThe private inbox deduplicates ideas and drafts the next evidence steps. It keeps source claims unmeasured. For an explicit Hugging Face file selection, resolve public metadata before considering a download:
fitr source resolve hf --repo owner/model --revision main --file model-Q4_K_M.gguf --out candidate.json
fitr source show candidate.json
fitr discover attach-source <idea-id> candidate.json
fitr discover plan <idea-id>The receipt pins a commit, preserves declared file sizes and hashes, and surfaces dependency gaps. It downloads no weights and does not qualify the model for a role. See source resolution, source attachments and investigation plans, discovery and the model library for the flow, and agent interoperability for the portable Agent Plugins package, read-only MCP tools, and the researched A2A, Hermes, Pi and OpenClaw integration boundaries.
For files already on disk, an explicit mapping can compare their local hashes with the pinned receipt before any runtime experiment:
fitr artifact bind --source candidate.json --mapping local-files.json --max-bytes 68719476736 --timeout 10m --out artifact.jsonRead artifact binding for the mapping and I/O bounds.
Matching local bytes still leaves runtime unbound, and capacity and quality unmeasured.
Declare the quality, context and memory requirements for a job, then attach canonical measurements. Each role keeps its own versioned preferences:
fitr role init coding --quality user_tasks --memory-gb 22 --ctx 8192
fitr role attach coding /path/to/canonical-result.json
fitr role review codingA candidate must clear every floor before preferences matter. Comparisons
retain uncertainty and check sensitivity to weight changes; missing evidence
and no-qualified-candidate are useful answers. With two to four compatible
candidates, fitr role confirm coding seals the policy and collects fresh
evidence after runtime and memory preflight. Explicit role adopt records the
selection; role status rechecks its evidence and expiry. A failed challenger
keeps the incumbent reference. See roles for the battery
screening scope and rollback limits. Automatic serving changes remain future work.
For a role whose requirements the text battery can measure, declare two to four installed candidates and a fixed runtime configuration:
fitr auto runtime C:/runtime/ollama.exe --models D:/ollama-models --out runtime.json
fitr auto start daily --mode establish --runtime runtime.json --candidate first:tag --candidate second:tag
fitr auto status auto-<id>
fitr auto adopt auto-<id>fitr checks feasibility and resources, starts its own local runtime, reserves each request, and collects comparable evidence. A preselected choice gets one fresh confirmation attempt before adoption. Quality floors stay fixed, and an uncertain result stays unresolved. Status explains each candidate's gaps.
Manual adoption is the default. --adoption confirmed-only can authorize
selection in fitr after confirmation and runtime cleanup. The first owner is
Windows-only, for CPU-only machines or one NVIDIA GPU, explicit installed text artifacts,
and finite time, request and output-cap allowances. See auto mode
for preparation, supported requirements, resume rules and limits.
fitr cleanup plan D:\modelsThe read-only plan shows apparent storage, large files and aged partial downloads worth investigating. It does not delete files or assume a large encoder, projector, shard or cache is unused. Links, incomplete scans and unresolved dependencies remain explicit. See cleanup planning.
| Capability | What it establishes | Details |
|---|---|---|
| Inventory and fit advice | Installed artifacts, evidence freshness, projected context fit, and an exact remedy when supported evidence says a configuration does not fit | Usage, choosing hardware |
| Local measurement | Structured output, instruction following, refusal, tool-channel behavior, degeneration, load state, TTFT, prefill, decode, context, placement, and allocation when the backend exposes them | Design, tasks |
| Capacity policy | A pre-load sealed resource domain, timestamped availability, explicit operator budget or reserve, exact usable-budget formula, component projection, and observed safe headroom | Choosing hardware, usage |
| Workload decisions | Constraint-based eligibility under a versioned declaration, with no universal weighted score | Decisions |
| Personal roles | Quality and resource floors, fixed preferences, fresh confirmation, explicit selection, evidence expiry and validated rollback | Roles, confirmation |
| Source resolution | Commit-pinned public file metadata, distinct declared hashes and unresolved dependencies, with no weight downloads | Source metadata |
| Discovery investigations | Private source attachments and separately stated metadata, dependency, runtime and quality gaps | Source attachments |
| Local artifact observations | Bounded whole-file hashes for explicit mappings, source comparisons and change detection without runtime promotion | Artifact binding |
| Agent interoperability | Read-only MCP 2026-07-28 candidate review and selected status, official SDK acceptance and an Agent Plugins 1.0.0 package | Protocol and client limits |
| Cleanup planning | Read-only bounded storage inventory and aged partial-download review candidates, with no inferred deletion authority | Cleanup |
| Document context tasks | A sealed document pack at one operating window, collected with fitr run --context-tiers; CLI and JSON only; not a role floor |
Context quality, usage |
| Context experiments | A predeclared exploratory context plan with shared task seeds, point-specific allocation, required-equal factors, and replayable bundles | Context experiment |
| Configuration tradeoffs | Conservative frontiers across sealed candidates, optional same-base conversion lineage, and no point-estimate winner when intervals overlap | Quant experiment, calibration |
| Fresh confirmation | A sealed candidate set, a fresh shared task seed, full paired runs, and confirmation only when requirements resolve and the objective separates | Confirmation |
| Validated work | A sealed workflow contract, typed proof classes, per-trial timing partitions, fixed policy-repair tools, independent verification, and signed terminal outcomes | Validated work, workload evidence |
The broader 1.0 thesis has seven evidence layers:
FIT Can this exact configuration run within the relevant capacity?
BEHAVIOR Does it perform the required primitives correctly?
PERFORMANCE How long do distinct load and inference phases take?
EXPLAIN What does the evidence support, contradict, or leave unresolved?
VALIDATED WORK
Does speed survive retries and independent verification?
TRADEOFFS Which configurations are dominated, and which remain choices?
COVERAGE Which declared workloads have earned local trust or need fallback?
FIT, behavior, burst performance, typed capacity policy, direct receipt diagnoses, typed context and configuration experiments, a collectable document-task scorecard, fresh confirmation, and one bounded validated-work contract are implemented. Broader causal explanation, operational experiments, and declared workload coverage remain pre-1.0 work. The roadmap distinguishes shipped slices from planned contracts.
README screenshots are deterministic fixtures rendered through the real
presentation paths with make screenshots. Host identity and local paths are
omitted. Full receipts and the other command surfaces live in the linked docs,
where they can be explained without turning the front page into a transcript.
- Fit is not performance. Artifact size, KV projection, addressable capacity, current availability, observed allocation, and sustained residency are different claims.
- Speed is not correctness. A fast model can emit malformed structure, call a tool in message content, repeat a call after the tool disappears, or degenerate into a loop.
- Capability is not competence. A runtime declaration can route a test. It cannot become a behavioral PASS.
- Exploration is not confirmation. The observations that selected an attractive context or quant cannot certify that choice. Confirmation uses a fresh sealed plan and fresh evidence.
- The worker is not the verifier. Validated work requires harness-owned state and independent proof of the final outcome.
- Uncertainty is an answer. Missing receipts, overlapping intervals, and blocked observations remain visible instead of becoming estimates.
One renderer-neutral analysis path rebuilds presentation claims from validated records. CLI, TUI, JSON, and HTML consume those facts; a renderer does not get to invent a verdict.
- It does not rank evidence across incompatible machines or configurations.
- It does not call an unmeasured model good or turn missing input into a number.
- It does not reuse a result after the runtime-bound artifact identity changes.
- It does not add isolated model measurements and call the sum co-residency.
- It does not infer a hardware root cause from a timing ratio.
- It does not certify an exploration winner on the data that selected it.
- It does not mutate or restart your serving runtime.
fitr applyprints a recipe. - It does not run generated code by default or silently score unavailable execution evidence.
These boundaries are part of the product. The detailed evidence model and known limits are in design.
OpenRouter and other OpenAI-compatible providers are optional. They can help develop adversarial cases, calibrate heuristic graders, compare task-pack behavior, or provide an explicitly labeled model-judged observation. They are never required for installation, local measurement, decisions, exports, or CI, and they cannot establish local fit, placement, residency, or performance.
The existing OpenAI-compatible backend can be used explicitly for protocol diagnostics. Credentials come from environment variables and are not written to a fitr result. Exact commands, current limits, and the future experiment-scoped provider receipt are documented in optional external validation.
Installed-model evaluation talks only to the selected endpoint. With a local backend and installed artifact, the measurement path can run offline. Network access otherwise occurs only for an explicit install, update, pull, or remote endpoint. Explicit source resolution also fetches public file metadata.
Ollama entries marked as remote are excluded from local measurement, even when the daemon is on localhost. Absent markers are not proof of local execution; see execution provenance for the checks and limits.
Results remain on your machine unless you explicitly export them. The HTML export omits raw model output, hostnames, local paths, the raw device fingerprint key, and arbitrary runtime configuration. Private workload bundles retain hashes and deterministic verifier output rather than raw prompts, replies, or tool contents. The resulting integrity receipt is intentionally not described as full replayability.
Local storage permissions matter: source and artifact receipts use mode 0600
on Unix and inherit the destination directory's ACL on Windows. Use a directory
restricted to your account. fitr does not encrypt receipts or replace Windows ACLs.
- Usage: installation, commands, flags, output modes, storage, and exit codes.
- Design: the evidence model, scoring principles, trust boundaries, and known limits.
- Decision specifications: workload requirements, eligibility, uncertainty, objectives, and evidence levels.
- Workload evidence: bounded workflows, independent verification, timing, coverage, authority, and retention.
- Choosing hardware: capacity, performance, workload fit, and honest hardware evidence.
- Tasks, statistics, and calibration: battery construction and inference.
- Backends, doctor, and TUI: runtime contracts, measurement readiness, and the optional terminal monitor.
- Optional external validation: the narrow, opt-in OpenRouter role and its evidence boundary.
- Discovery and agent interoperability: ideas, role-specific choices, bounded automation and integration contracts.
- Terminal design language: color, layout, live activity and reduced-motion behavior.
- Roadmap and release acceptance: what comes next and the receipts required to publish.
Apache License 2.0. Dependency licenses are in THIRD_PARTY_NOTICES.md.