Skip to content

Repository files navigation

ModelMonitor

ModelMonitor is a static dashboard for two related jobs: following the history of large language models and checking what changed in the current model ecosystem.

The site keeps historical milestones separate from live observations. A paper publication date, a model launch, a Hugging Face repository creation date, and the first time ModelMonitor saw an endpoint are different facts and are stored as such.

What the site shows

  • Overview — an automatically generated intelligence brief: recent changes, current totals, source health, and a browser-local “since your last visit” summary.
  • LLM History — a curated timeline from important pre-Transformer work through current reasoning, multimodal, open-weight, and agentic systems.
  • Releases — the broader release/observation feed with filters and provenance.
  • Free Models — provider free-tier/credit programs plus endpoints that have recent evidence for zero input and output token pricing. These are kept separate so a provider-level allowance is never mistaken for a permanently free model.
  • Open Models — tracked open-source and open-weight models, with licenses and available metadata kept distinct.
  • Analysis — transparent comparisons and derived metrics based on the data ModelMonitor actually has.
  • Glossary — plain-language explanations of LLM terminology used around the dashboard.

Run locally

Node.js 20+ is required for the collectors and checks.

npm ci
npm test
npm run lint
npm run validate
python -m http.server 8000 --bind 127.0.0.1

Open http://localhost:8000.

The frontend is intentionally static-host friendly. There is no database, account system, or application server.

Refreshing data

npm run fetch

That command collects live public data, normalizes it, updates generated datasets, compares the result with the previous successful state, and regenerates the intelligence brief.

To republish curated history/glossary/report data without making network requests:

node scripts/fetch-models.js --offline

The collector never runs Git commands. GitHub Actions handles validation and commits generated public/data/ changes after a successful scheduled run.

Optional provider API secrets

OpenRouter is checked from its public model catalog and can optionally use OPENROUTER_API_KEY if you want authenticated catalog access. Groq and Gemini model catalogs require credentials, so their live endpoint inventories are enabled only when these GitHub Actions repository secrets are present:

  • GROQ_API_KEY
  • GEMINI_API_KEY
  • optionally OPENROUTER_API_KEY

If those secrets are missing, the refresh still succeeds. ModelMonitor continues to verify the providers' public free-tier/allowance documentation and shows those programs separately from model-level $0 pricing. Never commit API keys into the repository.

Data flow

public APIs / official model catalogs
                +
       data/curation.json
       data/milestones.json
       data/glossary.json
                |
                v
     collectors + pipeline
                |
                v
        public/data/*.json
                |
                v
            index.jsx

Curated inputs

data/curation.json is a small manual verification/configuration layer. It exists for cases automation should not guess: explicit identity overrides, sourced alias/reveal relationships, public usage-limit facts, and collector source configuration. It is not the primary model database and the application does not give special UI treatment to records listed there.

data/milestones.json contains the reviewed LLM-history backbone and its eras. The main history is intentionally curated rather than an attempt to list every paper, fine-tune, quantization, or repository.

data/glossary.json contains the user-facing terminology reference.

Generated datasets

File Purpose
public/data/models.json Canonical/provider-scoped model entities and their evidence
public/data/timeline.json Machine-observed releases, availability events, reveals, archive records, and other dated events
public/data/availability.json Current and removed provider endpoints, prices, context, and capabilities
public/data/history.json Compact provider availability snapshots/deltas
public/data/report.json Structured intelligence brief generated by comparing the current and previous state
public/data/snapshot.json Compact current aggregate snapshot used by report/history logic
public/data/milestones.json Published copy of the curated LLM-history milestones
public/data/glossary.json Published glossary
public/data/metadata.json Build/source health, limitations, and archive metadata
public/data/limits.json Source-backed provider free-tier, credit, allowance, and usage-limit records
public/data/signals.json Separate unverified/community signals when present
public/data/audit.json Disposition of the preserved legacy archive
public/data/cache.json Collector fallback metadata; not a primary UI dataset

public/models.json is the preserved legacy archive. Its dates and claims are not silently promoted to verified release facts.

Intelligence brief

public/data/report.json is deterministic. It is generated from dataset differences rather than by asking an LLM to write commentary.

It can report changes such as:

  • model added/removed/updated
  • endpoint added/removed
  • paid to free / free to paid
  • token-price changes
  • context or capability changes
  • open-model/license changes
  • historical milestone additions/corrections
  • deprecations and retirements

The Overview also stores a last-seen timestamp in browser localStorage. That powers “Since your last visit” on that browser only; it is not synced to an account.

Evidence and dates

ModelMonitor uses four evidence labels:

  • official — the claim is directly supported by an official provider, paper, model card, or first-party source.
  • confirmed — the relationship is supported by strong, explicit evidence but is not represented as a first-party live field.
  • observed — the value comes from a secondary catalog or direct observation and is not promoted beyond what that source establishes.
  • unverified — retained for visibility/history but not treated as established fact.

Important date fields are kept separate:

  • releaseDate — an explicitly supported release/launch date.
  • repositoryCreatedAt — repository creation, not a launch date.
  • firstSeen — first local observation by ModelMonitor.
  • lastChecked — most recent evidence check.
  • legacyReportDate — date stored by the old archive; unverified unless separately sourced.

Open source vs open weights

ModelMonitor does not treat these as synonyms.

A model may expose downloadable weights while using a license with restrictions. Those records are classified as open weights. Recognized permissive weight licenses may be classified as open source within ModelMonitor's narrower tracking convention; the dashboard does not claim broader OSI/Open Source AI certification from weight availability alone.

Automation

.github/workflows/update-models.yml runs the model-intelligence refresh every 6 hours at minute 17 (00:17, 06:17, 12:17, and 18:17 UTC; 08:17, 14:17, 20:17, and 02:17 in Philippine Time). GitHub may start scheduled jobs a little late during platform load. The workflow can also be started manually from the Actions tab. The job installs dependencies, runs tests and linting, collects data, validates the result, and commits generated public/data/ changes only when something actually changed.

Source failures are isolated. A failed source should not erase valid data from other sources, and stale evidence is excluded from current availability/free-paid totals where the freshness rules require it. OpenRouter is collected automatically; Groq and Gemini live catalogs are optional authenticated sources. Cloudflare Workers AI and Hugging Face free access are represented as provider-level allowances/credits from official documentation rather than mislabeled as universally free models.

Development rules

When adding new logic:

  • keep collectors/provider adapters separate from the generic model/timeline UI;
  • do not hardcode named models into generic components;
  • prefer explicit unknown values over guesses;
  • keep historical dates and observation dates separate;
  • add curated historical milestones only when there is a useful, source-backed reason;
  • keep user-facing copy concise and factual;
  • run npm test, npm run lint, and npm run validate before publishing.

Releases

Packages

Contributors

Languages