Skip to content

feat(proxy): use tenant catalog token limits for VS Code models - #608

Open
bohdanmaliar wants to merge 10 commits into
codemie-ai:mainfrom
bohdanmaliar:feat/vscode-model-token-limits
Open

bohdanmaliar wants to merge 10 commits into
codemie-ai:mainfrom
bohdanmaliar:feat/vscode-model-token-limits

Conversation

@bohdanmaliar

@bohdanmaliar bohdanmaliar commented Oct 5, 2026 •

Copy link
Copy Markdown

Summary

codemie proxy connect vscode writes the languageModels config for VS Code from a hardcoded capability table. Those token limits were static, so models whose limits differ per tenant got the wrong maxInputTokens / maxOutputTokens. This PR reads the limits from the tenant model catalog and uses the table only as a fallback.

Changes

  • Catalog parsing (tenant-catalog.ts): TenantModelDescriptor now carries maxInputTokens and maxOutputTokens from max_input_tokens / max_output_tokens. Values are kept only when they are finite numbers > 0. Zero, negative, string, null and missing values are dropped.
  • Limit resolution (vscode-models.ts): new resolveVsCodeTokenLimits(entry, descriptor).
    • Output limit: catalog value, else the table entry (or the default for untabled models).
    • Input limit: catalog input minus the resolved output limit. The catalog value is the full context window, whereas VS Code's maxInputTokens is a prompt budget, so input plus output has to fit.
    • If the catalog input is missing, or the subtraction is not positive, the entry's own maxInputTokens is used.
  • Config writer (vscode.ts): buildManagedModel takes the descriptor and applies the resolved limits to both table-matched and default (untabled) models.
  • sso.http-client.ts: small type tweak (LlmModel) so the catalog type can be reused.
  • docs/COMMANDS.md: documents the resolution order (tenant catalog, then built-in table, then defaults) and the input-minus-output rule.

Examples

Model Catalog Result (in / out)
claude-4-5-sonnet (tabled) input 200000 136000 / 64000
gpt-6-sol (untabled) input 922000 913808 / 8192
any input 200000, output 16000 184000 / 16000
claude-4-5-sonnet input 64000 entry input / 64000 (fallback, since 64000 - 64000 is not positive)
claude-4-5-sonnet no limits table values unchanged

Tests

Added unit tests for catalog parsing (valid and invalid values), resolveVsCodeTokenLimits (all cases above) and an end-to-end config-write test covering both tabled and untabled models.

The docs/superpowers/tasks/2026-10-02-vscode-model-token-limits/ files are planning artifacts from the SDLC run that produced this change.

Bohdan Maliar and others added 10 commits October 2, 2026 17:57
Refs: EPMCDME-15572

Generated with AI

Co-Authored-By: codemie-ai <codemie.ai@gmail.com>
Resolve each limit as tenant catalog, then built-in table, then defaults.

Refs: EPMCDME-15572

Generated with AI

Co-Authored-By: codemie-ai <codemie.ai@gmail.com>
Refs: EPMCDME-15572

Generated with AI

Co-Authored-By: codemie-ai <codemie.ai@gmail.com>
vscode-model-token-limits task 1

Generated with AI

Co-Authored-By: codemie-ai <codemie.ai@gmail.com>
vscode-model-token-limits task 3

Generated with AI

Co-Authored-By: codemie-ai <codemie.ai@gmail.com>
vscode-model-token-limits task 2

Generated with AI

Co-Authored-By: codemie-ai <codemie.ai@gmail.com>
…rule

Generated with AI

Co-Authored-By: codemie-ai <codemie.ai@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant