diff --git a/src/content/docs-lite/en/features.md b/src/content/docs-lite/en/features.md index 0ffeead..bb4dff6 100644 --- a/src/content/docs-lite/en/features.md +++ b/src/content/docs-lite/en/features.md @@ -2,7 +2,7 @@ A page-by-page reference to what ThinkWatch Lite shows and does. The [Lite page](/lite) gives the short version. -The main window has nine pages: Overview, Traffic, Clients, Keys, Upstreams, Routing, Security, MCP and Settings. +The main window has ten pages: Overview, Traffic, Clients, Keys, Upstreams, Routing, Security, MCP, Plugins and Settings. ## Usage and cost @@ -14,7 +14,7 @@ Until the first request has gone to an upstream, the Overview page shows a get-s ## Traffic and sessions -The Traffic page lists requests as they arrive: status, key, model, upstream, time to first token and total time (with the generation speed on hover), tokens and cost, with marks for a converted API format, redacted keys and a blocked or suspicious tool call. The list can be searched (path, key, upstream, model and error message), filtered by key, upstream and model, or narrowed to failed or unpriced requests. It holds the latest 2,000 requests, and a search or filter also runs over the stored history, further back on request. With content search on, it also covers what each request newly sent (the last user turn, tool results included) and its answer (tool calls included), for as long as payloads are kept; a match shows its excerpt under the request. The Sessions view groups the requests of one conversation into turns, with the input tokens and the cost of each turn. +The Traffic page lists requests as they arrive: status, key, model, upstream, time to first token and total time (with the generation speed on hover), tokens and cost, with marks for a converted API format, redacted keys and a blocked or suspicious tool call. The list can be searched (path, key, upstream, model and error message), filtered by key, upstream and model, or narrowed to failed or unpriced requests. It holds the latest 2,000 requests, and a search or filter also runs over the stored history, further back on request. With content search on, it also covers what each request newly sent (the last user turn, tool results included) and its answer (tool calls included), for as long as payloads are kept; a match shows its excerpt under the request. The Sessions view groups the requests of one conversation into turns, with the input tokens and the cost of each turn. A session's Conversation tab replays it turn by turn: the messages each turn added and its answer, with text, folded thinking, tool calls beside their results, and images by type and size. Changes to the system prompt and restarts after compaction are marked, and a turn whose bodies are past retention, were too large to keep whole or cannot be read says so. Each turn opens its request. A request opens into its timeline, its routing (the rule it matched, the group it went through and each attempt with its status and duration), the request and response bodies, and its usage and cost. A request from DeepSeek Harness also shows the size of the session log it carried, the whole conversation the client attaches to every request; the gateway removes it before a request goes to an upstream other than DeepSeek. A finished request can be sent again, unchanged, to another upstream after an estimate of its cost, and the two responses are shown side by side. @@ -73,6 +73,10 @@ The MCP page covers what clients load from their own configuration files, which The app watches these files while it runs, and a new finding raises a system notification. +## Plugins + +Plugins are short JavaScript files that change requests before they go to an upstream and answers before they reach the client, such as asking for answers in a chosen language or converting file paths in tool calls between WSL and Windows. They run after routing, in a sandbox inside core with no network, files or memory between requests, see placeholders instead of the keys in a request, and pass through the same protections afterwards. Two plugins ship with the app, both off until turned on. Each plugin is one file that also holds its scope, its behavior on errors and its settings. It is edited in one editor with a Settings tab, whose changes are written into the file, and a Code tab; adding a plugin opens the same editor. Routine changes are saved directly; for a plugin that may change tool calls, installing it, turning it on, changing its code and approving a changed file are confirmed in a system dialog. A plugin whose file changes outside the app stops running until the new version is approved. The page lists each plugin with its status, permissions, scope and statistics, and offers a trial run on a recent request and its log; requests changed by plugins are marked in Traffic. The API, permissions, limits and security model are described in [Plugins](/docs/lite/plugins). + ## Settings Settings has six sections. Connection lists the local core and the saved remote cores, described in [Connecting to a remote core](/docs/lite/remote-core). General sets the language, the appearance, what the menu bar item shows on macOS, launch at login, whether notices arrive as system notifications, in the app only or not at all, and shows hidden guidance hints again. Listening sets who can reach the gateway (this machine only, the local network of a chosen interface, or every interface), its port and the allowed address ranges. Log retention sets how long request payloads and request records are kept, and a size cap for payloads. About shows the version, checks for updates and produces a diagnostics bundle with keys and addresses masked. Uninstall restores every connected client and removes the autostart entry, and is meant to be run before the app is deleted. diff --git a/src/content/docs-lite/en/overview.md b/src/content/docs-lite/en/overview.md index 9f624be..59d3ec1 100644 --- a/src/content/docs-lite/en/overview.md +++ b/src/content/docs-lite/en/overview.md @@ -9,6 +9,7 @@ ThinkWatch Lite is a local gateway for Claude Code, Codex and other AI clients, - **Connect once, switch freely.** Twelve clients are pointed at the gateway in one step, with the change previewed and the original backed up; Cursor, Continue and Antigravity CLI come with instructions. - **Protection against relays.** A relay sees every request and can rewrite every answer. Outbound redaction can replace credentials, ID numbers and bank card numbers before a request leaves, and tool-call inspection can cut off an answer that carries a dangerous tool call, such as download-and-run or sending out credential files, before the client runs it. Hidden-character detection, a content filter and an output limit complete the five protections, each in Off, Observe or Enforce. - **MCP servers, skills and hooks, scanned.** The MCP servers of thirteen clients side by side, and a scan of client configuration for hidden characters, prompt injection, dangerous commands and overly broad permissions. +- **Plugins.** Short JavaScript plugins adjust requests and answers, such as asking for answers in a chosen language or converting file paths between WSL and Windows. They run in a sandbox, see placeholders instead of keys, and every change they make is recorded. - **Every request traceable.** The matched rule, each attempt, any format conversion and the cost, with replay against another upstream; the whole history can be searched, including the text of requests and answers. - **Routing and failover.** Rules by model, tools, images and more; groups that fail over before the answer begins and keep each session on one upstream. - **Any upstream.** API keys, Amazon Bedrock, ChatGPT and Z.ai accounts, relays and local models, with conversion between the Anthropic, OpenAI and Gemini APIs. @@ -26,6 +27,7 @@ ThinkWatch Lite is a local gateway for Claude Code, Codex and other AI clients, | Routing | Routes, rules and groups, auxiliary requests, and the dry run | | Security | The security log and the five protections with their rules | | MCP | MCP servers, skills and hooks in each client, and the configuration scan | +| Plugins | JavaScript plugins that change requests and answers, with their permissions, trial runs and logs | | Settings | Connection, language, appearance, menu bar, notifications, listening, retention, updates and uninstall | Each page is described in [Features](/docs/lite/features). @@ -34,6 +36,7 @@ Each page is described in [Features](/docs/lite/features). - [Install and update](/docs/lite/install) - [Connecting to a remote core](/docs/lite/remote-core): the gateway is [ThinkWatch Core](/docs/core), which runs beside the app or on a Linux server. +- [Plugins](/docs/lite/plugins): writing a plugin, its permissions and the sandbox it runs in. - [Architecture](/docs/lite/architecture): Lite holds no routing, forwarding or accounting logic; it controls Core over an encrypted control channel. - [Build from source](/docs/lite/run-from-source) diff --git a/src/content/docs-lite/en/plugins.md b/src/content/docs-lite/en/plugins.md new file mode 100644 index 0000000..fffa344 --- /dev/null +++ b/src/content/docs-lite/en/plugins.md @@ -0,0 +1,341 @@ +# Plugins + +Plugins adapt requests and answers to a particular setup: adding instructions to the system prompt, removing a parameter that one upstream rejects, asking for answers in a chosen language, or rewriting file paths in tool calls between WSL and Windows. A plugin is a short JavaScript file. It runs in a sandbox inside core, sees placeholders instead of the keys it would otherwise find, and every change it makes is recorded on the request and checked by the same protections as anything a client sends. + +This page covers what plugins can do, where they run, how one is added and edited, the API for writing one, permissions, limits and the security model. Two plugins ship with the app, both off by default; they are described at the end. + +## What a plugin can change + +| Side | What can be changed | +|---|---| +| Request, before it goes to an upstream | The system prompt; the messages, including tool results and the arguments of earlier tool calls; the tool definitions; the model, `max_tokens`, `temperature`, `top_p` and `stop`. A plugin can also refuse the request. | +| Answer, before it reaches the client | The text of the answer, as a whole block or as it streams; the tool calls in the answer, which can be changed, removed or added. | + +A plugin sees the same structure whatever API format the client uses (Anthropic Messages, OpenAI Chat Completions, OpenAI Responses or Gemini). By default a plugin handles conversations: requests that generate an answer, and the token counts and Responses compaction requests that carry the same conversation. A plugin that declares them also handles embeddings and legacy completions, where it can change the text of each input; see [Embeddings and legacy completions](#embeddings-and-legacy-completions). Other endpoints, such as images and audio, pass every plugin untouched. + +Plugins also run on the WebSocket connection Codex uses for the Responses API: each `response.create` goes through the request hooks, and each answer through the answer hooks. A connection keeps the plugins that were in place when it opened. Other WebSocket connections, such as the Realtime API, are not handled by plugins. + +Core writes the changes back in the client's own format and touches only the items that changed; cache markers, signatures, images and fields it does not know are kept. A request that no plugin changes is forwarded byte for byte, so the upstream's prompt cache is unaffected. + +Request headers, upstream addresses and credentials are not available to plugins. Images are passed as their media type only, without their data, and thinking blocks can be read but not changed. + +## Where plugins run + +A request is routed first, on what the client sent: the routing rules, a rule's model rename, failover groups, keeping a session on one upstream and the models the key may use all apply to the client's original request, and plugins cannot change where it goes. Then, for each attempt to send it to an upstream: + +1. Keys in the request are replaced with placeholders. +2. The request hooks in scope for this attempt run in list order, starting from the request as the client sent it. The placeholders are restored after them. +3. If a plugin changed the request, the content filter and the hidden-character check look at it again and report only what the plugins added. A block refuses the whole request; it is not tried on another upstream. +4. The request is converted to the upstream's format if needed, outbound redaction applies, and it is sent. + +When an attempt fails and the request moves to another upstream, the hooks run again from the request as the client sent it, so changes made for one upstream never reach another. A retry to the same upstream, such as the one after a sign-in token is refreshed, reuses what the hooks produced. + +A plugin that changes the model renames only what is sent to the upstream of this attempt, like a routing rule's rename, and replaces any name a rule set. The request is not routed again and the upstream's model list is not checked again, but the models the key may use still apply: a model outside them refuses the request. + +On the answer side, the hooks run after the answer is converted to the client's format and before the tool-call inspection and the output limit. + +## Adding and editing a plugin + +Each plugin is a single file that holds everything about it: the code, the requests it handles, what happens when it fails and the value of each setting. The file alone describes the plugin, so a copy of it carries the settings too. Only whether the plugin is on is kept outside the file, as a switch in the app. + +**Settings** on a plugin's row opens its editor, which has two tabs and a single **Save**: + +- **Settings** holds the **Enabled** switch, the scope (**Applies to**), the behavior on errors (**On error**) and the plugin's own settings. Apart from **Enabled**, changes made here are written into the manifest in the code: the app rewrites only the manifest and leaves every other byte of the file as it is. +- **Code** holds the whole file. When typing pauses, core reads the code again and the Settings tab follows it. Code that cannot be loaded is marked at the line and column of the error, and the Settings tab waits until it loads again. + +Changes on both tabs are kept while switching between them and saved together. + +**Add plugin** opens the same editor on the **Code** tab, with a small working plugin to start from. The code is edited or pasted there, or loaded with **Import from file…** at the top right, and **Install** installs the plugin. For a new plugin the Settings tab also has the **Plugin ID**: lowercase letters, digits and hyphens, up to 40, suggested from the file name or the plugin's name. The ID cannot be changed after installing. A new plugin is installed turned off unless **Enabled** is switched on first. Plugins are installed only from a local file or pasted code; there is no installation from a link, no plugin marketplace and no automatic update. + +On the Plugins page: + +- **Order.** Plugins run in the order of the list, which **Reorder** changes by dragging or with the arrows. Each plugin sees the result of the one before it and is checked against its own permissions. +- **Trial run.** Runs the plugin on a recent request from the history, with the routing that request had: the upstream that answered and the model sent to it. It shows the request or answer before and after, with the plugin's log. Nothing is sent to an upstream, and a trial run does not count in the statistics. +- **Logs.** The latest 500 lines the plugin wrote with `console`. +- **Statistics** since the gateway started: runs, changes, rejections, errors and the average CPU time. + +On the Traffic page, requests changed by a plugin carry a mark. A request's detail lists every plugin run, grouped by attempt, with its outcome (unchanged, changed, rejected, error or skipped) and its CPU time, and shows the request as the client sent it and as it was sent to the upstream that answered. A failing plugin raises a notification. + +When the app is connected to a [remote core](/docs/lite/remote-core), plugins are installed on the server and run there. + +### Confirmation in a system dialog + +Routine changes are saved directly. A system dialog is needed only for four steps, and only for a plugin that can change tool calls (one with the `reply.tool_calls` permission): installing it, turning it on, saving a change to its code, and approving a change made to its file outside the app. A change to the code counts when the permission is there before or after it, so adding `reply.tool_calls` is confirmed as well, and a plugin whose permissions cannot be read is treated as one that has it. + +The dialog is raised by the app itself, outside the page. It names the plugin, what the plugin can do and what is changing, and for new code the beginning of its SHA-256 hash, to compare with the one shown in the app. Core accepts these four steps only through the dialog, so a script injected into the page cannot take them on its own. + +No system dialog is needed for the settings, scope and behavior on errors of any plugin, for turning a plugin off, reordering or deleting, or for installing, turning on and editing a plugin that cannot change tool calls. Editing the config file in the app or restoring a version from its version history cannot install a plugin that can change tool calls, turn it on or change its code; such a change is refused there and made on the Plugins page. + +### When the file changes outside the app + +A plugin runs only as it was approved. Core keeps an approved copy of each plugin together with its SHA-256 hash, and a save in the app updates the file, the copy and the hash at once. When the file is changed or removed outside the app, the plugin stops running within seconds and its status shows **File changed**. **Review changes** shows the differences from the approved copy, and **Approve changes** lets the plugin run again; saving the plugin in its editor instead writes the editor's code back to the file. + +## Writing a plugin + +A plugin is a single ES module file in UTF-8, at most 1 MiB. It exports a `manifest`, which describes the plugin and holds its configuration, and one or more hooks. + +```js +export const manifest = { + name: "Project notes", + api: 1, + description: "Appends the notes set here to the system prompt.", + permissions: ["system"], + match: { clients: ["claude-code"], models: [], upstreams: [] }, + on_error: "reject", + settings: { + notes: { type: "string", label: "Notes", value: "Use pnpm, not npm." }, + }, +}; + +export function onRequest(req, ctx) { + const notes = String(ctx.settings.notes).trim(); + if (notes === "") return; // nothing to add: the request is sent as it is + req.system = req.system ? `${req.system}\n\n${notes}` : notes; + return req; +} +``` + +### Manifest + +| Field | Required | Value | +|---|---|---| +| `name` | Yes | 1 to 64 characters. | +| `api` | Yes | `1`, the only version supported. | +| `description` | No | Up to 500 characters. | +| `permissions` | Yes | One or more of `system`, `messages`, `tools`, `params`, `reply.text` and `reply.tool_calls`; see [Permissions](#permissions). | +| `requests` | No | The kinds of request the plugin handles: one or more of `conversation`, `embeddings` and `completions`. Without it, conversations only; see [Embeddings and legacy completions](#embeddings-and-legacy-completions). | +| `match` | No | The [scope](#scope): `clients`, `models` and `upstreams`, each a list of up to 100 patterns in which `*` matches any run of characters. A missing or empty list matches everything. Shown as **Applies to** on the Settings tab. | +| `on_error` | No | `"reject"` (the default) or `"skip"`: what happens when the plugin fails, described in [When a plugin fails](#when-a-plugin-fails). Shown as **On error** on the Settings tab. | +| `reply` | No | `"block"` (the default) or `"stream"`: how `onReplyText` receives text. | +| `settings` | No | Up to 20 entries, each named with letters, digits and `_` and not starting with a digit, with `type` (`string`, `number` or `boolean`), `label` (plain text, up to 100 characters) and `value`, the current value (`""`, `0` or `false` when left out). Values are edited on the Settings tab and passed in `ctx.settings`. A string value may span several lines, up to 10,000 characters. | + +Whether the plugin is on is not part of the manifest; it is the **Enabled** switch in the app. + +The manifest is plain data, which lets core read it straight from the source and the app rewrite it without running the code. It is an object literal whose values are strings, numbers, `true`, `false`, `null`, lists and nested objects; keys are names or quoted strings, and trailing commas are allowed. Expressions, variables, function calls, spreads, computed keys, getters and template literals with `${}` are not data: a file that uses them in the manifest does not load, and neither does one whose code changes the manifest after declaring it. + +A rewrite replaces only the manifest literal, and every byte outside it stays as it is. The literal is written back in one fixed style, the one used in the examples on this page, and comments inside it are not kept, so notes about the settings belong above the manifest. + +The file is checked when it is added and whenever core loads it. It must export a valid manifest, with no fields other than these, and at least one hook. Each exported hook needs its permission, and each permission must be used by an exported hook, so that a plugin requests nothing it does not use; `onReplyTextEnd` also needs `onReplyText` and stream mode. A file that fails any check is not installed, and the error names the line and column where it can. + +### Scope + +The scope is the manifest's `match`. It is edited under **Applies to** on the Settings tab, which suggests the clients, models and upstreams the gateway knows. + +| List | Matches | +|---|---| +| `clients` | The client the request came from, as recorded on the request, such as `claude-code`. A request from a client that is not recognized matches only an empty list. | +| `models` | The model sent to the upstream, after a routing rule renamed it. | +| `upstreams` | The upstream of the attempt. | + +All three lists apply to request hooks and answer hooks alike, and patterns ignore case. A request hook is in scope per attempt: after a failover, a plugin limited to one upstream runs only for the attempts that go there. The request hooks of an attempt are chosen before any of them runs, so a model renamed by one plugin does not change which request hooks run; answer hooks match the model the answering upstream received. A plugin also handles only the kinds of request it declares in `requests`. + +### Hooks + +| Hook | Permission | Called | +|---|---|---| +| `onRequest(req, ctx)` | `system`, `messages`, `tools` or `params` | Before each attempt to send the request to an upstream, after routing. | +| `onReplyText(text, ctx)` | `reply.text` | For the text of the answer. | +| `onReplyTextEnd(ctx)` | `reply.text` | In stream mode, at the end of each text block. Optional. | +| `onToolCall(call, ctx)` | `reply.tool_calls` | For each tool call in the answer, once it is complete. | + +**`onRequest`** receives the [request view](#the-request-view) and returns it changed, or returns nothing to leave the request as it is. Calling `reject("reason")` refuses the request, and the client receives an error that names the plugin; `reject` ends the hook by throwing, and the refusal stands even if the plugin catches what it throws. The hook runs once for each attempt to send the request to an upstream, as described in [Where plugins run](#where-plugins-run). Besides requests that generate an answer, it runs on token counts (`/v1/messages/count_tokens`, Gemini's `countTokens` and the Responses API's `input_tokens`) and on Responses compaction, which carry the same conversation, so that what a plugin removes is not sent through them either. On these, a changed model is applied but changes to the other `params` are not, because these endpoints do not accept them. A token count that the gateway estimates itself, without contacting an upstream, runs no plugin. + +**`onReplyText`** in block mode, the default, is called once for each text block with the whole text of the block, and the text reaches the client after the call. In stream mode it is called for each piece of streamed text and returns what to send now; returning `""` holds the text back, and `onReplyTextEnd` returns whatever is still held when the block ends. An answer that is not streamed is passed in one call, followed by `onReplyTextEnd` in stream mode. Returning nothing leaves the text unchanged. + +**`onToolCall`** receives `{ id, name, input }`. Streamed arguments are collected until the call is complete, and the result is sent in the client's format. The hook returns nothing to leave the call unchanged, an object to replace it, `null` to remove it, or an array to replace it with several calls. An `id` is generated where one is missing. + +Answer hooks run only on the answers to conversations. Thinking blocks are not passed to plugins and reach the client unchanged. + +A hook may be an `async` function or return a promise; it is settled within the same call and the same limits, and a promise that never settles counts as an error. + +### The request view + +```ts +type RequestView = { + format: "anthropic" | "openai_chat" | "openai_responses" | "gemini"; // read-only + model: string; // read-only; the model is changed through params.model + system?: string; // with "system"; "" when there is none + messages?: Message[]; // with "messages" + tools?: Tool[]; // with "tools" + params?: Params; // with "params" +}; +type Message = { key?: string; role: "user" | "assistant" | "tool" | "system"; parts: Part[] }; +type Part = + | { key?: string; type: "text"; text: string } + | { key?: string; type: "thinking"; text: string } // read-only + | { key?: string; type: "tool_call"; id: string; name: string; input: unknown } // input can be changed + | { key?: string; type: "tool_result"; call_id: string; text: string; is_error: boolean } // text can be changed + | { key?: string; type: "image"; media_type: string | null } // read-only, no data + | { key?: string; type: "other"; label: string }; // read-only +type Tool = { key?: string; name: string; description: string; input_schema: unknown }; +type Params = { model: string; max_tokens?: number; temperature?: number; top_p?: number; stop?: string[] }; +``` + +`model` and `params.model` hold the model that will be sent to the upstream of this attempt, the same as `ctx.model`. `system` is the instruction at the start of the conversation: the `system` field in Anthropic Messages, the leading system or developer messages in Chat Completions, `instructions` in Responses and `systemInstruction` in Gemini. Setting it to `""` removes it. Messages that carry only tool results have the role `tool`, and system or developer messages later in the conversation have the role `system`. + +Core assigns a `key` to every message, part and tool. The rules for changes: + +- A message, part or tool keeps its `key` to be changed and has no `key` when it is new. An unknown or repeated `key` is an error. +- Messages can be removed, changed or added. Added messages have the role `user`, `assistant` or `system` and contain text parts only. Messages that are kept keep their order and their role. +- Within a kept message, parts can be removed, their editable fields changed, and text parts added. `type`, `id`, `name` and `call_id` cannot be changed, and neither can thinking, image and other parts. +- Tools can be removed or added, and their `description` and `input_schema` changed. Tool names stay unique. The view lists the tools the client defines itself; tools whose definition comes from the API, such as web search, are not shown, are kept as they are, and a new tool cannot take their names. +- `max_tokens`, `temperature`, `top_p` and `stop` can be changed or removed, and `params.model` can be changed. A changed `params.model` renames the model sent to the upstream of this attempt, as described in [Where plugins run](#where-plugins-run). + +A result that breaks a rule, or that contains a section the plugin was not granted, counts as an error. + +### Embeddings and legacy completions + +A plugin that lists them in `requests` also handles embeddings (OpenAI `/v1/embeddings`, Gemini `:embedContent` and `:batchEmbedContents`) and legacy completions (OpenAI `/v1/completions`). Their view has the same shape, without a system prompt or tools: + +```ts +type InputsView = { + format: "openai_embeddings" | "openai_completions" | "gemini_embed"; // read-only + model: string; // read-only; the model is changed through params.model + messages?: Message[]; // with "messages": one user message per input + params?: Params; // with "params": the model; for completions also max_tokens, temperature, top_p, stop +}; +``` + +- Each input is one `user` message. For OpenAI, each element of an `input` (embeddings) or `prompt` (completions) array is one message, and a single string is one message; for Gemini, each content is one message, with one part for each of its parts. +- Only the text of text parts can be changed. An input given as token ids is a read-only `other` part labelled `tokens`, and other parts that are not text are read-only too. +- Messages and parts cannot be added, removed or reordered, because the answer comes back input by input. +- `params` holds the model, and for completions also `max_tokens`, `temperature`, `top_p` and `stop`. Other fields, such as `suffix` and `dimensions`, are not shown and stay as they are. + +Everything else works as for conversations: the placeholders, a run for each attempt after routing, a changed model renaming what is sent to this upstream, recording and trial runs. A changed request is checked again by the content filter and the hidden-character check, on the text of its inputs. Answer hooks do not run on these requests: an embeddings answer carries no text, and a legacy completions answer is passed through as it is. + +`messages` and `params` apply to all three kinds; `system`, `tools`, `reply.text` and `reply.tool_calls` apply to conversations only. A plugin does not load when `requests` is empty, names an unknown kind or names one twice, when a kind it declares is reached by none of its permissions (embeddings and completions need `messages` or `params`), or when it holds a permission that applies to none of its kinds, such as `system` without `conversation`. + +A request of a kind that a plugin does not declare is outside its scope. It passes untouched and nothing is recorded, whatever the plugin's behavior on errors, and even while the plugin's file has changed or it fails to load. + +### `ctx` + +```ts +type Ctx = { + client: string | null; // the client, as recorded on the request, e.g. "claude-code" + model: string; // the model sent to the upstream, after a routing rule or an earlier plugin renamed it + requested_model: string; // the model the client asked for + format: "anthropic" | "openai_chat" | "openai_responses" | "gemini" + | "openai_embeddings" | "openai_completions" | "gemini_embed"; // the format of the client's request + upstream: string; // the upstream of this attempt, or the one that served the answer + settings: Record; +}; +``` + +`ctx` is frozen and cannot be changed. + +### Available JavaScript + +Plugins have the standard JavaScript built-ins, such as `JSON`, `RegExp`, `Map`, `Date` and `Math`; `console.log`, `console.info`, `console.warn` and `console.error`, which write to the plugin's log; and `reject`, which is valid only in `onRequest`. There is no `fetch`, `require` or timer, `import` and `import()` load nothing, and there is no access to files, the network, environment variables or processes. The sandbox's clock is UTC. Nothing a plugin stores in a variable outlives the request attempt or answer it runs for. + +## Permissions + +A permission decides both what a plugin sees and what it may change. Sections that are not granted are not passed to the plugin, and a result that changes them counts as an error. + +| Permission | Sees and may change | Note in the app | +|---|---|---| +| `system` | The system prompt | | +| `messages` | Messages: text, tool results and the arguments of earlier tool calls; thinking is read-only and images are passed as their type only. For embeddings and legacy completions, the text of each input | Can add instructions to the conversation | +| `tools` | Tool definitions | Changes which tools the model can use | +| `params` | Model, `max_tokens`, `temperature`, `top_p`, `stop`; for embeddings, the model only | Can change the model sent to the upstream, and so what the request costs | +| `reply.text` | The text of answers | | +| `reply.tool_calls` | The tool calls in answers: change, remove, add | High risk, shown in red: a plugin can change what the client runs. The tool-call inspection still checks the result, and installing such a plugin, turning it on, changing its code and approving a change to its file are confirmed in a system dialog. | + +## When a plugin fails + +A plugin fails when it throws an error, exceeds a [limit](#limits), returns something that is not valid, or changes a section it was not granted. A request of a kind the plugin handles whose body cannot be read counts as a failure as well. The manifest's `on_error`, shown as **On error** on the Settings tab, sets what happens then: + +- **Reject the request** (the default): the request is refused, or the answer ends, with an error that names the plugin. A refused request is not tried on another upstream. +- **Skip this plugin**: the plugin is left out of this request, or of the rest of the answer, and the request continues as if it were not installed. + +The same choice applies while a plugin cannot run: when its file has changed and has not been approved, or when it fails to load. Its behavior on errors and its scope then come from the approved file, which core reads as data even when the code cannot run, so such a plugin refuses only the attempts within its scope, matched by request kind, client, model and upstream like a working one. A plugin that fails to load counts as handling conversations only, and one whose manifest cannot be read at all rejects every conversation. Every failure is recorded on the request. + +## Limits + +| Hook | CPU time | Memory | Output | +|---|---|---|---| +| Request | 200 ms per call, including the module's top-level code | 128 MiB | Twice the size of the request view, plus 1 MiB | +| Answer | 20 ms per call; 2 s for the whole answer | 64 MiB | Twice the size of the text or tool call passed in, plus 1 MiB; text released by `onReplyTextEnd`: 1 MiB | + +Each call can write 100 lines to the log, and a line longer than 4 KiB is cut short. The plugin file can be up to 1 MiB. A call that exceeds a limit, including a 101st log line, is stopped and counts as an error. + +Across the gateway, at most 32 answer-hook instances run at the same time. Each plugin with answer hooks takes one for every answer it handles and gives it back when the answer ends. When none is free, the plugin's behavior on errors decides: **Reject the request** refuses the request before any of the answer is sent, and **Skip this plugin** lets that answer through without the plugin. Either way the run is recorded as an error and raises a notification. + +## Security model + +- **A sandbox with nothing in it.** Plugins run in QuickJS compiled to WebAssembly and executed by Wasmtime inside core. Plugin code never runs in the app's window and never runs as native code. The sandbox has no network, files, environment variables or processes; even a flaw in the JavaScript engine reaches only the sandbox's own memory, not the keys and tokens in core's memory. +- **Nothing is kept.** Every request-hook call runs in a new instance. For each answer, a plugin gets one instance, shared by its hooks for that answer and discarded when the answer ends. Plugins share nothing with each other. +- **Placeholders instead of keys.** Keys that the outbound redaction rules recognize are replaced with placeholders before a plugin sees them and restored after it, on the request and on the answer, whatever mode the protection is in. The answer side matters as much as the request: answers become part of the conversation and are sent upstream again with the next request, so a plugin that could see a real key could hide it, encoded, in an answer. +- **The protections still apply.** Request hooks run after routing, and a request they change is checked again by the content filter and the hidden-character check before outbound redaction; answer hooks run before the tool-call inspection and the output limit. Whatever a plugin writes is checked like anything else, and plugins cannot change where a request is routed or switch to a model the key may not use. +- **Approved code only.** A plugin runs only while its file matches the approved SHA-256 hash, and a save in the app updates the file and the hash together. Installing a plugin that may change tool calls, turning it on, changing its code and approving a change to its file are confirmed in a system dialog outside the page, as described in [Confirmation in a system dialog](#confirmation-in-a-system-dialog). +- **Every change is visible.** Each run is recorded on the request, a changed request is stored as it was sent to the upstream that answered (with keys replaced, like every stored request), and requests changed by plugins are marked on the Traffic page. +- **Bounded.** Every call has limits on CPU time, memory, output and log volume, and the number of answer instances alive at once is capped. Plugins run in a separate thread pool, so a slow plugin does not hold up the gateway's own work. +- **Plain text.** The app shows a plugin's name, description, setting labels, logs and errors as plain text. + +## What the sandbox cannot prevent + +A plugin is allowed to change content, and the sandbox cannot judge whether a change is honest. A plugin with `system` or `messages` can write instructions into the prompt, and one with `reply.text` can make an answer misleading. A plugin with `params` can switch to a more expensive model among those the key may use. A plugin with `reply.tool_calls` can change what the client runs; the tool-call inspection catches only the patterns in its rules. Keys that the outbound redaction rules do not recognize are visible to plugins; they are also sent to the upstream as they are, and a custom rule on the Security page covers such a format. + +These risks are limited by granting a plugin only the permissions it needs, by reading its code before installing it, and by comparing the request before and after the plugins in the request's detail. + +## Plugins that ship with the app + +Two plugins come with the app. They appear in the plugin list like any other, both off, and are turned on and configured the same way; the app shows their names and setting labels in the interface language. One of them, `wsl-paths`, may change tool calls, so turning it on is confirmed in a system dialog, while its setting, scope and behavior on errors change without one. While no plugin is on, core does not start the sandbox, so plugins that stay off take no memory. The plugins ship with core, so a remote core has them too. Both handle conversations only, and their code is in the ThinkWatch Core repository, in [`crates/tw-gateway/src/plugin/defaults`](https://github.com/ThinkWatchProject/ThinkWatch-Core/tree/main/crates/tw-gateway/src/plugin/defaults). + +### Answer in a chosen language (`reply-language`) + +Adds a fixed line to the end of the system prompt that asks the model to answer in the chosen language, unless the user explicitly asks for another. The line is the same on every request, so the upstream's prompt cache keeps hitting. + +- Permission: `system`. +- Setting: `language`, the answer language, shipped as `简体中文` (Simplified Chinese). Only a language name is accepted (letters, spaces, parentheses and hyphens, up to 40 characters), so the setting cannot add other instructions; with any other value the plugin fails on every request. + +The manifest as shipped: + +```js +export const manifest = { + name: "Answer in a chosen language", + api: 1, + description: "Adds a fixed line to the end of the system prompt that asks the model to answer in the language set here.", + permissions: ["system"], + on_error: "reject", + settings: { + language: { type: "string", label: "Answer language", value: "简体中文" }, + }, +}; +``` + +### Convert WSL and Windows paths (`wsl-paths`) + +A client running in WSL cannot open `C:\Users\…`, and a client running on Windows cannot open `/mnt/c/Users/…`. The plugin rewrites every tool-call argument whose whole value is a drive path into the form the client can open, in the tool calls of answers and in earlier tool calls in the conversation, so the model keeps seeing one form. Paths inside command lines, the text of the conversation and tool results are left as they are. + +- Permissions: `messages`, `reply.tool_calls`. +- Setting: `windows_client`, whether the client runs on Windows; shipped as `false`, for a client in WSL. The conversion is fixed in the code, and the setting cannot express any other rewrite. + +The manifest as shipped: + +```js +export const manifest = { + name: "Convert WSL and Windows paths", + api: 1, + description: "Rewrites drive paths in tool-call arguments to the form the client can open (WSL /mnt/c/… or Windows C:\\…), in answers and in the conversation history.", + permissions: ["messages", "reply.tool_calls"], + on_error: "reject", + settings: { + windows_client: { + type: "boolean", + label: "The client runs on Windows (otherwise WSL)", + value: false, + }, + }, +}; +``` + +A built-in plugin that is deleted is not added again, and one whose code was edited is left as it is; changes made on its Settings tab do not count as editing the code. Otherwise, when an update ships a newer version of a built-in plugin, the plugin is updated and keeps its on/off state, its behavior on errors, its scope and the value of each setting the new version still declares with the same type; a new version that asks for more permissions or more kinds of request is turned off. + +## Upgrading + +The first launch after updating to the version that introduces plugins clears the request history, including stored requests and answers. The configuration, keys and upstreams are kept, and the two built-in plugins are added to the list, turned off. + +## Next steps + +- [Features](/docs/lite/features): the other pages of the app, including the Security page and its protections. +- [Connecting to a remote core](/docs/lite/remote-core): plugins on a core running on a server. diff --git a/src/content/docs-lite/zh-CN/features.md b/src/content/docs-lite/zh-CN/features.md index 45f5cce..65a2d85 100644 --- a/src/content/docs-lite/zh-CN/features.md +++ b/src/content/docs-lite/zh-CN/features.md @@ -2,7 +2,7 @@ 逐页说明 ThinkWatch Lite 展示的内容与提供的功能。简要介绍见 [Lite 产品页](/zh-CN/lite)。 -主窗口共有九个页面:概览、流量、客户端、密钥、上游、路由、安全、MCP 和设置。 +主窗口共有十个页面:概览、流量、客户端、密钥、上游、路由、安全、MCP、插件和设置。 ## 用量与费用 @@ -14,7 +14,7 @@ ## 流量与会话 -流量页实时列出请求:状态、密钥、模型、上游、首 token 时间与总耗时(悬停时显示生成速度)、token 和费用,并标出格式转换、被脱敏的密钥,以及被拦截或可疑的工具调用。列表可以搜索(路径、密钥、上游、模型与错误信息),可以按密钥、上游和模型筛选,或只看失败、无法计价的请求。列表装有最近 2,000 条请求,搜索与筛选同时在全部请求记录中进行,可以继续向更早的记录搜索。打开「搜索内容」后,还会搜索每个请求新发送的内容(最后一轮用户消息,含工具结果)及其回答(含工具调用),范围以报文仍保留的请求为限;命中的片段显示在对应请求的下方。「会话」视图把同一段对话的请求归为若干轮次,给出每一轮的输入 token 与费用。 +流量页实时列出请求:状态、密钥、模型、上游、首 token 时间与总耗时(悬停时显示生成速度)、token 和费用,并标出格式转换、被脱敏的密钥,以及被拦截或可疑的工具调用。列表可以搜索(路径、密钥、上游、模型与错误信息),可以按密钥、上游和模型筛选,或只看失败、无法计价的请求。列表装有最近 2,000 条请求,搜索与筛选同时在全部请求记录中进行,可以继续向更早的记录搜索。打开「搜索内容」后,还会搜索每个请求新发送的内容(最后一轮用户消息,含工具结果)及其回答(含工具调用),范围以报文仍保留的请求为限;命中的片段显示在对应请求的下方。「会话」视图把同一段对话的请求归为若干轮次,给出每一轮的输入 token 与费用。会话的「对话」页按轮还原整段对话:每一轮新加入的消息和回答,包括文字、折叠的思考、工具调用及其结果,图片只显示类型和大小。系统提示词的变化和压缩上下文之后的重新开始会标出;报文已过保留期、过大未能完整保存或无法读取的轮次会注明原因。每一轮都可以打开对应的请求。 打开一个请求可以查看时间线、路由(命中的规则、经过的策略组,以及每一次尝试的状态与耗时)、请求与响应正文、用量与费用。DeepSeek Harness 发出的请求还会显示所带会话日志的大小:这是客户端随每个请求附带的整段对话记录,发往 DeepSeek 以外的上游之前由网关去除。已结束的请求可以在预估费用后原样发送到另一个上游,两次的响应并排对照。 @@ -73,6 +73,10 @@ MCP 页管理客户端从自己的配置文件中加载的内容,这些内容 应用运行期间会监视这些文件,出现新的发现时发送系统通知。 +## 插件 + +插件是简短的 JavaScript 文件,在请求发往上游之前改写请求、在回答交给客户端之前改写回答,例如要求用指定的语言回答、在 WSL 与 Windows 之间转换工具调用里的路径。插件在路由之后、于 core 内部的沙箱中运行,没有网络和文件,两次请求之间不留任何数据;看到的是占位符而不是请求里的密钥,改动之后照常经过各项防护。应用自带两个插件,出厂均为关闭。每个插件是一个文件,适用范围、出错时的处理方式和设置项也写在其中;编辑器分为「设置」与「代码」两页,「设置」页的更改写回文件,添加插件时打开的也是同一个编辑器。日常的更改直接保存,只有能改动工具调用的插件,安装、打开、更改代码和确认文件变更才要在系统对话框中确认。文件在应用之外被改动后,插件停止运行,确认新版本后恢复。插件页列出每个插件的状态、权限、适用范围与统计,提供用最近一条请求试运行和日志;被插件改动过的请求在「流量」页带有标记。接口、权限、限额与安全模型见[插件](/zh-CN/docs/lite/plugins)。 + ## 设置 设置页分为六节。「连接」列出本机 core 和已保存的远程 core,详见[连接远程 core](/zh-CN/docs/lite/remote-core)。「通用」设置语言、外观、菜单栏显示的内容(仅 macOS)、开机启动、提醒以系统通知发送、仅在应用内显示还是关闭,并可让设为不再显示的引导提示重新显示。「网关监听」设置网关的访问范围(仅本机、所选网卡所在的局域网或所有网卡)、端口和放行网段。「日志保留」分别设置请求报文与请求记录的保留天数,以及报文的空间上限。「关于」显示版本、检查更新,并可生成诊断包,其中的密钥与地址均已脱敏。「卸载」还原所有已接管的客户端并取消开机启动,应在删除应用之前执行。 diff --git a/src/content/docs-lite/zh-CN/overview.md b/src/content/docs-lite/zh-CN/overview.md index bb3279a..49fa2fd 100644 --- a/src/content/docs-lite/zh-CN/overview.md +++ b/src/content/docs-lite/zh-CN/overview.md @@ -9,6 +9,7 @@ ThinkWatch Lite 是 Claude Code、Codex 等 AI 客户端的本地网关,支持 - **一次接入,随时切换。** 十二款客户端可一键指向网关,写入前预览改动并备份原文件;Cursor、Continue 与 Antigravity CLI 提供配置说明。 - **防范中转站。** 中转站能看到每个请求,也能改写每一次回答。出站脱敏可在请求发出前替换其中的凭据、身份证号与银行卡号;回答中出现下载即执行、外发凭据文件之类的危险工具调用时,工具调用审查可在客户端执行前切断回答。另有隐藏字符检测、内容过滤与输出长度限制,共五项防护,每项可设为关闭、观察或拦截。 - **扫描 MCP、技能与钩子。** 十三款客户端的 MCP 服务器并列显示,并扫描客户端配置中的隐藏字符、提示注入、危险命令与过宽权限。 +- **插件。** 用简短的 JavaScript 插件调整请求与回答,例如要求用指定的语言回答、在 WSL 与 Windows 之间转换路径。插件在沙箱中运行,只看到占位符而看不到密钥,每一处改动都有记录。 - **每个请求都可追溯。** 命中的规则、每一次尝试、格式转换与费用都有记录,也可以重放到另一个上游对比;全部请求记录都可以搜索,包括请求与回答的内容。 - **路由与故障转移。** 按模型、工具、图片等条件分流;回答开始前上游出错时换用下一个,同一会话固定使用同一上游。 - **多种上游。** API 密钥、Amazon Bedrock、ChatGPT 与 Z.ai 账号、中转服务与本机模型,Anthropic、OpenAI、Gemini 接口之间自动转换。 @@ -26,6 +27,7 @@ ThinkWatch Lite 是 Claude Code、Codex 等 AI 客户端的本地网关,支持 | 路由 | 路由、规则与策略组,辅助请求,以及试算 | | 安全 | 安全日志,以及五项防护及其规则 | | MCP | 各客户端的 MCP 服务器、技能与钩子,以及配置扫描 | +| 插件 | 改写请求与回答的 JavaScript 插件,及其权限、试运行与日志 | | 设置 | 连接、语言、外观、菜单栏、提醒、网关监听、日志保留、更新与卸载 | 各页的详细说明见[功能详解](/zh-CN/docs/lite/features)。 @@ -34,6 +36,7 @@ ThinkWatch Lite 是 Claude Code、Codex 等 AI 客户端的本地网关,支持 - [安装与更新](/zh-CN/docs/lite/install) - [连接远程 core](/zh-CN/docs/lite/remote-core):网关本体是 [ThinkWatch Core](/zh-CN/docs/core),随应用在本机运行,也可以部署在 Linux 服务器上。 +- [插件](/zh-CN/docs/lite/plugins):编写插件、插件的权限及其运行的沙箱。 - [架构](/zh-CN/docs/lite/architecture):Lite 不含路由、转发与计费逻辑,通过加密的控制通道控制 Core。 - [从源码构建](/zh-CN/docs/lite/run-from-source) diff --git a/src/content/docs-lite/zh-CN/plugins.md b/src/content/docs-lite/zh-CN/plugins.md new file mode 100644 index 0000000..1220ee9 --- /dev/null +++ b/src/content/docs-lite/zh-CN/plugins.md @@ -0,0 +1,341 @@ +# 插件 + +插件让请求和回答适应具体的使用场景:在系统提示词中附加说明、去掉某个上游不接受的参数、要求用指定的语言回答,或者在 WSL 与 Windows 之间改写工具调用里的路径。插件是一段简短的 JavaScript,在 core 内部的沙箱中运行,看到的是占位符而不是请求里的密钥;它做的每一处改动都记录在请求上,并和客户端发来的内容一样经过各项防护的检查。 + +本页说明插件能改什么、在哪一步运行、如何添加和编辑、编写插件的接口、权限、限额与安全模型。应用自带两个插件,出厂均为关闭,末尾逐一说明。 + +## 插件能改什么 + +| 方向 | 可以改动的内容 | +|---|---| +| 请求,发往上游之前 | 系统提示词;对话消息,含工具结果和此前工具调用的参数;工具定义;模型、`max_tokens`、`temperature`、`top_p` 与 `stop`。插件也可以拒绝这次请求。 | +| 回答,交给客户端之前 | 回答的文字,可以整段处理,也可以随流式输出逐段处理;回答里的工具调用,可以修改、删除或新增。 | + +无论客户端使用哪种接口格式(Anthropic Messages、OpenAI Chat Completions、OpenAI Responses 或 Gemini),插件看到的都是同一种结构。插件默认处理对话:生成回答的请求,以及携带同一段对话的计算 token 数请求与 Responses 压缩请求。声明了它们的插件还处理嵌入与旧版补全,可以修改每一项输入的文字,见[嵌入与旧版补全](#嵌入与旧版补全)。其他接口(例如图片、音频)不经过任何插件,原样转发。 + +Codex 通过 WebSocket 使用 Responses 接口时,插件同样生效:每个 `response.create` 经过请求钩子,每次回答经过回答钩子。一条连接沿用它建立时的插件。其他 WebSocket 连接(例如 Realtime 接口)不经过插件。 + +core 按客户端自己的格式写回改动,只动改过的条目,缓存标记、签名、图片和不认识的字段都保持原样。没有被任何插件改动的请求逐字节原样转发,不影响上游的提示词缓存。 + +请求头、上游地址与凭据不对插件开放。图片只提供媒体类型,不提供内容;思考内容只读,不能修改。 + +## 插件在哪一步运行 + +请求先按客户端发来的原样路由:路由规则、规则对模型的改名、策略组的故障转移、同一会话固定使用同一上游,以及密钥的可用模型,依据的都是客户端的原始请求,插件改变不了请求的去向。之后,每次把请求发往一个上游之前: + +1. 把请求里的密钥替换为占位符; +2. 从客户端发来的原始请求开始,按列表顺序运行适用范围覆盖这一次发送的请求钩子,运行之后换回占位符; +3. 插件改动了请求时,内容过滤和隐藏字符检测再检查一遍,只报告插件加入的内容;拦截时整个请求被拒绝,不会转到别的上游; +4. 需要时转换为上游的格式,经过出站脱敏后发出。 + +一次发送失败、请求转到另一个上游时,请求钩子从客户端的原始请求重新运行,为一个上游做的改动不会带到另一个上游。向同一个上游重发(例如登录令牌刷新后的重试)时,沿用请求钩子已有的结果。 + +插件修改模型时,只改变这一次发往上游的模型名,效果与路由规则改名相同,并取代规则改过的名字。请求不会重新路由,也不再核对上游的模型列表,但密钥的可用模型依然有效:改成密钥不可用的模型时,请求被拒绝。 + +回答一侧,回答钩子在回答转换为客户端的格式之后、工具调用审查和输出长度限制之前运行。 + +## 添加与编辑插件 + +每个插件是一个文件,插件的一切都写在其中:代码、处理哪些请求、出错时如何处理,以及每个设置项的取值。单凭这个文件就能完整描述插件,复制文件时设置也随之带走。只有插件是否启用不写在文件里,它是应用中的一个开关。 + +在插件所在行选择「设置」即打开它的编辑器。编辑器分为两页,只有一个「保存」: + +- **设置**:「启用」开关、「适用范围」、「出错时」的处理方式,以及插件自己的设置项。除「启用」外,这一页的更改都写回代码中的 manifest:应用只改写 manifest,文件中的其他字节保持原样。 +- **代码**:完整的文件内容。输入停顿时,core 重新读取代码,「设置」页随之更新。代码无法加载时,标出出错的行号与列号,「设置」页等到代码能够加载之后才能修改。 + +在两页之间切换时,未保存的更改都会保留,保存时一并写入。 + +「添加插件」打开同一个编辑器,从「代码」页开始,里面是一个可以直接安装的简单插件。在这里修改或粘贴代码,或者用右上角的「从文件导入…」载入文件,选择「安装」即安装插件。新插件的「设置」页多一项「插件 ID」:小写字母、数字和连字符,最多 40 个,按文件名或插件名预先填好,安装后不可修改。新插件安装后处于关闭状态,除非安装前打开了「启用」。插件只能从本地文件或粘贴的代码安装,不支持通过链接安装,没有插件市场,也不会自动更新。 + +「插件」页还提供: + +- **顺序**:插件按列表顺序依次运行,可在「调整顺序」中拖动或用箭头更改。后一个看到的是前一个改过的结果,各自按自己的权限核对。 +- **试运行**:用请求历史中最近的一条请求运行插件,按这条请求当时的路由(作答的上游和发给它的模型名),显示改动前后的请求或回答以及插件日志。试运行不发往上游,也不计入统计。 +- **日志**:插件通过 `console` 写下的最近 500 行。 +- **统计**:网关启动以来的运行次数、改写次数、拒绝次数、出错次数和平均 CPU 时间。 + +「流量」页中,被插件改动过的请求带有标记。请求详情按每一次发送分组列出运行过的插件、各自的结果(未改动、已改写、已拒绝、出错或已跳过)与 CPU 时间,并可以对照客户端发来的原始请求和发往作答上游的请求。插件出错时发送提醒。 + +应用连接[远程 core](/zh-CN/docs/lite/remote-core) 时,插件安装在服务器上,也在服务器上运行。 + +### 在系统对话框中确认 + +日常的更改直接保存。只有能改动工具调用的插件(拥有 `reply.tool_calls` 权限)需要在系统对话框中确认,而且只限四种操作:安装、打开、保存对代码的更改,以及确认在应用之外对文件所做的更改。代码更改前后任何一版拥有这项权限,都要确认,因此新增 `reply.tool_calls` 同样要确认;读不出权限的插件按拥有这项权限对待。 + +系统对话框由应用自身弹出,不属于页面。对话框写明插件名称、它能做什么和这次改动的内容;涉及新代码时,还给出代码 SHA-256 哈希的开头部分,可与应用中显示的对照。core 只接受经过系统对话框的这四种操作,注入页面的脚本无法自行完成。 + +以下更改都无须系统对话框:任何插件的设置项、适用范围和出错时的处理方式,关闭插件、调整顺序和删除,以及安装、打开和编辑不能改动工具调用的插件。在应用中编辑配置文件,或者从版本历史恢复某个版本,都不能安装能改动工具调用的插件、打开它或更改它的代码;这类更改在那里会被拒绝,须在「插件」页进行。 + +### 文件在应用之外被改动时 + +插件只按确认过的样子运行。core 为每个插件保留一份确认过的副本及其 SHA-256 哈希,在应用中保存时,文件、副本与哈希一同更新。文件在应用之外被改动或删除后,插件在几秒内停止运行,状态显示为「文件已更改」。「审核更改」列出与确认过的副本之间的差异,选择「确认更改」后插件恢复运行;也可以在插件的编辑器中保存,用编辑器里的代码写回文件。 + +## 编写插件 + +插件是一个 UTF-8 编码的 ES 模块文件,不超过 1 MiB,导出 `manifest` 和一个或多个钩子函数;manifest 描述插件,同时保存它的配置。 + +```js +export const manifest = { + name: "附加项目说明", + api: 1, + description: "在系统提示词末尾附上这里填写的说明。", + permissions: ["system"], + match: { clients: ["claude-code"], models: [], upstreams: [] }, + on_error: "reject", + settings: { + notes: { type: "string", label: "说明内容", value: "使用 pnpm,不用 npm。" }, + }, +}; + +export function onRequest(req, ctx) { + const notes = String(ctx.settings.notes).trim(); + if (notes === "") return; // 没有要附加的内容:请求原样发出 + req.system = req.system ? `${req.system}\n\n${notes}` : notes; + return req; +} +``` + +### manifest + +| 字段 | 必填 | 取值 | +|---|---|---| +| `name` | 是 | 1 至 64 个字符。 | +| `api` | 是 | `1`,目前唯一支持的版本。 | +| `description` | 否 | 不超过 500 个字符。 | +| `permissions` | 是 | `system`、`messages`、`tools`、`params`、`reply.text`、`reply.tool_calls` 中的一项或多项,见[权限](#权限)。 | +| `requests` | 否 | 插件处理的请求种类:`conversation`(对话)、`embeddings`(嵌入)、`completions`(旧版补全)中的一项或多项。缺省时只处理对话,见[嵌入与旧版补全](#嵌入与旧版补全)。 | +| `match` | 否 | [适用范围](#适用范围):`clients`、`models`、`upstreams` 三个列表,每个最多 100 项,`*` 匹配任意一串字符;列表缺省或为空表示全部。对应「设置」页的「适用范围」。 | +| `on_error` | 否 | `"reject"`(默认)或 `"skip"`:插件出错时如何处理,见[插件出错时](#插件出错时)。对应「设置」页的「出错时」。 | +| `reply` | 否 | `"block"`(默认)或 `"stream"`,决定 `onReplyText` 以何种方式接收文字。 | +| `settings` | 否 | 最多 20 项,名称由字母、数字和 `_` 组成,不以数字开头,每项包含 `type`(`string`、`number` 或 `boolean`)、`label`(纯文本,不超过 100 个字符)与 `value`,即当前的取值(缺省时为 `""`、`0` 或 `false`)。取值在「设置」页填写,通过 `ctx.settings` 传给插件。字符串的取值可以有多行,不超过 10,000 个字符。 | + +插件是否启用不在 manifest 中,它是应用中的「启用」开关。 + +manifest 是纯数据,因此 core 可以直接从源码中读出它,应用也无须运行代码就能改写它。它是一个对象字面量,取值只能是字符串、数字、`true`、`false`、`null`、列表和嵌套的对象;键写成名称或带引号的字符串,允许结尾的逗号。表达式、变量、函数调用、展开、计算出的键、getter,以及带 `${}` 的模板字符串都不是数据:manifest 中用到它们的文件无法加载,声明之后又在代码中改动 manifest 的文件同样无法加载。 + +改写只替换 manifest 字面量,其余每个字节都保持原样。字面量按固定的格式写回,即本页示例所用的格式;其中的注释不会保留,关于设置项的说明应写在 manifest 上方。 + +文件在添加时和每次被 core 加载时都会检查:必须导出有效的 manifest(不能有上表以外的字段)和至少一个钩子。导出的每个钩子都要有对应的权限,申请的每项权限也都要有用到它的钩子,插件不会申请用不到的权限;`onReplyTextEnd` 还要求同时导出 `onReplyText` 并使用逐段模式。未通过检查的文件不会安装,错误信息尽可能给出行号和列号。 + +### 适用范围 + +适用范围即 manifest 中的 `match`,在「设置」页的「适用范围」中编辑,编辑时会提示网关已知的客户端、模型和上游。 + +| 列表 | 匹配的对象 | +|---|---| +| `clients` | 请求来自的客户端,即请求记录上的客户端,例如 `claude-code`。识别不出客户端的请求只与空列表匹配。 | +| `models` | 发往上游的模型名,即路由规则改名之后的名字。 | +| `upstreams` | 这一次发送的上游。 | + +三个列表对请求钩子和回答钩子都适用,匹配时不区分大小写。请求钩子按每一次发送判断是否在适用范围内:发生故障转移时,限定了上游的插件只在发往这些上游时运行。一次发送要运行哪些请求钩子,在其中任何一个运行之前就已确定,前面的插件改了模型名,不影响后面运行哪些请求钩子;回答钩子按作答上游实际收到的模型名匹配。此外,插件只处理它在 `requests` 中声明的请求种类。 + +### 钩子 + +| 钩子 | 权限 | 调用时机 | +|---|---|---| +| `onRequest(req, ctx)` | `system`、`messages`、`tools` 或 `params` | 路由之后,每次把请求发往一个上游之前。 | +| `onReplyText(text, ctx)` | `reply.text` | 处理回答中的文字。 | +| `onReplyTextEnd(ctx)` | `reply.text` | 逐段模式下,每段文字结束时调用。可选。 | +| `onToolCall(call, ctx)` | `reply.tool_calls` | 回答中的每个工具调用完整之后调用。 | + +**`onRequest`** 接收[请求视图](#请求视图),返回改过的视图;不返回则请求保持原样。调用 `reject("原因")` 拒绝这次请求,客户端收到注明插件名称的错误;`reject` 通过抛出异常结束钩子,即使插件自己接住这个异常,拒绝依然成立。每次把请求发往一个上游之前调用一次,详见[插件在哪一步运行](#插件在哪一步运行)。除了生成回答的请求,计算 token 数的请求(`/v1/messages/count_tokens`、Gemini 的 `countTokens` 与 Responses 接口的 `input_tokens`)和 Responses 的压缩请求携带的是同一段对话,也经过它,插件删去的内容不会经由这些接口发出。这些请求只采用插件对模型的修改,不采用对其他 `params` 的修改,因为这些接口不接受那些参数。网关自行估算、不发往上游的 token 计数不运行插件。 + +**`onReplyText`** 在整段模式(默认)下,每段文字到齐后调用一次,参数是整段文字,调用结束后文字才交给客户端。逐段模式下随流式输出逐段调用,返回值是此刻要发出的文字;返回 `""` 表示先扣住,这段文字结束时 `onReplyTextEnd` 的返回值把扣住的内容放出。非流式的回答一次传入全部文字,逐段模式下随后再调用 `onReplyTextEnd`。不返回表示文字不变。 + +**`onToolCall`** 接收 `{ id, name, input }`。流式输出的参数先收集完整,处理后按客户端的格式发出。不返回表示保持原样;返回一个对象替换这个调用;返回 `null` 删除它;返回一个数组则替换为多个调用。缺少 `id` 的由 core 生成。 + +回答钩子只在对话的回答上运行。思考内容不交给插件,原样交给客户端。 + +钩子可以是 `async` 函数,也可以返回 Promise:Promise 在同一次调用内、同样的限额之下落定;始终不落定的 Promise 按出错处理。 + +### 请求视图 + +```ts +type RequestView = { + format: "anthropic" | "openai_chat" | "openai_responses" | "gemini"; // 只读 + model: string; // 只读;修改模型要通过 params.model + system?: string; // 授予 "system" 时提供;没有系统提示词时为 "" + messages?: Message[]; // 授予 "messages" 时提供 + tools?: Tool[]; // 授予 "tools" 时提供 + params?: Params; // 授予 "params" 时提供 +}; +type Message = { key?: string; role: "user" | "assistant" | "tool" | "system"; parts: Part[] }; +type Part = + | { key?: string; type: "text"; text: string } + | { key?: string; type: "thinking"; text: string } // 只读 + | { key?: string; type: "tool_call"; id: string; name: string; input: unknown } // input 可改 + | { key?: string; type: "tool_result"; call_id: string; text: string; is_error: boolean } // text 可改 + | { key?: string; type: "image"; media_type: string | null } // 只读,不含图片数据 + | { key?: string; type: "other"; label: string }; // 只读 +type Tool = { key?: string; name: string; description: string; input_schema: unknown }; +type Params = { model: string; max_tokens?: number; temperature?: number; top_p?: number; stop?: string[] }; +``` + +`model` 与 `params.model` 是这一次将要发往上游的模型名,与 `ctx.model` 相同。`system` 是对话开头的系统指令:Anthropic Messages 的 `system` 字段、Chat Completions 开头的 system 或 developer 消息、Responses 的 `instructions`、Gemini 的 `systemInstruction`。设为 `""` 即删除。只含工具结果的消息角色为 `tool`,对话中途出现的 system 或 developer 消息角色为 `system`。 + +core 为每条消息、每个片段和每个工具分配一个 `key`。改动规则如下: + +- 要修改的消息、片段或工具保留原来的 `key`,新增的不带 `key`。出现未知或重复的 `key` 按出错处理。 +- 消息可以删除、修改或新增。新增的消息角色只能是 `user`、`assistant` 或 `system`,且只含文字片段。保留下来的消息保持原有的先后顺序和角色。 +- 保留下来的消息里,片段可以删除,可编辑的字段可以修改,也可以新增文字片段。`type`、`id`、`name` 与 `call_id` 不可修改,思考、图片和其他片段整体不可修改。 +- 工具可以删除或新增,`description` 与 `input_schema` 可以修改,工具名称不能重复。视图只列出客户端自己定义的工具;由接口提供定义的工具(例如网页搜索)不在视图中,保持原样,新增的工具也不能与它们重名。 +- `max_tokens`、`temperature`、`top_p` 与 `stop` 可以修改或删除,`params.model` 可以修改。修改 `params.model` 会改变这一次发往上游的模型名,见[插件在哪一步运行](#插件在哪一步运行)。 + +违反规则的结果,或者带有未授权部分的结果,都按出错处理。 + +### 嵌入与旧版补全 + +在 `requests` 中声明了它们的插件,还处理嵌入(OpenAI 的 `/v1/embeddings`,Gemini 的 `:embedContent` 与 `:batchEmbedContents`)和旧版补全(OpenAI 的 `/v1/completions`)。这两种请求的视图结构相同,只是没有系统提示词和工具: + +```ts +type InputsView = { + format: "openai_embeddings" | "openai_completions" | "gemini_embed"; // 只读 + model: string; // 只读;修改模型要通过 params.model + messages?: Message[]; // 授予 "messages" 时提供:每项输入一条 user 消息 + params?: Params; // 授予 "params" 时提供:模型;补全另有 max_tokens、temperature、top_p、stop +}; +``` + +- 每项输入是一条 `user` 消息。OpenAI 的 `input`(嵌入)或 `prompt`(补全)是数组时,每个元素一条,是单个字符串时就是一条;Gemini 的每个 content 一条,content 的每个部分对应一个片段。 +- 只有文字片段的文字可以修改。以 token ID 给出的输入是只读的 `other` 片段,`label` 为 `tokens`;其他不是文字的部分同样只读。 +- 消息和片段不能新增、删除或调换顺序,因为回答按输入逐项对应。 +- `params` 中有模型,补全另有 `max_tokens`、`temperature`、`top_p` 与 `stop`。`suffix`、`dimensions` 等其他字段不提供给插件,保持原样。 + +其余与对话相同:占位符、路由之后按每一次发送运行、修改模型只改变这一次发往上游的模型名、运行记录与试运行。改动过的请求按各项输入的文字,再经过一次内容过滤和隐藏字符检测。这两种请求不运行回答钩子:嵌入的回答不含文字,旧版补全的回答原样转发。 + +`messages` 与 `params` 适用于全部三种请求;`system`、`tools`、`reply.text` 与 `reply.tool_calls` 只适用于对话。出现以下情况时插件无法加载:`requests` 为空、含有未知的种类或同一种类写了两次;声明的某种请求没有任何一项权限用得上(嵌入和旧版补全需要 `messages` 或 `params`);或者某项权限不适用于声明的任何一种请求,例如申请了 `system` 却没有声明 `conversation`。 + +插件没有声明的请求种类不在它的适用范围内:这类请求原样通过,不留记录;无论出错时的处理方式如何,即使插件文件已更改或加载失败,也不会因它被拒绝。 + +### ctx + +```ts +type Ctx = { + client: string | null; // 请求记录上的客户端,例如 "claude-code" + model: string; // 发往上游的模型名,即路由规则或前面的插件改名之后的名字 + requested_model: string; // 客户端请求的模型 + format: "anthropic" | "openai_chat" | "openai_responses" | "gemini" + | "openai_embeddings" | "openai_completions" | "gemini_embed"; // 客户端请求的格式 + upstream: string; // 这一次发送的上游,或作答的上游 + settings: Record; +}; +``` + +`ctx` 已冻结,不能修改。 + +### 可用的 JavaScript + +插件可以使用 JavaScript 标准内置对象,例如 `JSON`、`RegExp`、`Map`、`Date` 与 `Math`;`console.log`、`console.info`、`console.warn` 与 `console.error` 写入插件日志;`reject` 只在 `onRequest` 中有效。没有 `fetch`、`require` 和定时器,`import` 与 `import()` 加载不了任何模块,也无法访问文件、网络、环境变量和进程。沙箱里的时钟是 UTC。插件存进变量的内容,不会留到它所处理的这一次发送或这个回答之后。 + +## 权限 + +权限同时决定插件看得到什么和能改什么。未授权的部分不会交给插件;返回的结果改动了未授权的部分,按出错处理。 + +| 权限 | 可以看到并修改 | 应用中的提示 | +|---|---|---| +| `system` | 系统提示词 | | +| `messages` | 对话消息:文字、工具结果和此前工具调用的参数;思考内容只读,图片只提供类型。嵌入与旧版补全:每项输入的文字 | 可以往对话中加入指令 | +| `tools` | 工具定义 | 会改变模型能使用的工具 | +| `params` | 模型、`max_tokens`、`temperature`、`top_p`、`stop`;嵌入只有模型 | 可能改变发往上游的模型,以及因此产生的费用 | +| `reply.text` | 回答中的文字 | | +| `reply.tool_calls` | 回答中的工具调用:修改、删除、新增 | 高风险,以红字提示:插件可以改动客户端要执行的操作。改动后的工具调用仍经过工具调用审查;安装这类插件、打开它、更改它的代码和确认它的文件变更,都要在系统对话框中确认。 | + +## 插件出错时 + +插件抛出错误、超出[限额](#限额)、返回无效的结果,或者改动了未授权的部分,都算出错;插件处理的那种请求,请求体无法读取时也算出错。出错时的处理方式由 manifest 的 `on_error` 决定,即「设置」页的「出错时」: + +- **拒绝这次请求**(默认):请求被拒绝,或回答在此处结束,并附上注明插件名称的错误。被拒绝的请求不会转到别的上游重试。 +- **跳过此插件**:这次请求(或这个回答余下的部分)不经过该插件,按未安装它的情况继续。 + +插件无法运行时(文件已更改而尚未确认,或者插件加载失败),按同一设置处理。这时出错时的处理方式和适用范围取自确认过的文件,即使代码无法运行,core 也能把它们当作数据读出;因此这样的插件与正常的插件一样,按每一次发送的请求种类、客户端、模型和上游判断是否在适用范围内,只拒绝范围内的发送。加载失败的插件按只处理对话对待,完全读不出 manifest 的插件拒绝所有对话。每次出错都记录在请求上。 + +## 限额 + +| 钩子 | CPU 时间 | 内存 | 输出 | +|---|---|---|---| +| 请求钩子 | 每次 200 毫秒,含模块顶层代码 | 128 MiB | 不超过请求视图大小的两倍加 1 MiB | +| 回答钩子 | 每次 20 毫秒,整个回答累计 2 秒 | 64 MiB | 不超过传入的文字或工具调用大小的两倍加 1 MiB;`onReplyTextEnd` 放出的文字不超过 1 MiB | + +每次调用最多写 100 行日志,超过 4 KiB 的行截短。插件文件不超过 1 MiB。超出限额的调用(包括写第 101 行日志)立即中止,按出错处理。 + +整个网关同时运行的回答钩子实例最多 32 个。带回答钩子的插件每处理一个回答占用一个,回答结束时释放。没有空余时,按该插件出错时的处理方式:「拒绝这次请求」在回答的任何内容发出之前拒绝请求,「跳过此插件」让这次回答不经过该插件。两种情况都记为一次出错,并发送提醒。 + +## 安全模型 + +- **沙箱里什么都没有。** 插件在 QuickJS 中运行,QuickJS 编译成 WebAssembly,由 core 内部的 Wasmtime 执行。插件代码从不在应用窗口中运行,也从不以原生代码运行。沙箱里没有网络、文件、环境变量和进程;即使 JavaScript 引擎本身有漏洞,也只能破坏沙箱自己的内存,碰不到 core 进程中的密钥与令牌。 +- **不留任何数据。** 每次调用请求钩子都使用新实例。每个回答中,一个插件使用一个实例,由它在这个回答上的各个钩子共用,回答结束即丢弃。插件之间不共享任何东西。 +- **只看到占位符。** 出站脱敏规则能识别的密钥,在交给插件之前替换为占位符,插件处理之后再换回;请求与回答两个方向都是如此,与这项防护处于哪一档无关。回答一侧同样重要:回答会进入对话历史,随下一次请求再次发往上游;插件如果看得到真实的密钥,就能把它编码后藏进回答里带出去。 +- **防护照常生效。** 请求钩子在路由之后运行,被它改动的请求在出站脱敏之前再经过内容过滤和隐藏字符检测;回答钩子在工具调用审查和输出长度限制之前运行。插件写入的内容和其他内容一样受检,插件既改变不了请求的路由,也不能换用密钥不可用的模型。 +- **只运行确认过的代码。** 插件文件与确认过的 SHA-256 哈希一致才会运行,在应用中保存时,文件与哈希一同更新。能改动工具调用的插件,安装、打开、更改代码和确认文件变更,都要在页面之外的系统对话框中确认,见[在系统对话框中确认](#在系统对话框中确认)。 +- **改动都有记录。** 每次运行都记录在请求上;被改动的请求另存一份发往作答上游的版本,和所有存下的请求一样替换掉了密钥;「流量」页为被插件改动的请求加上标记。 +- **用量有上限。** 每次调用都有 CPU 时间、内存、输出和日志量的上限,同时存在的回答实例也有数量上限。插件在单独的线程池中运行,运行缓慢的插件不会拖住网关本身的工作。 +- **只显示纯文本。** 插件的名称、说明、设置项标签、日志与错误信息,在应用中一律按纯文本显示。 + +## 沙箱无法防范的情况 + +插件本来就有权改动内容,沙箱无法判断一处改动是否正当。拥有 `system` 或 `messages` 权限的插件可以往提示词里写入指令;拥有 `reply.text` 的插件可以把回答改得误导人;拥有 `params` 的插件可以在密钥的可用模型中换用更贵的模型;拥有 `reply.tool_calls` 的插件可以改动客户端要执行的操作,而工具调用审查只能拦下其规则覆盖的写法。出站脱敏规则识别不了的密钥,插件看得到;这类密钥本来也会原样发往上游,可以在安全页添加自定义规则覆盖它的格式。 + +降低这些风险的办法:只授予插件所需的权限,安装前通读代码,并在请求详情中对照经过插件前后的请求。 + +## 应用自带的插件 + +应用自带两个插件。它们和其他插件列在同一张表里,出厂均为关闭,打开和设置的方式与其他插件相同;应用按界面语言显示它们的名称和设置项标签。其中「WSL 路径转换」能改动工具调用,打开它要在系统对话框中确认,修改它的设置项、适用范围和出错时的处理方式则无须确认。所有插件都关闭时,core 不启动沙箱,保持关闭的插件不占内存。这些插件随 core 发布,远程 core 上同样有。两个插件都只处理对话,代码在 ThinkWatch Core 仓库的 [`crates/tw-gateway/src/plugin/defaults`](https://github.com/ThinkWatchProject/ThinkWatch-Core/tree/main/crates/tw-gateway/src/plugin/defaults) 目录中。 + +### 指定回答语言(`reply-language`) + +在系统提示词末尾附上一句固定的要求:用选定的语言回答,用户明确要求其他语言时除外。这句要求每次都相同,上游的提示词缓存照常命中。 + +- 权限:`system`。 +- 设置:`language`,回答语言,出厂为 `简体中文`。只接受语言名称(字母、空格、括号和连字符,不超过 40 个字符),设置写不进其他指令;填入其他内容时,插件在每个请求上出错。 + +出厂时的 manifest: + +```js +export const manifest = { + name: "Answer in a chosen language", + api: 1, + description: "Adds a fixed line to the end of the system prompt that asks the model to answer in the language set here.", + permissions: ["system"], + on_error: "reject", + settings: { + language: { type: "string", label: "Answer language", value: "简体中文" }, + }, +}; +``` + +### WSL 路径转换(`wsl-paths`) + +在 WSL 中运行的客户端打不开 `C:\Users\…` 这样的路径,在 Windows 上运行的客户端同样打不开 `/mnt/c/Users/…`。插件把工具调用中整个取值为盘符路径的参数,统一改写为客户端一侧能打开的写法;回答中的工具调用和对话历史中此前的工具调用都会改写,模型看到的始终是同一种写法。命令行中夹带的路径、对话文字和工具结果不做改动。 + +- 权限:`messages`、`reply.tool_calls`。 +- 设置:`windows_client`,客户端是否运行在 Windows 上,出厂为 `false`,即客户端运行在 WSL 中。转换方式写死在代码里,设置无法表达其他改写。 + +出厂时的 manifest: + +```js +export const manifest = { + name: "Convert WSL and Windows paths", + api: 1, + description: "Rewrites drive paths in tool-call arguments to the form the client can open (WSL /mnt/c/… or Windows C:\\…), in answers and in the conversation history.", + permissions: ["messages", "reply.tool_calls"], + on_error: "reject", + settings: { + windows_client: { + type: "boolean", + label: "The client runs on Windows (otherwise WSL)", + value: false, + }, + }, +}; +``` + +删除的自带插件不会再次添加;代码被修改过的自带插件保持原样,在「设置」页所做的更改不算修改代码。除此之外,应用更新带来自带插件的新版本时,插件随之更新,保留开关状态、出错时的处理方式、适用范围,以及新版本仍然声明且类型不变的设置项的取值;新版本申请了更多权限或更多请求种类时,插件改为关闭。 + +## 升级须知 + +升级到加入插件的版本后,首次启动会清空请求记录,包括保存的请求与回答报文。配置、密钥与上游不受影响,两个自带插件以关闭状态加入列表。 + +## 延伸阅读 + +- [功能详解](/zh-CN/docs/lite/features):应用的其他页面,包括安全页及其各项防护。 +- [连接远程 core](/zh-CN/docs/lite/remote-core):在服务器上的 core 中使用插件。 diff --git a/src/content/docs/_meta.ts b/src/content/docs/_meta.ts index 62f5844..0b7d647 100644 --- a/src/content/docs/_meta.ts +++ b/src/content/docs/_meta.ts @@ -196,6 +196,16 @@ export const products: Product[] = [ "zh-CN": "Tauri 2 外壳与 React 19 前端,负责托管 Core,在 macOS 与 Linux 上通过 unix socket、在 Windows 上通过回环端口、连接服务器时通过 TCP 端口控制它,每种通道都经过加密握手。", }, }, + { + slug: "plugins", + label: { en: "Plugins", "zh-CN": "插件" }, + locales: both, + group: "reference", + summary: { + en: "JavaScript plugins that change requests and answers: where they run, adding and editing one, the API and the kinds of request, permissions, limits, the sandbox and what it cannot prevent, and the two that ship with the app.", + "zh-CN": "改写请求与回答的 JavaScript 插件:运行位置、添加与编辑、接口与请求种类、权限、限额、沙箱及其无法防范的情况,以及应用自带的两个插件。", + }, + }, { slug: "import-links", label: { en: "Import links", "zh-CN": "导入链接" },