diff --git a/docs-mintlify/admin/ai/bring-your-own-model.mdx b/docs-mintlify/admin/ai/bring-your-own-model.mdx
index a4340a6acd634..674801ed93468 100644
--- a/docs-mintlify/admin/ai/bring-your-own-model.mdx
+++ b/docs-mintlify/admin/ai/bring-your-own-model.mdx
@@ -25,6 +25,7 @@ how you manage AI costs.
| **GCP Vertex AI** | Yes | No |
| **Databricks** | Yes | No |
| **Snowflake Cortex** | Yes | No |
+| **LLM Gateway** | Yes | No |
## Configuration
@@ -69,13 +70,17 @@ embedding model that created them.
## Network configuration
-When using BYOM, Cube connects to your model provider from its control plane.
-If your provider requires IP allowlisting, ensure the Cube outbound IP
-addresses are added to your allowlist.
+For every provider except **LLM Gateway**, Cube connects to your model
+provider from its control plane. If your provider requires IP allowlisting,
+ensure the Cube outbound IP addresses are added to your allowlist.
For agents running in dedicated regions, additional per-region IP addresses
may also need to be allowlisted.
+The **LLM Gateway** provider works the other way around: the request is made
+by your own deployment, so nothing needs to be reachable from Cube and there
+is no allowlist to maintain. See [LLM Gateway](#llm-gateway) below.
+
## Billing
When using a BYOM model, **Cube AI tokens are not consumed**. You are billed
@@ -119,6 +124,208 @@ Supports two authentication methods:
- JWT authentication
- Key-pair authentication (requires an encrypted PKCS#8 PEM private key)
+### LLM Gateway
+
+Use this provider to route agent traffic through your own LLM gateway or
+proxy — including one that is only reachable from inside your own network.
+
+Unlike every other provider, Cube holds no credentials and never calls your
+model. Your deployment does: Cube sends the conversation to your deployment's
+runtime, which invokes a `chatCompletion` hook you write in your configuration
+file — `cube.js`, `cube.ts`, or `cube.py` under its snake_case name
+`chat_completion`. The gateway endpoint, its API key and all LLM egress stay
+inside your network.
+
+
+
+Because the request originates from your own deployment, a gateway with no
+public ingress — for example one reachable only within your VPC — works with
+no peering, allowlisting or inbound access from Cube.
+
+
+
+#### Writing the hook
+
+The hook may return a [LangChain](https://js.langchain.com/) chat model, which
+is the shortest path if your gateway already has a LangChain integration:
+
+```javascript
+const { ChatOpenAI } = require("@langchain/openai");
+
+module.exports = {
+ chatCompletion: new ChatOpenAI({
+ model: "gpt-5",
+ apiKey: process.env.LLM_GATEWAY_API_KEY,
+ configuration: {
+ baseURL: "https://llm-gateway.internal.example.com/v1",
+ },
+ }),
+};
+```
+
+Any LangChain chat model works the same way — `ChatAnthropic`,
+`ChatBedrockConverse`, `ChatVertexAI`, or your own subclass. Add the
+integration package to your project's `package.json`; Cube does not need to
+know which one you use.
+
+The model must support tool calling. Cube agents call tools on every turn, so
+a model without it cannot serve an agent.
+
+To choose the model per request — to route by user, or to act on the model
+name configured in **Admin > Models** — export a function instead. It is
+called once per turn:
+
+```javascript
+const { ChatOpenAI } = require("@langchain/openai");
+
+module.exports = {
+ chatCompletion: ({ model, securityContext }) => new ChatOpenAI({
+ model: model ?? "gpt-5",
+ apiKey: process.env.LLM_GATEWAY_API_KEY,
+ configuration: {
+ baseURL: "https://llm-gateway.internal.example.com/v1",
+ defaultHeaders: {
+ "x-cube-tenant": securityContext?.tenantId,
+ },
+ },
+ }),
+};
+```
+
+The function receives:
+
+| Field | Description |
+| --- | --- |
+| `model` | The **Model Name** configured for the model in **Admin > Models**, if any. A routing hint — Cube does not interpret it |
+| `messages` | The conversation so far, in LangChain's message format |
+| `tools` | Tool definitions for this turn, as JSON Schema |
+| `toolChoice` | How the model should use tools, when the agent constrains it |
+| `metadata` | Attribution for cost tracking on your side. Carries no query results or model output |
+| `securityContext` | The security context of the user whose turn triggered the call |
+| `signal` | An `AbortSignal`, aborted if the user cancels the turn |
+
+If you do not use LangChain, return an async iterable of chunks instead:
+
+```javascript
+module.exports = {
+ chatCompletion: async function* ({ messages, signal }) {
+ const response = await callYourGateway(messages, { signal });
+
+ for await (const token of response) {
+ yield { content: token };
+ }
+ },
+};
+```
+
+#### Chunk format
+
+A LangChain model produces these for you. You only need this section if your
+hook builds chunks by hand — an async iterable in `cube.js`, or either of the
+`cube.py` forms below.
+
+| Field | Description |
+| --- | --- |
+| `content` | Text produced by this chunk. Chunks are concatenated in order |
+| `tool_call_chunks` | Tool-call fragments (see below) |
+| `usage_metadata` | `input_tokens`, `output_tokens`, `total_tokens`. Send once, on the last chunk |
+| `response_metadata` | Free-form; `model_name` shows up in agent traces |
+
+
+
+Cube agents call tools on every turn, so a hand-written hook that only ever
+emits `content` cannot serve an agent. Your hook must translate the tool calls
+your gateway returns into `tool_call_chunks`.
+
+
+
+Each entry in `tool_call_chunks` carries an `index`, and the `args` of entries
+sharing an index are **concatenated as a JSON string** — that is how a streamed
+tool call arrives one fragment at a time. Send the `id` and `name` on the first
+fragment of each call:
+
+```javascript
+yield {
+ content: '',
+ tool_call_chunks: [
+ { index: 0, id: 'call_1', name: 'run_query', args: '{"limit"' },
+ ],
+};
+yield { content: '', tool_call_chunks: [{ index: 0, args: ': 10}' }] };
+```
+
+If your gateway hands you whole tool calls rather than fragments, emit each one
+as a single chunk whose `args` is the complete JSON string.
+
+#### Writing the hook in Python
+
+`cube.py` supports the same hook under its snake_case name, `chat_completion`,
+but what it can return is narrower. Values crossing between Python and
+JavaScript are limited to strings, numbers, booleans, lists, dicts and plain
+functions, so a Python hook **cannot** return a model object or an async
+generator. Return a list of chunks instead:
+
+```python
+from cube import config
+
+
+@config
+async def chat_completion(request):
+ response = await call_your_gateway(request["messages"])
+
+ return [{"content": response.text}]
+```
+
+To stream, return a `next` function that yields one chunk per call and `None`
+when the response is complete. A closure crosses the bridge, so it can hold
+whatever iterator your gateway call produced:
+
+```python
+from cube import config
+
+
+@config
+def chat_completion(request):
+ tokens = iter(call_your_gateway(request["messages"]))
+
+ async def next_chunk():
+ try:
+ return {"content": next(tokens)}
+ except StopIteration:
+ return None
+
+ return {"next": next_chunk}
+```
+
+Chunk dicts use the same fields as the [chunk format](#chunk-format) above,
+with Python naming — `content`, `tool_call_chunks`, `usage_metadata`.
+
+Three further differences from the JavaScript form:
+
+- The request is a dict with the same keys, but **without `signal`** —
+ cancellation cannot be delivered across the bridge, so a Python hook is not
+ told when the user cancels a turn. Enforce your own timeout if that matters
+- The hook must be a plain `def` or `async def`. A bound method, a
+ `functools.partial` or a callable class instance will not be picked up
+- There is no LangChain shortcut, so a Python hook always builds
+ `tool_call_chunks` by hand. If your gateway speaks the OpenAI API, a
+ `cube.js` hook with `ChatOpenAI` is considerably less work
+
+If you want to hand back a LangChain model directly, the deployment has to be
+configured with `cube.js` rather than `cube.py` — a deployment uses one
+configuration file, and `cube.py` takes precedence when both are present.
+
+#### Configuring the model in Cube
+
+1. Add a model in **Admin > Models** and choose the **LLM Gateway** provider
+2. Leave **Model Name** blank if your hook always serves one model, or set it
+ to a name your hook routes on
+3. Optionally set **Small Model Name** for the lighter follow-up calls agents
+ make; it defaults to the main model name
+
+There are no credential fields — the credentials belong in your deployment's
+environment variables, next to the hook that uses them.
+
## Troubleshooting
### Rate limit errors
@@ -131,6 +338,29 @@ not by Cube. Check your provider's rate limits and usage quotas.
Verify that the API key or credentials configured for the model are valid and
have the necessary permissions.
+### LLM Gateway errors
+
+Errors mentioning the LLM Gateway come from your own deployment, not from
+Cube's control plane:
+
+- **No `chatCompletion` hook is configured** — the model is assigned to an
+ agent but the deployment's configuration file does not export the hook, or
+ the deployment has not restarted since it was added
+- **Does not support tool calling** — the chat model the hook returned has no
+ `bindTools`. Cube agents require a tool-calling model
+- **Must return a LangChain chat model, a list of chunks, ...** — a `cube.js`
+ hook returned something with no stream in it
+- **Unable to represent PyObject in JS** — a `cube.py` hook returned a value
+ with no bridge representation, most often a model object or an async
+ generator. This is raised by the bridge itself, before Cube sees the return
+ value, which is why it reads nothing like the message above. Use one of the
+ two Python forms above
+- **Stream ended before the response was complete** — the connection to your
+ gateway dropped mid-answer. Check your gateway's timeouts and any proxy
+ between it and the deployment
+
+Your gateway's own errors are passed through with their original message.
+
### Model not found
Ensure the model ID configured in Cube matches a valid model offered by your
diff --git a/packages/cubejs-backend-native/js/index.ts b/packages/cubejs-backend-native/js/index.ts
index 0178709a8f5b7..0348ecabc420e 100644
--- a/packages/cubejs-backend-native/js/index.ts
+++ b/packages/cubejs-backend-native/js/index.ts
@@ -537,6 +537,14 @@ export interface PyConfiguration {
scheduledRefreshContexts?: (ctx: unknown) => Promise
scheduledRefreshTimeZones?: (ctx: unknown) => Promise
contextToGroups?: (ctx: unknown) => Promise
+ /**
+ * LLM Gateway hook; Core ignores it. Resolves to a list of chunks, or to
+ * `{ next }` for a stream — where `next()` resolves to `undefined`, NOT
+ * `null`, at the end: the bridge converts a Python `None` to `undefined`, so
+ * `if (chunk === null) break` never terminates. Python values with no bridge
+ * representation (a model object, an async generator) throw instead.
+ */
+ chatCompletion?: (request: unknown) => Promise
}
function simplifyExpressRequest(req: ExpressRequest) {
diff --git a/packages/cubejs-backend-native/python/cube/src/__init__.py b/packages/cubejs-backend-native/python/cube/src/__init__.py
index c793af0869056..d8ef6769b0076 100644
--- a/packages/cubejs-backend-native/python/cube/src/__init__.py
+++ b/packages/cubejs-backend-native/python/cube/src/__init__.py
@@ -78,6 +78,10 @@ class Configuration:
orchestrator_options: Union[Dict, Callable[[RequestContext], Dict]]
context_to_groups: Callable[[RequestContext], list[str]]
fast_reload: bool
+ # LLM Gateway hook. Must be a plain function, and its return value has to
+ # cross the Python->JS bridge: a list of chunk dicts, or {"next": fn} for a
+ # stream. Model objects and async generators cannot cross — use cube.js.
+ chat_completion: Callable[[Dict], Any]
def __init__(self):
self.web_sockets = None
@@ -126,6 +130,7 @@ def __init__(self):
self.pre_aggregations_schema = None
self.orchestrator_options = None
self.context_to_groups = None
+ self.chat_completion = None
self.fast_reload = None
def __call__(self, func):
diff --git a/packages/cubejs-backend-native/src/python/cube_config.rs b/packages/cubejs-backend-native/src/python/cube_config.rs
index aa9a5f38afe25..56e60f3c6b367 100644
--- a/packages/cubejs-backend-native/src/python/cube_config.rs
+++ b/packages/cubejs-backend-native/src/python/cube_config.rs
@@ -46,6 +46,9 @@ impl CubeConfigPy {
"web_sockets_base_path",
// functions
"can_switch_sql_user",
+ // Unlisted names are dropped silently, with no error to point at
+ // the omission — so this entry is what makes the hook reach JS.
+ "chat_completion",
"check_auth",
"check_sql_auth",
"context_to_api_scopes",
diff --git a/packages/cubejs-backend-native/test/config-chat-completion-object.py b/packages/cubejs-backend-native/test/config-chat-completion-object.py
new file mode 100644
index 0000000000000..bc460e03b6adb
--- /dev/null
+++ b/packages/cubejs-backend-native/test/config-chat-completion-object.py
@@ -0,0 +1,18 @@
+from cube import config
+
+config.schema_path = "models"
+
+
+class NotBridgeable:
+ """An arbitrary Python object, standing in for a LangChain chat model."""
+
+ def stream(self, messages):
+ return []
+
+
+# Documents the boundary the two supported forms exist to work around: an
+# arbitrary Python object has no cross-language representation, so it reaches
+# JavaScript as an unrepresentable reference and throws there.
+@config
+def chat_completion(request):
+ return NotBridgeable()
diff --git a/packages/cubejs-backend-native/test/config-chat-completion-stream.py b/packages/cubejs-backend-native/test/config-chat-completion-stream.py
new file mode 100644
index 0000000000000..5176ff3a34360
--- /dev/null
+++ b/packages/cubejs-backend-native/test/config-chat-completion-stream.py
@@ -0,0 +1,20 @@
+from cube import config
+
+config.schema_path = "models"
+
+
+# Streaming form of the `chat_completion` hook. A Python async generator cannot
+# cross the bridge to JavaScript, so a streamed response is handed back as a
+# `next` function that yields one chunk per call and None when complete —
+# closing over whatever iterator the gateway call produced.
+@config
+def chat_completion(request):
+ tokens = iter(["strea", "med ", request["model"]])
+
+ async def next_chunk():
+ try:
+ return {"content": next(tokens)}
+ except StopIteration:
+ return None
+
+ return {"next": next_chunk}
diff --git a/packages/cubejs-backend-native/test/config.py b/packages/cubejs-backend-native/test/config.py
index 1ad05ca00bcd4..c78efbcaaf6b3 100644
--- a/packages/cubejs-backend-native/test/config.py
+++ b/packages/cubejs-backend-native/test/config.py
@@ -107,3 +107,16 @@ def context_to_groups(ctx):
"dev",
"analytics",
]
+
+
+@config
+def chat_completion(request):
+ print("[python] chat_completion request=", request)
+
+ # A list of chunks: the whole response at once. Simplest form, and the one
+ # to use when the gateway call is not streamed.
+ return [
+ {"content": "Hello from "},
+ {"content": request["model"]},
+ {"usage_metadata": {"input_tokens": 3, "output_tokens": 4, "total_tokens": 7}},
+ ]
diff --git a/packages/cubejs-backend-native/test/python.test.ts b/packages/cubejs-backend-native/test/python.test.ts
index 30b3375aae073..264b1d9ec0a09 100644
--- a/packages/cubejs-backend-native/test/python.test.ts
+++ b/packages/cubejs-backend-native/test/python.test.ts
@@ -55,6 +55,94 @@ suite('Python Config', () => {
config = await loadConfigurationFile('config.py');
});
+ // `chat_completion` only reaches JavaScript if it is BOTH declared on the
+ // Python `Configuration` class AND present in the Rust allow-list in
+ // `cube_config.rs`. Miss either and the attribute is dropped in silence —
+ // no error, no warning, just a hook that never runs. These two cases are
+ // what makes that omission loud.
+ test('chat_completion returning a list of chunks', async () => {
+ if (!config.chatCompletion) {
+ throw new Error('chatCompletion was not defined in config.py');
+ }
+
+ expect(await config.chatCompletion({ model: 'gateway-model', messages: [] })).toEqual([
+ { content: 'Hello from ' },
+ { content: 'gateway-model' },
+ { usage_metadata: { input_tokens: 3, output_tokens: 4, total_tokens: 7 } },
+ ]);
+ });
+
+ test('chat_completion returning a next() pull stream', async () => {
+ const streamConfig = await loadConfigurationFile('config-chat-completion-stream.py');
+
+ if (!streamConfig.chatCompletion) {
+ throw new Error('chatCompletion was not defined in config-chat-completion-stream.py');
+ }
+
+ const stream = (await streamConfig.chatCompletion({
+ model: 'gateway-model',
+ messages: [],
+ })) as { next: () => Promise };
+
+ // A Python closure survives the bridge as a callable, which is the whole
+ // reason the streaming form is shaped as a pull function rather than as an
+ // async generator (an async generator object has no bridge representation).
+ expect(typeof stream.next).toEqual('function');
+
+ const chunks: unknown[] = [];
+ for (;;) {
+ // eslint-disable-next-line no-await-in-loop
+ const chunk = await stream.next();
+ if (chunk === undefined) {
+ break;
+ }
+ chunks.push(chunk);
+ }
+
+ expect(chunks).toEqual([
+ { content: 'strea' },
+ { content: 'med ' },
+ { content: 'gateway-model' },
+ ]);
+ });
+
+ // `CLRepr::Null` converts to `cx.undefined()`, so the Python `None` that ends
+ // a stream arrives as `undefined` and never as `null`. A consumer that stops
+ // on `null` alone loops forever on an already-finished stream — pinned here
+ // because the Python side says `return None` and nothing about writing it
+ // suggests the JavaScript side sees anything else.
+ test('a Python None crosses as undefined, not null', async () => {
+ const streamConfig = await loadConfigurationFile('config-chat-completion-stream.py');
+ const stream = (await streamConfig.chatCompletion!({
+ model: 'gateway-model',
+ messages: [],
+ })) as { next: () => Promise };
+
+ await stream.next();
+ await stream.next();
+ await stream.next();
+
+ const terminator = await stream.next();
+ expect(terminator).toBeUndefined();
+ expect(terminator).not.toBeNull();
+ });
+
+ // The boundary the two supported forms exist to work around. A customer's
+ // first instinct is to hand back a model object the way `cube.js` can; this
+ // pins that it fails loudly at the bridge rather than producing an empty
+ // response.
+ test('chat_completion returning an arbitrary Python object is rejected', async () => {
+ const objectConfig = await loadConfigurationFile('config-chat-completion-object.py');
+
+ if (!objectConfig.chatCompletion) {
+ throw new Error('chatCompletion was not defined in config-chat-completion-object.py');
+ }
+
+ await expect(
+ objectConfig.chatCompletion({ model: 'gateway-model', messages: [] })
+ ).rejects.toThrow(/PyObject/);
+ });
+
test('async checkAuth', async () => {
expect(config).toEqual({
schemaPath: 'models',
@@ -71,6 +159,7 @@ suite('Python Config', () => {
contextToGroups: expect.any(Function),
scheduledRefreshContexts: expect.any(Function),
scheduledRefreshTimeZones: expect.any(Function),
+ chatCompletion: expect.any(Function),
});
if (!config.checkAuth) {
diff --git a/packages/cubejs-server-core/src/core/optionsValidate.ts b/packages/cubejs-server-core/src/core/optionsValidate.ts
index 2824253895f29..4c1e3ba26d26e 100644
--- a/packages/cubejs-server-core/src/core/optionsValidate.ts
+++ b/packages/cubejs-server-core/src/core/optionsValidate.ts
@@ -158,6 +158,11 @@ const schemaOptions = Joi.object().keys({
serverless: Joi.boolean(),
allowNodeRequire: Joi.boolean(),
fastReload: Joi.boolean(),
+ // LLM Gateway hook. Core does not read it, but this schema rejects unknown
+ // keys, so without an entry every deployment declaring it fails to start.
+ // An object is allowed because the hook may be a model instance, not a
+ // factory. See docs-mintlify/admin/ai/bring-your-own-model.mdx.
+ chatCompletion: Joi.alternatives().try(Joi.func(), Joi.object()),
});
export function validateOptions(options: T): T {
diff --git a/packages/cubejs-server-core/src/core/types.ts b/packages/cubejs-server-core/src/core/types.ts
index 2e4084159a551..d1f219cd492c3 100644
--- a/packages/cubejs-server-core/src/core/types.ts
+++ b/packages/cubejs-server-core/src/core/types.ts
@@ -241,6 +241,13 @@ export interface CreateOptions {
allowNodeRequire?: boolean;
semanticLayerSync?: (context: RequestContext) => Promise | BiToolSyncConfig[];
fastReload?: boolean;
+ /**
+ * LLM Gateway hook — a chat model instance, a factory returning one, or a
+ * function returning a stream of chunks. Core accepts it and never reads it;
+ * declared so a `cube.ts` config that sets it type-checks. See
+ * docs-mintlify/admin/ai/bring-your-own-model.mdx.
+ */
+ chatCompletion?: unknown;
}
export interface DriverDecoratedOptions extends CreateOptions {
diff --git a/packages/cubejs-server-core/test/unit/optionsValidate.test.ts b/packages/cubejs-server-core/test/unit/optionsValidate.test.ts
index fd657744fcc82..3ff8a0e75c343 100644
--- a/packages/cubejs-server-core/test/unit/optionsValidate.test.ts
+++ b/packages/cubejs-server-core/test/unit/optionsValidate.test.ts
@@ -45,3 +45,46 @@ describe('validateOptions sanitized result', () => {
expect(validated).not.toBe(options);
});
});
+
+describe('validateOptions chatCompletion', () => {
+ // The schema rejects unknown keys, so without an entry for `chatCompletion`
+ // every deployment declaring Cube Cloud's LLM Gateway hook fails to start
+ // with "Invalid cube-server-core options" — a boot failure, not a warning.
+ test('accepts a function', () => {
+ expect(() => validateOptions({ chatCompletion: () => [] })).not.toThrow();
+ });
+
+ test('accepts a model instance rather than a factory', () => {
+ expect(() => validateOptions({ chatCompletion: { stream: () => [] } })).not.toThrow();
+ });
+
+ test('preserves a factory by reference', () => {
+ const chatCompletion = () => [];
+
+ expect(validateOptions({ chatCompletion }).chatCompletion).toBe(chatCompletion);
+ });
+
+ // The branch that can actually break. `validateOptions` returns Joi's `value`,
+ // so were Joi ever to clone the matched object, a chat model instance would
+ // arrive at the consumer as a plain object — no prototype, no `bindTools`,
+ // which the LLM Gateway reports as "does not support tool calling". An object
+ // literal cannot show that; a class instance can.
+ test('preserves a model instance by reference, prototype intact', () => {
+ class Model {
+ public bindTools() {
+ return this;
+ }
+ }
+ const chatCompletion = new Model();
+ const validated = validateOptions({ chatCompletion }).chatCompletion;
+
+ expect(validated).toBe(chatCompletion);
+ expect(validated).toBeInstanceOf(Model);
+ expect(typeof (validated as Model).bindTools).toBe('function');
+ });
+
+ test('still rejects an unknown option', () => {
+ expect(() => validateOptions({ chatCompletionn: () => [] } as any))
+ .toThrow(/chatCompletionn/);
+ });
+});