Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions src/lib/model-artifacts.ts
Original file line number Diff line number Diff line change
Expand Up @@ -207,11 +207,20 @@ export async function getModelsIndexMarkdown(): Promise<string> {
export function renderModelsIndexIntroMarkdown(): string {
return `Doubleword Batch API is priced per model based on token usage. Costs are calculated separately for input tokens (the content you send) and output tokens (the content generated by the model).

We can offer custom pricing for bulk discounts, large workloads, and dedicated deployments - reach out to [hello@doubleword.ai](mailto:hello@doubleword.ai).

The table below outlines the models we have available and their pricing per 1M tokens. If you are interested in understanding pricing for a model not listed below or if you'd like to request a new model - please reach out to support@doubleword.ai.

:::info{title="Prompt caching"}
Prompt-caching availability and rates are model-specific. Use **Cache read** to compare each supported model's reduced cached-input price with its standard input price. See the [prompt caching guide](/inference-api/prompt-caching) for setup, TTLs, and write pricing.
:::

:::warning{title="Async TTFT guarantees"}
We target a Time to First Token (TTFT) of under one minute for individual model calls, with two exceptions:

- **Tier availability:** the one-minute target does not apply to the flex tier for models that aren't also available on the realtime tier — currently mostly OCR models (as of Aug 2026), and subject to change as we expand the catalog.
- **High-volume bursts:** during large concurrent bursts to the same model and tier, the target applies to the first call, not the whole batch; remaining start times scale with queue depth.
:::
`;
}

Expand Down