Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
48 commits
Select commit Hold shift + click to select a range
411c7f9
CLOUD-4725: document latency p50 fields, mark mean fields as aliases
burak-upstash Sep 9, 2026
0b71e4c
CLOUD-4725: document the traffic and backup analysis endpoints, the n…
burak-upstash Sep 16, 2026
2c4fcfa
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 16, 2026
84ad723
CLOUD-4725 document the Insights tab and insights endpoint
burak-upstash Sep 16, 2026
a083600
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 16, 2026
b40d8e3
CLOUD-4725 document command counts in place of the protocol split
burak-upstash Sep 16, 2026
13cb2a5
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 16, 2026
ac78137
CLOUD-4725 auth failures stay zero for proxied databases
burak-upstash Sep 16, 2026
0aaa049
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 16, 2026
a61c7e2
CLOUD-4725 describe proxy rejections in the limits section
burak-upstash Sep 16, 2026
de5b84c
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 16, 2026
20062cf
CLOUD-4725 explain what counts as a cross-region route
burak-upstash Sep 16, 2026
6168d6e
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 16, 2026
f906cee
CLOUD-4725 integration pages: REST command count and proxy-side auth …
burak-upstash Sep 17, 2026
245b7fe
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 17, 2026
d77bba8
CLOUD-4725 five-view Insights tab, period-scoped traffic and counters…
burak-upstash Sep 21, 2026
691f9c2
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 21, 2026
cfb31b3
CLOUD-4725 auth failures also count commands sent before authenticating
burak-upstash Sep 21, 2026
5cc4dc3
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 21, 2026
f59bb6e
CLOUD-4725 keyspace view shows the latest daily backup analysis, prod…
burak-upstash Sep 21, 2026
8768cac
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 21, 2026
55e115b
CLOUD-4725 prod pack daily backup is included at no charge and turned…
burak-upstash Sep 21, 2026
fbd5699
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 21, 2026
815f2d6
CLOUD-4725 daily backups stay on while prod pack is enabled
burak-upstash Sep 21, 2026
02b433f
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 21, 2026
f2d7d10
CLOUD-4725 performance charts follow the updated insights mockup
burak-upstash Sep 21, 2026
5c4c8a3
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 21, 2026
14eb912
CLOUD-4725 first prod pack backup right away, latest daily backup kep…
burak-upstash Sep 21, 2026
4af61b3
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 21, 2026
3137e4e
CLOUD-4725 bandwidth chart, entry and serving region, entry region of…
burak-upstash Sep 22, 2026
1ab73b3
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 22, 2026
74bc7a0
CLOUD-4725 serving region of client networks
burak-upstash Sep 22, 2026
9bea5f1
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 22, 2026
fe6d82b
CLOUD-4725 revert serving region of client networks
burak-upstash Sep 23, 2026
5d5fbd8
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 23, 2026
bff2aab
CLOUD-4725 insights anomalies from the API and clearer traffic, clien…
burak-upstash Sep 23, 2026
7858315
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 23, 2026
cd115e1
CLOUD-4725 insights stats in the Prometheus and Datadog integration p…
burak-upstash Sep 24, 2026
cb73d69
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 24, 2026
393587d
CLOUD-4725 datadog scope, simple service latency wording and lua tabl…
burak-upstash Sep 25, 2026
7a7680e
CLOUD-4725 anomaly threshold wording in the openapi spec
burak-upstash Sep 25, 2026
4639e0d
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 25, 2026
e4ce036
CLOUD-4725 disabling prod pack keeps daily backups it did not turn on…
burak-upstash Sep 25, 2026
221b210
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 25, 2026
cc43427
CLOUD-4725 insights page without the slow commands table and the acce…
burak-upstash Sep 28, 2026
0110037
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 28, 2026
8851213
CLOUD-4725: backup analysis is not available for regional databases
burak-upstash Sep 29, 2026
b51a41b
chore(llms): regenerate llms.txt and llms-full.txt
github-actions[bot] Sep 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
601 changes: 597 additions & 4 deletions devops/developer-api/openapi.yaml

Large diffs are not rendered by default.

4 changes: 4 additions & 0 deletions docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -2261,6 +2261,8 @@
"GET /redis/databases",
"GET /redis/database/{id}",
"GET /redis/stats/{id}",
"GET /redis/traffic/{id}",
"GET /redis/insights/{id}",
"POST /redis/database",
"POST /redis/rename/{id}",
"POST /redis/reset-password/{id}",
Expand All @@ -2278,6 +2280,8 @@
"group": "Backup",
"pages": [
"GET /redis/list-backup/{id}",
"GET /redis/backup-analysis/{id}/{backupid}",
"POST /redis/backup-analysis/{id}/{backupid}",
"POST /redis/create-backup/{id}",
"POST /redis/restore-backup/{id}",
"PATCH /redis/enable-dailybackup/{id}",
Expand Down
183 changes: 162 additions & 21 deletions llms-full.txt

Large diffs are not rendered by default.

4 changes: 4 additions & 0 deletions llms.txt
Original file line number Diff line number Diff line change
Expand Up @@ -126,10 +126,12 @@
- [Authentication](https://upstash.com/docs/devops/developer-api/authentication.md): Authentication for the Upstash Developer API
- [HTTP Status Codes](https://upstash.com/docs/devops/developer-api/http_status_codes.md): The Upstash API uses the following HTTP Status codes:
- [Getting Started](https://upstash.com/docs/devops/developer-api/introduction.md)
- [Analyze Backup](https://upstash.com/docs/devops/developer-api/redis/backup/analyze_backup.md): This endpoint starts the offline analysis of a completed backup. The analysis scans the backup without touching the running database and takes minutes for large backups; poll the GET endpoint until the state is completed. Not available for RDB exports, HIPAA databases or regional databases.
- [Create Backup](https://upstash.com/docs/devops/developer-api/redis/backup/create_backup.md): This endpoint creates a backup for a Redis database.
- [Delete Backup](https://upstash.com/docs/devops/developer-api/redis/backup/delete_backup.md): This endpoint deletes a backup of a Redis database.
- [Disable Daily Backup](https://upstash.com/docs/devops/developer-api/redis/backup/disable_dailybackup.md): This endpoint disables daily backup for a Redis database.
- [Enable Daily Backup](https://upstash.com/docs/devops/developer-api/redis/backup/enable_dailybackup.md): This endpoint enables daily backup for a Redis database.
- [Get Backup Analysis](https://upstash.com/docs/devops/developer-api/redis/backup/get_backup_analysis.md): This endpoint returns the state of a backup's analysis and, once completed, the report. Daily backups of these databases are analyzed automatically when they complete. Available for databases on a Pro plan or with the Production Pack, except regional databases.
- [List Backup](https://upstash.com/docs/devops/developer-api/redis/backup/list_backup.md): This endpoint lists all backups for a Redis database.
- [Restore Backup](https://upstash.com/docs/devops/developer-api/redis/backup/restore_backup.md): This endpoint restores data from an existing backup.
- [Change Database Plan](https://upstash.com/docs/devops/developer-api/redis/change_plan.md): This endpoint changes the plan of a Redis database.
Expand All @@ -141,7 +143,9 @@
- [Enable Eviction](https://upstash.com/docs/devops/developer-api/redis/enable_eviction.md): This endpoint enables eviction for given database.
- [Enable TLS](https://upstash.com/docs/devops/developer-api/redis/enable_tls.md): This endpoint enables tls on a database.
- [Get Database](https://upstash.com/docs/devops/developer-api/redis/get_database.md): This endpoint gets details of a database.
- [Get Database Insights](https://upstash.com/docs/devops/developer-api/redis/get_database_insights.md): This endpoint returns the Insights series of a database: service latency percentiles, read, write, engine, replication, lock-wait and Lua latency, expired entries, replication lag, throughput and disk size, for the whole database, one replica or every region, with each series' latest window compared to the one before it, and how much the limit, security and Lua counters grew over the period. Latencies are in milliseconds. Available for databases on a Pro plan or with the Production Pack.
- [Get Database Stats](https://upstash.com/docs/devops/developer-api/redis/get_database_stats.md): This endpoint gets detailed stats of a database.
- [Get Database Traffic](https://upstash.com/docs/devops/developer-api/redis/get_database_traffic.md): This endpoint reports where a database's traffic comes from over a period. Available for databases on a Pro plan or with the Production Pack.
- [List Databases](https://upstash.com/docs/devops/developer-api/redis/list_databases.md): This endpoint list all databases of user.
- [Move To Team](https://upstash.com/docs/devops/developer-api/redis/moveto_team.md): This endpoint moves database under a target team
- [Rename Database](https://upstash.com/docs/devops/developer-api/redis/rename_database.md): This endpoint renames a database.
Expand Down
34 changes: 33 additions & 1 deletion redis/howto/datadog.mdx
Comment thread
muhammetssen marked this conversation as resolved.
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ This guide will walk you through the steps to seamlessly connect your Datadog ac
<Check >
**Integration Scope**

Upstash Datadog Integration only covers Pro databases or those included in the Enterprise Plan.
Upstash Datadog Integration only covers Pro databases, databases with the Production Pack, and those included in the Enterprise Plan.

</Check>

Expand Down Expand Up @@ -69,6 +69,38 @@ Upstash will suspend all metric publishing process after the you remove Datadog

After removing the integration on the Upstash side, it's crucial to go to your Datadog account and remove any related API keys or configurations associated with the integration.

## Metrics Published

Metrics arrive in Datadog under `upstash.db.<metric>` with `team` and `database` tags, refreshed every few minutes. Latencies are in milliseconds over the last ten minutes; a latency with no samples in that window is not sent.

| Metric | What it measures |
|---|---|
| `dailyprocessedcommands`, `dailyreadcommands`, `dailywritecommands` | Commands today, and the read and write share |
| `monthlycommands` | Commands this month |
| `readpersecond`, `writepersecond`, `throughput` | Commands per second by type, and in total |
| `hitrate`, `missrate` | Key hits and misses per second |
| `connections`, `restconnections` | Open TCP and REST connections |
| `totaldatasize`, `keyspace` | Data size in bytes, and number of keys |
| `dailybandwidth`, `dailybandwidthin`, `dailybandwidthout` | Bandwidth today, total and by direction, bytes |
| `monthlycost` | Cost this month |
| `overalllatencyp50`, `overalllatencyp99` | Simple service latency: server time from reading a request to writing its response, for requests where every command is a single-key read or write; multi-key commands, scripts and cross-region requests are excluded |
| `readlatencyp50`, `readlatencyp99`, `writelatencyp50`, `writelatencyp99` | Read and write request latency |
| `overalllatencyp90` | Simple service latency, p90 |
| `enginelatencyp99` | Time the engine spent executing commands, excluding network and queueing |
| `lockwaitlatencyp99` | Time commands waited for the execution lock |
| `replicationlatencyp999` | Time writes were held because of replication, p99.9 |
| `lualatencyp50`, `lualatencyp99` | Lua script execution latency |
| `throttledcommands` | Writes delayed today because persistence was falling behind |
| `authfailures` | Refused authentications per second over TCP and REST. REST requests over the rate limit are counted here as well |
| `restcommands` | Commands received over REST today |
| `expiredentriespermin` | Keys removed per minute because their TTL passed |
| `replicationlag` | Updates written on the primary that have not yet been sent to the replica |
| `oversizerequests`, `oversizerecordaccesses` | Requests over the request size limit, and accesses to keys over the record size limit, today |
| `keyspacenotifications`, `connectionsopened` | Keyspace notifications sent and TCP connections opened, per second |
| `luaexecutions`, `luacommands` | Script executions today, and the commands they ran |
| `luafailures`, `luatimeouts`, `luamemorylimits` | Script runs today that failed, timed out or hit the memory limit |
| `undeliverable` | Requests and connections the Upstash proxies could not hand to any replica today |

## Pricing

If you choose to integrate Datadog via Upstash, there will be an additional cost of $5 per month.
Expand Down
79 changes: 61 additions & 18 deletions redis/howto/metrics-and-charts.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -41,63 +41,106 @@ Total number of requests per day, over the last 5 days.

## Database Usage

If you click on the "Usage" tab, you can see more detailed charts about the usage of your database.
The "Usage" tab of a database shows detailed charts. Pick the period at the top of the tab: the past hour, 3 hours, 12 hours, day, 3 days or week. Databases on a Pro plan or with the Production Pack can also select the past month. Charts refresh every minute.

### Top Commands

The ten most executed commands over the period, one line each.

### Throughput

<Frame>
<img src="/img/metrics/throughput.png" alt="Redis throughput chart" width="100%" />
</Frame>

Throughput chart shows throughput values for reads, writes and commands (all
commands including reads and writes) per second. The chart covers the last 1
hour and it is updated every 10 seconds.
Reads, writes and all commands per second.

### Service Time Latency

<Frame>
<img src="/img/metrics/latency.png" alt="Redis latency chart" width="100%" />
</Frame>

This chart shows the processing time of the request between it is received by
the server and the response is sent to the caller. It shows the times in max,
mean, min, 99.9 percentile and 99.99 percentile. The chart covers the last 1
hour and it is updated every 10 seconds.
The time between a request reaching the server and its response leaving it, for reads and writes separately, as the median (p50) and the 99th percentile (p99). Every latency chart describes a sliding ten-minute window, so it reacts within minutes; an idle database shows no data rather than zero.

### Data Size

<Frame>
<img src="/img/metrics/datasize.png" alt="Redis data size chart" width="100%" />
</Frame>

This chart shows the data size of your database. The chart covers the last 24
hours and it is updated every 10 seconds.
The data size of your database on disk.

### Connections

<Frame>
<img src="/img/metrics/connections.png" alt="Redis connection count chart" width="100%" />
</Frame>

This chart shows the number of active client connections. It shows the number of
open connections plus the number of short-lived connections that started and
terminated in 10 seconds period. The chart covers the last 1 hour and it is
updated every 10 seconds.
Open TCP connections and open REST connections.

### Key Space

<Frame>
<img src="/img/metrics/keyspace.png" alt="Redis keyspace chart" width="100%" />
</Frame>

This chart shows the number of keys. The chart covers the last 24 hours and it
is updated every 10 seconds.
The number of keys in the database.

### Hits / Misses

<Frame>
<img src="/img/metrics/hitsmisses.png" alt="Redis cache hits and misses chart" width="100%" />
</Frame>

This chart shows the number of hits per second and misses per second. The chart
covers the last 1 hour and it is updated every 10 seconds.
Key hits and misses per second.

## Insights

Databases with the Production Pack have an **Insights** tab next to Usage; databases on a Pro plan have it too. It has five views: Performance, Traffic, Limits, Lua and Keyspace. One period selector at the top right, from the past hour to the past month, drives every view except Keyspace, and every number on a view covers that period. Hover the info icon next to a title for what the figure measures.

### Performance

One chart per signal. A global database gets one line per region, always in the same color, and a dashed line with their average. Each chart shows the current value and how the latest window compares with the window before it: the whole period for the past hour and the past three hours, the last three hours for longer periods. The comparison is left out when the earlier window has no data; a change of ten times or more is shown as a multiple, and Disk size shows the difference in bytes instead of a percentage. Charts also flag up to three anomalies: stretches where the value stayed well above its typical level for the period, which is the median, by at least three times the typical spread and at least 30%. Each metric has a small floor for the typical level, so a nearly idle chart does not turn a small bump into a huge percentage; an empty stretch at the start of the period, such as before the database existed, is ignored, and a rise that lasts more than a third of the period counts as a new level rather than a spike. The largest anomaly is highlighted with its peak against the typical value. Disk size is not checked, because it grows with your data. The same anomalies are returned by the [Developer API](/devops/developer-api/redis/get_database_insights), so they can be used outside the console.

- **Throughput** (commands per second): commands the engine processed per second, summed over replicas. Counts every command, including each command inside a pipeline or a script. Console and non-billable commands are excluded.
- **Bandwidth** (bytes per second): bytes moved between clients and the database per second, in and out combined, summed over replicas. The same counters that feed the daily bandwidth on the Usage tab, as a rate.
- **Disk size** (bytes): size of the data persisted on disk, largest replica. This is stored bytes, not memory in use.
- **Simple service latency** (ms): server time from reading a request to writing its response, for requests where every command is a single-key read or write served by this replica, as p50, p90 or p99. Multi-key commands, scripts and cross-region requests are excluded. Blocking time is subtracted.
- **Read service latency** and **Write service latency** (p99, ms): the same, for requests where every command is read-only, and for requests that include at least one write. All command types count.
- **Command execution latency** (p99, ms): time spent executing the command inside the engine, measured after the key lock is held and the key is loaded. Excludes parsing, lock wait, disk loads and the response.
- **Replication latency** (p99.9, ms): time a write was held before its response was sent because of replication: waiting for the backup to acknowledge, or backpressure applied when a replica falls behind.
- **Lock wait latency** (p99, ms): time a command waited to acquire its key's lock before executing. Grows with contention on hot keys and with long-running writes on the same key.
- **Lua execution latency** (p99, ms): wall time of EVAL, EVALSHA and FCALL, including the commands the script runs.
- **Expired entries** (per minute): keys removed because their TTL passed, whether they expired in memory, on access, by the expiry scan, or were purged from disk.
- **Replication lag** (entries): updates written on the primary that have not yet reached the replica in each region. A rising line means the replica is falling behind.

The **Server** selector narrows every chart to one replica. Latencies, expired entries and replication lag combine replicas by taking the highest value; throughput and bandwidth add up.

### Traffic

- **Request routing**: where requests entered the proxies and which replica served them, on a world map and in a table with REST requests, their share, TCP connections, bytes in and out and errors per route. The entry region is the region of the Upstash proxy a request came in through; clients connect to the nearest proxy, so it is usually close to where the client runs. The serving region is the region of the replica that executed the request, usually the nearest healthy replica; writes go to the primary region. A route is cross-region when the proxy had to connect to a replica in another region, which happens when requests enter a region where the database has no replica; writes that a read replica forwards to the primary are not visible here. A large cross-region share means a read region close to your clients would remove a network round trip from every request. The database's primary region is marked on the map and in the table. Above the table are the REST requests and TCP connections of the period and the share of bytes served cross-region, and below them the commands those REST requests carried and the commands run by Lua scripts: a REST request can carry several commands, since every command in a pipeline or transaction counts, and a Lua script counts each command it runs, whether it was called over REST or TCP.
- **Client networks**: the networks your clients connect from, as seen by the Upstash proxies. Each address is shortened to its /24 (IPv4) or /48 (IPv6) network, so single client addresses are never stored. The cloud provider and region are shown when the network is inside an IP range that AWS, Google Cloud or Cloudflare publishes; other networks, such as office or home connections, show as Other. Each row also shows the entry region the requests came in through.
- **Client types**: SDK, platform and runtime from the telemetry headers Upstash SDKs send with REST requests, as shares of those requests. A request sent through two SDKs, such as `@upstash/react-redis-browser` on top of `@upstash/redis`, counts once for each. TCP clients do not send telemetry and are not counted.

### Limits

Three meters compare the database with its plan: storage, open connections with the connections opened in the period, and the peak commands per second of the period. **Rejections** shows what a limit or a server guard refused or delayed on one stacked chart over the period, and lists each kind with its count and the time of the last hit:

- **Oversize requests**: requests over the request size limit.
- **Oversize keys**: accesses to keys over the record size limit, counted per access.
- **Throttled writes**: writes delayed because persistence was falling behind.
- **Undeliverable at proxy**: requests and connections the proxy could not hand to any replica. A request a replica answers with an error, such as 401 or 413, was delivered and does not count here.
- **Auth failures**: refused credentials over TCP and REST, and commands sent before authenticating. REST requests over the rate limit are counted here as well.

### Lua

Scripts loaded and run, executions with the commands they ran, failures split into timeouts and memory-limit hits, and the current p99 next to the script timeout of the plan. The latency chart shows p50 and p99 of EVAL, EVALSHA and FCALL including the commands they run, as the highest value across replicas. The table has the ten most executed scripts of each replica since midnight UTC: executions, commands and failures are summed across replicas, while the average and slowest duration and the memory use are the highest replica's value. A script is identified by the SHA1 of its source, the same hash `EVALSHA` takes; the table shows its first 12 characters and copies the full hash. To give a script a name in the table, start its source with `#!lua script_name=<name>`.

### Keyspace

The latest analysis of the database's daily backup. Every daily backup is scanned offline after it completes, without touching the running database, and the view shows the most recent report: the number of keys, elements and bytes, the distribution of key sizes and element counts per type, the biggest keys overall and per type with their expiry, and an expiration profile that shows how much data is persistent, already expired but not yet removed, or about to expire. Sizes are on-disk logical bytes. Turning on the Production Pack turns on daily backups, and the latest one is included at no extra charge; a database without daily backups, or a regional database, has no keyspace report.

## Backup Insights

The keyspace report of the latest daily backup is on the [Keyspace](#keyspace) view of the Insights tab. Other backups can be analyzed through the [Developer API](/devops/developer-api/redis/backup/get_backup_analysis).
Loading