Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion app_files/Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,6 @@ COPY app.py .
COPY gconf.py .
COPY test_json_1MB.json .
COPY start_services.sh .
ENV GUNICORN_CMD_ARGS="-c gconf.py --reuse-port"
ENV GUNICORN_CMD_ARGS="-c gconf.py"
EXPOSE 8000 8080-8082
ENTRYPOINT /src/start_services.sh
8 changes: 5 additions & 3 deletions app_files/gconf.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,12 +9,14 @@
else:
bind = '0.0.0.0:8000'

reuse_port = not os.getenv('SOCKET')

workers = os.getenv('WORKERS', 2)
threads = os.getenv('THREADS', 1)
# backlog - The number of pending connections.
backlog = 64
backlog = 1024
# Workers silent for more than this many seconds are killed and restarted.
timeout = 60
timeout = 300
# Timeout for graceful workers restart.
graceful_timeout = 30
# The number of seconds to wait for requests on a Keep-Alive connection.
Expand All @@ -27,7 +29,7 @@

if os.getenv('KEEPALIVE'):
# Workers silent for more than this many seconds are killed and restarted.
timeout = 60
timeout = 300
# Timeout for graceful workers restart.
graceful_timeout = 30
# The number of seconds to wait for requests on a Keep-Alive connection.
Expand Down
4 changes: 3 additions & 1 deletion bin/run_test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -2,4 +2,6 @@
tag=$1
echo "============================= Running tests with $tag ============================="
pytest -x --html=./reports/"$tag".html --self-contained-html --show-capture=stdout -vv -rP test_files/ -m "$tag"
echo "============================= Test run with $tag finished ========================="
status=$?
echo "============================= Test run with $tag finished ========================="
exit $status
77 changes: 42 additions & 35 deletions index.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ filename: index.md
This document is intended to provide some tips and ideas to get the most out of it


# Use these techniques to achieve 100-300% performance increase from your FastAPI application
## Use these techniques to achieve up to ~2.2x performance from your FastAPI application

> All tested on the same sized Docker containers (2 CPU cores). The performance numbers below represent real CI-verified measurements.

Expand All @@ -25,7 +25,7 @@ Each optimization below has a measurable, compounding effect. When applied toget
| **Gunicorn (1 worker) → FastAPI CLI (2 workers)** | +50-65% throughput | [Server runners](https://kisspeter.github.io/fastapi-performance-optimization/server_runners) |
| **JSON → ORJSON response class** | +4-13% throughput | [JSON response classes](https://kisspeter.github.io/fastapi-performance-optimization/json_response_class) |
| **BaseHTTPMiddleware → Starlette ASGI** | +35-44% throughput (avoiding BaseHTTPMiddleware cost) | [Middleware](https://kisspeter.github.io/fastapi-performance-optimization/middleware) |
| **Best vs Worst combo** | **+100% throughput** | [Measured below](#sync-endpoint-syncbig_json_response) |
| **Best vs Worst combo** | **+116% throughput (2.2x)** | [Measured below](#sync-endpoint-syncbig_json_response) |

### Async endpoints

Expand All @@ -35,11 +35,11 @@ Each optimization below has a measurable, compounding effect. When applied toget
| **Gunicorn (1 worker) → FastAPI CLI (2 workers)** | +50-65% throughput | [Server runners](https://kisspeter.github.io/fastapi-performance-optimization/server_runners) |
| **JSON → ORJSON response class** | +4-13% throughput | [JSON response classes](https://kisspeter.github.io/fastapi-performance-optimization/json_response_class) |
| **BaseHTTPMiddleware → Starlette ASGI** | +35-44% throughput (avoiding BaseHTTPMiddleware cost) | [Middleware](https://kisspeter.github.io/fastapi-performance-optimization/middleware) |
| **Best vs Worst combo** | **+297% throughput** | [Measured below](#async-endpoint-asyncbig_json_response) |
| **Best vs Worst combo** | **+118% throughput (2.2x)** | [Measured below](#async-endpoint-asyncbig_json_response) |

## Measured best vs worst configurations

> CI run [30018301964](https://github.com/KissPeter/fastapi-performance-optimization/actions/runs/30018301964) — all 4 jobs passed.
> CI run [36235373898](https://github.com/KissPeter/fastapi-performance-optimization/actions/runs/36235373898) — all 4 jobs passed, fully green, no non-2xx responses.

Using the big JSON response endpoint (1MB payload) across all tested combinations:

Expand All @@ -53,24 +53,24 @@ Using the big JSON response endpoint (1MB payload) across all tested combination

| Configuration | RPS | Latency |
|--------------|-----|---------|
| **Best**: FastAPI CLI (2 workers) + async + ORJSON + Starlette ASGI | 3405 | 29.4 ms |
| **Worst**: Gunicorn (1 worker, 0 threads) + sync + JSON + BaseHTTPMiddleware | 1695 | 59.1 ms |
| **Improvement** | **+100.85%** | **-29.7 ms** |
| **Best**: FastAPI CLI (2 workers) + async + ORJSON + Starlette ASGI | 27.90 | 3583.59 ms |
| **Worst**: Gunicorn (1 worker, 0 threads) + sync + JSON + BaseHTTPMiddleware | 12.94 | 7729.63 ms |
| **Improvement** | **+115.69%** | **-4146 ms** |

### Async endpoint (`/async/big_json_response`)

| Configuration | RPS | Latency |
|--------------|-----|---------|
| **Best**: FastAPI CLI (2 workers) + async + ORJSON + Starlette ASGI | 3433 | 29.4 ms |
| **Worst**: Gunicorn (1 worker, 0 threads) + sync + JSON + BaseHTTPMiddleware | 864 | 180.1 ms |
| **Improvement** | **+297%** | **-150.7 ms** |
| **Best**: FastAPI CLI (2 workers) + async + ORJSON + Starlette ASGI | 27.86 | 3589.66 ms |
| **Worst**: Gunicorn (1 worker, 0 threads) + sync + JSON + BaseHTTPMiddleware | 12.79 | 7820.47 ms |
| **Improvement** | **+117.86%** | **-4231 ms** |

### Small payload endpoints

| Endpoint | Best RPS | Worst RPS | Improvement |
|----------|----------|-----------|-------------|
| `/sync/items/` | 2081 | 974 | **+113.6%** |
| `/async/items/` | 2656 | 867 | **+206.4%** |
| `/sync/items/` | 3102.39 | 1391.80 | **+122.91%** |
| `/async/items/` | 4304.53 | 1766.31 | **+143.70%** |

### The compounding effect

Expand All @@ -82,33 +82,44 @@ When you combine ALL optimizations:
| Endpoint type | Sync | Async |
| Response class | JSON | ORJSON |
| Middleware | BaseHTTPMiddleware | Starlette ASGI (or none) |
| Nginx transport | TCP port | Unix socket |
| **Combined RPS (big JSON)** | **864** | **3433** |
| **Combined improvement** | | **~4x throughput** |
| **Combined RPS (big JSON)** | **12.79** | **27.86** |
| **Combined improvement** | | **~2.2x throughput** |

> The difference between a default FastAPI setup and a properly optimized one is **4x throughput** on big JSON responses and **3x on small payloads**. These are not theoretical numbers — they are measured in CI on identical Docker infrastructure (2 CPU cores per container).
> The difference between a default FastAPI setup and a properly optimized one is **~2.2x throughput** on big JSON responses and **~2.2-2.4x on small payloads**. These are not theoretical numbers — they are measured in CI on identical Docker infrastructure (2 CPU cores per container).

## All optimization topics
## Topics by category

## [Fastapi Middleware performance tuning](https://kisspeter.github.io/fastapi-performance-optimization/middleware)
## [Fastapi JSON response classes comparison](https://kisspeter.github.io/fastapi-performance-optimization/json_response_class)
## [Gunicorn workers and threads](https://kisspeter.github.io/fastapi-performance-optimization/workers_and_threads)
## [Nginx in front of FastAPI](https://kisspeter.github.io/fastapi-performance-optimization/nginx_port_socket)
## [Connection keepalive](https://kisspeter.github.io/fastapi-performance-optimization/keepalive)
## [Server Runners: Gunicorn vs Uvicorn vs FastAPI CLI](https://kisspeter.github.io/fastapi-performance-optimization/server_runners)
## [Sync / Async API Endpoints](https://kisspeter.github.io/fastapi-performance-optimization/sync_vs_async)
## [Connection Pool Size of External Resources](https://kisspeter.github.io/fastapi-performance-optimization/connection_pool)
## [Thread Pool Sizing (anyio tokens)](https://kisspeter.github.io/fastapi-performance-optimization/thread_pool_sizing)
## [Per-Worker Connection Pool](https://kisspeter.github.io/fastapi-performance-optimization/per_worker_connection_pool)
## [Pool Sizing Calculator](https://kisspeter.github.io/fastapi-performance-optimization/pool_sizing_calculator)
### Concurrency & workers

# Robustness & Reliability (Generic, not FastAPI-specific)
- **[Server Runners](https://kisspeter.github.io/fastapi-performance-optimization/server_runners)** — Gunicorn vs Uvicorn vs FastAPI CLI as a production runner; the single biggest lever (+50-65%).
- **[Workers & Threads](https://kisspeter.github.io/fastapi-performance-optimization/workers_and_threads)** — how many Gunicorn workers/threads to run on a 2-core container.
- **[Sync vs Async Endpoints](https://kisspeter.github.io/fastapi-performance-optimization/sync_vs_async)** — when async endpoints are worth it (+20-33%).
- **[Thread Pool Sizing](https://kisspeter.github.io/fastapi-performance-optimization/thread_pool_sizing)** — tuning anyio token semaphores for sync endpoints.
- **[Per-Worker Connection Pool](https://kisspeter.github.io/fastapi-performance-optimization/per_worker_connection_pool)** — pool isolation and the thread ceiling.

### Transport & networking

- **[Nginx in Front of FastAPI](https://kisspeter.github.io/fastapi-performance-optimization/nginx_port_socket)** — TCP port vs Unix socket transport.
- **[Keepalive](https://kisspeter.github.io/fastapi-performance-optimization/keepalive)** — HTTP connection reuse.
- **[Connection Pool](https://kisspeter.github.io/fastapi-performance-optimization/connection_pool)** — pool sizing for external resources.

### Response & middleware

- **[Middleware](https://kisspeter.github.io/fastapi-performance-optimization/middleware)** — BaseHTTPMiddleware vs native Starlette ASGI (+35-44%).
- **[Response Class](https://kisspeter.github.io/fastapi-performance-optimization/json_response_class)** — JSONResponse vs ORJSONResponse (+4-13%).

### Tools & debugging

- **[Pool Sizing Calculator](https://kisspeter.github.io/fastapi-performance-optimization/pool_sizing_calculator)** — an interactive calculator for the numbers above.
- **[Profiling](https://kisspeter.github.io/fastapi-performance-optimization/profiling)** — step-by-step cProfile investigation of a slow endpoint.

## Robustness & Reliability (Generic, not FastAPI-specific)

These patterns apply to any Python web application. They contribute to a robust and reliable application but are not FastAPI performance optimizations.

## [Retry Patterns and Circuit Breaker](https://kisspeter.github.io/fastapi-performance-optimization/retry_circuit_breaker)
- **[Retry and Circuit Breaker](https://kisspeter.github.io/fastapi-performance-optimization/retry_circuit_breaker)** — generic resilience patterns for any Python web app.

# Test environment
## Test environment

* All the tests were run on [GitHub Actions](https://github.com/KissPeter/fastapi-performance-optimization/actions/workflows/performance_tuning_measurements.yml)
* Application is built into a container, you can build it like this:
Expand Down Expand Up @@ -138,7 +149,3 @@ docker-compose build
pytest -vv -rP test_files/
```

# Stay tuned for new ideas:
## FastAPI application profiling
### Arbitrary place of code
### Profiling middleware
92 changes: 65 additions & 27 deletions keepalive.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
---
title: Keepalive support
description: "HTTP connection keepalive for FastAPI behind Nginx: +9-12% on small payloads, negligible on 1MB responses. Cheap, safe default — matters more when connections are expensive (HTTPS)."
layout: template
filename: keepalive.md
---
Expand All @@ -21,8 +22,14 @@ http_client = urllib3.PoolManager(
num_pools=10,
)
```
Let's see how to support HTTP connection keep-alive from FastAPI

> **TL;DR** Enable HTTP keepalive between Nginx and your Python app — it is a cheap, safe default. Measured **+8.98% on sync** and **+12.30% on async** small endpoints. On 1MB responses it makes no difference (within noise). The win grows when connection setup is expensive, e.g. **HTTPS**.

## Verdict

* Keepalive is a **small but near-free** win on small payloads: **+8.98%** sync, **+12.30%** async, simply by reusing upstream connections
* On **1MB responses** keepalive is a wash (+0.13% sync, -1.98% async — noise)
* If you use **HTTPS**, connection creation has even higher overhead due to the additional SSL layer, so keepalive matters more

## What is HTTP keepalive?

Expand All @@ -31,58 +38,89 @@ https://en.wikipedia.org/wiki/HTTP_persistent_connection)

<img src="https://upload.wikimedia.org/wikipedia/commons/thumb/d/d5/HTTP_persistent_connection.svg/600px-HTTP_persistent_connection.svg.png" alt="HTTP Keepalive">

Let's see how to support HTTP connection keep-alive from FastAPI as well.

## Measurements

> CI run [29770319196](https://github.com/KissPeter/fastapi-performance-optimization/actions/runs/29770319196) — Python 3.14, Ubuntu latest. Individual run data available in CI logs.
> CI run [36235373898](https://github.com/KissPeter/fastapi-performance-optimization/actions/runs/36235373898) — Python 3.14, Ubuntu latest. All jobs green, no non-2xx responses. Tested on the default app config (w3t1) over a Unix socket: **8009** = no upstream keepalive, **8017** = `keepalive 8` in the nginx upstream.

### Synchronous API endpoint with small request / response

#### Nginx - APP connection, but no keepalive

| **Test attribute** | **Average** |
|-----------------------|---------------|
| Requests per second | 6957.84 |
| Time per request [ms] | — |

| **Test attribute** | **Test run 1** | **Test run 2** | **Test run 3** | **Average** |
|-----------------------|------------------|------------------|------------------|---------------|
| Requests per second | 2716.62 | 2910.12 | 2950.52 | 2859.09 |
| Time per request [ms] | 36.81 | 34.363 | 33.892 | 35.0217 |

#### Nginx - APP connection with keepalive

| **Test attribute** | **Average** | Difference to baseline |
|-----------------------|---------------|--------------------------|
| Requests per second | 7082.61 | +1.79% |
| Time per request [ms] | — | — |

| **Test attribute** | **Test run 1** | **Test run 2** | **Test run 3** | **Average** | Difference to baseline |
|-----------------------|------------------|------------------|------------------|---------------|--------------------------|
| Requests per second | 3147.31 | 3141.11 | 3059.43 | 3115.95 | +8.98% |
| Time per request [ms] | 31.773 | 31.836 | 32.686 | 32.0983 | 2.92 ms |

### Observations
* +1.79% improvement only because we reuse our existing connections
* **+8.98%** throughput, **2.92 ms** lower latency — we simply reuse our existing connections instead of establishing a fresh upstream connection per request

### Asynchronous API endpoint with small request / response

#### Nginx - APP connection, but no keepalive

| **Test attribute** | **Average** |
|-----------------------|---------------|
| Requests per second | 7204.52 |
| Time per request [ms] | — |
| **Test attribute** | **Test run 1** | **Test run 2** | **Test run 3** | **Average** |
|-----------------------|------------------|------------------|------------------|---------------|
| Requests per second | 3445.89 | 3437.83 | 3489.34 | 3457.69 |
| Time per request [ms] | 29.02 | 29.088 | 28.659 | 28.9223 |

#### Nginx - APP connection with keepalive

| **Test attribute** | **Test run 1** | **Test run 2** | **Test run 3** | **Average** | Difference to baseline |
|-----------------------|------------------|------------------|------------------|---------------|--------------------------|
| Requests per second | 3865.38 | 3837.63 | 3945.8 | 3882.94 | +12.30% |
| Time per request [ms] | 25.871 | 26.058 | 25.343 | 25.7573 | 3.16 ms |

### Observations
* **+12.30%** throughput, **3.16 ms** lower latency. Keepalive helps async endpoints at least as much as sync ones by avoiding connection churn on the busy nginx↔app channel

### Synchronous API endpoint with 1MB response

#### Nginx - APP connection, but no keepalive

| **Test attribute** | **Test run 1** | **Test run 2** | **Test run 3** | **Average** |
|-----------------------|------------------|------------------|------------------|---------------|
| Requests per second | 18.08 | 17.79 | 17.96 | 17.9433 |
| Time per request [ms] | 5529.88 | 5619.71 | 5568.45 | 5572.68 |

#### Nginx - APP connection with keepalive

| **Test attribute** | **Average** | Difference to baseline |
|-----------------------|---------------|--------------------------|
| Requests per second | 7039.15 | -2.3% |
| Time per request [ms] | — | — |
| **Test attribute** | **Test run 1** | **Test run 2** | **Test run 3** | **Average** | Difference to baseline |
|-----------------------|------------------|------------------|------------------|---------------|--------------------------|
| Requests per second | 17.89 | 17.99 | 18.02 | 17.9667 | +0.13% |
| Time per request [ms] | 5588.77 | 5557.51 | 5550.14 | 5565.47 | 7.2 ms |

### Observations
* -2.3% regression for the async endpoint with keepalive — needs further investigation
* **+0.13%** — connection setup is nothing compared to moving a megabyte per request, keepalive is irrelevant here

## Verdict
### Asynchronous API endpoint with 1MB response

* Regardless of the use case sync / async endpoint we can improve our overall performance with this tiny change.
* If you use HTTPS connection creation has even higher overhead due the the additional SSL layer
#### Nginx - APP connection, but no keepalive

| **Test attribute** | **Test run 1** | **Test run 2** | **Test run 3** | **Average** |
|-----------------------|------------------|------------------|------------------|---------------|
| Requests per second | 18.14 | 18.11 | 18.24 | 18.1633 |
| Time per request [ms] | 5512.51 | 5521.74 | 5481.62 | 5505.29 |

# Pro tip:
#### Nginx - APP connection with keepalive

| **Test attribute** | **Test run 1** | **Test run 2** | **Test run 3** | **Average** | Difference to baseline |
|-----------------------|------------------|------------------|------------------|---------------|--------------------------|
| Requests per second | 17.94 | 17.78 | 17.69 | 17.8033 | -1.98% |
| Time per request [ms] | 5574.23 | 5623.66 | 5653.73 | 5617.21 | -111.92 ms |

### Observations
* **-1.98%** — same magnitude as the run-to-run noise on this scenario, not a real regression. Keepalive neither helps nor hurts large bodies

## Pro tip

* This is a full **Nginx config for FastAPI** with keepalive support:

Expand Down Expand Up @@ -152,4 +190,4 @@ http {
```html
<hr>

```
```
Loading
Loading