Skip to content

Let callers bound or skip the intermediate bodies read while following redirects and auth flows #1249

Description

@lesnik512

Raised as an issue directly: the template points to "Ideas" and "Potential Issue" discussion categories, but this repo's discussions only have Announcements and Q&A.

When Client.send() follows redirects or runs a multi-step auth flow, it reads every intermediate response body in full before moving on. This happens even with stream=True:

  • _send_handling_redirects: response.read() on each redirect when follow_redirects=True (_client.py:1056, async :1895 in 2.12.0)
  • _send_handling_auth: response.read() on each response the auth flow answers, such as a DigestAuth 401 challenge or a token-refresh retry (_client.py:1020, async :1858)

Most callers never look at those bodies. But a client that talks to untrusted or misbehaving servers, and bounds memory by streaming and capping the final body, can't bound these reads. A server can send a huge body on a 3xx or 401 and the client buffers all of it before the caller sees anything.

Reproduction (httpx2 2.12.0):

import httpx2

pulled = 0


def huge_body():
    global pulled
    for _ in range(1024):
        pulled += 1024
        yield b"x" * 1024


def handler(request: httpx2.Request) -> httpx2.Response:
    if request.url.path == "/redirect":
        return httpx2.Response(302, headers={"location": "/digest"}, content=huge_body())
    if not request.headers.get("authorization", "").startswith("Digest "):
        challenge = 'Digest realm="api", nonce="abc", qop="auth"'
        return httpx2.Response(401, headers={"www-authenticate": challenge}, content=huge_body())
    return httpx2.Response(200, content=b"ok")


with httpx2.Client(
    transport=httpx2.MockTransport(handler),
    auth=httpx2.DigestAuth("user", "pass"),
    follow_redirects=True,
) as client:
    with client.stream("GET", "https://example.test/redirect") as response:
        print(response.status_code, f"{pulled // 1024} KiB of intermediate bodies read before the stream opened")

Output: 200 3072 KiB of intermediate bodies read before the stream opened. That is the 302, the 401, and the 302 again when DigestAuth re-sends the original request.

The only workaround today is to re-implement both loops outside the client: send with follow_redirects=False and a no-op auth, follow response.next_request, and drive auth.sync_auth_flow() by hand. httpware, a client library built on httpx2, does this to enforce its response body cap. It handles redirects in modern-python/httpware#152 and auth flows in modern-python/httpware#153. That works with public API, but it means copying redirect counting, history handling and the private _build_request_auth fallback to URL credentials. Those copies can drift from httpx2.

Possible directions, in increasing scope:

  1. Close intermediate bodies unread when nothing needs them. Redirect bodies are never used by httpx2. Auth-step bodies are only needed when auth.requires_response_body is set. The catch: response.history[i].content would raise ResponseNotRead, so this probably has to be opt-in.
  2. An opt-in client option for (1), keeping today's default.
  3. A general response-size limit on the client, applied to these reads and to read(). This is related to Easy way to retrieve just the first X bytes from a URL encode/httpx#1392 (Easy way to retrieve just the first X bytes from a URL #521 here), which asked for an easy way to bound body size.

Would any of these be welcome?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions