Repository navigation
Keep Headers._encoding when constructing from another Headers instance #3763
Replies: 2 comments
|
You have identified a subtle behavior in Here is a breakdown of why this happens, how to ensure UTF-8 header encoding today, and how it relates to client design: 1. Why
|
|
There are two separate properties here:
import httpx
source = httpx.Headers({"x": "é"}, encoding="utf-8")
with httpx.Client(
headers=source,
transport=httpx.MockTransport(lambda request: httpx.Response(204)),
) as client:
request = client.build_request("GET", "https://example.test")
assert (b"x", b"\xc3\xa9") in request.headers.raw
assert request.headers.encoding == "utf-8"This works unchanged in HTTPX 0.28.1 and current master. Therefore, for normal valid UTF-8 values already constructed through The real inconsistency is that explicit policy is not preserved. For example: source = httpx.Headers({"x": "plain"}, encoding="utf-8")
copied = httpx.Headers(source)
assert copied.raw == source.raw
assert source.encoding == "utf-8"
assert copied.encoding == "ascii"That distinction matters when bytes are ambiguous under multiple encodings. For example, UTF-8 bytes explicitly interpreted as ISO-8859-1 will be reinterpreted as UTF-8 after copying, although the raw bytes remain identical. So preserving |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I’m trying to build a custom client that defaults header encoding to UTF-8 (related: #3238).
While looking at
httpx.Headers.__init__, it seems that when you pass an existing Headers instance, it copies ._list but ignores ._encoding.httpx/httpx/_models.py
Lines 151 to 164 in ae1b9f6
If I do Headers(existing_headers) (what BaseClient._merge_headers does), I expected it to behave like a copy, including keeping the same encoding; unless I explicitly pass encoding=.
I believe it should carry over its encoding (unless encoding is provided):
If this is intentional, what’s the recommended way to make a client default to UTF-8 header encoding without losing it when headers get copied around?
All reactions