You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Normalize non-ASCII errors in Base64 bytes loaders - #456
The bytes-like Base64 loaders encode string input as ASCII before decoding. Non-ASCII input therefore leaks UnicodeEncodeError instead of using the provider's ValueLoadError contract.
Fix
Convert UnicodeEncodeError to ValueLoadError("Bad base64 string", data) in the shared bytes loader. This covers bytes, bytearray, BytesIO, and IO[bytes] loaders. Padding behavior is unchanged.
Validation
Python 3.11 concrete-provider tests: 306 passed
Targeted public-seam matrix: 24 passed across result type, coercion, and debug-trail combinations
Mutation check: removing the Unicode handler makes all 24 targeted cases fail
The Coverage job failed before reading any coverage data: its first request, GET https://api.github.com/repos/reagento/adaptix, returned HTTP 503 with GitHub's Unicorn HTML page. All seven test jobs that produced the coverage artifacts passed. I attempted to rerun only the failed job, but fork contributors do not have the required repository permission.
@binggao1230 sorry for the long delay. Thanks for your interest in the project. Catching the UnicodeEncodeError exception looks like a very important improvement, but I still can't quite see the benefit of the additional padding checks. Does extra padding actually cause any problems during decoding?
You are right to ask. I rechecked this, and there is no decoded-byte corruption in these cases. AAA= and AAA== both decode to the same two bytes; AAAA, AAAA=, and AAAA== all decode to the same three bytes; and padding-only strings decode to empty bytes. The additional check therefore only enforces a canonical representation instead of fixing the decoded result. RFC 4648 also permits decoders to ignore excess terminal padding, so without a project requirement for canonical input, I do not think that part has enough benefit. I inferred stricter validation from the existing alphabet and missing-padding checks, but that is not a concrete decoding problem. I can remove the padding checks, related tests, and changelog wording, and keep the UnicodeEncodeError to ValueLoadError fix only.
Removed the excess-padding validation, its tests, and the padding-specific changelog wording. The PR now only converts non-ASCII input from UnicodeEncodeError to ValueLoadError; the 24-case public-seam matrix and all 306 concrete-provider tests pass on the signed follow-up commit.
Fixed the patch-level ParamSpec bound-source expectation on signed commit 67eb34f. The new run is fully green across CPython 3.10-3.14, PyPy 3.10/3.11, lint, coverage, SonarCloud, and docs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The bytes-like Base64 loaders encode string input as ASCII before decoding. Non-ASCII input therefore leaks
UnicodeEncodeErrorinstead of using the provider'sValueLoadErrorcontract.Fix
Convert
UnicodeEncodeErrortoValueLoadError("Bad base64 string", data)in the shared bytes loader. This coversbytes,bytearray,BytesIO, andIO[bytes]loaders. Padding behavior is unchanged.Validation
git diff --check: passed