Version
main at 5d6e287, and every released version back to v1.3.1. The store was added in #1630.
Client
Other (not client-specific; the failure is between NativeLink and the object store)
OS / Architecture
Linux x86_64
Deployment method
Other (any deployment using the ONTAP store)
Storage backend
S3 (provider: "ontap")
Configuration
{
stores: [
{
name: "CAS",
experimental_cloud_object_store: {
provider: "ontap",
endpoint: "https://s3.example.com",
vserver_name: "us-east-1",
bucket: "some-bucket",
},
},
],
}
Expected behavior
A PutObject under 5MB stores the object.
The S3 API specifies x-amz-checksum-sha256 as the base64 encoding of the 32-byte
SHA-256 digest. Base64 of 32 bytes is 44 characters and ends in a single = pad.
Amazon S3 validates this and rejects anything else, and S3-compatible
implementations follow suit.
Actual behavior
The request carries a 43-character unpadded value, the server rejects it with
HTTP 400, and the object is never stored.
Objects at or above 5MB are fine, because the multipart path does not set the header.
So a deployment can look healthy while silently failing every small blob, which for a
CAS is most of them.
The cause is a single call on the small-object upload path:
|
.set_checksum_algorithm(Some(aws_sdk_s3::types::ChecksumAlgorithm::Sha256)) |
|
.set_checksum_sha256(Some(BASE64_STANDARD_NO_PAD.encode(hash))) |
BASE64_STANDARD_NO_PAD strips the pad that the specification requires.
The test suite asserts the malformed value, so CI holds the bug in place:
|
assert_eq!( |
|
headers.get("x-amz-checksum-sha256"), |
|
Some("ZAbm15cMyBMkq2sUXRgXHNb3az8dLCm3tnyixpdxF+o") |
|
); |
Logs / error output
Here is a reproducer against Amazon S3 itself, to show this is the protocol rejecting
the value rather than one vendor being strict. It sends the same object twice, once
with the padded encoding and once with the padding stripped.
import base64
import hashlib
import boto3
from botocore.exceptions import ClientError
BUCKET = "your-bucket"
CONTENT = b"Demonstrating base64 padding enforcement in the S3 checksum header."
s3 = boto3.client("s3")
digest = hashlib.sha256(CONTENT).digest()
padded = base64.b64encode(digest).decode()
unpadded = padded.rstrip("=")
print(f"padded: {padded} ({len(padded)} chars)")
print(f"unpadded: {unpadded} ({len(unpadded)} chars)\n")
for label, checksum in (("unpadded", unpadded), ("padded", padded)):
try:
s3.put_object(
Bucket=BUCKET,
Key=f"checksum-{label}.txt",
Body=CONTENT,
ChecksumAlgorithm="SHA256",
ChecksumSHA256=checksum,
)
print(f"{label}: stored")
except ClientError as err:
error = err.response["Error"]
print(f"{label}: rejected -- {error['Code']}: {error['Message']}")
Output:
padded: jMemPDf6TOWab3EJ+6oBwisHwRT5iwXsMmR+uEU/bFI= (44 chars)
unpadded: jMemPDf6TOWab3EJ+6oBwisHwRT5iwXsMmR+uEU/bFI (43 chars)
unpadded: rejected -- InvalidRequest: Value for x-amz-checksum-sha256 header is invalid.
padded: stored
The unpadded value NativeLink sends is not accepted by Amazon S3, so the ONTAP store
is not usable against a spec-compliant endpoint for objects under 5MB.
Additional context
The deeper problem is that this store exists at all.
ontap_s3_store.rs is a 790-line reimplementation of StoreDriver that is
structurally a fork of s3_store.rs: the same constants, the same make_s3_path,
the same has / has_with_results / update / get_part layout, the same
HealthStatusIndicator tail. Hand-rolling the checksum is one of the few places the
two files differ, and it is the place the bug lives. s3_store.rs never had this
problem, because it lets the AWS SDK compute the header.
The fork has drifted in other ways too. S3Store implements optimized_for and
register_health; the ONTAP copy implements neither. Its retry buffer defaults to
20MB against S3Store's 5MB, with no stated reason. Every future fix to s3_store.rs
has to be remembered a second time here, and #2747 is an example of one that was not.
The repository already has a better pattern for this, added after #1630. Both the
Cloudflare R2 store and the OCI store are thin config adapters that build an S3 SDK
client and hand it to S3Store:
r2_store.rs — 102 lines
oci_store.rs — 144 lines, and it already handles the two things ONTAP needs,
force_path_style(true) and explicit checksum configuration through
RequestChecksumCalculation / ResponseChecksumValidation
Everything genuinely ONTAP-specific is client construction, not store behavior:
path-style addressing, a region taken from vserver_name, longer timeouts, and a
custom CA bundle. All of it fits the adapter shape.
Proposal
- Add the missing knobs to
ExperimentalAwsSpec and S3Store so it can address a
generic S3-compatible endpoint: endpoint, force_path_style, explicit
credentials, and checksum configuration. ExperimentalAwsSpec uses
deny_unknown_fields, so adding optional fields keeps existing configs working.
- Rewrite
OntapS3Store as an adapter over S3Store, the way R2Store and
OciStore already are.
- Remove the duplicated
StoreDriver implementation.
That fixes the checksum bug as a side effect, since the SDK emits a correctly padded
header, and it stops the two files from drifting further apart.
If keeping the fork is preferred, the one-line fix is BASE64_STANDARD in place of
BASE64_STANDARD_NO_PAD, plus updating the test expectation to
ZAbm15cMyBMkq2sUXRgXHNb3az8dLCm3tnyixpdxF+o=. I would rather do the consolidation,
and I am happy to send that PR. I will open one for step 1 shortly.
The root_certificates handling is left out of step 1 on purpose, since #2747 already
proposes moving it onto CommonObjectSpec for every S3-compatible store.
Version
mainat 5d6e287, and every released version back to v1.3.1. The store was added in #1630.Client
Other (not client-specific; the failure is between NativeLink and the object store)
OS / Architecture
Linux x86_64
Deployment method
Other (any deployment using the ONTAP store)
Storage backend
S3 (
provider: "ontap")Configuration
Expected behavior
A
PutObjectunder 5MB stores the object.The S3 API specifies
x-amz-checksum-sha256as the base64 encoding of the 32-byteSHA-256 digest. Base64 of 32 bytes is 44 characters and ends in a single
=pad.Amazon S3 validates this and rejects anything else, and S3-compatible
implementations follow suit.
Actual behavior
The request carries a 43-character unpadded value, the server rejects it with
HTTP 400, and the object is never stored.
Objects at or above 5MB are fine, because the multipart path does not set the header.
So a deployment can look healthy while silently failing every small blob, which for a
CAS is most of them.
The cause is a single call on the small-object upload path:
nativelink/nativelink-store/src/ontap_s3_store.rs
Lines 400 to 401 in 5d6e287
BASE64_STANDARD_NO_PADstrips the pad that the specification requires.The test suite asserts the malformed value, so CI holds the bug in place:
nativelink/nativelink-store/tests/ontap_s3_store_test.rs
Lines 309 to 312 in 5d6e287
Logs / error output
Here is a reproducer against Amazon S3 itself, to show this is the protocol rejecting
the value rather than one vendor being strict. It sends the same object twice, once
with the padded encoding and once with the padding stripped.
Output:
The unpadded value NativeLink sends is not accepted by Amazon S3, so the ONTAP store
is not usable against a spec-compliant endpoint for objects under 5MB.
Additional context
The deeper problem is that this store exists at all.
ontap_s3_store.rsis a 790-line reimplementation ofStoreDriverthat isstructurally a fork of
s3_store.rs: the same constants, the samemake_s3_path,the same
has/has_with_results/update/get_partlayout, the sameHealthStatusIndicatortail. Hand-rolling the checksum is one of the few places thetwo files differ, and it is the place the bug lives.
s3_store.rsnever had thisproblem, because it lets the AWS SDK compute the header.
The fork has drifted in other ways too.
S3Storeimplementsoptimized_forandregister_health; the ONTAP copy implements neither. Its retry buffer defaults to20MB against
S3Store's 5MB, with no stated reason. Every future fix tos3_store.rshas to be remembered a second time here, and #2747 is an example of one that was not.
The repository already has a better pattern for this, added after #1630. Both the
Cloudflare R2 store and the OCI store are thin config adapters that build an S3 SDK
client and hand it to
S3Store:r2_store.rs— 102 linesoci_store.rs— 144 lines, and it already handles the two things ONTAP needs,force_path_style(true)and explicit checksum configuration throughRequestChecksumCalculation/ResponseChecksumValidationEverything genuinely ONTAP-specific is client construction, not store behavior:
path-style addressing, a region taken from
vserver_name, longer timeouts, and acustom CA bundle. All of it fits the adapter shape.
Proposal
ExperimentalAwsSpecandS3Storeso it can address ageneric S3-compatible endpoint:
endpoint,force_path_style, explicitcredentials, and checksum configuration.
ExperimentalAwsSpecusesdeny_unknown_fields, so adding optional fields keeps existing configs working.OntapS3Storeas an adapter overS3Store, the wayR2StoreandOciStorealready are.StoreDriverimplementation.That fixes the checksum bug as a side effect, since the SDK emits a correctly padded
header, and it stops the two files from drifting further apart.
If keeping the fork is preferred, the one-line fix is
BASE64_STANDARDin place ofBASE64_STANDARD_NO_PAD, plus updating the test expectation toZAbm15cMyBMkq2sUXRgXHNb3az8dLCm3tnyixpdxF+o=. I would rather do the consolidation,and I am happy to send that PR. I will open one for step 1 shortly.
The
root_certificateshandling is left out of step 1 on purpose, since #2747 alreadyproposes moving it onto
CommonObjectSpecfor every S3-compatible store.