Skip to content

[Bug]: ONTAP S3 store sends an unpadded x-amz-checksum-sha256, which S3 rejects #2767

Description

@harshavardhana

Version

main at 5d6e287, and every released version back to v1.3.1. The store was added in #1630.

Client

Other (not client-specific; the failure is between NativeLink and the object store)

OS / Architecture

Linux x86_64

Deployment method

Other (any deployment using the ONTAP store)

Storage backend

S3 (provider: "ontap")

Configuration

{
  stores: [
    {
      name: "CAS",
      experimental_cloud_object_store: {
        provider: "ontap",
        endpoint: "https://s3.example.com",
        vserver_name: "us-east-1",
        bucket: "some-bucket",
      },
    },
  ],
}

Expected behavior

A PutObject under 5MB stores the object.

The S3 API specifies x-amz-checksum-sha256 as the base64 encoding of the 32-byte
SHA-256 digest. Base64 of 32 bytes is 44 characters and ends in a single = pad.
Amazon S3 validates this and rejects anything else, and S3-compatible
implementations follow suit.

Actual behavior

The request carries a 43-character unpadded value, the server rejects it with
HTTP 400, and the object is never stored.

Objects at or above 5MB are fine, because the multipart path does not set the header.
So a deployment can look healthy while silently failing every small blob, which for a
CAS is most of them.

The cause is a single call on the small-object upload path:

.set_checksum_algorithm(Some(aws_sdk_s3::types::ChecksumAlgorithm::Sha256))
.set_checksum_sha256(Some(BASE64_STANDARD_NO_PAD.encode(hash)))

BASE64_STANDARD_NO_PAD strips the pad that the specification requires.

The test suite asserts the malformed value, so CI holds the bug in place:

assert_eq!(
headers.get("x-amz-checksum-sha256"),
Some("ZAbm15cMyBMkq2sUXRgXHNb3az8dLCm3tnyixpdxF+o")
);

Logs / error output

Here is a reproducer against Amazon S3 itself, to show this is the protocol rejecting
the value rather than one vendor being strict. It sends the same object twice, once
with the padded encoding and once with the padding stripped.

import base64
import hashlib

import boto3
from botocore.exceptions import ClientError

BUCKET = "your-bucket"
CONTENT = b"Demonstrating base64 padding enforcement in the S3 checksum header."

s3 = boto3.client("s3")

digest = hashlib.sha256(CONTENT).digest()
padded = base64.b64encode(digest).decode()
unpadded = padded.rstrip("=")

print(f"padded:   {padded} ({len(padded)} chars)")
print(f"unpadded: {unpadded} ({len(unpadded)} chars)\n")

for label, checksum in (("unpadded", unpadded), ("padded", padded)):
    try:
        s3.put_object(
            Bucket=BUCKET,
            Key=f"checksum-{label}.txt",
            Body=CONTENT,
            ChecksumAlgorithm="SHA256",
            ChecksumSHA256=checksum,
        )
        print(f"{label}: stored")
    except ClientError as err:
        error = err.response["Error"]
        print(f"{label}: rejected -- {error['Code']}: {error['Message']}")

Output:

padded:   jMemPDf6TOWab3EJ+6oBwisHwRT5iwXsMmR+uEU/bFI= (44 chars)
unpadded: jMemPDf6TOWab3EJ+6oBwisHwRT5iwXsMmR+uEU/bFI (43 chars)

unpadded: rejected -- InvalidRequest: Value for x-amz-checksum-sha256 header is invalid.
padded: stored

The unpadded value NativeLink sends is not accepted by Amazon S3, so the ONTAP store
is not usable against a spec-compliant endpoint for objects under 5MB.

Additional context

The deeper problem is that this store exists at all.

ontap_s3_store.rs is a 790-line reimplementation of StoreDriver that is
structurally a fork of s3_store.rs: the same constants, the same make_s3_path,
the same has / has_with_results / update / get_part layout, the same
HealthStatusIndicator tail. Hand-rolling the checksum is one of the few places the
two files differ, and it is the place the bug lives. s3_store.rs never had this
problem, because it lets the AWS SDK compute the header.

The fork has drifted in other ways too. S3Store implements optimized_for and
register_health; the ONTAP copy implements neither. Its retry buffer defaults to
20MB against S3Store's 5MB, with no stated reason. Every future fix to s3_store.rs
has to be remembered a second time here, and #2747 is an example of one that was not.

The repository already has a better pattern for this, added after #1630. Both the
Cloudflare R2 store and the OCI store are thin config adapters that build an S3 SDK
client and hand it to S3Store:

  • r2_store.rs — 102 lines
  • oci_store.rs — 144 lines, and it already handles the two things ONTAP needs,
    force_path_style(true) and explicit checksum configuration through
    RequestChecksumCalculation / ResponseChecksumValidation

Everything genuinely ONTAP-specific is client construction, not store behavior:
path-style addressing, a region taken from vserver_name, longer timeouts, and a
custom CA bundle. All of it fits the adapter shape.

Proposal

  1. Add the missing knobs to ExperimentalAwsSpec and S3Store so it can address a
    generic S3-compatible endpoint: endpoint, force_path_style, explicit
    credentials, and checksum configuration. ExperimentalAwsSpec uses
    deny_unknown_fields, so adding optional fields keeps existing configs working.
  2. Rewrite OntapS3Store as an adapter over S3Store, the way R2Store and
    OciStore already are.
  3. Remove the duplicated StoreDriver implementation.

That fixes the checksum bug as a side effect, since the SDK emits a correctly padded
header, and it stops the two files from drifting further apart.

If keeping the fork is preferred, the one-line fix is BASE64_STANDARD in place of
BASE64_STANDARD_NO_PAD, plus updating the test expectation to
ZAbm15cMyBMkq2sUXRgXHNb3az8dLCm3tnyixpdxF+o=. I would rather do the consolidation,
and I am happy to send that PR. I will open one for step 1 shortly.

The root_certificates handling is left out of step 1 on purpose, since #2747 already
proposes moving it onto CommonObjectSpec for every S3-compatible store.

Activity

  1. MarcusSorealheis commented on Sep 16, 2026

    @MarcusSorealheis
    Member

    Wow. Incredible issue to read. Thank you much for opening it and we will support you in whatever way we can beginning with a review for Step 1.

  2. harshavardhana commented on Sep 16, 2026

    @harshavardhana
    ContributorAuthor

    One finding from #2768 that matters for step 2 above, so it does not have to be rediscovered.

    With flexible checksums enabled, the AWS SDK uploads a single-shot PutObject as Content-Encoding: aws-chunked with x-amz-content-sha256: STREAMING-UNSIGNED-PAYLOAD-TRAILER and the checksum in a trailer, rather than as a plain PUT with an x-amz-checksum-* header. That is the default and it is fine for AWS.

    It is not fine everywhere. oci_store.rs already sets RequestChecksumCalculation::WhenRequired because OCI rejects the chunked form, and the x-amz-content-sha256: UNSIGNED-PAYLOAD override in ontap_s3_store.rs looks like it is there for the same reason. So the hand-rolled checksum header in the ONTAP store was probably an attempt to keep a checksum while avoiding aws-chunked, and the encoding is what went wrong.

    That means rewriting the ONTAP store as an adapter needs a way to express "checksums on, chunked upload off", and whether ONTAP accepts the chunked form should be confirmed against a real endpoint first. #2768 deliberately leaves this alone and changes nothing on the wire for existing users.

  3. palfrey commented on Sep 18, 2026

    @palfrey
    Member

    Thanks. Now #2768 is merged, the primary item of this is done, so I'll be closing it. I've opened #2779 as a future item regarding actually rewriting the ONTAP store.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions