Expected Behavior
Archiving to Google Cloud Storage should allocate memory in proportion to the archived file, which is usually a few KB.
Actual Behavior
Every upload made by the gcloud archiver allocates a 16 MiB buffer, whatever the file size.
common/archiver/gcloud/connector/client.go creates the writer with object.NewWriter(ctx) and never sets ChunkSize. The storage client then uses its default of 16 MiB (googleapi.DefaultUploadChunkSize), and gensupport.NewMediaBuffer allocates a buffer of that size for every writer so it can retry the request. The storage.Writer.ChunkSize docs recommend the opposite for small objects:
If you upload small objects (< 16MiB), you should set ChunkSize to a value slightly larger than the objects' sizes to avoid memory bloat. This is especially important if you are uploading many small objects concurrently.
Each closed execution leads to up to three uploads (one history blob and two visibility records), so the history service allocates 16 MiB per file at the rate executions close.
In our production cluster (v1.32.0, history and visibility archival to GCS, tens of uploads per second at peak), an allocation profile of a history pod showed:
- 87% of allocated bytes came from
gensupport.NewMediaBuffer (called by gensupport.PrepareUpload), about 4 GB per minute per pod;
- about 64% of the CPU samples at peak were GC mark workers. Most were idle-priority workers, but they still count as container CPU, so the HPA scaled history out without more real load.
Setting GOGC=200 and GOMEMLIMIT cut CPU per request by about 28%, but the allocation rate is the same. There is no config to avoid it: GstorageArchiver only has credentialsPath.
Steps to Reproduce the Problem
- Enable history and visibility archival with the
gstorage provider.
- Run workflows that close at a steady rate.
- Take an allocation profile of a history pod (
/debug/pprof/allocs) and look for gensupport.NewMediaBuffer.
The benchmark below isolates the storage client: one upload per op to a fake GCS JSON API server, with the same cloud.google.com/go/storage v1.62.1 as main. The time column comes from a local fake server and says nothing about real GCS latency.
| File |
ChunkSize |
B/op |
ns/op |
| 20 KB |
default (16 MiB) |
16,875,766 |
425,853 |
| 20 KB |
file size + 1 |
322,519 |
138,966 |
| 2 MB |
default (16 MiB) |
16,874,510 |
1,135,904 |
| 2 MB |
file size + 1 |
2,448,968 |
1,208,443 |
Benchmark code
package chunkbench
import (
"context"
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"cloud.google.com/go/storage"
"google.golang.org/api/googleapi"
)
// Fake GCS JSON API: accepts any upload and answers with minimal object metadata.
func newClient(b *testing.B) *storage.Client {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = io.Copy(io.Discard, r.Body)
w.Header().Set("Content-Type", "application/json")
_, _ = io.WriteString(w, `{"bucket":"b","name":"o","size":"1"}`)
}))
b.Cleanup(srv.Close)
b.Setenv("STORAGE_EMULATOR_HOST", strings.TrimPrefix(srv.URL, "http://"))
c, err := storage.NewClient(context.Background())
if err != nil {
b.Fatal(err)
}
b.Cleanup(func() { _ = c.Close() })
return c
}
func bench(b *testing.B, size int, chunkSize func(int) (int, bool)) {
c := newClient(b)
file := make([]byte, size)
ctx := context.Background()
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
w := c.Bucket("b").Object("o").NewWriter(ctx)
if cs, ok := chunkSize(size); ok {
w.ChunkSize = cs
}
if _, err := w.Write(file); err != nil {
b.Fatal(err)
}
if err := w.Close(); err != nil {
b.Fatal(err)
}
}
}
var (
defaultCS = func(int) (int, bool) { return 0, false }
fileCS = func(n int) (int, bool) { return min(n+1, googleapi.DefaultUploadChunkSize), true }
)
func BenchmarkUpload20KB_Default(b *testing.B) { bench(b, 20<<10, defaultCS) }
func BenchmarkUpload20KB_FileSize(b *testing.B) { bench(b, 20<<10, fileCS) }
func BenchmarkUpload2MB_Default(b *testing.B) { bench(b, 2<<20, defaultCS) }
func BenchmarkUpload2MB_FileSize(b *testing.B) { bench(b, 2<<20, fileCS) }
Run with go test -run '^$' -bench . -benchtime 300x.
Specifications
- Version: v1.32.0. The same code is on
main (4db93cd).
- Platform: GKE, archival to GCS (
gstorage provider).
Proposed fix
Set ChunkSize to the file size plus one byte, capped at the 16 MiB default:
- The writer still keeps the whole file in a buffer, so retries work as today.
- Files under 16 MiB still go in a single request. The extra byte lets the client reach EOF before the buffer fills. Without it, a file whose size is an exact multiple of 256 KiB would switch to a resumable upload, because the client rounds
ChunkSize up to a multiple of 256 KiB.
- Files of 16 MiB or more keep the current behavior.
ChunkSize = 0 would avoid the buffer completely, but it turns off the client's retries, so I left it out.
I have a PR with this change and will link it here.
Expected Behavior
Archiving to Google Cloud Storage should allocate memory in proportion to the archived file, which is usually a few KB.
Actual Behavior
Every upload made by the gcloud archiver allocates a 16 MiB buffer, whatever the file size.
common/archiver/gcloud/connector/client.gocreates the writer withobject.NewWriter(ctx)and never setsChunkSize. The storage client then uses its default of 16 MiB (googleapi.DefaultUploadChunkSize), andgensupport.NewMediaBufferallocates a buffer of that size for every writer so it can retry the request. Thestorage.Writer.ChunkSizedocs recommend the opposite for small objects:Each closed execution leads to up to three uploads (one history blob and two visibility records), so the history service allocates 16 MiB per file at the rate executions close.
In our production cluster (v1.32.0, history and visibility archival to GCS, tens of uploads per second at peak), an allocation profile of a history pod showed:
gensupport.NewMediaBuffer(called bygensupport.PrepareUpload), about 4 GB per minute per pod;Setting
GOGC=200andGOMEMLIMITcut CPU per request by about 28%, but the allocation rate is the same. There is no config to avoid it:GstorageArchiveronly hascredentialsPath.Steps to Reproduce the Problem
gstorageprovider./debug/pprof/allocs) and look forgensupport.NewMediaBuffer.The benchmark below isolates the storage client: one upload per op to a fake GCS JSON API server, with the same
cloud.google.com/go/storagev1.62.1 asmain. The time column comes from a local fake server and says nothing about real GCS latency.Benchmark code
Run with
go test -run '^$' -bench . -benchtime 300x.Specifications
main(4db93cd).gstorageprovider).Proposed fix
Set
ChunkSizeto the file size plus one byte, capped at the 16 MiB default:ChunkSizeup to a multiple of 256 KiB.ChunkSize = 0would avoid the buffer completely, but it turns off the client's retries, so I left it out.I have a PR with this change and will link it here.