Skip to content

Fix compression in Dataset.to_csv - #8730

Open
PDGGK wants to merge 1 commit into
huggingface:mainfrom
PDGGK:fix/csv-export-compression-6043
Open

PDGGK wants to merge 1 commit into
huggingface:mainfrom
PDGGK:fix/csv-export-compression-6043

Conversation

@PDGGK

@PDGGK PDGGK commented Oct 2, 2026

Copy link
Copy Markdown

Summary

Fixes #6043.

Dataset.to_csv() currently renders each batch with pandas.DataFrame.to_csv(path_or_buf=None), which returns CSV text and ignores compression. This change applies compression to the output handle once, then writes all serial or multiprocessing-produced batches through that shared handle.

  • Honor explicit codecs, compression dictionaries, and filename inference.
  • Keep fsspec responsible for opening paths/URIs and handling storage options.
  • Finalize compression before closing an owned destination, while leaving caller-owned binary buffers open.
  • Write one member for ZIP/TAR exports across multiple batches.
  • Preserve the uncompressed encoded-byte return count for built-in compression writers, including older pandas ZIP wrappers whose write() returns None.

Why use pandas I/O helpers?

This uses the internal pandas.io.common.get_compression_method, infer_compression, and get_handle helpers to retain pandas' codec dictionary handling, including archive_name, gzip options, and TAR support. They are private APIs and carry a compatibility/maintenance risk; the version-specific tests are evidence for the tested runtimes, not a public API stability guarantee.

Using only fsspec's compression layer would require translating those options and archive semantics. In the inspected fsspec implementation, TAR is absent from the compression registry and TarFileSystem is read-only. fsspec's ZIP compression adapter uses filename rather than pandas' archive_name, defaults its member name to file, and uses ZipFile's stored default unless configured; pandas' wrapper defaults to deflated ZIP output. The existing fsspec destination opening is retained, with pandas wrapping that binary handle for top-level CSV compression.

Behavior changes and limitations

  • With no compression entry in storage_options, the default compression="infer" now produces compressed output for destinations such as out.csv.gz and out.zip. These paths previously received plain CSV bytes. Callers that intentionally need plain bytes with a compressed-looking filename can pass compression=None.
  • For path destinations, specifying a top-level codec or compression dictionary together with a compression key in storage_options now raises ValueError before opening or truncating the destination. This also applies when the storage-options value is None. Top-level compression omitted, None, or "infer" continues to let fsspec own the storage-options compression stream without adding a second layer. Caller-owned buffers continue to ignore storage_options.
  • pandas' ZIP and TAR wrappers buffer the entire uncompressed CSV payload in memory before writing the archive on close. batch_size limits batch conversion but does not bound this archive buffer. This can be significant for large exports; no large-memory stress claim is made. gzip/bz2/xz/zstd use streaming wrappers instead.
  • The returned count describes uncompressed encoded CSV bytes accepted by built-in compression wrappers, not compressed file size. Existing numeric custom-writer return values remain respected.
  • Path mode="a" behavior is unchanged: the existing destination-opening path still uses wb. Append coverage uses caller-owned binary append handles. This change does not claim full compatibility with every pandas CSV option.

Tests

The regression module covers explicit/inferred/dictionary/buffer compression, gzip/bz2/xz/ZIP/TAR/zstd, multiple batch sizes, serial and two-process generation, byte signatures and decoding, CSV round-trips, one archive member, gzip mtime, archive_name, compressed TAR, uncompressed controls, buffer ownership, storage-options compatibility, and conflict preflight without truncation.

Exact pandas 3.0.6 validation on Python 3.12.14 ran these two modules with the same tests in all three phases:

python -m pytest tests/io/test_csv_compression.py tests/io/test_csv.py -q
Production state Result
Candidate 231 passed
Only production writer reverted to base 151 failed, 80 passed
Candidate restored 231 passed

All 31 existing CSV I/O tests pass in every phase.

Prior investigation

Thanks to exs-avianello for the original report, and to aryanxk02 for reproducing it in July 2023 and discussing the path_or_buf=None cause and a possible direction.

@PDGGK
PDGGK marked this pull request as ready for review October 3, 2026 06:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Compression kwargs have no effect when saving datasets as csv

1 participant