Conversation
PDGGK
marked this pull request as ready for review
October 3, 2026 06:48
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #6043.
Dataset.to_csv()currently renders each batch withpandas.DataFrame.to_csv(path_or_buf=None), which returns CSV text and ignores compression. This change applies compression to the output handle once, then writes all serial or multiprocessing-produced batches through that shared handle.write()returnsNone.Why use pandas I/O helpers?
This uses the internal
pandas.io.common.get_compression_method,infer_compression, andget_handlehelpers to retain pandas' codec dictionary handling, includingarchive_name, gzip options, and TAR support. They are private APIs and carry a compatibility/maintenance risk; the version-specific tests are evidence for the tested runtimes, not a public API stability guarantee.Using only fsspec's compression layer would require translating those options and archive semantics. In the inspected fsspec implementation, TAR is absent from the compression registry and
TarFileSystemis read-only. fsspec's ZIP compression adapter usesfilenamerather than pandas'archive_name, defaults its member name tofile, and usesZipFile's stored default unless configured; pandas' wrapper defaults to deflated ZIP output. The existing fsspec destination opening is retained, with pandas wrapping that binary handle for top-level CSV compression.Behavior changes and limitations
compressionentry instorage_options, the defaultcompression="infer"now produces compressed output for destinations such asout.csv.gzandout.zip. These paths previously received plain CSV bytes. Callers that intentionally need plain bytes with a compressed-looking filename can passcompression=None.compressionkey instorage_optionsnow raisesValueErrorbefore opening or truncating the destination. This also applies when the storage-options value isNone. Top-level compression omitted,None, or"infer"continues to let fsspec own the storage-options compression stream without adding a second layer. Caller-owned buffers continue to ignorestorage_options.batch_sizelimits batch conversion but does not bound this archive buffer. This can be significant for large exports; no large-memory stress claim is made. gzip/bz2/xz/zstd use streaming wrappers instead.mode="a"behavior is unchanged: the existing destination-opening path still useswb. Append coverage uses caller-owned binary append handles. This change does not claim full compatibility with every pandas CSV option.Tests
The regression module covers explicit/inferred/dictionary/buffer compression, gzip/bz2/xz/ZIP/TAR/zstd, multiple batch sizes, serial and two-process generation, byte signatures and decoding, CSV round-trips, one archive member, gzip
mtime,archive_name, compressed TAR, uncompressed controls, buffer ownership, storage-options compatibility, and conflict preflight without truncation.Exact pandas 3.0.6 validation on Python 3.12.14 ran these two modules with the same tests in all three phases:
All 31 existing CSV I/O tests pass in every phase.
Prior investigation
Thanks to exs-avianello for the original report, and to aryanxk02 for reproducing it in July 2023 and discussing the
path_or_buf=Nonecause and a possible direction.