Skip to content

Report pages shared copy-on-write as Shared in /proc/[pid]/smaps. - #15430

Open
copybara-service[bot] wants to merge 1 commit into
masterfrom
test/cl994851982
Open

copybara-service[bot] wants to merge 1 commit into
masterfrom
test/cl994851982

Conversation

@copybara-service

@copybara-service copybara-service Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Report pages shared copy-on-write as Shared in /proc/[pid]/smaps.

smaps reported every resident page of a writable vma as Private_Dirty and
never reported Shared_Clean or Shared_Dirty. On Linux, a private anonymous page
that a fork()ed child still shares with its parent has mapcount 2 and is
Shared_Dirty in both (fs/proc/task_mmu.c:smaps_account()); it becomes
Private_Dirty in each once one of them writes it and breaks copy-on-write.

Programs measure copy-on-write this way: Valkey's and Redis's fork()ed RDB/AOF
child sums Private_Dirty from its own smaps (zmalloc_get_private_dirty) and
reports it as current_cow_size and rdb_last_cow_size. Under gVisor the child's
COW size starts at the whole heap and never grows as the parent writes, so
Valkey's "Test child sending info" fails with "COW info wasn't reported".

For private copy-on-write pmas, smaps now reports pages on which the
MemoryFile holds more than one reference as shared (dirty if the vma is
writable, clean otherwise), and divides each such page among its references in
Pss, as Linux divides by mapcount. A copy-on-write page's references are the
pmas mapping it. Other private pmas are skipped because their extra references
are pins, which Linux reports as private.

Reading smaps now walks the MemoryFile's reference counts for copy-on-write
pmas. In the worst layouts measured (1 GiB forked, then every other page
written by the parent, so that each page is its own pma or reference-count
segment) a read takes 4-32 ms instead of 0.2-12 ms; Linux, which walks every
page table entry, takes 11-57 ms for the same layouts on the same hosts.

Tests: proc_pid_smaps_test ForkCopyOnWrite fails before this change on runsc
(Shared_Dirty 0, Private_Dirty 64 kB and Pss 64 kB right after fork) and
passes after it and on native.
FUTURE_COPYBARA_INTEGRATE_REVIEW=#15437 from tamird:tuntap-leak-check 3b35968

@copybara-service copybara-service Bot added the exported Issue was exported automatically label Oct 7, 2026
@copybara-service
copybara-service Bot force-pushed the test/cl994851982 branch 2 times, most recently from d937d65 to 2fc6e6f Compare October 7, 2026 16:23
@copybara-service
copybara-service Bot force-pushed the test/cl994851982 branch 3 times, most recently from 6936ff5 to 42fde64 Compare October 7, 2026 21:59
@copybara-service
copybara-service Bot requested a review from relkochta as a code owner October 7, 2026 21:59
smaps reported every resident page of a writable vma as Private_Dirty and
never reported Shared_Clean or Shared_Dirty. On Linux, a private anonymous page
that a fork()ed child still shares with its parent has mapcount 2 and is
Shared_Dirty in both (fs/proc/task_mmu.c:smaps_account()); it becomes
Private_Dirty in each once one of them writes it and breaks copy-on-write.

Programs measure copy-on-write this way: Valkey's and Redis's fork()ed RDB/AOF
child sums Private_Dirty from its own smaps (zmalloc_get_private_dirty) and
reports it as current_cow_size and rdb_last_cow_size. Under gVisor the child's
COW size starts at the whole heap and never grows as the parent writes, so
Valkey's "Test child sending info" fails with "COW info wasn't reported".

For private copy-on-write pmas, smaps now reports pages on which the
MemoryFile holds more than one reference as shared (dirty if the vma is
writable, clean otherwise), and divides each such page among its references in
Pss, as Linux divides by mapcount. A copy-on-write page's references are the
pmas mapping it. Other private pmas are skipped because their extra references
are pins, which Linux reports as private.

Reading smaps now walks the MemoryFile's reference counts for copy-on-write
pmas. In the worst layouts measured (1 GiB forked, then every other page
written by the parent, so that each page is its own pma or reference-count
segment) a read takes 4-32 ms instead of 0.2-12 ms; Linux, which walks every
page table entry, takes 11-57 ms for the same layouts on the same hosts.

Tests: proc_pid_smaps_test ForkCopyOnWrite fails before this change on runsc
(Shared_Dirty 0, Private_Dirty 64 kB and Pss 64 kB right after fork) and
passes after it and on native.
FUTURE_COPYBARA_INTEGRATE_REVIEW=#15437 from tamird:tuntap-leak-check 3b35968
PiperOrigin-RevId: 994851982
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

exported Issue was exported automatically

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant