Skip to content

Accept single-bit ranges in bitmap lists - #15446

Merged
copybara-service[bot] merged 1 commit into
google:masterfrom
tamird:bitmap-inclusive-ranges
Oct 7, 2026
Merged

copybara-service[bot] merged 1 commit into
google:masterfrom
tamird:bitmap-inclusive-ranges

Conversation

@tamird

@tamird tamird commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Writing "1-1" to cpuset.cpus failed with EINVAL because the bitmap
parser rejected ranges with equal endpoints. CpusetCgroup.SetMask
reaches this case on a two-CPU system when it removes CPU 0 from the
initial mask. Linux accepts equal endpoints as a single-bit range 1.

Reject only descending ranges and correct the existing parser case that
expected an equal-endpoint range to fail. This applies to both cgroup
versions, whose CPU and memory-node masks use the same parser.

Assisted-by: OpenAI Codex

Writing "1-1" to cpuset.cpus failed with EINVAL because the bitmap parser
rejected ranges with equal endpoints. CpusetCgroup.SetMask reaches this
case on a two-CPU system when it removes CPU 0 from the initial mask.
Linux accepts equal endpoints as a single-bit range [1].

Reject only descending ranges and correct the existing parser case that
expected an equal-endpoint range to fail. This applies to both cgroup
versions, whose CPU and memory-node masks use the same parser.

[1]: https://github.com/torvalds/linux/blob/e5f0a698b/lib/bitmap-str.c#L220-L231

Assisted-by: OpenAI Codex
copybara-service Bot pushed a commit that referenced this pull request Oct 7, 2026
Writing "1-1" to cpuset.cpus failed with EINVAL because the bitmap parser rejected ranges with equal endpoints. CpusetCgroup.SetMask reaches this case on a two-CPU system when it removes CPU 0 from the initial mask. Linux accepts equal endpoints as a single-bit range [1].

Reject only descending ranges and correct the existing parser case that expected an equal-endpoint range to fail. This applies to both cgroup versions, whose CPU and memory-node masks use the same parser.

[1]: https://github.com/torvalds/linux/blob/e5f0a698b/lib/bitmap-str.c#L220-L231

Assisted-by: OpenAI Codex

<!-- codex-thread: 01a0680f-bbf5-7511-95f7-4414f4e10d2e -->

FUTURE_COPYBARA_INTEGRATE_REVIEW=#15446 from tamird:bitmap-inclusive-ranges 01d6bdf
PiperOrigin-RevId: 995164008
copybara-service Bot pushed a commit that referenced this pull request Oct 7, 2026
Writing "1-1" to cpuset.cpus failed with EINVAL because the bitmap parser rejected ranges with equal endpoints. CpusetCgroup.SetMask reaches this case on a two-CPU system when it removes CPU 0 from the initial mask. Linux accepts equal endpoints as a single-bit range [1].

Reject only descending ranges and correct the existing parser case that expected an equal-endpoint range to fail. This applies to both cgroup versions, whose CPU and memory-node masks use the same parser.

[1]: https://github.com/torvalds/linux/blob/e5f0a698b/lib/bitmap-str.c#L220-L231

Assisted-by: OpenAI Codex

<!-- codex-thread: 01a0680f-bbf5-7511-95f7-4414f4e10d2e -->

FUTURE_COPYBARA_INTEGRATE_REVIEW=#15446 from tamird:bitmap-inclusive-ranges 01d6bdf
PiperOrigin-RevId: 995164008
copybara-service Bot pushed a commit that referenced this pull request Oct 7, 2026
smaps reported every resident page of a writable vma as Private_Dirty and
never reported Shared_Clean or Shared_Dirty. On Linux, a private anonymous page
that a fork()ed child still shares with its parent has mapcount 2 and is
Shared_Dirty in both (fs/proc/task_mmu.c:smaps_account()); it becomes
Private_Dirty in each once one of them writes it and breaks copy-on-write.

Programs measure copy-on-write this way: Valkey's and Redis's fork()ed RDB/AOF
child sums Private_Dirty from its own smaps (zmalloc_get_private_dirty) and
reports it as current_cow_size and rdb_last_cow_size. Under gVisor the child's
COW size starts at the whole heap and never grows as the parent writes, so
Valkey's "Test child sending info" fails with "COW info wasn't reported".

For private copy-on-write pmas, smaps now reports pages on which the
MemoryFile holds more than one reference as shared (dirty if the vma is
writable, clean otherwise). Other private pmas are skipped because their extra
references are pins, which Linux reports as private. Pss is still Rss.

Tests: proc_pid_smaps_test ForkCopyOnWrite fails before this change on runsc
(Shared_Dirty 0, Private_Dirty 64 kB right after fork) and passes after it and
on native.
FUTURE_COPYBARA_INTEGRATE_REVIEW=#15446 from tamird:bitmap-inclusive-ranges 01d6bdf
PiperOrigin-RevId: 994851982
copybara-service Bot pushed a commit that referenced this pull request Oct 7, 2026
The fd and fdinfo directories and the fd/N symlinks took their owner from the
task's credentials when the inode was created. /proc/PID entries are cached,
so a task that was looked up while root and then switched to another uid kept
a root-owned, mode 0500 /proc/PID/fd, and processes running as the new uid got
EACCES listing its fds. setpriv does exactly this (libcap-ng reads
/proc/self/status before setresuid), which breaks tools that map sockets to
processes through socket:[N] links (ss -p, lsof, fuser) when run as the
service user.

Linux recomputes the owner of /proc/PID entries on every revalidation
(fs/proc/base.c:pid_revalidate -> pid_update_inode -> task_dump_owner), and of
fd/N in fs/proc/fd.c:tid_fd_revalidate. The fd/N symlinks and the fdinfo
directory are now taskOwnedInodes like the rest of /proc/PID. The fd directory
keeps its same-thread-group permission exception and computes its owner the
same way.

Wrapping these inodes made them implement kernel.TaskOwnedInode, so
pidfd_send_signal and setns on them took the /proc/PID pidfd path and failed
with ESRCH. Any proc file other than /proc/PID itself now returns EBADF there,
as on Linux, which also fixes the errno for files such as /proc/PID/status.

Tests: proc_test ProcPid.FdOwnerFollowsSetuid fails on runsc before this
change and passes on Linux. pidfd_test ProcPidFdinfoIsNotAPidfd,
ProcPidFdLinkIsNotAPidfd and ProcPidStatusIsNotAPidfd check the errnos.
FUTURE_COPYBARA_INTEGRATE_REVIEW=#15446 from tamird:bitmap-inclusive-ranges 01d6bdf
PiperOrigin-RevId: 994853358
copybara-service Bot pushed a commit that referenced this pull request Oct 7, 2026
Linux picks the SIGSEGV si_code from the vma lookup: no vma covering the
address is SEGV_MAPERR, a vma that forbids the access is SEGV_ACCERR
(arch/x86/mm/fault.c:do_user_addr_fault(): bad_area_nosemaphore() vs
bad_area_access_error()). After a failed HandleUserFault the sentry corrected
only the ACCERR case (EPERM) and otherwise kept the platform's code. The KVM
platform derives that code from the access type alone (vCPU.fault reports any
write or instruction fetch as SEGV_ACCERR), so on KVM a write to or a jump into
an unmapped page raised SEGV_ACCERR. systrap and ptrace forward the host's
siginfo and are unaffected.

Valgrind extends a client stack only on SEGV_MAPERR inside its stack
reservation (coregrind/m_signals.c:extend_stack_if_appropriate), so on KVM
programs run under valgrind die at their first stack extension with "Bad
permissions for mapped region".

The sentry now also sets SEGV_MAPERR when the vma lookup fails with EFAULT.
This only runs for faults the sentry could not handle.

Only page faults carry an access type. The ptrace platform treated every
SIGSEGV as a page fault, and systrap derived an access type from SigError for
any SIGSEGV, so a general protection fault (SI_KERNEL, e.g. from hlt) reached
the fault path and would now have been reported as SEGV_MAPERR. Both platforms
now pass SIGSEGVs whose code is not SEGV_MAPERR or SEGV_ACCERR through as
plain signals, as KVM already does.

Tests: mmap_test MMapTest.UnmappedIsMapErr fails on runsc_kvm before this
change for writes and instruction fetches, and passes on all platforms.
WriteReadOnlyIsAccErr covers the SEGV_ACCERR case. exceptions_test
HaltIsSiKernel and NonCanonicalAccessIsSiKernel check that general protection
faults keep SI_KERNEL; HaltIsSiKernel fails on runsc_ptrace without the
platform changes.
FUTURE_COPYBARA_INTEGRATE_REVIEW=#15446 from tamird:bitmap-inclusive-ranges 01d6bdf
PiperOrigin-RevId: 994856758
copybara-service Bot pushed a commit that referenced this pull request Oct 7, 2026
Linux picks the SIGSEGV si_code from the vma lookup: no vma covering the
address is SEGV_MAPERR, a vma that forbids the access is SEGV_ACCERR
(arch/x86/mm/fault.c:do_user_addr_fault(): bad_area_nosemaphore() vs
bad_area_access_error()). After a failed HandleUserFault the sentry corrected
only the ACCERR case (EPERM) and otherwise kept the platform's code. The KVM
platform derives that code from the access type alone (vCPU.fault reports any
write or instruction fetch as SEGV_ACCERR), so on KVM a write to or a jump into
an unmapped page raised SEGV_ACCERR. systrap and ptrace forward the host's
siginfo and are unaffected.

Valgrind extends a client stack only on SEGV_MAPERR inside its stack
reservation (coregrind/m_signals.c:extend_stack_if_appropriate), so on KVM
programs run under valgrind die at their first stack extension with "Bad
permissions for mapped region".

The sentry now also sets SEGV_MAPERR when the vma lookup fails with EFAULT.
This only runs for faults the sentry could not handle.

Only page faults carry an access type. The ptrace platform treated every
SIGSEGV as a page fault, and systrap derived an access type from SigError for
any SIGSEGV, so a general protection fault (SI_KERNEL, e.g. from hlt) reached
the fault path and would now have been reported as SEGV_MAPERR. Both platforms
now pass SIGSEGVs whose code is not SEGV_MAPERR or SEGV_ACCERR through as
plain signals, as KVM already does.

Tests: mmap_test MMapTest.UnmappedIsMapErr fails on runsc_kvm before this
change for writes and instruction fetches, and passes on all platforms.
WriteReadOnlyIsAccErr covers the SEGV_ACCERR case. exceptions_test
HaltIsSiKernel and NonCanonicalAccessIsSiKernel check that general protection
faults keep SI_KERNEL; HaltIsSiKernel fails on runsc_ptrace without the
platform changes.
FUTURE_COPYBARA_INTEGRATE_REVIEW=#15446 from tamird:bitmap-inclusive-ranges 01d6bdf
PiperOrigin-RevId: 994856758
@copybara-service
copybara-service Bot merged commit f59c2c5 into google:master Oct 7, 2026
14 of 15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants