Skip to content

Build without std, behind a default-on std feature - #8

Merged
pathscale merged 5 commits into
masterfrom
feat/no-std
Sep 7, 2026
Merged

pathscale merged 5 commits into
masterfrom
feat/no-std

Conversation

@pathscale

Copy link
Copy Markdown
Owner

Arctic builds without std. The default feature set gains std, so nothing
changes for any existing user: same code paths, same 41 tests, none modified.

What it actually needed std for

Less than the 15,820 lines suggest. Four things, and only one of them was in the
algorithm:

was is
the SIMD level std::sync::LazyLock spin::Once
a Sync marker PhantomData<std::sync::Mutex<V>> a private MutexLike<V> with the two unsafe impls spelled out
the error impl std::error::Error core::error::Error, stable since 1.81
everything else std::fmt, std::cmp, Box, Vec core:: and alloc::

The marker is the one worth looking at in review. core has no type whose auto
traits are "Send when V: Send, Sync when V: Send", which is what
Mutex<V> gave. Cell and UnsafeCell are the near misses and both are never
Sync, so it is written out with the two impls and a SAFETY note rather than
approximated.

What you lose without std, stated rather than buried

Runtime SIMD detection. is_aarch64_feature_detected! and its x86 twin are
std macros, so fearless_simd::Level::new is gated on std. Without it the
level comes from cfg!(target_feature) at compile time.

On aarch64 that costs nothing: Neon is mandatory in the architecture, so
detection had exactly one possible answer, and it was reaching it through a
Once, an atomic load, and a match at each of the ten dispatch! sites. On
x86 it means the scalar path unless the build asks, with -C target-feature=+avx2
or a target-cpu that implies it. If runtime detection is ever wanted without
std, cpufeatures does it with core::arch::__cpuid and no std.

I checked the compiled artifact rather than trusting the source: 1,049 Neon
instructions
, cmeq for the key compares, ushll/ushll2 widening for the
u16x16 sort. The SIMD path is real and present.

Reclamation. smr-ps-reclaim and stat-garbage now imply std, because
ps-reclaim is std only and the garbage counter batches per thread, which is
this crate's only thread_local!. NoOp is what remains available.

Dependencies

fearless_simd needs libm without std, for float maths a radix tree never
calls. That is the sanctioned path rather than a workaround: upstream has an
open draft adding no-std libm tests to their CI (linebender/fearless_simd#229).

Every crate in the no_std graph already declares no_std, so nothing needs
forking
:

arctic-wt
├── fearless_simd → libm
└── ribbit        → arbitrary-int, portable-atomic
                  → ribbit-derive (proc-macro, host only)

Verified

cargo check and cargo check --no-default-features both clean, zero warnings
on either. cargo fmt --check clean. 41 tests pass.

Not verified: a genuinely bare target. This repo pins toolchain 1.96.0 in
rust-toolchain.toml and only aarch64-apple-darwin is installed for it, so
cargo check --target thumbv7em-none-eabi fails on a missing core rather than
on anything in this change. One rustup target add closes that, and it is worth
doing before this is relied on.

While in here, one thing that is not in this PR

Scalable vectors are a dead end for this crate, and I would rather write down
why than have it re-derived. This M4 Max reports sme_max_svl_b: 64, a 512-bit
streaming vector length, and node_47's 48-key search is four Neon compares
that would be one. Three separate walls stop it: SVE intrinsics are nightly only
and blocked on an unaccepted RFC for unsized types (rust-lang/rust#145052);
Apple has no non-streaming SVE, and -C target-feature=+sve produces a binary
that dies with SIGILL on this machine, which I measured; and fearless_simd has
refused scalable vectors by design (linebender/fearless_simd#339), because its
whole API rests on a compile-time lane count.

meh added 2 commits September 7, 2026 01:04
The default feature set gains `std`, so nothing changes for any existing user:
the same code paths, the same 41 tests passing, none modified.

With it off the crate does not link `std`. That is not the same as not having an
operating system, and Arctic barely wants one: what it needed `std` for was a
lazy static, a marker type, an error impl, and a prelude.

What changed:

`std::sync::LazyLock` held the detected SIMD level. It is a `spin::Once` now.
That was the only thing in the core that required `std`.

Runtime SIMD detection is a `std` facility: `is_aarch64_feature_detected!` and
its x86 twin are `std` macros, and `fearless_simd::Level::new` is gated on the
`std` feature accordingly. Without it the level comes from `cfg!(target_feature)`
at compile time. On aarch64 that costs nothing measurable, because Neon is
mandatory in the architecture, so detection had one possible answer. On x86 it
means the scalar path unless the build asks for the instructions, with
`-C target-feature=+avx2` or a `target-cpu` that implies them.

`fearless_simd` needs `libm` without `std`, for float maths that a radix tree
never calls. That is the sanctioned path rather than a workaround: upstream has
an open draft adding no-std libm tests to their CI.

`PhantomData<std::sync::Mutex<V>>` was the marker asserting `ConcurrentMap` is
`Sync` when `V` is `Send`. `core` has no type with those auto traits, and the
near misses are not near: `Cell` and `UnsafeCell` are never `Sync`. It is a
private `MutexLike<V>` with the two `unsafe impl`s spelled out.

`impl std::error::Error` became `core::error::Error`, stable since 1.81, so it
needs no feature gate.

`smr-ps-reclaim` and `stat-garbage` now imply `std`. ps-reclaim is `std` only,
and the garbage counter batches per thread, which is this crate's only
`thread_local!`. `NoOp` is what remains available without `std`.

The rest is mechanical: `std::` to `core::` and `alloc::`, and the `alloc`
imports where `Box`, `Vec` and `ToOwned` lost their prelude.

The four crates in the no_std dependency graph, `fearless_simd`, `libm`,
`ribbit` and its `arbitrary-int` and `portable-atomic`, all already declare
`no_std`. Nothing needs forking.
The idea occurs to everyone who reads node_47: 48 sorted keys is four 16-byte
Neon compares, this chip reports a 512-bit streaming vector length, and four
instructions become one. It is the most obvious win in the crate and it is a
34x loss.

Two million searches of 48 keys on an M4 Max, all three implementations checked
against each other on every one of the 256 possible target bytes first:

    scalar                       3.64 ns
    neon, 4x128-bit              0.93 ns
    sve 512-bit, smstart once   31.62 ns
    sve 512-bit, per call       43.44 ns

The third row is the one that matters: streaming mode is entered once for all
two million iterations, so SMSTART is amortised to nothing and it is still
thirty-four times slower. The instructions themselves are slow here, because
Apple's SME unit is built to feed a ZA tile with matrix outer products and
streaming SVE rides along because the architecture requires it, not because
there is a fast general-purpose vector engine underneath.

docs/M4-SVE-perf-review.md carries the harness, the three toolchain and library
walls behind the number, and the note that none of this transfers to hardware
with real non-streaming SVE. A comment above SIMD_LEVEL points at it, because
that is where the question gets asked.
meh added 2 commits September 7, 2026 01:05
A patch bump rather than a minor one. The change is additive: nothing is
removed, no signature moves, and the feature combination WorkTable pins,
`default-features = false` with `smr-ps-reclaim`, still resolves and still
compiles, which I checked rather than assumed.

That keeps the three consumers in this house on `^0.1` picking it up rather than
stranding them behind a manifest edit. What they gain automatically is `spin`,
which has no dependencies of its own.
AGENTS.md and CLAUDE.md were on this branch's base and are not on master: the
commit carrying them was dropped when master was rewritten. They are not part of
the no_std change and are here only because they are otherwise lost.
Two defects, both of which CI caught and I did not, because every check I ran
was on aarch64 with default features.

Level::fallback does not exist on an x86-64 build. fearless_simd compiles the
Fallback variant and its constructor only when the target does not statically
support something better, and on x86-64 SSE4.2 is the floor, so the no_std
detect arm referenced a function that was cfg'd out. force_support_fallback
makes the variant exist on every target. This is the right knob rather than a
cfg matrix mirroring upstream's, because without std there is no runtime
detection to fall back from and detect names the variant directly.

smr-hazard does not link std but uses it. smr/hazard/thread.rs keys ID_FAST and
ID_SLOW off thread_local! and holds its free list in a std::sync::Mutex. Both
have no_std answers and neither is written yet, so the feature carries the
requirement, the way smr-ps-reclaim already does.

Reproduced and fixed against x86_64-unknown-linux-gnu, which is what the runners
use:

  smr-seize   x86_64: OK
  smr-epoch   x86_64: OK
  smr-hazard  x86_64: OK

plus default and no-default on both architectures, fmt and clippy.
@pathscale
pathscale merged commit d0f916b into master Sep 7, 2026
7 checks passed
@pathscale
pathscale deleted the feat/no-std branch September 7, 2026 09:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant