Repository navigation
Add required info for Anthropic's OSS Scanner program - #14628
tschneidereit wants to merge 1 commit into
Conversation
This adds a `Dockerfile` and a threat model description to be consumed by the [OSS Scanner program](https://red.anthropic.com/oss-scanner/), following the [program's template](https://github.com/anthropics/oss-scanner/tree/main/templates). The `Dockerfile` builds quite a bunch of stuff, which is by design: the image build process happens on much larger machines than where they're then run (16 instead of 2 cores, and something more than the 8GB it gets at runtime.) The threat model largely describes our stability tiers and gives some information on what to focus on. The one place where I decided to deviate a bit is to say that issues with the aarch64 backend should be considered vulnerabilities and reported as such, even though that backend is tier 2. The reason is that a lack of continuous fuzzing is what's holding that backend back from tier 1, and I think with this initiative (plus other auditing we're doing right now) we should be able to reconsider this soon.
|
Oh, and I tested the |
alexcrichton
left a comment
There was a problem hiding this comment.
I'm fine with these being follow-ups, but I'm hesitant to duplicate so much of our documentation in these docs because it seems like it'll inevitable get pretty far out of sync and we'll forget to update it
| # 2. `clif-util` at `target/debug/clif-util`, for Cranelift's `.clif` tests. | ||
| RUN cargo build --locked -p cranelift-tools | ||
| # 3. The full test suite as CI runs it, with all features enabled. | ||
| RUN ./ci/run-tests.py --locked --no-run |
There was a problem hiding this comment.
I'd probably say the tests here and cargo test above can be skipped
There was a problem hiding this comment.
See my other comments on these things: this is meant to prime the builds for running on the smaller machines after the image has been built.
There was a problem hiding this comment.
(and note that it doesn't run the tests, just build them)
| # `target/x86_64-unknown-linux-gnu/release/`. A failure here should not make | ||
| # the rest of the image unusable, so only warn about it. | ||
| RUN cargo "+$FUZZ_NIGHTLY" fuzz build --debug-assertions \ | ||
| || echo "WARNING: building the fuzz targets failed" |
There was a problem hiding this comment.
I suspect nothing is going to read these warning messages, but also the fuzzers should always build, so perhaps just skip this?
There was a problem hiding this comment.
The template says "Run the cheap tests so you know the image works, but do not let a failing test abort the build", which is why it's done this way.
Also, apparently the primary contact will be emailed, so we'll have to see whether having that be a mailing list of all core maintainers will be the right call or not
| # Smoke-test the build: run the spec test suite on the CLI's test harness. A | ||
| # failure here should be visible in the build log, but not abort the build. | ||
| RUN cargo test --locked --test wast -- --quiet \ | ||
| || echo "WARNING: wast tests reported failures" | ||
|
|
||
| # No network from here on. | ||
| ENV CARGO_NET_OFFLINE=true |
There was a problem hiding this comment.
These seem like they can be dropped
| # Install the same toolchains as CI (see `.github/actions/install-rust`): stable | ||
| # is MSRV + 2, and the nightly is the one pinned to match OSS-Fuzz, used for the | ||
| # fuzz targets. | ||
| ARG FUZZ_NIGHTLY=nightly-2026-07-09 | ||
| RUN msrv=$(grep 'rust-version.*1' Cargo.toml | sed 's/.*\.\([0-9]*\)\..*/\1/') \ | ||
| && curl -sSf https://sh.rustup.rs | sh -s -- -y --no-modify-path \ | ||
| --profile minimal --default-toolchain "1.$((msrv + 2)).0" \ | ||
| && rustup target add wasm32-wasip1 wasm32-wasip2 wasm32-unknown-unknown \ | ||
| && rustup toolchain install "$FUZZ_NIGHTLY" --profile minimal \ | ||
| && cargo install --locked cargo-fuzz --version 0.13.2 \ | ||
| && cargo install --locked wasm-tools |
There was a problem hiding this comment.
Could this perhaps try to read
and use that version of rustc? That can be the default toolchain for this whole build, I don't think there's any need to switch between stable/nightly| * `cargo test --test wast [filter]` runs the spec tests plus | ||
| `tests/misc_testsuite/**/*.wast` under each compiler (Cranelift, Winch, | ||
| Pulley), with and without the pooling allocator, and with each GC collector. | ||
| `;;! name = true` lines at the top of a `.wast` file enable features (see | ||
| `crates/test-util/src/wast.rs`). |
There was a problem hiding this comment.
This is probably better modeled as ./target/debug/wasmtime wast *.wast
| * `cargo test --test all [filter]` runs the integration tests in `tests/all/`, | ||
| which use the public `wasmtime` API. |
There was a problem hiding this comment.
Is this useful to include because all tests are always passing?
There was a problem hiding this comment.
it's useful to include because it primes the build cache, so modifications to individual files go more quickly. The machines building the image are much larger than the ones running it, so this kind of front-loading is apparently advisable.
| against the spec interpreter and V8), `instantiate`, `component_api`, | ||
| `gc_ops`, and `call_async` are good starting points. | ||
| * `wasm-tools` is installed, for `wasm-tools shrink` (test-case reduction), | ||
| `wasm-tools smith`, `print`, and `validate`. |
There was a problem hiding this comment.
We've git a skill in-repo for reduction, could this point there?
| ## How you rate severity | ||
|
|
||
| The impact on an embedder running an untrusted guest is what counts. | ||
| Reproduce in release mode where possible: a bug that needs debug assertions to | ||
| show up has no release-mode impact unless you show one. | ||
|
|
||
| * **Critical**: a guest-controlled sandbox escape with the default | ||
| configuration or a commonly used one (pooling allocator, async with | ||
| fuel/epochs, Winch): arbitrary read/write of host memory, control of host | ||
| execution, or a use-after-free that a guest can trigger and steer. | ||
| Typical sources: miscompiled or elided bounds checks, broken CFI, type | ||
| confusion in GC or `funcref`s, freed host memory reachable from Wasm. | ||
| * **High**: a sandbox escape that needs a less common but supported | ||
| configuration; a guest reading host memory or another instance's data | ||
| (including stale data in a reused pooling-allocator slot); WASI capability | ||
| bypasses, such as reading or writing files outside the preopened | ||
| directories; host memory unsafety reachable from safe embedder code | ||
| driven by guest-controlled values. | ||
| * **Medium**: denial of service against the host that a guest triggers at | ||
| run time: a host panic or abort, an infinite loop that fuel or epoch | ||
| interruption does not stop when configured, or memory or other resource | ||
| exhaustion that escapes configured limits (`ResourceLimiter`, `StoreLimits`, | ||
| pooling limits) or grows without bound in the host (e.g. in WASI | ||
| implementations). Memory unsafety that needs unusual but safe embedder API | ||
| use, with no guest control, also belongs here. | ||
| * **Low**: issues that need unlikely configurations or embedder behavior and | ||
| have limited impact; defense-in-depth weaknesses (e.g. a guard region that | ||
| is smaller than documented) with no demonstrated exploit; serious bugs in | ||
| Tier 2 features (marked as such). | ||
|
|
||
| **Not vulnerabilities** but worth reporting as normal bugs: | ||
|
|
||
| * Anything during *compilation* that is not memory unsafety: panics, slow | ||
| compilation, excessive memory use while compiling, infinite loops in the | ||
| register allocator, and so on. | ||
| * Wasm executing with the wrong semantics but staying inside the sandbox: | ||
| wrong results, spurious or missing traps. | ||
|
|
||
| **Not vulnerabilities** and not worth reporting at all: | ||
|
|
||
| * Behavior that Wasm allows to differ between engines: NaN bit patterns, | ||
| relaxed-SIMD results, how deep recursion can go before a stack overflow, | ||
| whether `memory.grow`/`table.grow` succeed. WASIp1 error codes that differ | ||
| from other engines. | ||
| * Memory or CPU use by a guest that stays within the limits the embedder | ||
| configured, or that is unlimited because the embedder set no limits. | ||
| * Bugs that need untrusted precompiled artifacts or cache contents, untrusted | ||
| CLI flags or `Config`, misuse of `unsafe` APIs, or bugs in host functions | ||
| written by the embedder. | ||
| * Spectre-style side channels beyond the mitigations described in | ||
| `docs/security.md`. | ||
|
|
There was a problem hiding this comment.
This is mostly just a duplication of our other docs, right?
| ## How reports and patches should look | ||
|
|
||
| * A minimal reproducer, in order of preference: a `.wast` file (with the | ||
| `wasmtime wast` flags needed listed in a comment at the top), a `.wat` or | ||
| `.wasm` for `wasmtime run` plus the exact command line, or a Rust `#[test]` | ||
| that uses only the public `wasmtime` API. For Cranelift-only issues, a | ||
| `.clif` file runnable with `clif-util` also helps, but should come with a Wasm | ||
| reproducer that shows the bug is reachable from Wasm. | ||
| * The exact configuration: CLI flags or `Config` calls, Cargo features, and | ||
| whether it reproduces in release mode. | ||
| * The commit you tested, and the expected and actual behavior (the trap, | ||
| panic message, or sanitizer and `gdb` output). | ||
| * One report per root cause. Fuzzers find the same bug in many shapes, so | ||
| deduplicate first. | ||
| * Patches should be minimal, fix the root cause rather than the symptom, and | ||
| add a regression test (usually a `.wast` file in `tests/misc_testsuite/` or a | ||
| test in `tests/all/`). |
There was a problem hiding this comment.
This is mostly a duplication of the wasmtime-auditor skill I think?
This adds a
Dockerfileand a threat model description to be consumed by the OSS Scanner program, following the program's template.The
Dockerfilebuilds quite a bunch of stuff, which is by design: the image build process happens on much larger machines than where they're then run (16 instead of 2 cores, and something more than the 8GB it gets at runtime.)The threat model largely describes our stability tiers and gives some information on what to focus on. The one place where I decided to deviate a bit is to say that issues with the aarch64 backend should be considered vulnerabilities and reported as such, even though that backend is tier 2. The reason is that a lack of continuous fuzzing is what's holding that backend back from tier 1, and I think with this initiative (plus other auditing we're doing right now) we should be able to reconsider this soon.