📖 Documentation: Concepts · Architecture · Protocol · MDS design · Mirroring
PREFIX=${HOME}/local
./autogen.sh
./configure --prefix=${PREFIX}
make -j$(nproc)
make install
OST_ADDR=192.168.0.1:7777
MDS_ADDR=192.168.0.1:7776
##
# OST Server
#
OST_DATADIR=/var/rawstor
mkdir -p ${OST_DATADIR}
rawstor-ost \
--bind ${OST_ADDR} \
file://${OST_DATADIR}
##
# MDS Server (optional: only for chunked mds:// objects)
#
MDS_DATADIR=/var/lib/rawstor/mds
mkdir -p ${MDS_DATADIR}
# One line per OST: <uuid> <location> <weight> [[[<dc>/]<row>/]<rack>/]<server>
cat > ${MDS_DATADIR}/topology.conf <<EOF
$(cat /proc/sys/kernel/random/uuid) ost://${OST_ADDR} 100 dc1/row1/rack1/host1
EOF
rawstor-mds \
--bind ${MDS_ADDR} \
--db ${MDS_DATADIR}/mds.db \
--topology ${MDS_DATADIR}/topology.conf
##
# Client
#
# An object split into 256M chunks, placed across the OSTs by the MDS...
OBJECT_TARGET=$(rawstor create mds://${MDS_ADDR} --size=1G --chunk-size=256M --mirrors=1)
# ...or a plain object stored whole on one OST, no MDS needed:
# OBJECT_TARGET=$(rawstor create ost://${OST_ADDR} --size=1G --mirrors=1)
VHOST_RUNDIR=${PREFIX}/var/run/rawstor
mkdir -p ${VHOST_RUNDIR}
rawstor-vhost \
--socket-path=${VHOST_RUNDIR}/rawstor1.sock \
${OBJECT_TARGET}
qemu-system-x86_64 \
-enable-kvm \
-m 4G \
-machine accel=kvm,memory-backend=mem \
-drive file=image.qcow2,if=none,id=drive1 \
-device virtio-blk-pci,drive=drive1 \
-object memory-backend-memfd,id=mem,size=4G,share=on \
-chardev socket,id=rawstor1,reconnect=1,path=${VHOST_RUNDIR}/rawstor1.sock \
-device vhost-user-blk-pci,chardev=rawstor1,num-queues=1,disable-legacy=on
The following environment variables can be used to tune the behavior of the Rawstor client and server. Default values are shown below.
| Variable | Default | Description |
|---|---|---|
RAWSTOR_LOCATION |
(none) | Fallback LOCATION for rawstor create/list/info when it's omitted from the command line. |
RAWSTOR_MDS_OPTS_INFO_INTERVAL |
300000 |
MDS backend INFO/health polling interval in milliseconds; each OST is polled once per interval at its own fixed offset, spreading probes evenly. Must be positive. |
RAWSTOR_MDS_OPTS_INFO_CONCURRENCY |
128 |
Maximum concurrent MDS backend probes, including retries (1–1024); see MDS health and location information. |
RAWSTOR_OPTS_IO_ATTEMPTS |
10 |
Number of attempts for an I/O operation before giving up, covering any failure from a broken connection to a well-formed rejection from a live backend (e.g. EBUSY/ENOSPC) -- every retry reconnects first (except a plain EBUSY, where the session is fine and just backed up against the remote server's own write-throttling). A rejection known to never succeed on retry (e.g. ENOENT) isn't retried at all, regardless of this value. |
RAWSTOR_OPTS_IO_RETRY_BACKOFF_BASE |
100 |
Base delay, in milliseconds (ms), before the first retry of an I/O operation; doubles with each further attempt, capped at RAWSTOR_OPTS_IO_RETRY_BACKOFF_MAX. |
RAWSTOR_OPTS_IO_RETRY_BACKOFF_MAX |
30000 |
Upper bound, in milliseconds (ms), on the exponential retry backoff delay above. |
RAWSTOR_OPTS_IO_RETRY_BACKOFF_JITTER |
50 |
Percentage (0-100, not a time value) of the computed retry backoff delay that is randomized, to avoid many clients retrying in lockstep. 0 disables jitter (a plain exponential backoff); 100 is "Full Jitter" (the whole delay is randomized); 50 is "Equal Jitter" (half the delay is fixed, half is randomized). |
RAWSTOR_OPTS_SO_SNDTIMEO |
5000 |
Socket send timeout, in milliseconds (ms). Sets SO_SNDTIMEO for network sockets. |
RAWSTOR_OPTS_SO_RCVTIMEO |
5000 |
Socket receive timeout, in milliseconds (ms). Sets SO_RCVTIMEO for network sockets. |
RAWSTOR_OPTS_TCP_USER_TIMEOUT |
5000 |
TCP user timeout, in milliseconds (ms) (Linux TCP_USER_TIMEOUT). Defines how long transmitted data may remain unacknowledged before the connection is closed. |
RAWSTOR_OPTS_LIST_LIMIT |
1000 |
Server-side page size cap for list operations: the maximum number of objects returned in a single call, regardless of the caller-requested limit. Larger listings are paginated across multiple calls. |
RAWSTOR_OPTS_WRITE_THROTTLE_LIMIT |
128 |
Per-session cap on writes dispatched to a file:// backing store without their completion arriving yet; writes past the cap wait for a dispatch slot instead of being sent immediately. |
RAWSTOR_OPTS_WRITE_BACKLOG_CAPACITY |
256M |
Per-session cap on writes queued behind RAWSTOR_OPTS_WRITE_THROTTLE_LIMIT but not yet dispatched to a file:// backing store; a write that would push the backlog over the cap fails with EBUSY instead of queuing. Takes a size with a mandatory unit suffix (B, K, M, G, T, P, E), e.g. 256M or 4096B; a bare number is rejected. |
Note: All timeout values are expressed in milliseconds unless stated otherwise.
An unset variable takes its default. A set one must be valid -- a plain
decimal integer within the variable's range (or, for a size, a number with a
unit) -- otherwise the program logs the variable and the accepted range and
exits with EX_CONFIG (78), which the shipped systemd units do not restart.
0 is accepted only where it means something: it disables
RAWSTOR_OPTS_SO_SNDTIMEO, RAWSTOR_OPTS_SO_RCVTIMEO,
RAWSTOR_OPTS_TCP_USER_TIMEOUT and the retry backoff, and allows no backlog in
RAWSTOR_OPTS_WRITE_BACKLOG_CAPACITY.
rawstor-ost implements the OST protocol (see Protocol), handling network connections and providing access to data stored in locations (as defined in the Concepts documentation).
file://scheme → serves data directly from the local filesystem.ost://scheme → acts as a proxy to an underlying OST backend.- Comma‑separated list → supports mirroring or data locality.
rawstor-ost [-h] -b ADDR LOCATION
| Option | Description |
|---|---|
-h, --help |
Show help message and exit. |
LOCATION |
Comma‑separated list of backend locations (e.g., file:///path, ost://host:port). |
-b, --bind ADDR |
Bind address in <ip>:<port> format (e.g., 127.0.0.1:7777). |
-w, --workers N |
Number of worker threads, each with its own client connections and I/O queue, all accepting on the same listening socket (default: 12). |
Serve local directory:
rawstor-ost -b 0.0.0.0:7777 file:///var/rawstor/dataProxy to remote OST:
rawstor-ost -b 0.0.0.0:7777 ost://192.168.1.100:7777Data locality (local cache + proxy):
rawstor-ost -b 0.0.0.0:7777 file:///var/rawstor/data,ost://remote:7777Mirroring between two OST backends:
rawstor-ost -b 0.0.0.0:7777 ost://left:7777,ost://right:7777The rawstor-ost deb/rpm package ships the rawstor-ost@.service
template: one instance per OST, named after the OST's id (the <uuid> the
MDS topology lists it under), configured by
/etc/rawstor/ost/<uuid>.conf. Installing the package starts nothing; an
instance without its config refuses to start. A host runs as many
instances as it has OSTs, each on its own port and store.
The config is an environment file (KEY=VALUE lines):
| Key | Default | Description |
|---|---|---|
BIND_ADDR |
— (required) | <ip>:<port> to listen on, unique per host. |
LOCATION |
file:///var/lib/rawstor/ost/<uuid> |
Backing store this instance serves. |
QUEUE_SIZE |
4096 |
--queue-size. |
WORKERS |
12 |
--workers. |
RAWSTOR_OPTS_* |
Tuning knobs, see Environment Variables. |
The instance runs as the rawstor user under ProtectSystem=strict. Its
default store, /var/lib/rawstor/ost/<uuid> (StateDirectory=), is
created on start and owned by rawstor; a disk mounted there needs no
further configuration. A store anywhere else needs a per-instance drop-in
granting the sandbox access to it, and must be writable by rawstor:
sudo systemctl edit rawstor-ost@<uuid>.service
# [Service]
# ReadWritePaths=/srv/disk1/rawstorTwo OSTs on one host:
OST1=$(cat /proc/sys/kernel/random/uuid)
OST2=$(cat /proc/sys/kernel/random/uuid)
echo "BIND_ADDR=0.0.0.0:7777" | sudo tee /etc/rawstor/ost/${OST1}.conf
echo "BIND_ADDR=0.0.0.0:7778" | sudo tee /etc/rawstor/ost/${OST2}.conf
sudo systemctl enable --now rawstor-ost@${OST1} rawstor-ost@${OST2}
systemctl status rawstor-ost@${OST1} # state
journalctl -u rawstor-ost@${OST1} # log
rawstor info ost://127.0.0.1:7777 # serving?
sudo systemctl restart rawstor-ost@${OST1}Removing an instance never touches its store: stop and disable it, then delete its config. Recreating the same config later serves the same data again; the store itself is removed by hand, if at all.
sudo systemctl disable --now rawstor-ost@${OST2}
sudo rm /etc/rawstor/ost/${OST2}.conf
# data stays in /var/lib/rawstor/ost/${OST2}A package upgrade restarts the running instances; removing the package stops them, leaving configs, stores and enablement in place.
rawstor-vhost is a userspace VirtIO block device backend implementing the
vhost-user protocol
for virtio-blk-pci/vhost-user-blk-pci devices. It gives a guest direct,
zero-copy access to a rawstor object's data: virtqueue descriptors are
resolved straight into the guest's shared memory regions and read from or
written to the backing object via librawstor's native io_uring-based
object I/O, with no host-side kernel block layer or copy in between.
The vhost-user protocol itself (feature/memory-region negotiation,
virtqueue kick/call handling, request dispatch) is implemented natively in
vhost/ — it does not depend on qemu's libvhost-user library. Every I/O
path, including the control socket, is asynchronous and non-blocking, and
multiple in-flight requests on a virtqueue may complete out of order.
rawstor-vhost [-h] -s SOCKET_PATH TARGET [--queue-size SIZE] [--write-cache on|off] [--readonly] [-v]
| Option | Description |
|---|---|
-h, --help |
Show help message and exit. |
-s, --socket-path PATH |
Location of the vhost-user Unix domain socket. |
TARGET |
Comma‑separated list of rawstor backend targets (see Concepts). |
--queue-size SIZE |
RawIO queue (io_uring) depth of each virtqueue's own queue. Default: 4096. |
--write-cache on|off |
Advertise a writeback (on) or write-through (off, default) cache to the guest; write-through makes every write durable on completion, writeback relies on the guest issuing an explicit flush. |
--readonly |
Export the object read-only: advertises VIRTIO_BLK_F_RO to the guest and opens TARGET with RAWSTOR_READONLY (no mirror write quorum needed; writes fail). Required to export a version target. |
-v, --version |
Print version and exit. |
PREFIX=${HOME}/local
OST_ADDR=192.168.0.1:7777
OBJECT_ID=...
VHOST_RUNDIR=${PREFIX}/var/run/rawstor
rawstor-vhost \
--socket-path=${VHOST_RUNDIR}/rawstor1.sock \
ost://${OST_ADDR}/${OBJECT_ID}
qemu-system-x86_64 \
-enable-kvm \
-m 4G \
-machine accel=kvm,memory-backend=mem \
-drive file=image.qcow2,if=none,id=drive1 \
-device virtio-blk-pci,drive=drive1 \
-object memory-backend-memfd,id=mem,size=4G,share=on \
-chardev socket,id=rawstor1,reconnect=1,path=${VHOST_RUNDIR}/rawstor1.sock \
-device vhost-user-blk-pci,chardev=rawstor1,num-queues=1,disable-legacy=on
rawstor-vhost negotiates VIRTIO_BLK_F_SIZE_MAX, VIRTIO_BLK_F_SEG_MAX,
VIRTIO_BLK_F_BLK_SIZE, VIRTIO_BLK_F_TOPOLOGY, VIRTIO_BLK_F_MQ,
VIRTIO_BLK_F_FLUSH, VIRTIO_BLK_F_CONFIG_WCE, VIRTIO_BLK_F_DISCARD,
VIRTIO_BLK_F_WRITE_ZEROES, VIRTIO_RING_F_INDIRECT_DESC and
VIRTIO_RING_F_EVENT_IDX, and services read (VIRTIO_BLK_T_IN), write
(VIRTIO_BLK_T_OUT), flush (VIRTIO_BLK_T_FLUSH), discard
(VIRTIO_BLK_T_DISCARD), write-zeroes (VIRTIO_BLK_T_WRITE_ZEROES) and
identify (VIRTIO_BLK_T_GET_ID) requests. A flush is device-wide: since
each virtqueue has its own connection to TARGET (see below), a
VIRTIO_BLK_T_FLUSH arriving on one makes durable every write issued
through any of them, not just its own. VIRTIO_BLK_F_MQ is negotiated
and genuinely serviced: the front-end alone decides how many virtqueues
to use (QEMU's own num-queues=, which defaults to the guest's vCPU
count; rawstor-vhost reports up to 1024 via VHOST_USER_GET_QUEUE_NUM),
and each one it sets up is processed on its own thread and its own
connection to TARGET in parallel, started the first time the front-end
configures that queue.
rawstor-vhost serves exactly one front-end connection per process
invocation: it accepts a connection on the socket, serves it until the
front-end disconnects (exiting cleanly), and then exits. A setup or
protocol error (e.g. failing to connect a new virtqueue to TARGET) is reported and also exits
the process, rather than silently waiting for another connection. Pair it
with reconnect=1 on the QEMU chardev and an external supervisor (e.g.
systemd with Restart=always, or a wrapper loop) if the backend needs to
survive guest-side reconnects. Send SIGINT/SIGTERM to stop it: it
first waits for every request already in flight to complete, so QEMU,
reconnecting to the restarted backend, carries on exactly where it left
off. It also negotiates VHOST_USER_PROTOCOL_F_INFLIGHT_SHMFD: every
request is logged in shared memory QEMU keeps across backend restarts, so
after a crash the restarted backend resubmits whatever was left in
flight, in its original order.
rawstor-vhost ships in its own rawstor-vhost deb/rpm package (separate
from librawstor), along with the rawstor-vhost@.service systemd
template unit. The package creates the same system user/group (rawstor)
that rawstor-ost uses, if it doesn't already exist — it does not
depend on libvirt, since rawstor-vhost has no need for it (it talks to
QEMU purely over a vhost-user Unix socket, negotiated when QEMU connects).
Whatever user actually runs QEMU (libvirt-qemu on Debian/Ubuntu, qemu
on Fedora/RHEL, or something else entirely if you invoke QEMU by hand)
needs permission to connect to the socket under RuntimeDirectory=rawstor
(/run/rawstor/*.sock). Add that user to the rawstor group rather than
running rawstor-vhost as it:
sudo usermod -aG rawstor libvirt-qemu # Debian/Ubuntu + libvirt
sudo usermod -aG rawstor qemu # Fedora/RHEL + libvirtThis works because rawstor-vhost chmod()s the socket to 0660 itself
right after creating it, regardless of the caller's umask — on Linux,
connect(2) to a UNIX stream socket requires write permission on the
socket file itself (not just directory access), so leaving it at whatever
bind(2) produced under the process's umask (commonly 0755, i.e.
group gets read+execute but no write) would silently prevent anyone but
the socket's owner from ever connecting. Since the daemon enforces this
itself, it holds regardless of how or by whom rawstor-vhost is invoked —
including if you override User=/Group= below via a drop-in.
If User=/Group=rawstor in the unit doesn't fit your setup (e.g. you'd
rather run rawstor-vhost as the same user QEMU runs as, instead of
sharing access via the group), override it with a drop-in instead of
editing the shipped unit file — systemctl edit rawstor-vhost@.service
(all instances) or systemctl edit rawstor-vhost@<uuid>.service (one
instance) opens an editor and saves the result under
/etc/systemd/system/…/override.conf, which survives package upgrades.
See the comment above User= in systemd/rawstor-vhost@.service and
systemd.unit(5) for details.
rawstor-vduse is a userspace VirtIO block device backend implementing the
VDUSE (vDPA Device in
Userspace) protocol for virtio-blk devices. Unlike rawstor-vhost
(vhost-user, a QEMU-only Unix-socket protocol), VDUSE creates a real kernel
vDPA device: once attached to the vDPA bus, it can be driven either by
virtio-vdpa (the guest's virtio-blk driver talks to the kernel vDPA
framework directly -- no VMM involved at all, e.g. for containers) or by
vhost-vdpa (a VMM such as QEMU drives it through /dev/vhost-vdpa-N).
The VDUSE control-plane protocol (device/virtqueue setup via ioctl(2) on
/dev/vduse/control and /dev/vduse/$NAME, IOTLB-backed memory mapping,
kick/interrupt handling) is implemented natively in vduse/ -- like
vhost/, it does not vendor or link against any third-party protocol
library (qemu's libvduse included); only the kernel uAPI struct/ioctl
definitions themselves are vendored (vduse/include/stdheaders/linux/,
trimmed the same way vhost/include/stdheaders/ vendors its own kernel
headers). Every I/O path, including the control channel, is asynchronous
and non-blocking via RawIO, and multiple in-flight requests on a virtqueue
may complete out of order -- the same design vhost/ uses for
vhost-user.
rawstor-vduse [-h] TARGET [--queue-size SIZE] [--num-queues N] [--write-cache on|off] [--readonly] [-v]
| Option | Description |
|---|---|
-h, --help |
Show help message and exit. |
TARGET |
Comma‑separated list of rawstor backend targets (see Concepts). Creates /dev/vduse/UUID, where UUID is the target object's own UUID -- there is no separate name to pick, since the UUID already uniquely and stably identifies it. |
--queue-size SIZE |
Virtqueue size, a power of two, of each virtqueue's own queue. Default: 256, max 1024. |
--num-queues N |
Number of virtqueues advertised to the guest. Each one the driver actually enables is serviced by its own thread and its own connection to TARGET, started at that point (Linux's virtio_blk enables up to its CPU count). Default: the host's number of CPUs. |
--write-cache on|off |
Advertise a writeback (on) or write-through (off, default) cache to the guest. |
--readonly |
Export the object read-only: advertises VIRTIO_BLK_F_RO to the guest and opens TARGET with RAWSTOR_READONLY (no mirror write quorum needed; writes fail). Required to export a version target. |
-v, --version |
Print version and exit. |
OST_ADDR=192.168.0.1:7777
OBJECT_ID=...
sudo modprobe vduse
sudo rawstor-vduse \
ost://${OST_ADDR}/${OBJECT_ID} &
# Attach the device to the vDPA bus once rawstor-vduse has created it
# (named after OBJECT_ID, i.e. /dev/vduse/${OBJECT_ID}):
sudo vdpa dev add name ${OBJECT_ID} mgmtdev vduse
# Either hand it to a guest's virtio-vdpa driver directly (no VMM), or
# drive it from QEMU over /dev/vhost-vdpa-N:
qemu-system-x86_64 \
-enable-kvm \
-m 4G \
-device vhost-vdpa-device-pci,vhostdev=/dev/vhost-vdpa-0
rawstor-vduse negotiates VIRTIO_BLK_F_SEG_MAX, VIRTIO_BLK_F_BLK_SIZE,
VIRTIO_BLK_F_TOPOLOGY, VIRTIO_BLK_F_FLUSH, VIRTIO_BLK_F_MQ,
VIRTIO_BLK_F_DISCARD, VIRTIO_BLK_F_WRITE_ZEROES, plus whatever baseline
virtio/ring features the kernel VDUSE driver itself requires
(VIRTIO_F_VERSION_1, VIRTIO_F_ACCESS_PLATFORM,
VIRTIO_F_NOTIFY_ON_EMPTY, VIRTIO_RING_F_EVENT_IDX,
VIRTIO_RING_F_INDIRECT_DESC), and services read (VIRTIO_BLK_T_IN), write
(VIRTIO_BLK_T_OUT), flush (VIRTIO_BLK_T_FLUSH), discard
(VIRTIO_BLK_T_DISCARD), write-zeroes (VIRTIO_BLK_T_WRITE_ZEROES) and
identify (VIRTIO_BLK_T_GET_ID) requests. A flush is device-wide: since
each virtqueue has its own connection to TARGET (see below), a
VIRTIO_BLK_T_FLUSH arriving on one makes durable every write issued
through any of them, not just its own. VIRTIO_BLK_F_MQ is negotiated
and genuinely serviced: rawstor-vduse advertises --num-queues
virtqueues, and each one the driver enables is processed on its own
thread and its own connection to TARGET in parallel.
Unlike vhost-user, VDUSE has no driver-writable config space at all --
there is no protocol message equivalent to VHOST_USER_SET_CONFIG -- and
the kernel VDUSE driver unconditionally rejects VIRTIO_BLK_F_CONFIG_WCE
for virtio-blk devices (device creation fails outright if it's
advertised), so rawstor-vduse doesn't negotiate it. --write-cache
still works: Linux's virtio_blk driver enables its write-cache
(flush-before-trusting-durability) assumption whenever
VIRTIO_BLK_F_FLUSH is negotiated regardless of CONFIG_WCE, so the
guest always issues FLUSH appropriately either way -- the flag purely
controls whether rawstor-vduse itself additionally makes every write
durable on completion or relies solely on that FLUSH.
rawstor-vduse creates one VDUSE device per process invocation and keeps
running until the process is stopped (SIGINT/SIGTERM) or the VDUSE
device itself is destroyed out from under it -- unlike rawstor-vhost,
there is no per-connection front-end to disconnect from, since the kernel
is always "connected". Attaching the created device to the vDPA bus (vdpa dev add name UUID mgmtdev vduse) and, if applicable, driving it from a
VMM, are separate, external steps.
On SIGINT/SIGTERM it first waits for every request already in flight
to complete. A device still bound to the vDPA bus outlives the process; a
restarted rawstor-vduse reattaches to it and, if the driver is already
using it, resumes every ready virtqueue where the previous instance left
off, without the driver noticing -- as long as it is back within the
kernel's VDUSE message timeout (/sys/class/vduse/UUID/msg_timeout, 30 s
by default), after which the kernel marks the device broken. Requests in
flight during a crash are resubmitted too: every request is logged, in
the same format as vhost-user's in-flight memory, in
/dev/shm/rawstor-vduse-UUID.inflight, a tmpfs file that outlives the
process but not a reboot (which takes the VDUSE device with it anyway),
and is removed along with the device. A reattached device keeps the
virtqueue count it was created with, so --num-queues must not change
across restarts.
rawstor-vduse ships in its own rawstor-vduse deb/rpm package (separate
from librawstor), along with the rawstor-vduse@.service systemd
template unit and a udev rule granting the rawstor system user/group
(shared with rawstor-ost/rawstor-vhost) access to /dev/vduse/control
and the device nodes it creates under /dev/vduse/.
Creating a VDUSE device requires the vduse kernel module (modprobe vduse), and attaching it to the vDPA bus requires the vdpa tool
(iproute2) and CAP_NET_ADMIN -- neither is something rawstor-vduse
itself does; both are external, one-time-per-device administrative steps.
rawstor-mds is the metadata server behind mds://<host>:<port>/<uuid>
targets: a logical object split into fixed-size chunks, each independently
placed (and optionally mirrored) across a static OST topology, opened/
read/written through the same rawstor_target_*()/rawstor CLI surface
as a plain object (see MDS design). A single instance owns the explicit
chunk map (SQLite, WAL journal) and the static topology config for its own
cluster; it is not in the data path -- rawstor_target_open() talks to
the placed OSTs directly once it has resolved an object's own chunk map.
rawstor-mds [options] -b ADDR -d DBPATH -t TOPOLOGY
| Option | Description |
|---|---|
-h, --help |
Show help message and exit. |
-b, --bind ADDR |
Bind address in <ip>:<port> format (e.g., 127.0.0.1:7776). |
-d, --db PATH |
SQLite database file holding the chunk map (created if missing). |
-t, --topology PATH |
Static topology config file: one <uuid> <location> <weight> [[[<dc>/]<row>/]<rack>/]<server> line per OST (failure domains from the root down; only <server> is required), <location> being a single location URI (ost://host:port; a client-local one such as file:// only makes sense on a single host; to put several stores under one entry, list a rawstor-ost serving them all) (see MDS design). |
--queue-size SIZE |
RawIO queue (io_uring) depth. Default: 4096. |
-w, --workers N |
Number of worker threads, each with its own client connections and I/O queue, all accepting on the same listening socket and sharing one database (default: 4). |
-r, --reconstruct |
Rebuild the chunk map from a LIST+META scan of every OST in the topology before serving -- for recovering from a lost or corrupted database. |
Serve a topology of two OSTs:
cat > topology.conf <<EOF
018f4e2a-1000-7000-8000-000000000001 ost://host1:7777 100 dc1/row1/rack1/host1
018f4e2a-1000-7000-8000-000000000002 ost://host2:7777 100 dc1/row1/rack1/host2
EOF
rawstor-mds -b 0.0.0.0:7776 -d /var/lib/rawstor/mds/mds.db -t topology.confCreate and grow an mds:// object:
rawstor create -t mds://127.0.0.1:7776/018f4e2a-2000-7000-8000-000000000001 --size=1G --chunk-size=256M --mirrors=1
rawstor resize mds://127.0.0.1:7776/018f4e2a-2000-7000-8000-000000000001 --size=2GAdd an OST: append its line to topology.conf and send SIGHUP
(systemctl reload rawstor-mds@<uuid>). New chunks may then be placed on
it; existing ones stay where they are. Client connections are kept. A
topology that doesn't parse, or drops an OST still holding chunks, is
refused: on reload the current one is kept and the reason logged
(Topology reload from ... failed, keeping the current one: ...), at
startup the MDS doesn't start.
rawstor-mds ships in its own rawstor-mds deb/rpm package (needs
sqlite3; skip building it with --without-sqlite3), along with the
rawstor-mds@.service systemd template and a commented topology example,
/usr/share/doc/rawstor-mds/topology.conf. One instance serves one
cluster, named after the cluster's id and configured by
/etc/rawstor/mds/<uuid>.conf. Installing the package starts nothing; an
instance without its config refuses to start.
The config is an environment file (KEY=VALUE lines):
| Key | Default | Description |
|---|---|---|
BIND_ADDR |
— (required) | <ip>:<port> to listen on, unique per host. |
TOPOLOGY_PATH |
/etc/rawstor/mds/<uuid>.topology |
Topology file; must exist (may be empty) and be readable by rawstor. |
DB_PATH |
/var/lib/rawstor/mds/<uuid>/mds.db |
Database; its directory must be writable by rawstor. |
QUEUE_SIZE |
4096 |
--queue-size. |
WORKERS |
4 |
--workers. |
RAWSTOR_MDS_OPTS_*, RAWSTOR_OPTS_* |
Tuning knobs, see Environment Variables. |
The instance runs as the rawstor user under ProtectSystem=strict;
/var/lib/rawstor/mds/<uuid> (StateDirectory=) is created on start.
A database anywhere else needs a drop-in (ReadWritePaths=), as for
rawstor-ost.
Two clusters on one host:
for C in ${CLUSTER1} ${CLUSTER2}; do
sudo cp my-${C}.topology /etc/rawstor/mds/${C}.topology
done
echo "BIND_ADDR=0.0.0.0:7776" | sudo tee /etc/rawstor/mds/${CLUSTER1}.conf
echo "BIND_ADDR=0.0.0.0:7786" | sudo tee /etc/rawstor/mds/${CLUSTER2}.conf
sudo systemctl enable --now rawstor-mds@${CLUSTER1} rawstor-mds@${CLUSTER2}
rawstor info mds://127.0.0.1:7776 # serving?
# edit /etc/rawstor/mds/${CLUSTER1}.topology, then
sudo systemctl reload rawstor-mds@${CLUSTER1}
journalctl -u rawstor-mds@${CLUSTER1} | grep Topology # applied?Removing an instance (systemctl disable --now, then deleting its config)
keeps its database; a package upgrade restarts the running instances and
removing the package stops them, leaving configs and databases in place.
make test
See CONTRIBUTING.md for the development environment, code style, commit conventions and pull request guidelines.
io_uring_queue_init() failed: Operation not permitted
First check if io_uring is disabled or not in sysctl:
sysctl -a | grep io_uringAccording to the documentation for the sysctl files in /proc/sys/kernel/:
io_uring_disabled:Prevents all processes from creating new
io_uringinstances. Enabling this shrinks the kernel’s attack surface.
0- All processes can createio_uringinstances as normal. This is the default setting.
1-io_uringcreation is disabled (io_uring_setup()will fail with-EPERM) for unprivileged processes not in theio_uring_groupgroup. Existingio_uringinstances can still be used. See the documentation forio_uring_groupfor more information.
2-io_uringcreation is disabled for all processes.io_uring_setup()always fails with-EPERM. Existingio_uringinstances can still be used.
So you need to set it to 0:
sysctl kernel.io_uring_disabled=0io_uring_register_buf_ring failed: Invalid argument
Seen from librawio's multishot recv path (uring_buffer.cpp), and from
anything that depends on it: rawstor-ost itself reads wire responses via
multishot recv with a registered provided buffer ring, so every ost://-backed
test (OstIOTest, OstLifecycleTest, MirrorOstTest, ...) and
librawio's own MultishotTest fail with this same symptom when it's present.
Confirm it's this, not something this project is doing wrong, with a minimal
reproduction independent of librawstor entirely (gcc repro.c -luring):
#include <liburing.h>
#include <sys/mman.h>
#include <string.h>
int main(void) {
struct io_uring ring;
io_uring_queue_init(16, &ring, 0);
size_t size = 4096;
void *buf = mmap(NULL, size, PROT_READ | PROT_WRITE,
MAP_ANONYMOUS | MAP_PRIVATE, -1, 0);
struct io_uring_buf_reg reg = {0};
reg.ring_addr = (unsigned long long)(uintptr_t)buf;
reg.ring_entries = 1; /* any power of 2 */
reg.bgid = 1;
return io_uring_register_buf_ring(&ring, ®, 0); /* -EINVAL here */
}If this fails identically regardless of ring_entries/size (including the
smallest possible registration), under sudo, with ulimit -l far above
what's being requested, and with Seccomp: 0 in /proc/self/status -- it
isn't a resource limit, a privilege issue, or a code bug: RLIMIT_MEMLOCK,
kernel.io_uring_disabled (see above), and process privilege can all be ruled
out this way.
This combination -- Ubuntu's own linux-image-*-generic kernel,
IORING_REGISTER_PBUF_RING (provided buffer ring registration), EINVAL --
has been reported before: see
axboe/liburing#1432,
where the io_uring maintainer notes Ubuntu's kernel "is not part of the stable
series and does not receive updates." Canonical backports its own set of
io_uring security fixes onto an otherwise-frozen base version, independently
of upstream -- this exact code path has had several
(CVE-2024-0582,
CVE-2024-35880, CVE-2025-21836; see
this write-up)
-- a combination where behavior diverging from upstream's own is plausible.
There's no known sysctl/ulimit fix for it; running the affected tests
against a different kernel (a genuine upstream stable release, or a different
distribution's) is the only known way around it so far.