Skip to content

Add internal/nodepage: the on-disk byte layout of a B+Tree node - #9

Merged
suhailopensource merged 1 commit into
mainfrom
step/0x09-nodes-as-pages
Sep 26, 2026
Merged

suhailopensource merged 1 commit into
mainfrom
step/0x09-nodes-as-pages

Conversation

@suhailopensource

Copy link
Copy Markdown
Member

Nodes as pages

Why this is needed

The B+Tree from Phase A works and disappears the moment the process exits. Every
link inside it — children []*node, next *node — is a Go pointer, which is a
RAM address. Write one to disk, restart, and it points at whatever happens to
occupy that address now.

in memory                          on disk
children []*node   8 bytes each    child core.PageID   4 bytes each
                   a RAM address                       a page number
                   valid until exit                    valid forever

Page 12 is at byte 12 * 4096 = 49152 — in this process, in the next one, and on
another machine. This change introduces internal/nodepage, the byte layout that
makes that swap possible. The tree does not use it yet; this is the format the
tree will be written in.

What is in the change

file lines what
internal/nodepage/nodepage.go 273 the format: constructors, accessors, append, search, validate
internal/nodepage/describe.go 179 the reader: hexdump, region map, byte-for-byte decode, fanout table
internal/nodepage/nodepage_test.go 456 17 tests, including a golden byte layout and 11 corruption cases
internal/nodepage/fuzz_test.go 141 1 fuzz target, 2 benchmarks
cmd/scutedb-demo/nodepage.go 220 the walkthrough behind make demo-nodepage

nodepage imports core and page and nothing else. It is a second importer of
page, alongside slots.

The layout

A slotted page — the same shape Postgres and SQLite use:

region         bytes         holds
page header    0..15         id, kind, key count, free start, free end
node header    16..23        next leaf / first child, level, 2 pad
slot array     24..          4 bytes per entry, in KEY order
free space                   shrinks from both ends
cells          ..4095        key + payload, in INSERTION order

The page header is the existing 16 bytes from 0x02, untouched. The node header
adds 8: a 4-byte link, a 2-byte level, and 2 bytes of padding that exist so the
slot array starts at byte 24 and stays 4-byte aligned.

The link field holds two different things. In a leaf it is the next leaf;
in an internal node it is the leftmost child. They are never both needed, the
kind byte says which one it is, and sharing saves 4 bytes on every node in the
tree. Next() and FirstChild() are two names for the same four bytes.

How the code is organised

Making a node. Both constructors set FreeStart to 24 rather than 16,
because the node header sits between the page header and the slot array.

function does
NewLeaf(id) an empty leaf page, level 0
NewInternal(id, level) an empty internal page; rejects level 0, because leaf and level-0 are the same statement and allowing both spellings allows them to disagree
Load(p page.Page) wraps a page read off disk, rejecting anything that is not a B+Tree node kind

Node is struct{ page.Page } — embedding, so ID(), Kind(), FreeSpace()
and the rest of the page methods are available on a node without redeclaring any
of them.

Reading a node.

function does
KeyCount() number of entries, from the page header's item count
Level() 0 for a leaf, higher going up
Next() / FirstChild() the shared link field, named for the kind
KeyRef(i) key i as a view into the page — no copy, used on every comparison
Key(i) key i copied — safe to keep after the page buffer is recycled
Row(i) leaf only: the RowID stored after key i
Child(i) internal only: child i, where Child(0) comes from the header
Search(key) binary search over the slot array, returning position and whether it was an exact hit
ChildFor(key) internal only: which child subtree a key belongs in

The Key / KeyRef split follows the same rule codec and slots already use,
and it matters more here: once a buffer pool recycles page buffers, a retained
view points at a different page's bytes.

Writing a node.

function does
AppendLeaf(key, rid) appends a cell of key ‖ rowPage(4) ‖ rowSlot(2)
AppendInternal(key, child) appends a cell of key ‖ childPage(4)
Fits(keyLen) / EntrySize(keyLen) ask before appending, rather than appending and handling failure
appendCell the shared body: writes the cell at the bottom of free space, writes the slot at the top, moves both pointers, bumps the count

Append is append, not insert: it writes the slot at the end of the array, so
the caller must supply keys in ascending order. Validate catches a caller that
does not. Ordered insertion into the middle belongs to the step that makes the
tree use this format, and building it now would be an untested guess at what that
step needs.

Checking a node. Validate() verifies twelve facts and is called by nearly
every test.

Reading it as a human. Describe() prints the region map, a hexdump of the
headers, every field decoded with its offset, the slot array, the cells, and for
an internal node the full child list. Hexdump(b, base) and FanoutTable(keyLen)
are the pieces it is built from, usable on their own.

The flow of one lookup

This is the path make demo-nodepage walks, reading pages out of a real file:

pf.Read(0)                      4096 bytes at byte offset 0 * 4096
  -> nodepage.Load(raw)         checks the kind byte says btree-internal
  -> n.Leaf()                   false, so keep descending
  -> n.ChildFor(key)            binary search the slot array via KeyRef
                                  -> n.Child(i)  reads a child page number
                                     i == 0 -> from the node header
                                     i  > 0 -> last 4 bytes of cell i-1
  -> returns core.PageID(2)

pf.Read(2)                      4096 bytes at byte offset 2 * 4096
  -> nodepage.Load(raw)         kind byte says btree-leaf
  -> n.Leaf()                   true, so stop descending
  -> n.Search(key)              binary search -> (slot 1, found)
  -> n.Row(1)                   last 6 bytes of that cell -> RowID{200, 30}

Every hop is a page number turned into a byte offset by page.Offset(id), which
is one multiplication. Nothing in that path is a pointer.

Why the slot array exists

Two requirements pull against each other: keys must be in sorted order so search
can be binary, and key bytes must not move on every insert.

The slot array resolves it by sorting 4-byte references instead of the keys.
Inserting in the middle shifts a handful of slots; the key bytes stay wherever
they were first written. That is why the slot array is sorted while the cells are
in arrival order, and why KeyRef(i) has to go through a slot rather than
indexing into the page directly.

Growing the two regions toward each other, from opposite ends, means neither
needs a size fixed in advance. They meet when the page is full, and FreeSpace()
— which is just FreeEnd - FreeStart from the existing page header — is the
entire bookkeeping. A fixed boundary between them would have to be guessed, and
would be wrong for every key length but one.

The key length is never stored, because the suffix is fixed per kind: a leaf key
is cellLen - 6, an internal key is cellLen - 4.

Fanout is computed, not chosen

  page header        16
  node header         8
  left for entries 4072

  internal entry = 4 slot + 8 key + 4 child  = 16 bytes -> 254 keys, 255 children
  leaf entry     = 4 slot + 8 key + 6 row id = 18 bytes -> 226 keys

  levels   reads per get  keys it holds
  1        1              226
  2        2              57,630
  3        3              14,695,650
  4        4              3,747,390,750

This corrects a figure in the README. 0x06 claimed roughly 340 keys per
internal node, from 4080 / (8 + 4). That ignored the 8-byte node header and the
4-byte slot each entry needs. The real number is 254 — bookkeeping costs 25%
of the fanout. The estimate erred in the usual direction: it counted the data and
forgot the structure that makes the data findable. MaxKeys and Fanout now
compute it, and TestMaxKeysMatchesWhatActuallyFits fills a real page for every
key length from 1 to 64, for both kinds, and compares.

An internal node with n keys has n+1 children; the extra one has no slot to
sit in, which is what the header link is for on that side.

Reading it back without any of this code

Three pages written to a file, then read with xxd:

$ xxd -s 4096 -l 32 tree.db
00001000: 0000 0001 0300 0002 0020 0fe4 0000 0000  ......... ......
00001010: 0000 0002 0000 0000 0ff2 000e 0fe4 000e  ................
bytes value meaning
0000 0001 1 page id
03 3 kind, btree-leaf
0002 2 key count
0020 32 free start, 24 + 2 slots x 4
0fe4 4068 free end, where the lowest cell begins
0000 0002 2 next leaf is page 2
0000 0 level, 0 means leaf
0ff2 000e 4082, 14 slot 0 points at a 14-byte cell

Slot 0 names byte 4082 of page 1, which is file offset 4096 + 4082 = 8178:

$ xxd -s 8178 -l 14 tree.db
00001ff2: 7fff ffff ffff ffff 0000 0064 0000       ...........d..

7F FF FF FF FF FF FF FF is the key -1, sign bit flipped by the ordered key
encoding from 0x03. Then 0000 0064 is row page 100, and 0000 is slot 0.
Every layer built so far is visible in those fourteen bytes.

How the real engines do it

Postgres calls a page number a BlockNumber, a uint32 index into the relation
file, and its B-Tree pages carry a btpo_level field matching the level here.
SQLite uses 1-based 32-bit page numbers, page 1 always holding the schema. InnoDB
uses 32-bit page numbers inside a tablespace with 16 KB pages. None of them store
an address.

How it is verified

Validate checks twelve structural facts. The strongest is that the cells
exactly tile the region from free end to the end of the page — sorted by
offset, the first must start at free end, each must end where the next begins,
and the last must end at 4096. Gaps and overlaps are both rejected. It also
checks that slots are in strictly ascending key order, that free start equals
24 + 4 x count, that no cell is too short to hold its own suffix, that the
padding bytes are still zero, and that leaf-ness and level-0 agree.

  • 11 deliberate corruptions are each asserted to fail it: a slot pointing into
    the slot array, two slots aliasing one cell, a cell too short for a row id,
    swapped slots, moved free pointers, a used padding byte, a leaf claiming a
    level, a flipped kind byte.
  • TestGoldenLeafLayout pins 14 exact byte ranges, so the format cannot drift
    without a test failing, and asserts the free space between the regions is still
    all zeros.
  • TestRoundTripThroughARealFile writes through page.File, reads back, and
    asserts the bytes are identical.
  • FuzzNodePageSurvivesArbitraryKeys ran 2,382,065 executions building pages
    from arbitrary keys at arbitrary widths, validating and reading back each one.

116 tests, 10 fuzz targets, go test -race clean, go vet clean.

Not in this change

internal/btree still uses pointers; wiring it to this format — splitting and
merging pages rather than structs — is the next step. No free list, no buffer
pool, and the reserved header bytes still hold no checksum.

Run the walkthrough with make demo-nodepage.

Every link in the Phase A tree is a Go pointer, which is a RAM address,
which means the tree cannot survive the process that built it. This adds
a byte-for-byte page format where a link is a page number instead: page
12 is at byte 12*4096, in this process and in the next one.

The layout is a slotted page, the shape Postgres and SQLite use. A
16-byte page header, then an 8-byte node header holding the next-leaf or
first-child link, the level, and two bytes of padding that keep the slot
array 4-byte aligned. Slots grow up from byte 24 at 4 bytes each, cells
grow down from byte 4095, and free space is whatever lies between.

The point of the indirection is that the slot array is sorted while the
cells are not. Inserting a key in the middle shifts a few 4-byte slots
and never moves a key byte, so binary search stays O(log n) while writes
stay cheap. Growing the two regions toward each other also avoids having
to decide a split between them up front, which would be wrong for every
key length but one.

A leaf cell is key bytes followed by a 6-byte row id; an internal cell
is key bytes followed by a 4-byte child. The key length is never stored
because the suffix is fixed per kind. An internal node with n keys has
n+1 children and the extra one has no slot to sit in, so it lives in the
node header, sharing four bytes with the leaf's next pointer and
disambiguated by the kind byte.

Fanout is now computed rather than estimated. With 4096-byte pages and
8-byte keys an internal node holds 254 keys and 255 children, a leaf
holds 226, and three levels reach 14,695,650 keys. This corrects the
README, which said roughly 340 per internal node from 4080/(8+4). That
ignored the node header and the 4-byte slot each entry needs; overhead
is 25% of fanout. MaxKeys and Fanout compute it, and a test fills a real
page for every key length from 1 to 64 to confirm the arithmetic.

Validate checks twelve structural facts, the strongest being that the
cells exactly tile the region from free end to the end of the page with
no gaps and no overlaps, and that slots are in strictly ascending key
order. Eleven deliberate corruptions are asserted to fail it.
TestGoldenLeafLayout pins fourteen exact byte ranges so the format
cannot drift silently. FuzzNodePageSurvivesArbitraryKeys ran 2,382,065
executions, and a round trip through a real page.File asserts the bytes
read back are identical to the bytes written.

The B+Tree does not use this yet. 116 tests, 10 fuzz targets, race
clean.
@suhailopensource
suhailopensource merged commit 1f991b0 into main Sep 26, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant