Skip to content

Release v0.9.7 - #2

Merged
libraz merged 77 commits into
mainfrom
develop
Aug 6, 2026
Merged

libraz merged 77 commits into
mainfrom
develop

Conversation

@libraz

@libraz libraz commented Aug 6, 2026

Copy link
Copy Markdown
Owner

Promotes develop to main for the v0.9.7 release.

Release gate

Check Result
Oracle (formula + CF + workbook) 4185 cases, 100% passing, 0 failures, 294 documented skips
Fast ctest tier 11680/11680
WASM uncompressed 2.26 MiB (hard ceiling 3.00 MiB)
WASM Brotli 609 KiB (hard ceiling 768 KiB)

Highlights

  • formulon paginate across the CLI, C ABI and every binding, backed by fm_workbook_paginate.
  • Workbook memory-footprint estimate (fm_workbook_memory_usage), reported to V8 as external memory by the Node addon.
  • XLSB writer emits native styles, the mandatory worksheet prefix and dynamic-array metadata, and retains worksheet-tail records verbatim.
  • OOXML writer interns literal text into a generated shared-string table and escapes XML-invalid controls per context.
  • GROUPBY / PIVOTBY honour a total_depth of +/-2; number-format colour names are read in the UI locale.
  • Recalc pools its layer workers, tracks whole-axis references as rectangle dependencies, and scopes Tarjan to the dirty set.

Full list: CHANGELOG.md

libraz added 30 commits July 19, 2026 02:46
ci.yml's native test job triggers only on main, so a commit that breaks
the C++ suite gets no signal until the develop->main fast-forward. Add a
native-fast job that does a Release build and runs the fast ctest labels,
gated by the same expected_flakes allowlist ci.yml uses, so a red develop
surfaces on the develop push instead of at promotion time.
A Release build of the heavy C++ TUs compiled in parallel OOM-kills the
2-core ubuntu-latest runner (SIGTERM / exit 143). Switch the native-fast
job to the debug preset, matching ci.yml's build type and `make build`;
the stale-test failures this job guards against reproduce identically
under -O0.
The first run on an empty build cache does a full Debug compile of the
whole test suite on the 2-core runner and does not fit in 45 minutes;
raise the timeout to 60 (matching ci.yml). Cached incremental runs stay
well under it.
`cmake --build --parallel` with no bound links the heavy Debug test and
bench binaries across all logical CPUs at once and OOM-kills the 7 GB
runner (SIGTERM / exit 143), regardless of build type. Pin the build to
2 concurrent jobs (the runner's physical core count).
- remove the one-line Apache-2.0 copyright banner from every source file; the license terms live in the top-level LICENSE
- stop tools/codegen/gen_bindings.py from emitting that banner into generated binding translation units
…tersection

- ROW/COLUMN over a multi-cell range and OFFSET over a multi-cell rectangle now spill elementwise instead of collapsing to a scalar, including negative height/width, which walks the rectangle in reverse
- add a shared implicit_intersection.h so the @ operator and the VM dispatcher take the top-left element (#VALUE! for an empty array) instead of passing the operand through
- move ISFORMULA, FORMULATEXT, ISREF, SHEET and SHEETS onto a shared reference-call resolver so OFFSET / CHOOSE / IF / INDIRECT results resolve uniformly
- teach range_args the intersection operator (#NULL! when the operands are disjoint) and array literals directly
- cap dynamic-array allocation centrally at 2^20 cells with an overflow-checked multiply; SEQUENCE with a zero dimension yields #CALC! and EXPAND validates against sheet bounds
- add a shared omitted_arg.h, fixing SORT's omitted by_col / sort_index handling and reusing it in XLOOKUP
- bind an erroring LAMBDA argument lazily instead of short-circuiting the whole call
- clear a stale spill region in SpillCommitter when a formerly array-producing formula evaluates to a scalar
- add text_format/rounding.h (round to 15 significant digits, display-decimal rounding) and use it in FIXED, DOLLAR, BAHTTEXT and the numeric renderer, replacing ad hoc rounding that broke ties-away-from-zero and leaked binary-residue artifacts
- stop doubling the sign for conditional format sections that carry an explicit literal minus
- suppress the fraction part of a "# ?/?" format for whole values
- honour the minimum digit width of the [hh], [mm] and [ss] elapsed-time tokens
- enforce Excel's 32,767 UTF-16 unit text cap in SUBSTITUTE and return #VALUE! past it
- deduplicate the half-to-full katakana voicing tables into jp_kana_table, replacing a modular-arithmetic derivation that mis-composed dakuten and handakuten
- scale the regularized incomplete gamma iteration cap with the shape parameter and return #NUM! on genuine non-convergence instead of truncating silently
- avoid an Inf/Inf intermediate in F.DIST for large x
- interpolate MEDIAN and PERCENTILE.INC/EXC via weighted endpoints so extreme finite pairs stay finite
- derive the ROUND tie-break bias from DBL_EPSILON instead of a hardcoded constant
- share the near-integer snap helper and apply it to CEILING.PRECISE and FLOOR.PRECISE as well
- use fmod in ISEVEN/ISODD instead of an int64 cast that was undefined past 2^63
…ate builtins

- reject serials outside Excel's 0..9999-12-31 range in WORKDAY, WORKDAY.INTL, NETWORKDAYS, NETWORKDAYS.INTL and the financial date-argument reader, including during day stepping
- broadcast array arguments cell by cell in the date and time builtins instead of requiring scalars
- fall back to return type 1 for WEEKNUM with an invalid return_type on every profile
…al cell

- resolve the "width", "color", "prefix" and "format" info types from the referenced cell's style and the sheet's column layout instead of returning fixed stub values
…eping

- replace per-layer thread spawn/join with LayerWorkerPool, a barrier-synchronized pool that lives for the whole parallel pass, and thread the iterative-solve progress callback through the parallel path
- fix node-count bookkeeping in clear_dependencies_of, which checked the reverse-edge bucket before erasing it; add a borrowing dependencies_of_ref for hot paths and move cyclic-component classification into DepGraph
- cap eager dependency-range expansion at 10M cells and fall back to marking the cell volatile
- use cached formula-cell values for ordinary dependency-ordered recalc, not only iterative mode, bounding recursion depth on long chains
- reject an oversized unknown name in FunctionRegistry before allocating a normalized key
- reject a structured reference whose table bottom row is not below its top row, preventing an unsigned wraparound
- static_assert that the lazy-dispatch table is sorted and look it up by binary search
- allow a process-wide structured-log sink and minimum level to be installed, dropping below-threshold records before serialization
- stop the default stderr path from flushing on every record
… syntax

- validate the completed AST against the depth limit in parse(), covering left-associative chains that never recurse, and guard formula formatting and structural ref shifting with a shared iterative depth walker
- inspect both endpoints in the range-operator validity check instead of only the right-hand side
- reject any $-bearing identifier run in the tokenizer, catching forms like A$$1
- consolidate sheet-name quoting and quote all-numeric sheet names as well
- format a blank literal as an omitted argument rather than ""
- apply the row/column shift transform to 3-D references during structural edits
- an equal-threshold DataBar rule drew no bar at all; fill the bar completely, matching Excel
- rewrite hyperlink locations, data-validation formulas and pivot-cache worksheet sources on sheet rename, not only defined names
- freeze referencing formulas and defined names to #REF! on sheet removal, and reindex or drop the affected tables and pivot caches
- rewrite conditional-format and table ranges on row/column insert and delete, unregister then reregister in two phases to avoid a dependency-graph key collision when coordinates swap, and validate the count against the remaining sheet bounds
- trigger a dependency reindex when a defined name is added so #NAME? formulas can resolve
- store opaque sheets and raw worksheet extensions for OOXML round-trip fidelity
- add set_cell_text, and synchronize set_cell_xf_index and set_cell_phonetic under the same mutex so they no longer no-op on a missing row
- drop the spill table's per-phantom reverse index in favour of anchor-keyed rectangle scans
- carry chart, dialog and macro sheets through as opaque sheets instead of failing the read, and capture and re-emit package-level and per-sheet relationships of unrecognised types with zip-slip validation
- preserve workbook <fileVersion>, <fileSharing>, <extLst> and worksheet <protectedRanges>, <scenarios>, <customSheetViews>, <phoneticPr>, <ignoredErrors>, <legacyDrawingHF>, <picture>, <oleObjects> and <controls> in schema order; emit <sheetFormatPr> from the model and strip a stale sheetPr@tabHidden
- add xfId, the apply* and quotePrefix attributes and explicit presence bits for font bold/italic/strike to styles so a differential format's val="0" survives, and round-trip <colors>, <tableStyles>, <extLst> and unrecognised styleSheet children
- preserve <autoFilter>, <sortState>, <extLst> and unmodelled attributes on tables
- match a comment sheet's VML part by source path so it survives renumbering
- escape root extra namespace attributes on re-emission, and report kIoZipEncrypted distinctly from kIoZipCorrupt in ZipReader
- read an empty <v/> on a numeric cell as blank, validate a dynamic-array ref as an ordered rectangle anchored at its own cell, and keep a cached formula value loaded from file until an explicit recalc
…mmands

- add a shared file_io with read_file and write_file_atomically, centralizing the slurp and atomic-replace paths with mkstemp, permission preservation via fchmod, an EINTR-safe write loop and an fsync before rename
- collapse exit codes to {0, 1, 64} via exit_code_for_status, replacing a masking scheme that could collide after 8-bit truncation, and accept a -- terminator in eval
- escape control characters in dump's line-oriented output and report write failures
- evaluate eval through the read-only array C API instead of writing into A1 and recalculating, so =A1+1 sees an empty A1; pre-validate syntax and render multi-cell array results in both plain and JSON form
- preserve the workbook's own iteration settings in recalc --iterative
…ulas

- add a native xl/styles.bin emitter and round-trip row/column layout, merged rectangles and the date1904 flag
- write a formula that cannot be lowered to the supported Ptg set as its cached literal and report it through a new XlsbWriteResult downgrade count instead of aborting the package write; count and report deferred sheet features the same way
- read and re-emit unknown package-level relationships with escaped attributes
- extend the Ptg codec with parenthesis tokens, whole-column and whole-row area encoding, a decode depth guard and eleven more function-id entries
- add an xlsb_fuzz libFuzzer target registered as a SLOW ctest entry, and register the new styles writer and CLI file_io sources in the build
…ection

- project the layout through the C API using the workbook's Excel profile and the selected Compact, Tabular or Outline report layout, so the default ja-JP profile emits localized labels instead of the legacy English grid
- honour each field's sort direction when building hierarchy levels, and emit one column-subtotal entry per custom subtotal function carrying its own aggregation
- render an empty row/column intersection as blank instead of the aggregation identity
- place row subtotals according to each field's subtotal_top, and reserve cells through a checked multiply
- drop per-record string allocation in label matching, skip unresolved cache items instead of matching them against genuine blanks, and format numeric labels with the shortest round-trip renderer so GETPIVOTDATA matches
- make PivotTable::last_result a mutex-guarded shared_ptr<const PivotResult> snapshot so a reader cannot race a live replacement
- record the change in CHANGELOG.md
…xports

- add fm_workbook_set_cell_phonetic / fm_workbook_get_cell_phonetic, fm_workbook_set_iterative_enabled / fm_workbook_get_iterative, fm_styles_get_font_ex / fm_styles_add_font_ex with a new fm_font_record_ex carrying vert_align, and fm_workbook_save_xlsb_with_result; the new struct and functions are additive, so the existing ABI is unchanged
- return kNotFound from fm_sheet_get_comment_at for an absent comment and reserve kInvalidArgument for a bad sheet index
- write fm_workbook_set_text through Workbook::set_cell_text instead of interning into a per-handle deque that was never pruned
- clear the out-param in fm_workbook_load before validating its arguments
- add a C11 -fsyntax-only compile check of the public header as its own ctest entry
…WASM, Node and Python

- expose setCellPhonetic / getCellPhonetic and font vertAlign through the _ex font records in WASM
- add getCommentResult to Node and WASM, returning a status alongside the comment so an absent comment is distinguishable from an invalid sheet, and raise from Python's get_comment on a genuine error instead of returning None
- rename the binding-layer handle-destroyed constant to kBindingInvalidHandle so it no longer collides by name with the C ABI's kBindingNullPointer
- evaluate WASM and Python eval_formula read-only instead of writing into Sheet1!A1 and recalculating, so the intersection operator surfaces #NULL!
- add a WASM32 struct-layout check to check_binding_drift.py that derives field offsets from formulon_c.h and diffs them against the Python _structs.py layouts
- describe the default win-365-ja_JP profile in the binding READMEs
…overnance

- recapture the goldens against Excel 16.111.2; the content changes reflect the spill, WEEKNUM and OFFSET fixes on this branch
- add the cross_sheet_refs, intersect_operator, iterative_calc and spill_collision suites, and re-close the secondary-oracle goldens with documented divergences in a new tests/ironcalc_divergence.yaml
- declare per-case tolerance and compare mode in divergence.yaml (keyed by suite as well as id) and bake them into the generated goldens, replacing hardcoded variant heuristics in the C++ runner
- add divergence_check.py, which validates every divergence entry against a real case, suite, alias or documented non-oracle scope, and golden_coverage_check.py, which fails when a declared case has no golden, superseding the informational catalog coverage test
- register one ctest entry per secondary-oracle golden file instead of a single monolithic entry, so one allowlist exception can no longer mask every failure
- switch the macOS driver to app.api.iteration, whose predecessor stopped resolving
…one channel

- return ctest's skip code 77 when fewer than two channels are active, so a degraded run no longer reports success
- assert each active channel against the fixture's expected result rather than only cross-channel agreement
- add unit tests for the harness itself
…exclusion

- feed the ctest JSON listing to check_test_failures.py and fail on expected_flakes.txt entries that no longer name a registered test, so a stale allowlist cannot hide silently
- move the ctest label exclusion from SLOW|LOAD|BENCH to SLOW|BENCH|TSAN and follow suit in the Makefile targets
- drop the blanket secondary-oracle skip and three fixed criteria-coercion entries from the allowlist
- add install and command-line sections for eval, recalc and dump to the READMEs
- restate the WASM size bullet as the CI-enforced 3.00 MiB hard and 2.50 MiB soft ceilings measured by make size-check
- clarify that the function-metadata merge helpers are native Node and Python helpers and that the raw WASM C ABI has none by design
…n arena byte cap

- add Sheet::set_cell_cached_value_borrowed so a reader can install a workbook-owned shared-string view without copying the Text bytes into every cell
- bump a cell_enumeration_revision counter on each stored-cell, spill and structural mutation so consumers can invalidate a cached address set
- add RowLayout::has_height so an explicit row height is distinguishable from the sheet default
- add utils/a1_column with the bijective base-26 column-name encoder that the writers, CLI and reference formatting can share
- give Arena an optional byte budget and a sticky exhausted() flag that stays set until reset(), so a caller can tell resource exhaustion from a legitimate empty result
- cover pivot-anchor shifting across row and column inserts and deletes
…tersection operands

- accept a whole-column (A:A) or whole-row (1:3) tail after a sheet span, quoted or not, and build it as a Ref3D carrying the same whole-axis flags as its single-sheet counterpart
- treat a space before a quoted sheet name or a parenthesized range as the intersection operator, while a spaced opening parenthesis after a cell-shaped name stays a function call unless it closes a range tail
- render a 3-D whole-axis range as one A:C tail in the AST dump and the formatter instead of repeating the axis on both sides
- take the column-letter encoding from utils/a1_column and check the grid bound when formatting an A1 reference
- correct the header comments that described parser diagnostics as an editor-facing position model; they are an internal C++ API
…pe Tarjan to the dirty set

- record a whole-column / whole-row reference as a compact rectangle dependency instead of promoting the formula to volatile, and dirty every rectangle watcher from the workbook mutation path
- run Tarjan over the induced subgraph of dirty cells, and over the dirty cells inside the viewport closure for a partial recalc, instead of the workbook-wide graph
- stop inferring divergence from three non-decreasing residuals and stop overwriting an unresolved cycle with #NUM!; Excel iterates to the cap and keeps the last approximation
- report arena exhaustion as kOutOfMemory from full and partial recalc and record it on EvalState, so an allocation failure is no longer indistinguishable from a worksheet error
- evaluate an array literal as a first-class dynamic array in the tree walker and route the remaining literal paths through resolve_range_arg
- expand a 3-D whole-column / whole-row span to each sheet's populated extent in range-aware dispatch, and reject a span whose endpoints disagree on the whole axis
- index GROUPBY and PIVOTBY groups by a canonical hashed key instead of rescanning existing groups, and build the MODE frequency table by sorting instead of quadratic duplicate detection
- take the column-letter encoding from utils/a1_column in reference formatting and drop the unused RecalcStats::cancelled field
- add a single-cell-edit mode to the recalc benchmark that measures one dirty cell in a large workbook
…ution inverses

- surface an infinite or NaN result as #NUM! in SUMIF / SUMIFS, AVERAGEIF / AVERAGEIFS, PERCENTRANK.INC / .EXC and the complex-number formatter instead of emitting a non-finite number
- invert T.INV, F.INV and BETA.INV by bisection over an expanding bracket against a directly evaluated right tail, so heavy df=1 quantiles past the old fixed brackets and extreme-tail beta roots resolve instead of returning NaN
- clamp YIELD and ODDFYIELD iterates to the non-negative yield domain PRICE accepts, damping the Newton step and taking a one-sided derivative at the boundary
- clamp the HYPGEOM.DIST cumulative sum to the unit interval and bound the PERCENTRANK significance argument to the representable integer range
- record the pre-1900 WEEKDAY policy in tests/divergence.yaml: Formulon returns the real proleptic-Gregorian weekday rather than reproducing the 1900 serial-calendar offset
… OOXML text

- intern literal text cells into a generated xl/sharedStrings.xml, write those cells as t="s", and reserve the part across the content types, workbook relationships and passthrough collision set
- escape XML-invalid C0 controls with OOXML _xHHHH_ notation in element text and attribute values, escape a literal escape-shaped run as _x005F_xHHHH_, and decode both when reading cell, shared-string and phonetic text
- replace an invalid UTF-8 byte with U+FFFD so a written part cannot become unreparseable
- reject a malformed <item x> attribute in the pivot reader instead of defaulting it to index 0, which silently pointed at unrelated cache data
- write ht for a row whose explicit height is zero, driven by RowLayout::has_height, and retain a one-cell dynamic-array anchor so its metadata survives the write
- install shared-string cell text through set_cell_cached_value_borrowed rather than copying the payload into every cell
- take the column-letter encoding from utils/a1_column in the conditional-format, comments and cell-reference writers
libraz added 29 commits August 6, 2026 02:29
- add load_xml_buffer_inplace to src/io/xml_utils.*, a destructive part loader that parses out of the caller's buffer instead of the private copy pugixml allocates, dropping peak parse memory from two copies of the part to one; document the stricter contract (bytes are rewritten in situ, the document aliases them) and when the copying overload is still required
- factor the shared kIoXmlParse envelope and the parse flags out of both loaders so the copying and in-place paths are indistinguishable to callers matching on the message or context
- turn read_shared_strings and read_pivot_cache_records into sink functions taking the buffer by value, so the two parts whose size scales with workbook content can be parsed in place
- cover tree equality against the copying loader, buffer aliasing, whitespace-only leaf text and identical failure envelopes in tests/unit/io/xml_utils_test.cpp
- drop unknown_parts from OoxmlReadResult and XlsbReadResult and hand the payload to Workbook::set_passthrough_parts by move; passthrough carries every unmodelled binary in the package, xl/media/* above all, so mirroring it doubled the resident cost of opening a workbook with large embedded images
- switch the OOXML worksheet parse and the metadata-shell parse to load_xml_buffer_inplace, so a large sheet no longer costs twice its own XML size to open; order the shell and byte buffers so they outlive the documents that alias them
- pass the SST and pivot-cache-record buffers to their readers by move now that both parse in place
- read passthrough through workbook.passthrough_parts() in the metadata, roundtrip, reader and XLSB writer tests, and add a case asserting the payload survives moving the workbook out of the read result and still round-trips
- load xl/sharedStrings.xml in read_via_sax and resolve the pending shared-string cells that read_sheet_data_sax queues in its context, including phonetic payloads
- without the pass the comparison put resolved text on the DOM side against Text("") placeholders on the SAX side and reported the reader paths as divergent when they are not
- note in the en/ja root READMEs that the WASM builds (npm and PyPI) read worksheet XML through the DOM parser only while the native CLI switches to a streaming parser past 256 KiB, so loading needs memory proportional to the largest single worksheet's XML
- state that the peak is per worksheet rather than per workbook, that the practical ceiling is the 32-bit WASM address space, and that results are identical on every surface
- add the same note to the npm and Python package READMEs next to their load APIs
- add src/utils/thread_launch.{h,cpp}: launch_thread wraps _beginthreadex on Windows and pthread_create elsewhere and returns Expected<Thread, Error>, because every target builds with exceptions disabled and std::thread's constructor terminates the process when the OS refuses a thread
- make Thread move-only and self-joining in its destructor, with no detach, so a worker cannot outlive the state it borrows; map EAGAIN and ENOMEM to kOutOfMemory since both a thread ceiling and an exhausted Emscripten pthread pool report them
- expose a launch-failure injection hook so the refusal paths are reachable from tests without real thread exhaustion
- treat SchedulerConfig::num_threads as an upper bound in LayerWorkerPool: keep whichever workers started, log the shortfall as recalc.worker.launch_failed, and dispatch each layer against the live worker count
- fall through to serial evaluation when no worker starts at all, counting those layers in SchedulerStats::serial_fallback_steps; results do not depend on how many workers were obtained
- keep the worker slots in a reserved vector so a launched worker's pointer into it can never dangle
- register the new TU in CMakeLists.txt and add unit coverage plus scheduler cases for the no-worker and partial-launch pools
- add Workbook::approximate_memory_bytes(), estimating the cell store (slots, formula text, owned cached text, phonetic text), the shared-string storage every Text value borrows from, the passthrough part payloads and the workbook-level metadata strings
- leave allocator overhead, the dependency graph, arenas, styles and pivot caches out of the estimate; the figure is a pressure signal for host runtimes, not an allocation ledger, and it is O(cells) so it belongs at load and recalc boundaries
- surface it as fm_workbook_memory_usage in the C ABI, rejecting a NULL handle or out-param with kBindingNullPointer so a binding can tell an empty workbook from a broken one
- correct the load comment in the C API workbook part: the workbook now owns the passthrough payload as well as the text storage, so moving it out of the read result takes everything the handle needs
- cover growth with the cell store, passthrough payload accounting, repeat-call stability, and the C ABI null rejections
- add Workbook::SyncExternalMemory, which reads fm_workbook_memory_usage and hands V8 the delta since the last report, so a collector that only sees a pointer-sized handle feels the real weight of a loaded workbook
- report on create, load, recalc and handle destruction; destroying reports the whole amount back, and doing it after the destroy keeps the accounting from going negative if the release path runs twice
- add a memoryUsage() instance method returning the current estimate and refreshing the report, which is what keeps a long run of cell writes from leaving the figure stale
- document the method in index.d.ts and add a memory-accounting section to the package README, including that the hint is advisory and dispose() remains the way to free at a chosen point
- correct the README method count to 179 and note the two native-only methods
- cover growth, the post-dispose zero, repeat-call stability and a recalc in the smoke suite
- add a BrtName record builder emitting the workbook-scoped prefix the reader actually consumes, and let WorkbookBin place name records after the sheet bundle where both name passes look for them
- assert that PtgNameX in a cell formula keeps the cached value, stores no formula and counts one undecoded formula, since reusing the name index against this workbook's own table would silently retarget the reference
- add a decodable PtgInt defined name as the positive control, so a zero counter means nothing was undecodable rather than no name record having been parsed
- assert that PtgNameX as a defined name's own body is dropped instead of registered with a fabricated formula, is counted, and still lets the read succeed
- decode the reference text of an INDIRECT call in count_areas and count each comma-separated area it names, so =AREAS(INDIRECT("A1,B2")) returns 2 instead of 1
- keep the decode textual only: commas inside a quoted sheet qualifier do not split, so 'Q1,Q2'!A1 stays one area
- fall back to the previous single-opaque-reference count for text the decoder does not recognise (R1C1, a defined name, malformed input), including the explicit a1 = FALSE form that parse_a1_ref rejects on its own
- drop INDIRECT from returns_single_reference, which now covers only OFFSET, IFS, SWITCH, INDEX and XLOOKUP
- cover unions mixing ranges and sheet qualifiers, the quoted-comma case and the decoder-rejects fallback in the AREAS unit tests
- no oracle regeneration was run: the arg_indirect_union golden still carries the skipped marker and needs a macOS Excel run
- stop short-circuiting BYROW / BYCOL on the first errored slice: the lambda result, error or not, lands in that slice's own output cell so the other slices still spill their results
- hand REDUCE and SCAN an errored cell verbatim instead of returning it, so an IFERROR-guarded body recovers and keeps folding
- let an unguarded body leave the error where it belongs: SCAN carries it forward on the accumulator into every later cell, REDUCE into the returned accumulator
- update the invoke_lambda_with_values contract comment, since no caller pre-filters errored arguments any more
- move the bycol_error_in_column, map_error_propagation and scan_error_in_element oracle formulas off 1/0 inside an array constant, an expression Excel rejects at formula entry, onto the #DIV/0! error literal
- cover the per-cell error placement, the guarded-recovery path and SCAN's accumulator poisoning in the helper unit tests
- no oracle regeneration was run: the three goldens still carry the skipped marker and need a macOS Excel run
- run the autocorrelation scan of detect_seasonality on the linear-trend residual instead of the raw series, so a perfectly linear (non-stationary) series reports no period rather than picking lag 2 off correlations that all clear the threshold
- fit the trend by ordinary least squares against the sample index, which closes in one pass because the resample grid is equally spaced by construction
- add a relative residual floor measured against the raw mean-centred variance, so a residual that is pure rounding noise reports no period instead of ranking arbitrary autocorrelations
- rank positive autocorrelation only: a negative correlation at lag k means the residual flips sign every k steps, so the repetition is at 2k and magnitude ranking would let that lag outrank the real period
- cover the linear-series case and a seasonal pattern riding a trend in the unit tests
- no oracle regeneration was run: the forecast_ets_seasonality_linear_no_period golden still carries the skipped marker and needs a macOS Excel run
- floor the residual sum of squares in build_stats_output at the squared rounding error the data's own scale can still resolve, so an exact fit with stats=TRUE reports a huge finite F like Excel instead of surfacing #N/A
- leave the floor at zero for a flat response (ss_total == 0), keeping F at #N/A where it is genuinely undefined
- note in the comment that only the finiteness at that floor is contractual; the magnitude carries no statistical meaning and will not match Excel digit for digit
- cover the perfect-fit and flat-response cases in the LINEST unit tests
- no oracle regeneration was run; linest_univariate_stats_true stays a divergence for the remaining magnitude difference
- add build_outer_grouping and subtotal_label to the shared common TU: the outer level is the first key column alone, built over the groups the flat composite-key pass already produced, with outer groups numbered in first-occurrence order
- emit one subtotal row per outer group in GROUPBY, placed on the same side of its block as the grand total, with the sort applying within the hierarchy so each outer group stays contiguous
- do the same for PIVOTBY rows, aggregating each outer group per (column group, value column) intersection, leaving empty intersections blank and filling the grand-total strip only in the single-value-column layout
- keep a single key column on the +/-1 layout, since the outer level would coincide with the groups themselves and every subtotal row would restate its group
- keep PIVOTBY subtotal columns (|col_total_depth| == 2) unimplemented: they insert an extra value-wide block per outer column group, reshaping every column index, so that request still falls back to the grand-total-only layout behind the eval.pivotby.subtotals_unsupported warning, now tagged with the axis
- drop the GROUPBY fallback warning and its structured-log include, and derive the ja-JP subtotal suffix from the grand-total label of the same locale, isolated in subtotal_label so one edit corrects every row once goldens exist
- cover both placements, the single-key-column shape, contiguity under a sort and the diagnostic expectations in the GROUPBY and PIVOTBY unit tests
- no oracle regeneration was run: the groupby_subtotal_depth_two and pivotby_row_subtotal_depth_two goldens still carry the skipped marker and need a macOS Excel run, which is also what confirms the ja-JP wording
- drop arg_indirect_union, now that AREAS decodes an INDIRECT reference string textually and counts every area it names
- drop bycol_error_in_column, map_error_propagation and scan_error_in_element, now that the lambda helpers spill an error into its own output slot instead of collapsing the call
- drop forecast_ets_seasonality_linear_no_period, now that the seasonality detector removes the linear trend before the autocorrelation scan
- drop groupby_subtotal_depth_two and pivotby_row_subtotal_depth_two, now that both functions build the outer/inner hierarchy and place a subtotal row per outer group
- record in their place that PIVOTBY subtotal columns remain unimplemented behind the grand-total-only fallback, with no oracle case covering that shape yet
- rewrite the linest_univariate_stats_true reason: both engines now report a huge finite F and the two magnitudes differ by a numerically unstable factor, which is the same class as logest_univariate_stats_true
- note that the affected goldens still carry the skipped marker and need a macOS regeneration run to go live, and that the ja-JP subtotal-row wording is derived from the grand-total label pending that run
- return from scan_ident_or_cellref_or_bool after consuming one byte when
  the identifier run decoded nothing, so byte_pos_ always advances
- a truncated or otherwise malformed lead byte previously left the position
  unmoved while still classifying as an identifier start, so the scanner and
  the dispatcher spun forever appending empty tokens until memory ran out
- report the byte as LexerErrorCode::InvalidCharacter, matching how the
  dispatcher treats any other byte it cannot begin a token with
- cover a truncated lead byte, a lead byte followed by a non-continuation
  byte, and that the remainder of the formula still tokenizes
- the TU already uses standard algorithms but relied on a transitive
  include for their declarations
- drop count_areas_in_reference_text / count_areas_in_indirect from
  areas_lazy.cpp: AREAS no longer decodes an INDIRECT reference string to
  count comma-separated areas
- INDIRECT builds at most one rectangle, so a successful resolution counts
  as 1 area and a failure travels out of AREAS as that error, carried by the
  new kAreasPropagate sentinel and its propagated ErrorCode out-param
- a union string such as "A1,B2" is not a reference INDIRECT can return,
  so =AREAS(INDIRECT("A1,B2")) is #REF! rather than 2
- the fresh Excel 365 ja-JP golden (2026-08-06) confirms #REF!, replacing
  the skipped arg_indirect_union case and the count behaviour introduced by
  "Count every area an INDIRECT reference string names"
- a total_depth of +/-2 subtotal row now copies its outer key verbatim into
  the first key column and leaves the inner key columns blank, instead of
  the derived "<key> Total" label; subtotal_label is removed
- add hierarchy_grand_total_label, which promotes the ja-JP grand total to
  its distinct top-level wording when subtotal rows share the same column,
  and reuses grand_total_label elsewhere
- apply the same rule to PIVOTBY's row axis; col_total_depth +/-2 stays
  unimplemented and still falls back to the +/-1 layout, so move the
  divergence note onto the column comment
- goldens regenerated from a live Excel 365 ja-JP run on 2026-08-06 confirm
  the layout, replacing the skipped groupby_subtotal_depth_two and
  pivotby_row_subtotal_depth_two cases
- regenerate on 2026-08-06: the FORECAST.ETS seasonality case and the
  BYCOL / MAP / SCAN inline-array error cases now carry real expectations
  instead of a skip
- record last_verified_excel_version 16.111.2 on the divergence entries the
  run confirms
- re-diagnose text_color_red_discarded: the read-back is stable, and the
  real cause is that ja-JP Excel localizes TEXT color tokens, accepting the
  Japanese name and rejecting [Red], while Formulon knows only the English
  names, so prefer flips from formulon to excel
- re-scope the GETPIVOTDATA page-axis entry: observing Excel's surface needs
  a PivotTable in the workbook, so only the workbook track's Windows primary
  can clear it
- add the FM_FUZZ_SANITIZERS cache variable, default unchanged at
  fuzzer,address,undefined, so a host can drop to fuzzer,undefined while the
  nightly runner keeps the full set
- LLVM 18's ASan deadlocks inside its own shadow-memory init against the
  macOS 26 dyld, so the harnesses never reach main there
- libFuzzer accepts either individual replay inputs or corpus directories
  but not a mix, so naming the XLSB fixture alongside seeds/xlsb made
  FuzzXlsbSmoke exit with "Not a directory" before running anything
- stage the fixture and the seeds into one build-tree corpus directory and
  point the harness at that single directory
- add --brotli-ceiling-bytes (768 KiB) and --brotli-soft-ceiling-bytes
  (640 KiB) to tools/bench/wasm_size_report.sh, with the same value/inline
  forms and integer validation as the uncompressed options
- report a per-measure status in both the text and JSON output, and derive
  the overall status from the worse of the two so a Brotli overage exits 1
  even while the uncompressed size still has headroom
- keep a host without brotli on PATH on a skipped Brotli status instead of
  a failure, so a missing tool cannot look like a size regression
- state both hard ceilings in README.md and README_ja.md
- match the eight localized colour names in is_color_specifier, so the
  English [Red] / [ColorN] spellings now fall through to the
  invalid-bracket path and surface #VALUE! exactly as Excel does
- match a name as a prefix, so a longer localized word that starts with a
  colour name still parses, while a leading blank does not
- accept blanks after the localized index prefix, leading zeros and
  full-width digits in the indexed form, and read its digit run greedily
  so an already overshot value cannot fall back into 1..56
- flag a second colour in one section as an invalid bracket: either
  bracket alone is inert, but Excel rejects the pair with #VALUE!
- cover the accepted and rejected spellings in
  tests/unit/eval/text_format_number_test.cpp and
  tests/oracle/cases/text_format.yaml
- oracle diff: the primary text_format golden was regenerated from Excel
  365 ja-JP 16.111.2, and the text_color_red_discarded skip is retired
  from tests/divergence.yaml now that the case is answerable
- divergence-skip the imported third-party engine corpus on its [Red]#.#
  column with prefer: mac, since that engine is en-US and discards the
  bracket
- flip the "[Red] colour bracket" probe from diverged to impl, matching
  the number-format colour behaviour it now describes
- point the BYCOL / MAP / SCAN error-propagation probes at #DIV/0!
  instead of 1/0, which is what the oracle cases actually contain, and
  flip them to impl
- flip FORECAST.ETS.SEASONALITY's linear-trend probe to impl for the
  same reason
- document the paginate CLI subcommand alongside eval / recalc / dump: it
  resolves the print area, page breaks and page count
- restate the pivot-cache non-goal: the stored pivotCacheRecords snapshot
  is preserved as-is rather than rebuilt from the source range, while
  PivotTable results are still evaluated on demand through the API
- note that the XLSB reader/writer covers styles and cross-sheet 3-D
  references, leaving array-constant literals and post-2007 future
  function IDs as the remaining gaps against the OOXML path
- replace the verification bullet list with a table: 11680/11680 on the
  PR gate, 11932/11932 with the SLOW tier, 3942/3942 primary formula
  oracle with 140 skips, 23/23 conditional formatting, 63/70 workbook
  with 7 skips, and 12510/12510 on the imported third-party engine
  corpus with 168 documented divergences
- raise the oracle category count to 97 and refresh the closure counts to
  518 of 522, with the remaining 4 (ARRAYTOTEXT, FILTERXML, GETPIVOTDATA,
  PHONETIC) failing only behaviors_declared
- keep README.md and README_ja.md in step
- add the 0.9.7 CHANGELOG section written from the 72 commits since v0.9.6,
  split across Added, Fixed, Performance, Changed, Testing, Build / CI and
  Documentation
- retarget the [Unreleased] compare link at v0.9.7 and add a [0.9.7] link
- add the v0.9.7 entry to the release index in docs/releases/README.md and
  drop the "Latest release" marker from v0.9.6
- bump the version string in lockstep across CMakeLists.txt, src/version.h,
  both npm package.json files and pyproject.toml, and regenerate uv.lock
- release gate: oracle 4185 cases, 100% passing, 0 failures, 294 documented
  skips; fast ctest tier 11680/11680; wasm 2373826 bytes uncompressed
  (2.26 MiB) and 623900 bytes brotli (609 KiB), both inside the 3.00 MiB /
  768 KiB hard ceilings
- default-construct in place in the XLSB styles writer's Normalized helper
  (storage.assign(1U, T{})) and drop the fallback reference parameter, so
  the five call sites no longer bind a T{} temporary to a reference
  parameter of a reference-returning function
- GCC's -Wdangling-reference reports that binding regardless of whether the
  reference can escape, which -Werror turned into a hard error on the Linux
  jobs while Apple Clang with libc++ stayed silent; behaviour is unchanged
- include <cmath> in eval/dynamic_array/filtering.cpp for std::isfinite,
  which the TU had been reaching through a transitive include that
  libstdc++ does not provide
- both translation units compile clean under GCC 15 with -Wall -Wextra
  -Wdangling-reference -Werror, and the pre-change helper signature still
  reproduces the diagnostic
- the local build succeeds and the 359 XLSB / styles / filtering ctest
  cases pass
- take the structured binding in BuildCache by const auto& so the range-for
  no longer copies each std::pair out of the braced initializer list
- silences -Wrange-loop-construct, which -Werror turned into a build failure
  for the concurrency test target under GCC; Apple Clang does not warn
- every translation unit in the build compile database, 567 of them, was
  syntax-checked with GCC 15 under the project warning and -Werror flags,
  and none reports a diagnostic
- saturate the coerced double in read_digits before the cast, so TRUNC,
  ROUND, ROUNDDOWN and ROUNDUP read the same digit count on every host
- converting a double outside int's range is undefined and the two
  architectures disagree: AArch64 saturates to INT_MAX while x86-64 yields
  INT_MIN, so a digit count such as 1E50 read as a huge positive on one
  host and a huge negative on the other
- =TRUNC(1E100, 1E50) was therefore a no-op on macOS and collapsed to 0 on
  Linux; every caller compares the value against a +/-308 threshold, which
  sits far inside int's range, so saturating carries the same meaning as
  the real magnitude
- cover TRUNC with a huge positive and a huge negative digit count, plus a
  case exercising ROUND, ROUNDDOWN and ROUNDUP through the same helper
- the imported third-party engine corpus caught this on the Linux CI runner
  via calc_tests__MROUND_TRUNC_INT__TRUNC cases C48 and C58; no golden
  changed, since the goldens already carried the correct values and only
  the Linux result was wrong
- drop the three SheetSpillTest entries: those cases compile only under NDEBUG, while CI builds Debug, where the same spill shape and footprint rejection paths are registered as SheetSpillDeathTest abort cases
- the names were never registered in the configuration CI runs, so the stale-entry check in check_test_failures.py failed the job on every push even with a passing suite
- both configurations cover the rejection paths, so the recorded debt is already paid and the entries are removed rather than renamed
- state in the header that an entry must name a test CTest registers in the Debug configuration
@libraz
libraz merged commit edd9c5e into main Aug 6, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant