onix is a Rust rewrite of Python DeepDiff's core: byte-compatible output, 37-4245x faster, with ignore_order support included. Install it as deepdiff-rs, a drop-in DeepDiff class for Python, or run the diff engine as the onix command-line tool.
deepdiff-rs reads live Python objects (or JSON) and produces the exact same report DeepDiff does at verbose_level=2, so it slots into code that already parses DeepDiff output while running dramatically faster on large or deeply nested inputs.
Status (September 2026): deepdiff-rs 0.x is live on PyPI (Python 3.9+, wheels for Linux x86_64/aarch64, macOS arm64/x86_64, and Windows x64, plus an sdist); the onix CLI builds from source, and nothing is on crates.io yet. Ordered and ignore_order diffing are complete, differentially tested against real DeepDiff 9.1.0, and benchmarked. It is 0.x, not stable or 1.0: the API may still change before 1.0.
- Install
- Quickstart
- Diffing tables
- Performance
- Reference
- Layout
- Known limitations
- Contributing
- License
Python (the deepdiff-rs package, import name deepdiff_rs):
pip install deepdiff-rsFrom source: with the Rust toolchain and maturin installed (Setup from scratch), run make python-test or maturin develop --release in crates/onix-py; the sequence is in CONTRIBUTING's Python bindings section.
CLI (the onix binary), from a clean clone:
cargo install --path crates/onix-cliLibrary crate (onix-core), a path dependency only (it sets publish = false):
[dependencies]
onix-core = { path = "crates/onix-core" }The drop-in DeepDiff class, on live Python objects:
from deepdiff_rs import DeepDiff
diff = DeepDiff({"a": 1}, {"a": 2})
if diff:
print(diff.to_json()) # byte-compatible with DeepDiff(...).to_json() at verbose_level=2
print(diff.to_dict()) # the same report as a native Python dict{"values_changed":{"root['a']":{"new_value":2,"old_value":1}}}
{'values_changed': {"root['a']": {'new_value': 2, 'old_value': 1}}}
A custom object diffs by its attributes:
from dataclasses import dataclass
from deepdiff_rs import DeepDiff
@dataclass
class Point:
x: int
y: int
print(DeepDiff(Point(1, 2), Point(1, 3)).to_json()){"values_changed":{"root.y":{"new_value":3,"old_value":2}}}
diff_json, the fast path when you already have JSON text (it parses, diffs, and serializes entirely in Rust, with no Python-object conversion):
from deepdiff_rs import diff_json
print(diff_json('{"a": 1}', '{"a": 2}')){"values_changed":{"root['a']":{"new_value":2,"old_value":1}}}
The onix CLI, diffing two JSON files (compact JSON to stdout, {} for no differences):
$ echo '{"a": 1}' > left.json
$ echo '{"a": 2}' > right.json
$ onix diff left.json right.json
{"values_changed":{"root['a']":{"new_value":2,"old_value":1}}}Pass --ignore-order to compare every list by value instead of by position, mirroring DeepDiff(..., ignore_order=True).
diff_tables compares two tables the way DeepDiff compares two objects. It takes any object implementing the Arrow PyCapsule interface — a pyarrow Table or RecordBatch, a polars DataFrame, a DuckDB relation — and imports it with no Python round trip. The two tables are matched on a required, non-empty set of key columns (the table's primary key).
It reports the schema diff (which columns were added, removed, or changed type), the keyed row diff (which rows were added, removed, or changed, and which keys are duplicated), and the per-cell diff (cells_changed): one row per changed cell, carrying the key columns, column, old_value/new_value, and change (became_null/became_non_null, type_changed, or value_changed), ordered by the canonical string rendering of the key columns (nulls first), then left-schema column order — the exact rendering and change-classification rules are in docs/design/row-diff.md's "Per-cell changes" section. Rows are matched by the key columns; rows_added, rows_removed, cells_changed, and duplicate_keys return Arrow tables, and summary() counts each outcome.
import pyarrow as pa
from deepdiff_rs import diff_tables
# 9 is a duplicate key
left = pa.table({"id": pa.array([1, 2, 3, 9, 9], pa.int64()), "amount": pa.array([10, 20, 30, 90, 91], pa.int32())})
right = pa.table({"id": pa.array([2, 3, 4], pa.int64()), "amount": pa.array([20, 31, 40], pa.int64()),
"note": pa.array(["a", "b", "c"], pa.string())})
diff = diff_tables(left, right, key=["id"])
print(diff.summary(), "added ids:", pa.table(diff.rows_added()).column("id").to_pylist(), "removed ids:", pa.table(diff.rows_removed()).column("id").to_pylist())
print("cells changed:", pa.table(diff.cells_changed()).to_pylist(), "duplicate keys:", pa.table(diff.duplicate_keys()).to_pylist()){'columns_added': 1, 'columns_removed': 0, 'columns_type_changed': 1, 'rows_added': 1, 'rows_removed': 1, 'rows_changed': 1, 'duplicate_keys': 1, 'null_keys': 0, 'cells_changed': 1} added ids: [4] removed ids: [1]
cells changed: [{'id': 3, 'column': 'amount', 'old_value': '30', 'new_value': '31', 'change': 'value_changed'}] duplicate keys: [{'id': 9, 'left_count': 2, 'right_count': 0}]
A key appearing more than once on either side is reported in duplicate_keys (left_count/right_count) and excluded from added/removed/changed; a null key matches its counterpart and is counted in null_keys. Rows are compared by the non-key columns present on both sides, with onix's value semantics (integers and integral floats fold together, all NaNs compare equal, 1.00 equals 1.0000, a timestamp compares by its instant and a time or duration by its value across units, dictionary-encoded values equal their plain form, and null equals null); a nested non-key column is skipped rather than compared. See the "Value semantics" section of docs/design/row-diff.md.
Type comparison uses the full logical Arrow type (timestamp unit and timezone, decimal precision and scale, and so on), but physical encodings sharing a logical type compare equal — a dictionary-encoded string equals a plain string, polars' Utf8View equals pyarrow's Utf8, the list variants normalize together, and a map compares equal however a library spells it — so the same table read through pyarrow, polars, or DuckDB reports no spurious type changes; nullability is ignored but reported in each record. The full rules are on normalized_type/map_entries in crates/onix-arrow/src/schema.rs. Column names must be unique on each side; a repeated name raises ValueError. diff.schema_arrow is the same result as an Arrow table: it implements __arrow_c_stream__, so polars.DataFrame(diff.schema_arrow) needs no pyarrow, and .to_pyarrow() returns a pyarrow.Table; pandas.api.interchange.from_dataframe(diff.schema_arrow) also works, but needs pyarrow installed regardless.
pyarrow is optional: pip install deepdiff-rs[arrow]. It's needed only for to_pyarrow() and passing pyarrow objects in — importing deepdiff_rs and diffing polars or DuckDB tables need it not at all. An object implementing neither Arrow protocol raises TypeError; to_pyarrow() without pyarrow installed raises ImportError naming the extra. diff.to_json() gives the whole diff — schema, summary, and rows_added/rows_removed/cells_changed/duplicate_keys in full, one JSON object per row — as a single string with no pyarrow, polars, or pandas needed; see Known limitations for its row cap.
Beyond a join-based diff, diff_tables reports duplicate and null keys rather than silently multiplying or dropping them, distinguishes type_changed from value_changed per cell, renders every value by one documented rule set (Python's, with Duration the one exception -- an ISO 8601 string, not Python's own str()), and produces byte-identical output at any thread count with memory proportional to row count from a streamed input. What that costs over a hand-rolled join is in perf/arrow/RESULTS.md's polars-backed ceiling section.
Two committed, regenerable reports back the numbers below; every figure here is copied verbatim from them.
The Python bindings against real deepdiff on live Python objects, the number a real caller pays (source: crates/onix-py/benchmarks/bench_bindings.py, macOS 26.5.1, Apple M5 Max, median of 11 isolated subprocess runs per side, run on 2026-09-23):
| Shape | deepdiff | deepdiff_rs | Speedup |
|---|---|---|---|
ignore_order, 10k shuffled ints, ~5% mutated (live objects) |
699.68ms | 66.23ms | 10.56x |
| peak RSS | 401.8 MB | 203.4 MB | 1.98x |
| CPU seconds | 0.699 s | 0.066 s | 10.56x |
| Heterogeneous API-payload records, n=20,000 (live objects) | 3486.50ms | 141.42ms | 24.65x |
| peak RSS | 228.1 MB | 274.6 MB | 0.83x |
| CPU seconds | 3.484 s | 0.141 s | 24.64x |
| Typed records (datetime/tuple/set fields), n=10,000 (live objects) | 795.03ms | 46.45ms | 17.11x |
| peak RSS | 169.8 MB | 175.6 MB | 0.97x |
| CPU seconds | 0.794 s | 0.046 s | 17.11x |
Same typed-records shape, ignore_order (live objects) |
60834.60ms | 864.82ms | 70.34x |
| peak RSS | 225.6 MB | 257.6 MB | 0.88x |
| CPU seconds | 60.786 s | 0.863 s | 70.47x |
Same ignore_order shape, via diff_json (JSON-string path) |
731.31ms | 74.54ms | 9.81x |
| peak RSS | 402.3 MB | 203.9 MB | 1.97x |
| CPU seconds | 0.730 s | 0.074 s | 9.81x |
Same API-payload shape, via diff_json (JSON-string path) |
5125.41ms | 95.33ms | 53.76x |
| peak RSS | 249.1 MB | 269.7 MB | 0.92x |
| CPU seconds | 4.971 s | 0.095 s | 52.22x |
| Same API-payload shape, both tools reading two JSON files from disk | 4774.33ms | 97.33ms | 49.05x |
| peak RSS | 249.0 MB | 269.7 MB | 0.92x |
| CPU seconds | 4.771 s | 0.097 s | 49.12x |
The engine's own diff-only time and peak resident memory against pinned deepdiff 9.1.0 (source: perf/RESULTS.md, same machine, median over tier-appropriate runs, diff time excluding process startup and JSON parsing on both sides):
| Fixture | onix diff-only (median, min-max) | deepdiff diff-only (median, min-max) | Speedup | onix peak RSS | deepdiff peak RSS | Memory ratio | ≥5x threshold |
|---|---|---|---|---|---|---|---|
flat_dict_10k |
3.154 ms (3.058 ms-3.220 ms) | 141.155 ms (140.606 ms-142.781 ms) | 44.75x | 5.78 MB | 39.29 MB | 6.79x | ✅ |
flat_dict_100k |
38.440 ms (38.000 ms-38.806 ms) | 1.594 s (1.581 s-1.602 s) | 41.47x | 40.57 MB | 110.82 MB | 2.73x | ✅ |
flat_dict_1m |
460.082 ms (454.820 ms-465.133 ms) | 17.061 s (16.926 s-17.162 s) | 37.08x | 478.15 MB | 753.65 MB | 1.58x | ✅ |
flat_list_100k |
82.907 ms (81.131 ms-84.654 ms) | 4.751 s (4.715 s-4.820 s) | 57.31x | 38.17 MB | 154.95 MB | 4.06x | ✅ |
nested_uniform_d6_b10 |
207.242 ms (205.249 ms-217.027 ms) | 71.458 s (70.959 s-71.599 s) | 344.80x | 227.41 MB | 868.32 MB | 3.82x | ✅ |
api_payloads |
162.489 ms (159.446 ms-178.736 ms) | 93.764 s (93.687 s-94.474 s) | 577.05x | 270.09 MB | 609.93 MB | 2.26x | ✅ |
deep_narrow_d120 |
0.029 ms (0.028 ms-0.030 ms) | 123.650 ms (123.230 ms-125.018 ms) | 4245.55x | 2.15 MB | 41.27 MB | 19.23x | ✅ |
startup_trivial |
0.001 ms (0.001 ms-0.001 ms) | 0.175 ms (0.169 ms-0.183 ms) | 147.75x | 2.15 MB | 32.67 MB | 15.22x | ✅ |
ignore_order_10k |
73.475 ms (71.672 ms-73.940 ms) | 12.976 s (12.900 s-13.005 s) | 176.60x | 60.11 MB | 345.19 MB | 5.74x | ✅ |
identical_1m |
9.254 ms (6.726 ms-9.875 ms) | 15.790 s (15.660 s-15.989 s) | 1706.35x | 315.41 MB | 503.19 MB | 1.60x | ✅ |
Both reports carry their full methodology, fairness rules, and the reproduce command. perf/RESULTS.md is an upper bound (JSON parsed straight into the engine, no Python-object conversion); the bindings table is the product-surface number. Regenerate them with perf/run_bench.sh and crates/onix-py/benchmarks/bench_bindings.py (see CONTRIBUTING.md). The Arrow table diff, on two seeded ~5 GB parquet pairs (a narrow and a wide fixture), takes 1.81x DuckDB's time and 1.62x polars' on the narrow pair and 3.01x DuckDB's and 5.24x polars' on the wide pair, in the Results tables of perf/arrow/RESULTS.md, regenerated with perf/arrow/bench_tables.py.
Python API. The public surface is DeepDiff, diff_json, diff_tables (returning a TableDiff), MaxDepthError, MAX_DEPTH_CEILING, and __version__ (a str matching the installed deepdiff-rs distribution version).
DeepDiff(t1, t2, ignore_order=False, max_depth=None): diffs two live Python objects of a supported type (see Known limitations);.to_json()returns the DeepDiff-compatible JSON string,.to_dict()the same report as a dict — with Python types preserved, so a value the diff found in atuple,setorfrozensetcomes back as one and adatetime/date/time/timedeltacomes back as a real one of those (a custom object, though, comes back as a plaindictof its attributes, since onix cannot reconstruct the instance — see Known limitations) — and the instance is falsy when there is no difference. Theset_item_added/set_item_removedcategories are lists of path strings, each ending in the item itself (root['a'][2],root['x'],root[(1, 2)]).diff_json(a, b, ignore_order=False, max_depth=None) -> str: diffs two JSON strings entirely in Rust and returns the report as a JSON string.MaxDepthError(aValueErrorsubclass) is raised when input exceedsmax_depth;MAX_DEPTH_CEILING(20,000) is the hard upper bound onmax_depth; a larger value raisesValueErrorbefore any diffing.diff_tables(left, right, key=[...], threads=None) -> TableDiff: diffs two Arrow tables (see Diffing tables).threadssets the row diff's worker count —Noneuses the machine's available parallelism,1runs single-threaded, and a value below 1 or above 1024 raisesValueError(the diff spawns one worker per thread); the result is byte-identical at any value.TableDiff.schemais the list of changed columns,.schema_arrowthe same as an Arrow table,.summary()the schema and row change counts,.to_json()the full diff (schema, summary, and every row-level member, capped at 10,000 embedded rows);.rows_added(),.rows_removed(),.cells_changed(), and.duplicate_keys()return Arrow tables.
CLI. onix diff <a.json> <b.json> [--max-depth N] [--ignore-order] [--timing] reads both files as JSON and prints a compact, single-line DeepDiff-compatible report to stdout ({} when there is no difference).
--max-depth Noverrides the recursion-depth bound (default: theONIX_MAX_DEPTHenvironment variable if set, else 512).--ignore-ordercompares every list by hash-based matching instead of by position, mirroringDeepDiff(..., ignore_order=True).--timingprints one line of JSON ({"parse_ns": N, "diff_ns": N}) to stderr.
Exit codes:
| Code | Meaning |
|---|---|
0 |
Diff computed successfully (whether or not the report is empty; differences are carried in the stdout JSON, not the exit code). |
1 |
Usage error (missing/unknown subcommand, wrong argument count, unknown flag, non-numeric --max-depth); details and a usage line go to stderr. |
2 |
I/O error (e.g. a missing input file) or a JSON-parse error on either input. |
3 |
max_depth exceeded; the path that tripped the bound goes to stderr. |
crates/onix-core # the diff engine (library, no I/O)
crates/onix-cli # the `onix` binary (thin CLI over the core)
crates/onix-arrow # Arrow table diffing (schema diff and keyed row diff)
crates/onix-py # PyO3 bindings, published as `deepdiff-rs`
docs/design/ # algorithm/invariant reference pages (list-diff, ignore-order, value-model, value-conversion, depth-budget, row-diff)
scripts/ # gen_goldens.py: regenerates tests/golden/ from real DeepDiff
tests/golden # DeepDiff-generated expected outputs (the compatibility corpus)
perf/ # cross-language benchmark harness and RESULTS.md
- Only the core diff is implemented:
exclude_paths,significant_digits, custom operators,verbose_level != 2, and delta/patch are not (yet) supported. - Value types: the supported types are
None,bool,int,float(NaN/Infinity/-Infinityincluded),str,dict,list,tuple,set,frozenset,datetime.datetime,datetime.date,datetime.time, anddatetime.timedelta; aset/frozensetmember may be any of these except alist,dictorset, or a subclass of those that defines__hash__, which raiseTypeError, matching Python's own hashability rule, transitively through whatever the member nests. Atzinfowhoseutcoffset()is not a whole number of seconds raisesValueError, fordatetimeandtimealike. A custom object is diffed by its attributes (see the Custom objects bullet below). A subclass of a supporteddict,list,tuple,set,frozenset,datetime.datetime,datetime.date,datetime.time, ordatetime.timedelta(includingnamedtupleand pandas'Timestamp) converts and compares as its base type but carries its own class name into atype_changesentry, matching DeepDiff'stype(obj).__name__reporting, with three exceptions: a calendar-type (datetime/date/time/timedelta) subclass held as aset/frozensetmember compares as its base with notype_changes, matching DeepDiff, since set membership has no pairwise comparison to report one against; atuple/frozensetsubclass, including anamedtuple, is not accepted as aset/frozensetmember at all and raisesTypeErrorwhere DeepDiff accepts and compares it by value (a documented divergence); andnamedtuplediffs positionally rather than by field (also documented). Atype_changesentry'sold_type/new_typeare type names into_dict(), where DeepDiff returns the type objects. The Numbers and strings, Datetimes and Sets bullets below cover the deliberate divergences for those types. Seetests/golden/README.md. - Dict keys and paths: a
dictkey may bestr,None,bool,int,float,datetime.datetime,datetime.date, or atupleof those (never nested), or atuple/datetime/datesubclass key (namedtupleincluded), accepted and matched against its base-type value the same way; adictkey outside that list, a nestedtuplekey included, raisesTypeErrornaming the key's type and path. Atuple/datetime/datesubclass dict key whose__eq__/__hash__is overridden matches by its base type's structural value, where DeepDiff matches by the subclass's own equality (seetests/golden/README.md), and adatetime/datesubclass key's path renders deterministically, where DeepDiff's collapses toNone(seetests/golden/README.md). A non-strkey's path renders via Python's ownrepr(), except atuplekey, which splits into one bracket group per element (root[1][2]for(1, 2), neverroot[(1, 2)], with one deliberate exception for a real DeepDiff bug on the empty tuple, and another on a non-finite-float key — seetests/golden/README.mdandtests/golden/README.md). A nested dict value's ownbool/None/int/floatkey stringifies the wayjson.dumpsdoes, but itsdatetime/date/tuplekey renders that samerepr()text where DeepDiff's ownto_json()raisesTypeErroron one — a superset, not a difference in the findings (seetests/golden/README.md).1/1.0/Truematch as the same key between two dicts (Pythondict/setequality). Two keys whose rendered paths collide collapse to one entry in both tools, but the surviving finding follows DeepDiff's dict insertion order and onix's alphabetical key order (seetests/golden/README.md). - Numbers and strings: an
intof any magnitude converts and keeps its exact value (an arbitrary-precision integer beyondi64/u64is compared, ordered, hashed and rendered by its full digits). Underignore_order, comparing an integer beyondf64::MAX(about2**1024) reports the change deterministically where real DeepDiff raisesOverflowError(its distance function callsfloat()on it unguarded), a documented crash-class divergence like the datetime one below (seetests/golden/README.md). Reading JSON text (diff_jsonand the CLI, not the Python-object path), an integer beyondi64/u64parses as the nearestfloat, so two documents whose integers differ only past that range compare equal and diff to{}— a limitation of the JSON number reader, not the value model, tracked in #92. Non-finite floats (NaN,Infinity,-Infinity) compare and hash like real Python: twoNaNs are never equal,Infinity == Infinity,to_json()renders the same bareNaN/Infinity/-Infinitytokens Python'sjson.dumpsdoes, andignore_ordermatching treats everyNaNas one shared item, matchingDeepHash. The one divergence, always deterministic: this crate's value model carries Python object identity for custom objects only, so twoNaNs outside one shared custom object always compare unequal here, where DeepDiff sometimes reports no difference (t1 is t2) or lets one collapse into another (asetmember, an ordered-list match) when the two objects, or their containers, happen to be the same one. Seetests/golden/README.md. Astr(ordictkey) containing a lone (unpaired) surrogate code point (e.g.'\udc80', legal in Python but not encodable as UTF-8) is compared and reported exactly like any otherstr, matching DeepDiff's plain==;to_json()renders it withjson.dumps's own single-backslash\uXXXXescape. The one accepted divergence, always deterministic: hashing one — aset/frozensetmember, or any value at all onceignore_order=True(DeepHashhashes every value there, not just a set's members) — crashes real DeepDiff with an unhandledUnicodeEncodeError, where onix hashes by code point and reports normally. Seetests/golden/README.md. Astrinside atupleorfrozensetset item is escaped exactly as Python'srepr()escapes it, against Unicode 16.0.0; on a Python older than 3.14 (an olderunicodedatatable), a code point assigned to Unicode after that Python's own version is escaped by DeepDiff and rendered literally by onix. Seetests/golden/README.md. - Datetimes: compare by instant, with a naive value read as UTC, matching DeepDiff. A changed pair is reported normalized to UTC (
to_json()renders...+00:00,to_dict()returns UTC-awaredatetimes); everywhere else a datetime keeps its raw value. Three deliberate departures:to_json()renders adateasYYYY-MM-DDwhere DeepDiff's ownto_json()raisesTypeError(a documented superset); azoneinfo/pytztzinfo comes back fromto_dict()as a fixed-offsetdatetime.timezonecarrying the offset it was in force at, not the original zone object (seetests/golden/README.md's "Normalized versus raw datetimes" section, its "Fixed-offsettzinforound-trip" point); and a set holding both a naive and an aware value at one instant reports both as members, where DeepDiff's own digest cache can report only one (seetests/golden/README.md's "Set iteration order" section, its "A naive and an aware datetime" point). Comparing two datetimes whose UTC form would leave year 1..=9999 raisesValueErrornaming the path, where DeepDiff raisesOverflowError; underignore_orderDeepDiff's hasher normalizes every datetime and so raises for such a value even when it is only added, removed, or shuffled, where onix hashes by instant and reports it normally, raising the sameValueErroronly where it is compared (seetests/golden/README.md).truncate_datetimeis not supported. Atime/timedelta, unlike a datetime, is never normalized for report (DeepDiff comparestime/date/timedeltawith a plain!=), and a naivetimeis never equal to an aware one;to_json()renders atimeastime.isoformat()'s bytes and atimedeltaasstr(timedelta)'s, both supersets (seetests/golden/README.md). Underignore_order,DeepHashhashes atimeby whole seconds-of-day only — dropping the microsecond and any offset, a confirmed upstream quirk — while atimedeltahashes exactly; seetests/golden/README.md's "Known DeepDiff quirks" section. - Sets: diffed deterministically, where DeepDiff's own answers depend on the order the running process happens to iterate a set in (hash order, and
PYTHONHASHSEED-dependent forstrmembers) or on how its digest cache/computation handles a tuple, frozenset, or calendar member independently of Python's own==. Each consequence — entry order, which member of an equality class is reported, set-versus-sequence coercion, and a tuple/frozenset member's own (positional, not order-/repetition-insensitive) matching rule — is shown with both tools' output intests/golden/README.md's "Set iteration order" section. A report holding afrozensetvalue also serializes to JSON here, where DeepDiff's ownto_json()raisesTypeError— a superset, not a difference in the findings. A set serialized into a report, and afrozensetrendered inside aset_item_*path, list their members in one fixed structural order, where DeepDiff lists them in the process's set iteration order (seetests/golden/README.md). - Custom objects: diff by their attributes, matching DeepDiff's
_diff_obj(attribute_added/attribute_removed,root.attrpaths,type_changesbetween two different class objects), and anEnummember by itsnameandvalue, matching_diff_enum. A type DeepDiff routes to a handler onix lacks —bytes,bytearray, a generator or any other iterable,complex,Decimal,Fraction, anumpyscalar,uuid,ipaddress, a class object, a module, or a bareobject()— raisesTypeErrorat the root; below it, such a value is equal only to the same object, as DeepDiff'st1 is t2check makes it, and raisesTypeErrornaming its path wherever the report would have to show it. Two identical Python objects at one position are never walked, a class attribute converts only when a report compares it with a shadowing value, and a class attribute is left out of a whole object in a report, as in DeepDiff (seetests/golden/README.md). Apydanticmodel raisesTypeErrorwhere DeepDiff diffs it (seetests/golden/README.md). A custom non-dictMapping, and are.Patternpair whose named groups differ, raiseTypeErrorwhere DeepDiff diffs them (seetests/golden/README.md). A@propertythat raisesAttributeError, or an unset slot beside a__dict__, raisesTypeErrorwhere DeepDiff reports the object asunprocessed(seetests/golden/README.md). A child that points back at an object on its path reports nothing when it is on the first side; on the second side only it is compared as the object it points back at, as DeepDiff'sparents_idstracks only the first side's ancestry (seetests/golden/README.md). Three attribute-view divergences (ignore_orderhashing, and the whole-object values into_json()andto_dict()) are tracked in #99, seetests/golden/README.md. - Adversarially deep input raises
MaxDepthErrorinstead of crashing: the defaultmax_depthis 512 and the hard ceiling isMAX_DEPTH_CEILING(20,000). A class attribute or an object a child points back at that nests past the budget can raiseMaxDepthErrorwhenignore_ordercompares two list items that contain such a value, at the path inside it where the budget runs out; not when the same pair was already compared at a shallower position (seedocs/design/ignore-order.md's "Distance memo" section). ADeepDiff/diff_jsoncall whose input nests past 32 levels runs on a worker thread that reservesmax(max_depth * 8,192 * 2, 512 KiB)bytes of virtual address space (committed lazily, so resident memory is unaffected): 8 MiB at the defaultmax_depth, 312.5 MiB at the ceiling (constants incrates/onix-py/src/guard.rs, per-level cost measured bycrates/onix-core/examples/stack_frame_cost.rs). Everydiff_tablescall reserves the ceiling's 312.5 MiB regardless of any depth, because the Arrow import recurses before the depth is known. Concurrent calls each reserve their own, and a process whose address-space limit (RLIMIT_AS, strict overcommit, a host with less RAM plus swap than the reservation) cannot hold it raisesRuntimeErrorinstead of diffing. JSON text read bydiff_jsonand the CLI is refused past 127 nested arrays or objects by the JSON parser (aValueErrorfromdiff_json, exit code 2 from the CLI) regardless ofmax_depth. Seedocs/design/value-model.mdfor the parser limit andcrates/onix-py/src/guard.rsfor the depth guard. ignore_orderpairing isO(N^2)in unpaired elements per side and carries a polynomial cost in both time and memory with input depth (about 7.5 s and 146 MB at depth 400 on a few-KB input, growing roughly cubically with depth, all under the defaultmax_depth); it has nomax_passes/max_diffscutoff, so bound the size and depth of untrusted input yourself. Seedocs/design/ignore-order.md's "Bounds" section.- A
values_changedbetween two multi-line strings runs adifflib-styleO(N*M)line diff on the default path with no opt-out, worst when changes are spread evenly through the text (about 35 s for a heavily edited 1 MB string and growing quadratically, so a few megabytes is minutes), so bound the size of untrusted strings yourself. Seecrates/onix-core/src/unified_diff.rs. - Ordered sequences (
listortuple) of scalars (null, bool, number, string, datetime, date, time, timedelta) run adifflib-styleO(N*M)matcher on the default path with its popular-element (autojunk) purge disabled forDeepDiffparity, worst for sequences of a few repeated values with dense edits (about 620 s for two 8,000-element lists of two repeated values with every other element changed, growing faster than quadratically), so bound the size of untrusted sequences yourself. Seecrates/onix-core/src/lcs.rs. diff_tablesinherits some Arrow-interchange quirks: a column name with an embedded NUL byte (\0) arrives truncated at the NUL through the C Data Interface (the report shows the truncated name); a list of structs named exactlykey/valuewith a nullable key is not distinguished from a real map, so a migration between the two is not reported as a type change (polars exports both as the same Arrow type — seemap_entriesincrates/onix-arrow/src/schema.rs); and a polars all-null (Null-typed) column fails at Arrow C import with aValueError(the datatype "Null" doesn't expect buffer at index 0), so give such a column a concrete type first (a pyarrow all-Nonecolumn, inferred asnull, works and compares as all-null).- In
diff_tables, DuckDB labels aTIMESTAMP WITH TIME ZONEcolumn with the connection's session time zone when it exports to Arrow (a UTC session asTimestamp(µs, "UTC"), anAmerica/New_Yorksession asTimestamp(µs, "America/New_York")), so on a non-UTC machine such a column can be reported as a type change against a UTC column from another library. RunSET TimeZone='UTC'on the DuckDB connection first for a deterministic, machine-independent result. diff_tablesrefuses a column whose Arrow type is nested deeper thanMAX_NESTING_DEPTH(128) with aMaxDepthError, because comparing arbitrarily deep nesting would overflow the native stack. Importing a schema nested many thousands of levels deep is also slow regardless, a cost of the Arrow C Data Interface itself.diff_tablesrow diff memory and temp disk, by term, in the table below (wall time isN·log Nacross workers). The cell pass is both sides' spilled rows × column width, plus the rendered output, 2xcells_changed(a per-cell term of changed cells × cell width), and also copies each column whose two sides share a type once per partition, aligned for comparison (changed rows ÷ partitions × that column's width; byte-view columns are compared as their spilled large-offset type, so they copy in full). Kept rows pin their input batch, per side: the whole batch when over half is kept, its view data otherwise. The size gate peeks up to 50,000 rows or 64 MB decoded per side, holding one whole producer batch per side beyond it.
| Term | Shape | Threads | Figure |
|---|---|---|---|
| Hash: 32 B per row per side (left only in parallel) | linear, 8M / 37M rows/side |
1 / 18 | 600 / 618 MB at 8M; 2.7 / 2.5 GB at 37M |
| Duplicate-key report | dup, 200k rows/side, 16 B / 1 KB keys |
18 | 39 MB / 0.83 GB |
| Cell pass, output-dominated | wide, 200k rows/side, 1 KB cell, every row changed |
2 / 18 / 64 | 1.45 / 0.96 / 1.25 GB |
| Cell pass, output-dominated | wide, 1M rows/side, 1 KB cell, every row changed |
18 | about 4.8 GB (4.77 to 4.91 GB over ten runs, two builds) |
| Cell pass, spill-dominated | manycols, 150k rows/side, eight 512 B columns, one differing |
2 / 18 / 64 | 2.39 / 1.56 / 1.96 GB |
| Cell pass, aligned copy | int64diff, 500k changed rows/side, three equal 1 KB columns (Utf8View, BinaryView, Utf8), only an Int64 differing |
2 / 18 | 6.24 / 2.84 GB |
| Right-side duplicate keys | 500k right keys the left holds once, repeated in a later batch (duprightonce), 1M rows/side, 2 KB of value columns |
18 | +952 MB (whole-process peak 2.32 GB) |
| Right-only keys | per 1M keys the left lacks, 8 B int64 keys (the entry is 32 B at any key width) |
18 | +85 MB |
| Kept rows pinned | 100 removed rows over a 1M-row side, two 1 KB Utf8View columns |
1 / 18 | ~2.0 GB |
| Size-gate peek | 49,999 rows/side, 8 KB cells; 100-row batches, then one whole-side batch | 1 / 18 | 12 / 65 MB; ~1.6 GB / 1.2 GB |
| Temp disk, repeated right keys | per 500k right keys the left holds once that a later batch repeats (one spilled row each), 2 KB of value columns, 1M rows/side | 2 / 64 | 1.03 GB |
| Temp disk, every spool resident at once | wide pair | 2 / 18 / 64 | ~23.8 GB |
| Whole-process peak | wide pair | 2 / 18 / 64 | 39.8 / 28.8 / 31.4 GB |
- Temp disk: each input spools to an anonymous file (
tempfile: unlinked, mode 0600, no name). Changed rows spill to anonymous per-partition files, byte-view columns cast to a large-offset type (i32 caps at ~2 GB) and dictionaries decoded, independent of the partition count (spilled undecoded, about 97 GB at 64 threads against 31.4 GB decoded); in parallel also one row per right key the left holds once that a later batch repeats. All resident until the cell pass ends; on Linux the spill may be a RAM-backedtmpfs. A full temp filesystem raisesValueErrornamingTMPDIR. - None of the row-diff memory, temp-disk or size-gate-peek costs in this section has a built-in cap. Bound: row count; changed fraction; right-side duplicate keys (on the parallel path, one full-width row per repeated key, spilled or held); column widths (key-column width, for duplicate-heavy data; for the cell pass, the total width of all common value columns, since every one spills in full per changed row whether it changed or not, plus the rendered changed cells; any column's width for the size-gate peek, which buffers decoded cells whether or not they change); the producer's batch size; the thread count: the per-partition terms are lowest at 18 threads (the default) and higher at 2 (two partitions each holding half the rows) and at 64; fewer partitions enlarge the aligned copy (the table's 2 / 18 / 64 rows).
- Figures are the peak resident set of
ROW_DIFF_THREADS=<n> ROW_DIFF_BATCH=<rows> cargo run -p onix-arrow --release --example row_diff_rss(method and shapes incrates/onix-arrow/examples/row_diff_rss.rs's module doc), macOS on an Apple M-series laptop, 2026-09-24 (int64diff: 2026-10-08), medians over 3 runs, except the 1-thread kept-rows (viewsparse) figure and the two 150kmanycolsrows at 18 threads over 5, the 8M hash row one run, the 37M rows four,int64diff8 at 2 threads and 3 at 18. Seeperf/arrow/RESULTS.mdand its column-wise compare section for these rows, its size-gate peek and disk usage sections for the peek and decode-vs-undecoded partition figures, andcrates/onix-arrow/src/row_diff.rsfor the implementation. TableDiff.to_json()is the one member with a built-in cap: it embedsrows_added,rows_removed,cells_changed, andduplicate_keysin full, one JSON object per row, so it refuses withValueError— naming the row count and the cap — once those four together hold more than 10,000 rows. The cap bounds row count only, the same way the cell-pass term above is stated as changed cells times cell width, not column count or cell width — so a table under the row cap but with wide or large cells can still be large; use the Arrow-returning accessors instead. Seecrates/onix-arrow/src/json_rows.rs.diff_tablescompares scalar columns by value (hashed: null, booleans, every integer, float and decimal width, strings and binary in every encoding, timestamps, dates, times, durations, intervals, and dictionaries of these; refused withValueError: run-end encoded columns and any type-and-unit combination Arrow itself cannot build; nested non-key columns skipped, nested key columns refused), with the exact enumeration inis_hashableandis_nestedincrates/onix-arrow/src/row_diff.rs. It also refuses a key column whose type differs across the two inputs after encoding normalization (the conservative choice: a primary key that changed type is refused rather than guessed, not coerced). A row whose only difference is a lossless type change — anInt32widened toInt64, a timestamp unit change at the same instant — hashes equal on both sides, so it is in neitherrows_changednorcells_changed; the column's type change is still reported inschema.- Output is byte-identical to DeepDiff except for the cases listed above and the path-rendering quirks in
tests/golden/README.md;tests/golden/README.mdenumerates every accepted exception, including integers past2^53(the limit of exactf64representation) inside ordered scalar lists andignore_orderpairing among naive datetimes, which DeepDiff ranks using the process's local timezone while onix reads a naive value as UTC everywhere.
Issues and pull requests are welcome. Open an issue to report a bug, a DeepDiff divergence (include both inputs and the report each engine produces), or a question. Building from source, the quality checks, the golden corpus, benchmarking, mutation testing, and publishing are all in CONTRIBUTING.md.
MIT: see LICENSE.
onix reimplements algorithms from CPython's difflib (PSF License) and
reproduces the behavior of DeepDiff (MIT); their notices and license texts are
in THIRD-PARTY-NOTICES.md.