Repository navigation
Table cell matching merges two columns when a PDF textline spans them; the resulting cell overlaps the neighbouring cell in the same row #4189
Description
Activity
Hi, I’d like to work on this issue. I can investigate the table cell matching logic and implement a fix to prevent matched cells from overlapping neighbouring columns, while preserving the existing correct structure prediction. I’ll also add regression tests covering the textline/word merge case described here and verify that the fix works across the relevant PDF backends. Please let me know if I can take this up.
Reacted by meinolf-piwekI measured how often the self-inconsistency you describe — a cell whose box overlaps the
next cell of its own row — occurs at docling 2.129.0 with default options, since it is
detectable from theDoclingDocumentalone.Rates (overlap ≥ 1 pt and ≥ 10% of the narrower cell; identical at 25% and 50% on
the first corpus):corpus tables with an overlap rows with an overlap 42 born-digital PDFs with a reference table (opendataloader-bench) 2 / 54 (3.7%) 3 / 372 (0.8%) 131 page images with tables (PubTabNet, OmniDocBench, DocLayNet slices) 12 / 152 (7.9%) 39 / 1,991 (2.0%) TEDS-S does not see any of it, because the grid is right in every case.
Mechanism, in the two born-digital hits. Since #1238 the matcher works on word-level
cells, so the textline fusion you observed is not the only route. What the boxes show is
the orphan-recovery fusion docling-ibm-models#188 describes:- 3-column table, a row whose column-0 value is the single word
Events. The column-2
cell came out asEvents presentations, webinars, …with box x = 76–524, i.e. starting
at column 0's x; column 0 is empty; the box overlaps the column-1 cell by 48 pt. The
next row splits the two-line cellPhysical or / digitalby word intoPhysical digital
(col 0) andor(col 1). - 4-column sheet, row
Highlight: the column-3 cell's box is x = 55–910, the full table
width — the row label from column 0 was fused into the last column's paragraph and the
other three cells shifted one column left.
A cheap guard. The overlap is computable inside
MatchingPostProcessorbefore the
table is emitted: refuse an orphan attachment whose resulting box would cross into a
neighbouring cell of the same row (keep the word as its own cell, or drop it with the
existing WARNING — never concatenate), and, second, re-split a matched cell whose box
crosses a column boundary by the column bands the good-IoU cells define. The first alone
clears both hits above. I have the measurement script and both documents; happy to
prototype the guard in docling-ibm-models and report the delta with the same check, if
that is a direction the maintainers would take.- 3-column table, a row whose column-0 value is the single word
@AbarnaaSree Sorry I got carried away chasing other issues. I'm interested in what you come up with here.
Correction to my comment above: the mechanism sentence is wrong for the two documents I described. Replaying
MatchingPostProcessor.processoffline on the dumped inputs (54 tables; output byte-identical to the live run) shows that orphan recovery adds nothing in either document — the fusion is complete after step 8, the final intersection-over-PDF-cell assignment, and it starts in TableFormer's own cell boxes:- 7×3 table: the column-2 box of five of the seven rows was predicted on top of column 0 (x ≈ 145–250, where the row labels sit). The label
Events(x 152–214) intersects the misplaced column-2 box at 1.0 and the column-0 box at 0.74, so it goes to column 2. The real column-2 words (x ≥ 526) then match nothing, arrive as orphans into a column-2 band that now starts at x = 145, and step 8.a aligns the cell box to the union of its words (152–1048) — the overlap. - 4×4 table: the column-3 box of one row was predicted at x 100–212, on the row label
Highlight; the row's other boxes are each one column too far left. The same rule then shifts the whole row.
So for these two cases the grid is right, the boxes are misplaced, and the assignment trusts the boxes without checking them against the columns they claim. docling-ibm-models#188 (the orphan fallback) is a different route to the same symptom; the textline merge in this issue's own example is a third.
@AbarnaaSree — you offered to take this on, so I'll leave the fix to you. If useful: the overlap check runs on the emitted
DoclingDocumentalone (a cell whose box crosses a same-row neighbour in another column, no declared span), and a replay harness that dumpsMatchingPostProcessor.processinputs per table and re-runs the post-processor without the model made the tracing above cheap. One observation from that replay that may matter for a fix insideMatchingPostProcessor:_deduplicate_cellsreturns the same structural cell several times, so any statistic overtable_cellsthere needs to dedupe bycell_idfirst. Happy to share the two documents and the harness.- 7×3 table: the column-2 box of five of the seven rows was predicted on top of column 0 (x ≈ 145–250, where the row labels sit). The label
Summary
When a PDF stores a row's leading field and the next field on one baseline with a
gap no wider than a word space,
docling-parsereturns them as a singletextline cell, and TableFormer's cell matching assigns that whole line to the
first column. The structure prediction is correct; the cell content is not.
The result is self-inconsistent in a way the library can detect: the cell
assigned to column 0 overlaps, horizontally, the cell assigned to column 1 in the
same row.
Observed
A four-column ruled table,
Datum | Erläuterung | Betrag Soll EUR | Betrag Haben EUR. Predicted shape15x4— correct.Header row, showing where the columns actually are:
'Datum''Erläuterung''Betrag Soll EUR''Betrag Haben EUR'First data row:
'07.01.2025 Basislastschrift CREDITOR NAME REDACTED S.C.A.''304-0000000-0000000 XXXXXXXXXXXXXX'The column-0 cell runs to x=265.2, well past the column-1 header at x=123.0, and
overlaps the column-1 cell of its own row across almost its entire width. Only
'07.01.2025'belongs in column 0.Where it originates
Not in the structure prediction, and not in the backend.
docling-parsereportsthe date and the first description word as one textline cell:
At word level they are separate — no word cell spans the two columns. The
merge is at line level, and it is metrically reasonable in isolation: the gap
between the date and the following word is about the width of a word space, so
no threshold on horizontal distance alone separates "word gap" from "column
gap" here.
All three PDF backends produce the identical fused cell, which is consistent
with the cause being the line merge rather than any backend's geometry:
^date\s+\Sthreaded_docling_parse(default)dlparse_v4pypdfium2The documented workaround makes it worse
do_cell_matching's own comment anticipates this case:Setting it to
Falseon this document:^date\s+\Sdo_cell_matching=Truedo_cell_matching=FalseMore fused cells, and a sixth of the text gone. On a second document (137 tables,
a technical export) it did reduce fusion 211 → 45, but lost 7% of the words and
emitted duplicated header cells. It is not a usable workaround in either
direction.
Reproducing without our file
The document is a personal bank statement and cannot be shared. The case should
reconstruct from:
Datum | Erläuterung | …text ends
dd.mm.yyyydate in column 0 and text in column 1 on thesame baseline, with the horizontal gap between the date and the first word of
column 1 close to the font's space width
only the first line of each row is affected
do_ocr=False,TableFormerMode.ACCURATE.Suggested fixes
falls inside the cell's bounding box. The prediction is already correct here,
and the inconsistency is detectable without it: a cell assigned to column n
should not horizontally overlap a cell assigned to column n+1 in the same
row.
The words are correctly separated already; only the line merge loses the
boundary.
docling_parse'sDecodeConfighashorizontal_cell_toleranceandword_space_width_factor_for_merge, butdoclingbuilds theDecodeConfiginternally and plumbs onlyenforce_same_font(
docling/backend/docling_parse_backend.py::_make_docling_parse_decode_config),so a caller cannot reach them from
PdfPipelineOptions— nor fromdocling-serve, whose convert options expose onlypdf_backendandtable_cell_matching.(1) or (2) would fix it; (3) would let callers whose corpus has this shape work
around it in the meantime.
Environment
docling-serve1.31.0,docling2.124.0,docling-core2.93.0,docling-ibm-models4.0.1,docling-parse7.16.0, imageghcr.io/docling-project/docling-serve-cu130:main, Python 3.12, CUDA 13.Scale
Measured across an archive of 6 599 PDFs (14 330 tables in 3 849 documents). A
narrow leading column of dates or reference numbers beside a wide description is
the most common table shape in it, so this affects a large share of the tables we
extract. In the 13-table statement above, up to 134 of 479 cells match the fused
pattern — an upper bound, since the pattern also matches cells that legitimately
begin with a date, such as a
"Kontostand am 03.01.2025, Auszug Nr. 1"heading.(Payment and mandate references above are redacted, with character counts
preserved so the quoted x-ranges remain consistent.)