Skip to content

Table cell matching merges two columns when a PDF textline spans them; the resulting cell overlaps the neighbouring cell in the same row #4189

Description

@meinolf-piwek

Summary

When a PDF stores a row's leading field and the next field on one baseline with a
gap no wider than a word space, docling-parse returns them as a single
textline cell, and TableFormer's cell matching assigns that whole line to the
first column. The structure prediction is correct; the cell content is not.

The result is self-inconsistent in a way the library can detect: the cell
assigned to column 0 overlaps, horizontally, the cell assigned to column 1 in the
same row.

Observed

A four-column ruled table, Datum | Erläuterung | Betrag Soll EUR | Betrag Haben EUR. Predicted shape 15x4 — correct.

Header row, showing where the columns actually are:

cell column x-range
'Datum' 0 70.9 – 100.8
'Erläuterung' 1 123.0 – 176.1
'Betrag Soll EUR' 2 394.2 – 464.0
'Betrag Haben EUR' 3 486.0 – 568.3

First data row:

cell column x-range
'07.01.2025 Basislastschrift CREDITOR NAME REDACTED S.C.A.' 0 69.9 – 265.2
'304-0000000-0000000 XXXXXXXXXXXXXX' 1 123.3 – 218.1

The column-0 cell runs to x=265.2, well past the column-1 header at x=123.0, and
overlaps the column-1 cell of its own row across almost its entire width. Only
'07.01.2025' belongs in column 0.

Where it originates

Not in the structure prediction, and not in the backend. docling-parse reports
the date and the first description word as one textline cell:

textline cell : '07.01.2025 Basislastschrift'   bbox x 69.9 .. 191.5

At word level they are separate — no word cell spans the two columns. The
merge is at line level, and it is metrically reasonable in isolation: the gap
between the date and the following word is about the width of a word space, so
no threshold on horizontal distance alone separates "word gap" from "column
gap" here.

All three PDF backends produce the identical fused cell, which is consistent
with the cause being the line merge rather than any backend's geometry:

backend tables words in table cells cells matching ^date\s+\S
threaded_docling_parse (default) 13 3 991 134
dlparse_v4 14 4 120 131
pypdfium2 14 4 452 133

The documented workaround makes it worse

do_cell_matching's own comment anticipates this case:

do_cell_matching: bool = (
    True
    # True:  Matches predictions back to PDF cells. Can break table output if PDF cells
    #        are merged across table columns.
    # False: Let table structure model define the text cells, ignore PDF cells.
)

Setting it to False on this document:

cells matching ^date\s+\S words in table cells
do_cell_matching=True 134 3 991
do_cell_matching=False 159 3 349 (−17.8%)

More fused cells, and a sixth of the text gone. On a second document (137 tables,
a technical export) it did reduce fusion 211 → 45, but lost 7% of the words and
emitted duplicated header cells. It is not a usable workaround in either
direction.

Reproducing without our file

The document is a personal bank statement and cannot be shared. The case should
reconstruct from:

  • a ruled table, four columns, header row Datum | Erläuterung | …
  • column 0 narrow (~30 pt of text), column 1 starting ~22 pt after column 0's
    text ends
  • each data row: a dd.mm.yyyy date in column 0 and text in column 1 on the
    same baseline
    , with the horizontal gap between the date and the first word of
    column 1 close to the font's space width
  • column 1's text wrapping onto further lines below, which are not fused —
    only the first line of each row is affected

do_ocr=False, TableFormerMode.ACCURATE.

Suggested fixes

  1. Split a matched PDF cell at a predicted column boundary when the boundary
    falls inside the cell's bounding box. The prediction is already correct here,
    and the inconsistency is detectable without it: a cell assigned to column n
    should not horizontally overlap a cell assigned to column n+1 in the same
    row.
  2. Prefer word cells over textline cells when matching inside a table region.
    The words are correctly separated already; only the line merge loses the
    boundary.
  3. Expose the merge thresholds. docling_parse's DecodeConfig has
    horizontal_cell_tolerance and word_space_width_factor_for_merge, but
    docling builds the DecodeConfig internally and plumbs only
    enforce_same_font
    (docling/backend/docling_parse_backend.py::_make_docling_parse_decode_config),
    so a caller cannot reach them from PdfPipelineOptions — nor from
    docling-serve, whose convert options expose only pdf_backend and
    table_cell_matching.

(1) or (2) would fix it; (3) would let callers whose corpus has this shape work
around it in the meantime.

Environment

docling-serve 1.31.0, docling 2.124.0, docling-core 2.93.0,
docling-ibm-models 4.0.1, docling-parse 7.16.0, image
ghcr.io/docling-project/docling-serve-cu130:main, Python 3.12, CUDA 13.

Scale

Measured across an archive of 6 599 PDFs (14 330 tables in 3 849 documents). A
narrow leading column of dates or reference numbers beside a wide description is
the most common table shape in it, so this affects a large share of the tables we
extract. In the 13-table statement above, up to 134 of 479 cells match the fused
pattern — an upper bound, since the pattern also matches cells that legitimately
begin with a date, such as a "Kontostand am 03.01.2025, Auszug Nr. 1" heading.

(Payment and mandate references above are redacted, with character counts
preserved so the quoted x-ranges remain consistent.)

Activity

  1. AbarnaaSree commented on Sep 9, 2026

    @AbarnaaSree

    Hi, I’d like to work on this issue. I can investigate the table cell matching logic and implement a fix to prevent matched cells from overlapping neighbouring columns, while preserving the existing correct structure prediction. I’ll also add regression tests covering the textline/word merge case described here and verify that the fix works across the relevant PDF backends. Please let me know if I can take this up.

  2. wittjeff commented on Sep 22, 2026

    @wittjeff
    Contributor

    I measured how often the self-inconsistency you describe — a cell whose box overlaps the
    next cell of its own row — occurs at docling 2.129.0 with default options, since it is
    detectable from the DoclingDocument alone.

    Rates (overlap ≥ 1 pt and ≥ 10% of the narrower cell; identical at 25% and 50% on
    the first corpus):

    corpus tables with an overlap rows with an overlap
    42 born-digital PDFs with a reference table (opendataloader-bench) 2 / 54 (3.7%) 3 / 372 (0.8%)
    131 page images with tables (PubTabNet, OmniDocBench, DocLayNet slices) 12 / 152 (7.9%) 39 / 1,991 (2.0%)

    TEDS-S does not see any of it, because the grid is right in every case.

    Mechanism, in the two born-digital hits. Since #1238 the matcher works on word-level
    cells, so the textline fusion you observed is not the only route. What the boxes show is
    the orphan-recovery fusion docling-ibm-models#188 describes:

    • 3-column table, a row whose column-0 value is the single word Events. The column-2
      cell came out as Events presentations, webinars, … with box x = 76–524, i.e. starting
      at column 0's x; column 0 is empty; the box overlaps the column-1 cell by 48 pt. The
      next row splits the two-line cell Physical or / digital by word into Physical digital
      (col 0) and or (col 1).
    • 4-column sheet, row Highlight: the column-3 cell's box is x = 55–910, the full table
      width — the row label from column 0 was fused into the last column's paragraph and the
      other three cells shifted one column left.

    A cheap guard. The overlap is computable inside MatchingPostProcessor before the
    table is emitted: refuse an orphan attachment whose resulting box would cross into a
    neighbouring cell of the same row (keep the word as its own cell, or drop it with the
    existing WARNING — never concatenate), and, second, re-split a matched cell whose box
    crosses a column boundary by the column bands the good-IoU cells define. The first alone
    clears both hits above. I have the measurement script and both documents; happy to
    prototype the guard in docling-ibm-models and report the delta with the same check, if
    that is a direction the maintainers would take.

  3. wittjeff commented on Sep 22, 2026

    @wittjeff
    Contributor

    @AbarnaaSree Sorry I got carried away chasing other issues. I'm interested in what you come up with here.

  4. wittjeff commented on Sep 22, 2026

    @wittjeff
    Contributor

    Correction to my comment above: the mechanism sentence is wrong for the two documents I described. Replaying MatchingPostProcessor.process offline on the dumped inputs (54 tables; output byte-identical to the live run) shows that orphan recovery adds nothing in either document — the fusion is complete after step 8, the final intersection-over-PDF-cell assignment, and it starts in TableFormer's own cell boxes:

    • 7×3 table: the column-2 box of five of the seven rows was predicted on top of column 0 (x ≈ 145–250, where the row labels sit). The label Events (x 152–214) intersects the misplaced column-2 box at 1.0 and the column-0 box at 0.74, so it goes to column 2. The real column-2 words (x ≥ 526) then match nothing, arrive as orphans into a column-2 band that now starts at x = 145, and step 8.a aligns the cell box to the union of its words (152–1048) — the overlap.
    • 4×4 table: the column-3 box of one row was predicted at x 100–212, on the row label Highlight; the row's other boxes are each one column too far left. The same rule then shifts the whole row.

    So for these two cases the grid is right, the boxes are misplaced, and the assignment trusts the boxes without checking them against the columns they claim. docling-ibm-models#188 (the orphan fallback) is a different route to the same symptom; the textline merge in this issue's own example is a third.

    @AbarnaaSree — you offered to take this on, so I'll leave the fix to you. If useful: the overlap check runs on the emitted DoclingDocument alone (a cell whose box crosses a same-row neighbour in another column, no declared span), and a replay harness that dumps MatchingPostProcessor.process inputs per table and re-runs the post-processor without the model made the tracing above cheap. One observation from that replay that may matter for a fix inside MatchingPostProcessor: _deduplicate_cells returns the same structural cell several times, so any statistic over table_cells there needs to dedupe by cell_id first. Happy to share the two documents and the harness.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions