Skip to content

Add recrop_figures.py: re-crop vector figures from saved GROBID TEI - #4

Open
LarsSchimmelpfennig wants to merge 1 commit into
griffithlab:mainfrom
LarsSchimmelpfennig:recrop-figures
Open

LarsSchimmelpfennig wants to merge 1 commit into
griffithlab:mainfrom
LarsSchimmelpfennig:recrop-figures

Conversation

@LarsSchimmelpfennig

Copy link
Copy Markdown

Adds an optional, standalone step that re-crops figure images so vector figures come out as figures instead of caption strips.

Problem. pdf_to_bioc.save_figure_images() crops from GROBID's <graphic> box, else the <figure> box. GROBID reports <graphic> boxes only for embedded bitmaps, and for vector figures (survival curves, bar charts, forest plots) the <figure> box covers the caption, not the drawing. So those PNGs are caption strips, or there is no PNG when GROBID gives no box at all. Example, PMID 10088816 Figure 1: the current crop is the three-line caption; the new crop is the scatter plot.

Changes

  • src/pipeline_steps/recrop_figures.py (new): re-crops from the saved 02_grobid/<stem>.tei.xml.gz + source PDF, writing only to a separate --out dir. It needs no GROBID server and never touches the pub dirs. Per figure:
    • <graphic> box: cropped as before, widened to touching vector panels and labels (axis titles, panel letters).
    • No <graphic> box: finds the caption in the PDF and crops the vector drawings/images in the zone next to it, bounded by body text, other captions and tables (source: "layout").
    • Layout search fails: keeps GROBID's <figure> box, as before (source: "figure_fallback"), unless it is mostly text (text blocks ≥ 35% of the box and graphics < 5%). Then no image (source: "none").
    • Tables are cropped exactly as in pdf_to_bioc.py. Output keeps the figures.json schema, plus grobid_box, reason and caption_bbox, alongside recrop_summary.tsv and an optional review.html.
  • README.md / CLAUDE.md: document it.

pdf_to_bioc.py, civic_pubtator.py, backfill_figures.py and the existing 02_grobid/figures/ outputs are unchanged.

Run

python3 src/pipeline_steps/recrop_figures.py /data/pub-data --out /data/recrop --review

Options:

  • --include-supplementary also re-crops 02_grobid/s/....
  • --review writes <out>/review.html with old and new crops side by side.
  • --workers N sets the number of documents processed in parallel.
  • --force redoes documents already done by the current cropper version (recrop_figures/2).
  • --out must be outside the input dirs.

Needs PyMuPDF ≥ 1.24.2 for page.cluster_drawings (it falls back to a naive merge on older versions); tested with 1.27.2.

Evaluation

100 papers, labelled by hand (98 with saved TEI): 274 figures. 50 GROBID crops were unusable (30 papers), all among the 55 figures without a <graphic> box. Recrop fixed or recovered 38 of them (21 papers). The 219 <graphic> figures were left alone (208 identical, 11 widened to include axis labels). Before the <figure>-box fallback there were 2 regressions: PMID 15256671 Figure 2, now fixed by the fallback, and one junk scanned-page item.

Full test corpus, unlabelled (588 PDFs with saved TEI, 2,369 figures in 564 papers, 0 failures):

  • <graphic> box (1,375): 1,127 identical, 248 widened (median 1.22× area), 0 lost. The 42 widened > 2× are almost all multi-panel figures where GROBID had boxed one bitmap panel; 4 are label-less journal cover-art images that pull in title/abstract text.
  • No <graphic> box (994):
    • 396 old crops replaced by a layout crop; a random sample of 36 had ~34 clean figure crops.
    • 116 figures with no old image now have one (82 papers).
    • 143 keep GROBID's crop via the fallback.
    • 262 old crops dropped as mostly text. On visual review of all 262: 2 were usable figures (both on OCR'd scanned pages: PMID 8730290 Fig 1, 9830390 Fig 1) and ~3 were whole text pages containing a small figure; the rest were caption strips, paragraphs, references, author lists and tables.
    • 77 had no image before or after.
  • With --include-supplementary: 1,957 more documents, 4,182 figures, 0 failures. 67 converted .xlsx/.docx supplements, whose prepared PDF civic_pubtator.py deleted after the run, are reported as "no PDF" and skipped. A missing main PDF is still an error.
Full corpus: old GROBID box → new source (reason)
  figure   -> figure_fallback fallback_caption_not_found               56
  figure   -> figure_fallback fallback_no_graphics_near_caption        68
  figure   -> figure_fallback fallback_same_region_as_fig_0             2
  figure   -> figure_fallback fallback_same_region_as_fig_2             6
  figure   -> figure_fallback fallback_same_region_as_fig_3             1
  figure   -> figure_fallback fallback_same_region_as_fig_4             3
  figure   -> figure_fallback fallback_same_region_as_fig_5             1
  figure   -> figure_fallback fallback_same_region_as_fig_6             3
  figure   -> figure_fallback fallback_same_region_as_fig_7             1
  figure   -> figure_fallback fallback_scanned_page                     2
  figure   -> layout          layout_above                            379
  figure   -> layout          layout_below                             21
  figure   -> layout          layout_facing_page                        4
  figure   -> layout          layout_frame                             40
  figure   -> layout          layout_side                              68
  figure   -> none            caption_not_found                        46
  figure   -> none            no_graphics_near_caption                238
  figure   -> none            same_region_as_fig_0                      6
  figure   -> none            same_region_as_fig_1                      4
  figure   -> none            same_region_as_fig_2                      7
  figure   -> none            same_region_as_fig_3                      6
  figure   -> none            same_region_as_fig_5                      3
  figure   -> none            same_region_as_fig_7                      2
  figure   -> none            same_region_as_fig_9                      1
  figure   -> none            scanned_page                             26
  graphic  -> graphic         graphic                                1127
  graphic  -> graphic         graphic_expanded                        248

The fallback's graphics threshold is 5% rather than 15%: at 15%, 15256671 Figure 2 is still rejected, because its panels are 34 small blot images covering 10% of the box and its caption wraps around them. Lowering the threshold only adds crops; every candidate below 5% graphics with ≥ 35% text was junk.

Known gaps

  • When GROBID merges two captions into one <figure> ("Figure 2 .Figure 3"), only the first figure is cropped.
  • Scanned PDFs have no vector layer to search. If the page has an OCR text layer, GROBID's caption-strip boxes are dropped as text.
  • Widening a <graphic> box can pull in text that sits on a coloured background (seen only on decorative cover art).
  • 11 papers have duplicate TEI xml:ids, which makes file names and the "old" column of review.html ambiguous there (as in pdf_to_bioc.py).
  • On Windows, keep --out short (260-character path limit).

Follow-up (maintainers' call): an orchestrator flag to run this after GROBID, or replacing save_figure_images with it.

🤖 Generated with Claude Code

GROBID only boxes embedded bitmaps, so vector figures fall back to the
<figure> box and come out as caption strips (or no PNG). The new standalone
step re-crops from 02_grobid/<stem>.tei.xml.gz + the source PDF into a
separate --out dir, finding each figure's drawings next to its caption in
the PDF, and keeps GROBID's <figure> box when that search fails unless the
box is mostly text. Pipeline steps and existing outputs are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant