Repository navigation
Add recrop_figures.py: re-crop vector figures from saved GROBID TEI - #4
Open
LarsSchimmelpfennig wants to merge 1 commit into
Open
LarsSchimmelpfennig wants to merge 1 commit into
LarsSchimmelpfennig wants to merge 1 commit into
Conversation
GROBID only boxes embedded bitmaps, so vector figures fall back to the <figure> box and come out as caption strips (or no PNG). The new standalone step re-crops from 02_grobid/<stem>.tei.xml.gz + the source PDF into a separate --out dir, finding each figure's drawings next to its caption in the PDF, and keeps GROBID's <figure> box when that search fails unless the box is mostly text. Pipeline steps and existing outputs are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds an optional, standalone step that re-crops figure images so vector figures come out as figures instead of caption strips.
Problem.
pdf_to_bioc.save_figure_images()crops from GROBID's<graphic>box, else the<figure>box. GROBID reports<graphic>boxes only for embedded bitmaps, and for vector figures (survival curves, bar charts, forest plots) the<figure>box covers the caption, not the drawing. So those PNGs are caption strips, or there is no PNG when GROBID gives no box at all. Example, PMID 10088816 Figure 1: the current crop is the three-line caption; the new crop is the scatter plot.Changes
src/pipeline_steps/recrop_figures.py(new): re-crops from the saved02_grobid/<stem>.tei.xml.gz+ source PDF, writing only to a separate--outdir. It needs no GROBID server and never touches the pub dirs. Per figure:<graphic>box: cropped as before, widened to touching vector panels and labels (axis titles, panel letters).<graphic>box: finds the caption in the PDF and crops the vector drawings/images in the zone next to it, bounded by body text, other captions and tables (source: "layout").<figure>box, as before (source: "figure_fallback"), unless it is mostly text (text blocks ≥ 35% of the box and graphics < 5%). Then no image (source: "none").pdf_to_bioc.py. Output keeps thefigures.jsonschema, plusgrobid_box,reasonandcaption_bbox, alongsiderecrop_summary.tsvand an optionalreview.html.README.md/CLAUDE.md: document it.pdf_to_bioc.py,civic_pubtator.py,backfill_figures.pyand the existing02_grobid/figures/outputs are unchanged.Run
Options:
--include-supplementaryalso re-crops02_grobid/s/....--reviewwrites<out>/review.htmlwith old and new crops side by side.--workers Nsets the number of documents processed in parallel.--forceredoes documents already done by the current cropper version (recrop_figures/2).--outmust be outside the input dirs.Needs PyMuPDF ≥ 1.24.2 for
page.cluster_drawings(it falls back to a naive merge on older versions); tested with 1.27.2.Evaluation
100 papers, labelled by hand (98 with saved TEI): 274 figures. 50 GROBID crops were unusable (30 papers), all among the 55 figures without a
<graphic>box. Recrop fixed or recovered 38 of them (21 papers). The 219<graphic>figures were left alone (208 identical, 11 widened to include axis labels). Before the<figure>-box fallback there were 2 regressions: PMID 15256671 Figure 2, now fixed by the fallback, and one junk scanned-page item.Full test corpus, unlabelled (588 PDFs with saved TEI, 2,369 figures in 564 papers, 0 failures):
<graphic>box (1,375): 1,127 identical, 248 widened (median 1.22× area), 0 lost. The 42 widened > 2× are almost all multi-panel figures where GROBID had boxed one bitmap panel; 4 are label-less journal cover-art images that pull in title/abstract text.<graphic>box (994):--include-supplementary: 1,957 more documents, 4,182 figures, 0 failures. 67 converted.xlsx/.docxsupplements, whose prepared PDFcivic_pubtator.pydeleted after the run, are reported as "no PDF" and skipped. A missing main PDF is still an error.Full corpus: old GROBID box → new source (reason)
The fallback's graphics threshold is 5% rather than 15%: at 15%, 15256671 Figure 2 is still rejected, because its panels are 34 small blot images covering 10% of the box and its caption wraps around them. Lowering the threshold only adds crops; every candidate below 5% graphics with ≥ 35% text was junk.
Known gaps
<figure>("Figure 2 .Figure 3"), only the first figure is cropped.<graphic>box can pull in text that sits on a coloured background (seen only on decorative cover art).xml:ids, which makes file names and the "old" column ofreview.htmlambiguous there (as inpdf_to_bioc.py).--outshort (260-character path limit).Follow-up (maintainers' call): an orchestrator flag to run this after GROBID, or replacing
save_figure_imageswith it.🤖 Generated with Claude Code