Skip to content

Feature request: reject known CAD-exported PDFs when partial text extraction is unacceptable #2667

Description

CAD-exported PDFs can produce successful but incomplete Markdown because drawing geometry is not represented in the output. Can MarkItDown offer an explicit rejection policy for callers who cannot accept this loss?

Observed behavior

Tested current main at 4cc9fa17653d695d64fb9eee5b33d4de55ff84e8 with synthetic, valid vector PDFs (not files exported by an installed CAD application):

  • A floor plan with vector room outlines and embedded ROOM A / ROOM B labels converts successfully, returning the labels without the room outlines or spatial relationships.
  • A geometry-only variant also returns successfully, with effectively empty output.
  • Enabling the OCR plugin does not recover the labeled vector plan: it contains extractable text but no embedded raster images, so the plugin follows its text extraction path. This was tested locally without invoking an external model.

Autodesk also documents that SHX text may be exported as geometry, comments, or hidden text. This creates a second way for plain text extraction to omit content:
https://help.autodesk.com/cloudhelp/2024/ENU/AutoCAD-Core/files/GUID-56EA988C-A1DA-4E85-8765-B3F31A01AB02.htm

Proposed bounded fix

Add --reject-cad-pdfs and the conversion keyword reject_cad_pdfs=True. Check PDF Creator/Producer metadata for explicitly named CAD exporters before selecting a converter. Reject known matches with an actionable exception, including when plugins or cloud converters are enabled. Preserve existing extraction behavior when the policy is not requested.

This is an exporter-based policy, not a drawing classifier or a guarantee of conversion fidelity. Missing/rewritten metadata and unrecognized exporters cannot be detected by it. It may also reject text-heavy documents exported by those applications; the caller explicitly chooses that conservative policy. A line-count heuristic would misclassify tables and ordinary charts.

Full drawing interpretation would require a separate rendered-page/vision workflow and validation; OCR alone cannot guarantee dimensions, topology, or geometric relationships. I am proposing the small rejection feature rather than claiming faithful drawing conversion.

Regression coverage

Use generated PDFs with correct xref offsets to cover labeled and geometry-only drawings, Creator/Producer identification, unknown metadata, unchanged default extraction, non-PDF inputs, plugins, and CLI file/stdin rejection before writing output.

Related: #296 discusses general PDF fidelity. I found no CAD/AutoCAD/SHX-specific issue or PR in the repository searches.

Activity

  1. cagdasyurekli commented on Oct 7, 2026

    @cagdasyurekli
    ContributorAuthor

    Adam Fourney (@afourney) Gagan Bansal (@gagb) — could you advise on the scope questions below?

    Following up with seven public PDFs (56 pages): five direct AutoCAD, Revit, SOLIDWORKS, and MicroStation exports identify a CAD exporter in metadata, and the opt-in policy rejects all five. Two processed drawing documents identify PowerPoint or a scanner and evade detection. Successful text extraction omits geometry and does not establish dimension relationships; the existing OCR path also misses labeled vector drawings.

    PR #2668 keeps rejection opt-in and preserves existing conversion behavior. Exporter detection can also flag text-heavy CAD reports, so these samples do not establish that default rejection is reliable enough. Any default change and drawing-aware extraction would be separate contributions.

    Could you advise on three questions?

    1. Should rejection remain opt-in, or should a future default change be explored with detection evaluation and a deprecation period?
    2. Is an optional drawing-aware mode (located labels, SHX text, and page views, without semantic dimension reconstruction) wanted in core, or better suited to a plugin?
    3. Would an optional Markdown sidecar describing conversion limitations be acceptable?

    Before proposing inline images, I would measure base64 Markdown size and usability on the 23-page AutoCAD export and dense Revit page. None of this would be described as lossless conversion. I will wait for your guidance before extending the implementation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions