CAD-exported PDFs can produce successful but incomplete Markdown because drawing geometry is not represented in the output. Can MarkItDown offer an explicit rejection policy for callers who cannot accept this loss?
Observed behavior
Tested current main at 4cc9fa17653d695d64fb9eee5b33d4de55ff84e8 with synthetic, valid vector PDFs (not files exported by an installed CAD application):
- A floor plan with vector room outlines and embedded
ROOM A / ROOM B labels converts successfully, returning the labels without the room outlines or spatial relationships.
- A geometry-only variant also returns successfully, with effectively empty output.
- Enabling the OCR plugin does not recover the labeled vector plan: it contains extractable text but no embedded raster images, so the plugin follows its text extraction path. This was tested locally without invoking an external model.
Autodesk also documents that SHX text may be exported as geometry, comments, or hidden text. This creates a second way for plain text extraction to omit content:
https://help.autodesk.com/cloudhelp/2024/ENU/AutoCAD-Core/files/GUID-56EA988C-A1DA-4E85-8765-B3F31A01AB02.htm
Proposed bounded fix
Add --reject-cad-pdfs and the conversion keyword reject_cad_pdfs=True. Check PDF Creator/Producer metadata for explicitly named CAD exporters before selecting a converter. Reject known matches with an actionable exception, including when plugins or cloud converters are enabled. Preserve existing extraction behavior when the policy is not requested.
This is an exporter-based policy, not a drawing classifier or a guarantee of conversion fidelity. Missing/rewritten metadata and unrecognized exporters cannot be detected by it. It may also reject text-heavy documents exported by those applications; the caller explicitly chooses that conservative policy. A line-count heuristic would misclassify tables and ordinary charts.
Full drawing interpretation would require a separate rendered-page/vision workflow and validation; OCR alone cannot guarantee dimensions, topology, or geometric relationships. I am proposing the small rejection feature rather than claiming faithful drawing conversion.
Regression coverage
Use generated PDFs with correct xref offsets to cover labeled and geometry-only drawings, Creator/Producer identification, unknown metadata, unchanged default extraction, non-PDF inputs, plugins, and CLI file/stdin rejection before writing output.
Related: #296 discusses general PDF fidelity. I found no CAD/AutoCAD/SHX-specific issue or PR in the repository searches.
CAD-exported PDFs can produce successful but incomplete Markdown because drawing geometry is not represented in the output. Can MarkItDown offer an explicit rejection policy for callers who cannot accept this loss?
Observed behavior
Tested current
mainat4cc9fa17653d695d64fb9eee5b33d4de55ff84e8with synthetic, valid vector PDFs (not files exported by an installed CAD application):ROOM A/ROOM Blabels converts successfully, returning the labels without the room outlines or spatial relationships.Autodesk also documents that SHX text may be exported as geometry, comments, or hidden text. This creates a second way for plain text extraction to omit content:
https://help.autodesk.com/cloudhelp/2024/ENU/AutoCAD-Core/files/GUID-56EA988C-A1DA-4E85-8765-B3F31A01AB02.htm
Proposed bounded fix
Add
--reject-cad-pdfsand the conversion keywordreject_cad_pdfs=True. Check PDF Creator/Producer metadata for explicitly named CAD exporters before selecting a converter. Reject known matches with an actionable exception, including when plugins or cloud converters are enabled. Preserve existing extraction behavior when the policy is not requested.This is an exporter-based policy, not a drawing classifier or a guarantee of conversion fidelity. Missing/rewritten metadata and unrecognized exporters cannot be detected by it. It may also reject text-heavy documents exported by those applications; the caller explicitly chooses that conservative policy. A line-count heuristic would misclassify tables and ordinary charts.
Full drawing interpretation would require a separate rendered-page/vision workflow and validation; OCR alone cannot guarantee dimensions, topology, or geometric relationships. I am proposing the small rejection feature rather than claiming faithful drawing conversion.
Regression coverage
Use generated PDFs with correct xref offsets to cover labeled and geometry-only drawings, Creator/Producer identification, unknown metadata, unchanged default extraction, non-PDF inputs, plugins, and CLI file/stdin rejection before writing output.
Related: #296 discusses general PDF fidelity. I found no CAD/AutoCAD/SHX-specific issue or PR in the repository searches.