Describe the bug
DOCXToDocument silently drops all text inside block-level content controls (w:sdt). Word writes these for Insert > Cover Page (title, subtitle, abstract), the table of contents, and content controls in templates and forms. Their paragraphs and tables live in w:sdt/w:sdtContent, but _extract_elements only handles the direct w:p and w:tbl children of the body, so a w:sdt element is skipped as a whole, without any warning. For RAG pipelines this means that the content of such sections can never be retrieved.
_extract_elements iterates document.element.body and only handles tags ending in p or tbl:
https://github.com/deepset-ai/haystack/blob/11c2c96fe/haystack/components/converters/docx.py#L244-L272
Error message
None, the text is silently dropped.
Expected behavior
The paragraphs and tables inside a content control are converted like the ones directly in the body, in reading order.
To Reproduce
import io
import docx
from docx.oxml import parse_xml
from haystack.components.converters import DOCXToDocument
from haystack.dataclasses import ByteStream
W = 'xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"'
doc = docx.Document()
doc.add_paragraph("Paragraph before the content control.")
sdt = parse_xml(
f"""<w:sdt {W}>
<w:sdtPr><w:alias w:val="Summary"/></w:sdtPr>
<w:sdtContent>
<w:p><w:r><w:t>Text inside a content control.</w:t></w:r></w:p>
</w:sdtContent>
</w:sdt>"""
)
doc.element.body.insert(1, sdt)
doc.add_paragraph("Paragraph after the content control.")
buf = io.BytesIO()
doc.save(buf)
print(repr(DOCXToDocument().run(sources=[ByteStream(data=buf.getvalue())])["documents"][0].content))
# 'Paragraph before the content control.\nParagraph after the content control.'
Additional context
I have a fix with tests ready and will open a PR. It expands content controls (including nested ones) into the elements they hold before the existing paragraph and table handling.
System:
- OS: macOS 26.3 (arm64)
- Haystack version: main @ 11c2c96
Describe the bug
DOCXToDocumentsilently drops all text inside block-level content controls (w:sdt). Word writes these for Insert > Cover Page (title, subtitle, abstract), the table of contents, and content controls in templates and forms. Their paragraphs and tables live inw:sdt/w:sdtContent, but_extract_elementsonly handles the directw:pandw:tblchildren of the body, so aw:sdtelement is skipped as a whole, without any warning. For RAG pipelines this means that the content of such sections can never be retrieved._extract_elementsiteratesdocument.element.bodyand only handles tags ending inportbl:https://github.com/deepset-ai/haystack/blob/11c2c96fe/haystack/components/converters/docx.py#L244-L272
Error message
None, the text is silently dropped.
Expected behavior
The paragraphs and tables inside a content control are converted like the ones directly in the body, in reading order.
To Reproduce
Additional context
I have a fix with tests ready and will open a PR. It expands content controls (including nested ones) into the elements they hold before the existing paragraph and table handling.
System: