Bug
In WebVTTDocumentBackend.convert, _extract_components caches the current paragraph in a local variable (par = cue_text[-1]) before iterating over a cue payload. When a nested span (voice, bold, italic, underline) wraps a line terminator, the recursive call starts a new paragraph in cue_text, but the caller keeps writing into its stale par. Everything that follows the span on the next line is therefore appended to the first line's paragraph, so the cue text comes out in the wrong order.
Steps to reproduce
sample.vtt:
WEBVTT
00:00:01.000 --> 00:00:05.000
<v Bob>Hello
there</v> and afterwards
from docling.document_converter import DocumentConverter
doc = DocumentConverter().convert("sample.vtt").document
print([t.text for t in doc.texts])
Actual
['Hello', ' and afterwards', 'there']
export_to_markdown() renders Hello and afterwards first and there after it.
Expected
['Hello', 'there', ' and afterwards']
and afterwards sits on the same cue line as there, so it belongs to that line's paragraph.
Docling version
main (2.x), Python 3.12
Bug
In
WebVTTDocumentBackend.convert,_extract_componentscaches the current paragraph in a local variable (par = cue_text[-1]) before iterating over a cue payload. When a nested span (voice, bold, italic, underline) wraps a line terminator, the recursive call starts a new paragraph incue_text, but the caller keeps writing into its stalepar. Everything that follows the span on the next line is therefore appended to the first line's paragraph, so the cue text comes out in the wrong order.Steps to reproduce
sample.vtt:Actual
export_to_markdown()rendersHello and afterwardsfirst andthereafter it.Expected
and afterwardssits on the same cue line asthere, so it belongs to that line's paragraph.Docling version
main (2.x), Python 3.12