Skip to content

webvtt: text after a multi-line span is attached to the wrong cue line #4104

Description

@businessarshgoyal

Bug

In WebVTTDocumentBackend.convert, _extract_components caches the current paragraph in a local variable (par = cue_text[-1]) before iterating over a cue payload. When a nested span (voice, bold, italic, underline) wraps a line terminator, the recursive call starts a new paragraph in cue_text, but the caller keeps writing into its stale par. Everything that follows the span on the next line is therefore appended to the first line's paragraph, so the cue text comes out in the wrong order.

Steps to reproduce

sample.vtt:

WEBVTT

00:00:01.000 --> 00:00:05.000
<v Bob>Hello
there</v> and afterwards
from docling.document_converter import DocumentConverter

doc = DocumentConverter().convert("sample.vtt").document
print([t.text for t in doc.texts])

Actual

['Hello', ' and afterwards', 'there']

export_to_markdown() renders Hello and afterwards first and there after it.

Expected

['Hello', 'there', ' and afterwards']

and afterwards sits on the same cue line as there, so it belongs to that line's paragraph.

Docling version

main (2.x), Python 3.12

Activity

  1. dosubot commented on Aug 29, 2026

    @dosubot

    Hi @businessarshgoyal! I'm Dosu and I'm helping the IBM Docling team.

    I can confirm this bug in the current main code. In _extract_components, par is bound to cue_text[-1] once at the top of the call. When a span like <v Bob>Hello\nthere</v> contains a line terminator, the recursive call into _extract_components(component.internal_text.components) appends a new AnnotatedPar to the shared cue_text list and starts writing into it. But when the recursive call returns, the outer loop's par still references the original paragraph. Any sibling component processed afterward (like the plain-text " and afterwards") therefore gets appended to the stale first paragraph instead of the current one — exactly the misordering you described.

    The existing tests don't cover this scenario: test_multi_line_cue_text_preserved only uses plain-text multi-line cues, and test_cue_ids_plus_multiline_with_voice_and_style has voice/style spans but no line break inside a span's payload. So your repro is a legitimate gap.

    The fix should be straightforward — re-fetch par = cue_text[-1] after the recursive _extract_components call returns (line 155), so the outer scope always writes into whatever paragraph is currently last in cue_text. Something like:

                        parents.append(item)
                        _extract_components(component.internal_text.components)
                        parents.pop()
    +                   par = cue_text[-1]  # the recursive call may have started a new paragraph

    I don't see an existing PR addressing this specific ordering bug yet.

    To reply, just mention @dosu.


    Share context across your team and agents. Try Dosu.

    Leave Feedback Learn about docling with Dosu Add Dosu to your team

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions