Skip to content

bug/html-text-inside-unrecognized-elements-dropped #4496

Description

@huuufu

Describe the bug

partition_html() silently drops the text inside any element the v1 HTML parser has no class for. Every unregistered tag gets a DefaultElement, and DefaultElement.iter_text_segments() (unstructured/partition/html/parser.py) emits only the element's tail, so the element's own text and everything nested in it are lost without a warning.

For <font> this is a regression: before the parser rewrite in #3218 (0.15.0), font was listed in TEXT_TAGS (git show c27e0d00^:unstructured/documents/html.py, line 34).

Content this loses in practice:

  • <font>, still common in email HTML (Outlook and Gmail reply headers), so .eml/.msg HTML bodies lose text;
  • Inline XBRL in SEC 10-K/10-Q filings: <ix:nonFraction> wraps the numbers inside sentences, and <ix:nonNumeric>/<ix:continuation> wrap whole notes, tables included;
  • custom elements (the <hauptteil_AB> case in bug/content within custom HTML tags is skipped #3708).

Everything routed through partition_html() is affected, including .md, .epub, .rst and .org.

To Reproduce

from unstructured.partition.html import partition_html

for html in [
    "<p><font face=Arial>Quarterly revenue</font> rose 12%.</p>",
    '<p>Revenue was $<ix:nonFraction name="us-gaap:Revenues">4,213</ix:nonFraction> million.</p>',
    "<my-card><h2>Pricing</h2><p>Plans start at $10.</p></my-card>",
]:
    print([str(e) for e in partition_html(text=html)])

Output on main (0ca5563):

['rose 12%.']
['Revenue was $ million.']
[]

On the repo's own fixtures, measured as the share of word tokens from a reference text (built from each fixture's HTML, with <head>, <script>, <style> and hidden elements removed) that appear in the output:

fixture elements token coverage
example-docs/example-10k.html 314 66.0%: the cover-page values and all of Notes 1–14 are missing
example-docs/eml/email-no-utf8-2008-07-16.062410.eml, HTML body 2 31.5%: the replies and their From:/Sent: headers are inside <font>

Expected behavior

The text of an unrecognized element is kept, the way a browser shows an unknown element (display: inline), with any block elements inside it partitioned as usual, as if the element were a <span>. Elements whose content isn't displayed as text, such as <select>, <svg> or the hidden Inline XBRL <ix:header>, should still be skipped.

Environment Info

  • unstructured: main at 0ca5563 (0.27.8)
  • Python 3.12.13, lxml 6.1.2 (libxml2 2.11.9)
  • Windows 11

Additional context

#3842 tracks this as a feature request ("custom HTML tag support"), with the open question of whether an unknown element should be treated as a block or as inline. The parser already turns block items nested in phrasing into their own elements, so an unknown element that is transparent phrasing handles both a few wrapped words and a wrapped section with headings and tables. I have a fix with tests and will open a PR that references this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions