Skip to content

Releases: Unstructured-IO/unstructured

0.16.11

10 Dec 00:51
b981d71
Compare
Choose a tag to compare

Enhancements

  • Enhance quote standardization tests with additional Unicode scenarios
  • Relax table segregation rule in chunking. Previously a Table element was always segregated into its own pre-chunk such that the Table appeared alone in a chunk or was split into multiple TableChunk elements, but never combined with Text-subtype elements. Allow table elements to be combined with other elements in the same chunk when space allows.
  • Compute chunk length based solely on element.text. Previously .metadata.text_as_html was also considered and since it is always longer that the text (due to HTML tag overhead) it was the effective length criterion. Remove text-as-html from the length calculation such that text-length is the sole criterion for sizing a chunk.

Features

Fixes

  • Fix ipv4 regex to correctly include up to three digit octets.

0.16.10

07 Dec 18:13
59e6cff
Compare
Choose a tag to compare

0.16.10

Enhancements

Features

Fixes

  • Fix original file doctype detection from cct converted file paths for metrics calculation.

0.16.9

02 Dec 21:52
0fb814d
Compare
Choose a tag to compare

What's Changed

Full Changelog: 0.16.8...0.16.9

0.16.8

26 Nov 19:39
0fe6ac6
Compare
Choose a tag to compare

0.16.8

Enhancements

  • Metrics: Weighted table average is optional

Features

Fixes

0.16.7

26 Nov 17:47
e48d79e
Compare
Choose a tag to compare

0.16.7

Enhancements

  • Add image_alt_mode to partition_html Adds an image_alt_mode parameter to partition_html() to control how alt text is extracted from images in HTML documents for html_parser_version=v2 . The parameter can be set to to_text to extract alt text as text from <img> html tags

Features

Fixes

0.16.6

22 Nov 02:09
626f73a
Compare
Choose a tag to compare

0.16.6

Enhancements

  • Every <table> tag is considered to be ontology.Table: Added special handling for tables in HTML partitioning. This change is made to improve the accuracy of table extraction from HTML documents.
  • Every HTML has default ontology class assigned: When parsing HTML to ontology, each defined HTML in the ontology has an assigned default ontology class. This allows assigning an ontology class instead of UncategorizedText when the HTML tag is predicted correctly but has no class assigned.
  • Use (number of actual table) weighted average for table metrics: In evaluating table metrics, the mean aggregation now uses the actual number of tables in a document to weight the metric scores.

Features

  • None added in this release.

Fixes

  • ElementMetadata consolidation: Now, text_as_html metadata is combined across all elements in CompositeElement when chunking HTML output.

0.16.5

07 Nov 20:32
a6aefee
Compare
Choose a tag to compare

What's Changed

Full Changelog: 0.16.4...0.16.5

0.16.4

31 Oct 19:00
df156eb
Compare
Choose a tag to compare

0.16.4

Enhancements

  • value attribute in <input/> element is parsed to OntologyElement.text in ontology
  • id and class attributes removed from Table subtags in HTML partitioning
  • cleaned to_html and newly introduced to_text in OntologyElement
  • Elements created from V2 HTML are less granular Added merging of adjacent text elements and inline html tags in the HTML partitioner to reduce the number of elements created from V2 HTML.

Features

  • Add support for link extraction in pdf hi_res strategy. The partition_pdf() function now supports link extraction when using the hi_res strategy, allowing users to extract hyperlinks from PDF documents more effectively.

Fixes

0.16.3

25 Oct 20:55
340a07f
Compare
Choose a tag to compare

0.16.3

Enhancements

Features

Fixes

  • V2 elements without first parent ID can be parsed
  • Fix missing elements when layout element parsed in V2 ontology
  • updated unstructured-inference to be 0.8.1 in requirements/extra-pdf-image.in

0.16.2

24 Oct 17:36
9835fe4
Compare
Choose a tag to compare

0.16.2

Enhancements

Features

  • Whitespace-invariant CCT distance metric. CCT Levenshtein distance for strings is by default computed with standardized whitespaces.

Fixes

  • Fixed retry config settings for partition_via_api function If the SDK's default retry config is not set the retry config getter function does not fail anymore.