Releases: Unstructured-IO/unstructured
Releases · Unstructured-IO/unstructured
0.16.11
Enhancements
- Enhance quote standardization tests with additional Unicode scenarios
- Relax table segregation rule in chunking. Previously a
Table
element was always segregated into its own pre-chunk such that theTable
appeared alone in a chunk or was split into multipleTableChunk
elements, but never combined withText
-subtype elements. Allow table elements to be combined with other elements in the same chunk when space allows. - Compute chunk length based solely on
element.text
. Previously.metadata.text_as_html
was also considered and since it is always longer that the text (due to HTML tag overhead) it was the effective length criterion. Remove text-as-html from the length calculation such that text-length is the sole criterion for sizing a chunk.
Features
Fixes
- Fix ipv4 regex to correctly include up to three digit octets.
0.16.10
0.16.9
What's Changed
- chore: fix CHANGELOG formatting by @cragwolfe in #3800
- Use native ntlk download by @vangheem in #3796
Full Changelog: 0.16.8...0.16.9
0.16.8
0.16.7
0.16.7
Enhancements
- Add image_alt_mode to partition_html Adds an
image_alt_mode
parameter topartition_html()
to control how alt text is extracted from images in HTML documents forhtml_parser_version=v2
. The parameter can be set toto_text
to extract alt text as text from<img>
html tags
Features
Fixes
0.16.6
0.16.6
Enhancements
- Every
<table>
tag is considered to be ontology.Table: Added special handling for tables in HTML partitioning. This change is made to improve the accuracy of table extraction from HTML documents. - Every HTML has default ontology class assigned: When parsing HTML to ontology, each defined HTML in the ontology has an assigned default ontology class. This allows assigning an ontology class instead of
UncategorizedText
when the HTML tag is predicted correctly but has no class assigned. - Use (number of actual table) weighted average for table metrics: In evaluating table metrics, the mean aggregation now uses the actual number of tables in a document to weight the metric scores.
Features
- None added in this release.
Fixes
- ElementMetadata consolidation: Now,
text_as_html
metadata is combined across all elements inCompositeElement
when chunking HTML output.
0.16.5
What's Changed
- Add max recursion limit and fix to_text() method by @plutasnyy in #3773
- Fix extracting value from field by @plutasnyy in #3774
- chore: remove dev and release as 0.16.5 by @badGarnet in #3775
Full Changelog: 0.16.4...0.16.5
0.16.4
0.16.4
Enhancements
value
attribute in<input/>
element is parsed toOntologyElement.text
in ontologyid
andclass
attributes removed from Table subtags in HTML partitioning- cleaned
to_html
and newly introducedto_text
inOntologyElement
- Elements created from V2 HTML are less granular Added merging of adjacent text elements and inline html tags in the HTML partitioner to reduce the number of elements created from V2 HTML.
Features
- Add support for link extraction in pdf hi_res strategy. The
partition_pdf()
function now supports link extraction when using thehi_res
strategy, allowing users to extract hyperlinks from PDF documents more effectively.
Fixes
0.16.3
0.16.2
0.16.2
Enhancements
Features
- Whitespace-invariant CCT distance metric. CCT Levenshtein distance for strings is by default computed with standardized whitespaces.
Fixes
- Fixed retry config settings for partition_via_api function If the SDK's default retry config is not set the retry config getter function does not fail anymore.