Skip to content

Fix TypeError when combining the text table strategy with text_layout - #1403

Open
MohammadHijjawi97 wants to merge 1 commit into
jsvine:developfrom
MohammadHijjawi97:fix/table-text-strategy-text-settings
Open

MohammadHijjawi97 wants to merge 1 commit into
jsvine:developfrom
MohammadHijjawi97:fix/table-text-strategy-text-settings

Conversation

@MohammadHijjawi97

Copy link
Copy Markdown

Problem

The table-extraction docs say that all text_-prefixed settings are passed to text extraction, and that "All possible arguments to Page.extract_text(...) are also valid here." With the "text" strategy, though, some of them crash:

page.extract_tables({
    "vertical_strategy": "text",
    "horizontal_strategy": "text",
    "text_layout": True,
})
# TypeError: WordExtractor.__init__() got an unexpected keyword argument 'layout'

The same happens with other extract_text-only arguments (e.g. text_x_density, text_layout_width). With the default "lines" strategy, the same settings work fine.

Cause

For the "text" strategy, TableFinder.get_edges() calls self.page.extract_words(**settings.text_settings). extract_words passes everything to WordExtractor, which only accepts word-extraction arguments.

Fix

When finding words for the text strategy, only pass on the text_settings keys that WordExtractor accepts. These are listed in utils.text.WORD_EXTRACTOR_KWARGS, which chars_to_textmap and extract_text already use for the same purpose. Cell text extraction still receives all text_* settings, so text_layout etc. keep working there.

Tests

  • New test_text_layout_with_text_strategy in tests/test_table.py, using the existing senate-expenditures.pdf. It fails on develop with the TypeError above and passes with this change. It also checks that the layout-extracted cells, once stripped, match the non-layout ones.
  • tests/test_table.py passes.
  • black --check, isort --check-only, flake8 and mypy --strict --implicit-reexport all pass.

CHANGELOG updated.

With the "text" table strategy, TableFinder passed every text_* setting
to Page.extract_words(), so settings that only apply to text extraction
(e.g., text_layout=True, as documented for table settings) raised
"TypeError: WordExtractor.__init__() got an unexpected keyword argument".
Only pass on the settings WordExtractor accepts.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant