refactor: batch Flash PDF page extraction - #40
Merged
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Flash opens each page through nine separate accessors and decodes visible paths again for drawing-line and path metadata extraction.
Modification
Add a private Python-only page snapshot collected inside one PDFium page lifecycle. Share decoded subpaths while isolating failures in each derived result. Retain independent accessors and move pdftext list/PageChars conversion into one adapter module.
This is stage 3/6 of the Flash PDF equivalence refactor, based on
codex/flash-pdf-02-hotspots. Review and integrate stages in order.Validation
31 documents / 298 pages remain completely equivalent. All 49 PDFDocument tests pass, including rotation/CropBox, Form, images, signatures, links, single-open cleanup and shared path decoding. Representative medians improve further over the preceding performance stage.
The original
test_demo_sparse_table_confidence_manifestbbox expectation failure is documented separately and its expected data is not modified.Compatibility
The existing
PdfModel.predict(), PDFDocument methods, ModelJson/MiddleJson output and renderer contracts remain unchanged. Recognition thresholds, candidate priority and fallback behavior are preserved.Checklist
git diff --checkpass.