2
[Feature Request] OCR rendition on linked documents
Source: paperless-ngx/paperless-ngx#13524 · opened by @R3dst4r
Description Imagine you have two documents: letter.docx and letter.pdf. The .docx-file unoffically represents the single source of truth whereas letter.pdf is either the exported version or - even worse - a scan of the printed .docx-file. Not to mention that this is likely to be modified, too. Now, running OCR on the .docx will create a highly accurate result. I think we should be able to reuse this precise content to apply it on our .pdf file. Using this method, we will drastically increase the quality of the PDF's OCR with almost no effort. A downside of this could be that my modified pdf-version is not 100% congruent with the docx. Maybe someone filled in PDF form fields or performed other minor changes to the printed and rescanned document (add inbox stamp, sign document, and so on). However, the rendition still can be used to verify most of the other OCR result, too and point to diffs as an additional step in OCR quality assurance. Adding a dialogue in /documents/$i…
No pledges yet. Be the first to back this.
Comments
Similar requests
[Feature] Optional removal of blank pages from the archive file (originals untouched)
1 vote · 0 comments
[Feature Request] Support a self-hostable remote OCR service
30 votes · 0 comments
[Feature Request] OCR: Extract specific data from zones into Fields [PR incoming]
12 votes · 0 comments
[Feature Request] Make Remote OCR respect `PAPERLESS_OCR_MODE` and local preprocessing options
5 votes · 0 comments
[Feature Request] Add alternative LLM based OCR
25 votes · 0 comments
No comments yet.