Declined
4
[Feature Request] Preserve original file when pre-consume script modifies document (garbled PDF text detection)
Source: paperless-ngx/paperless-ngx#12344 · opened by @croneter
Description [Feature Request] Preserve original file when pre-consume script modifies document (garbled PDF text detection) Problem Description Many digitally-created PDFs — especially from European banks, insurance companies, and telecom providers — use CID-encoded fonts or custom font encodings. When Paperless-ngx ingests these PDFs in the default skip OCR mode, pdftotext cannot decode the text layer. The result: the "Content" tab shows only garbled characters (squares, symbols, gibberish), and the document is not searchable. <img width="457" height="670" alt="image" src=" /> The PDF *looks* fine visually (the rendered output is correct), but the underlying text data is garbage. This affects a significant portion of digitally-generated documents in the D/A/CH region and likely elsewhere. I do NOT want to force-OCR every PDF I ingest, but only the ones where I need to (no loss of e.g. data quality). Current …
No pledges yet. Be the first to back this.
Comments
Similar requests
Extend recoverable PDF handling to attachments
1 vote · 0 comments
[Feature Request] [GoBD] Optionally disable PDF editor (to protect original document)
7 votes · 0 comments
[Feature Request] OCR rendition on linked documents
2 votes · 0 comments
Change .eml file file type detection
1 vote · 0 comments
[Feature Request] Append printed E-mail and/or attach e-mail body to PDF attachments
8 votes · 0 comments
No comments yet.