FeatureFuel
4

[Feature Request] Preserve original file when pre-consume script modifies document (garbled PDF text detection)

Source: paperless-ngx/paperless-ngx#12344 · opened by @croneter
Description [Feature Request] Preserve original file when pre-consume script modifies document (garbled PDF text detection) Problem Description Many digitally-created PDFs — especially from European banks, insurance companies, and telecom providers — use CID-encoded fonts or custom font encodings. When Paperless-ngx ingests these PDFs in the default skip OCR mode, pdftotext cannot decode the text layer. The result: the "Content" tab shows only garbled characters (squares, symbols, gibberish), and the document is not searchable. <img width="457" height="670" alt="image" src=" /> The PDF *looks* fine visually (the rendered output is correct), but the underlying text data is garbage. This affects a significant portion of digitally-generated documents in the D/A/CH region and likely elsewhere. I do NOT want to force-OCR every PDF I ingest, but only the ones where I need to (no loss of e.g. data quality). Current …

No pledges yet. Be the first to back this.

Make a pledge

Pledge your monetary support if this feature is added.

$

Comments

No comments yet.

Replying to

Add a comment

What do you think about this feature request?


Similar requests