1
[Feature Request] Re-run OCR automatically when the extracted text is poor
Source: paperless-ngx/paperless-ngx#14320 · opened by @Ollornog
Description Since #13633, remote OCR can be used selectively per document. What's missing is the trigger: deciding automatically *which* documents need it. Suggestion: after consumption, check the text quality, and if it's poor, run OCR again (with remote OCR, or locally with force). Typical cases: PDFs with a broken or garbage text layer, phone photos, faxes. 1. Rule-based check (cheap, no AI): text too short, too few real words, too many odd characters. 2. AI check (optional, when AI is enabled): the AI already reads the text for suggestions; let it flag "this text is unreadable" and re-run OCR before the suggestions are applied. As a workflow trigger/condition this would fit well: "if text quality is poor → re-OCR with engine X". Related: #13389 (selective remote OCR, solved by #13633), #13264, #12029. Other <details> <summary>Details and suggestions</summary> Quality rules that work well in practice (all configurable): •…
No pledges yet. Be the first to back this.
Comments
Similar requests
[Feature Request] Support a self-hostable remote OCR service
35 votes · 0 comments
[Feature] Optional removal of blank pages from the archive file (originals untouched)
1 vote · 0 comments
[Feature Request] OCR: Extract specific data from zones into Fields [PR incoming]
12 votes · 0 comments
[Feature Request] Selective / Conditional Remote OCR Routing (Azure OCR)
5 votes · 0 comments
[Feature Request] Make Remote OCR respect `PAPERLESS_OCR_MODE` and local preprocessing options
5 votes · 0 comments
No comments yet.