1
Make the Azure Document Intelligence model configurable (prebuilt-layout in addition to prebuilt-read)
Source: paperless-ngx/paperless-ngx#14299 · opened by @isi4d
Description Description The remote OCR parser currently calls the Azure Document Intelligence prebuilt-read model. I would like to request an optional setting to select the model, e.g. PAPERLESS_REMOTE_OCR_MODEL, defaulting to prebuilt-read so nothing changes for existing users. Why 1. Table structure is lost. With prebuilt-read, tabular documents come back as a flat list of cell values, one per line, with no relationship between them. A line-item table from a repair quote arrives as: description, part number, quantity, unit price — each on its own line, in reading order. The text is searchable, but the structure needed for any later extraction (totals, amounts, quantities) is gone. prebuilt-layout returns tables as tables. 2. Layout appears to recognise characters better, not just structure. Microsoft's own changelog states that the Layout model received "improvements to the OCR model for scanned text targeting improvements for single characters, boxed text, an…
No pledges yet. Be the first to back this.
Comments
Similar requests
[Feature Request] Self-hosted MinerU as a remote OCR engine
4 votes · 0 comments
[Feature Request] Selective / Conditional Remote OCR Routing (Azure OCR)
5 votes · 0 comments
[Feature Request] AI chat/RAG: cache the embedding model instead of reloading it on every query
1 vote · 0 comments
[Feature Request] "Document Deleted" / "Document Trashed" Workflow Trigger
6 votes · 0 comments
[Feature Request] Chat: show the scope of each question, and optionally show the model's reasoning
1 vote · 0 comments
No comments yet.