1
Allow AI classification strategy to be selected independently of embeddings
Source: paperless-ngx/paperless-ngx#14358 · opened by @GregDog
Feature request Allow the AI metadata classification strategy to be selected independently of whether embeddings / the LLM index are enabled. This is not a request to disable embeddings, RAG, similar-document retrieval, or document chat. Those remain useful. The request is that turning on the vector index should not force classification to use neighbour-derived taxonomy candidates and neighbour document text. Problem On current paperless_ai, classification is retrieval-assisted: 1. Retrieve similar document chunks (TAXONOMY_CANDIDATE_TOP_K = 15 in ai_classifier.py). 2. Aggregate tag / type / correspondent / storage-path weights from those neighbours (taxonomy.py). 3. Cap tags at MAX_TAG_CANDIDATES = 10. 4. Include RAG text from up to 5 similar documents. 5. Ask the LLM to classify. 6. Accept existing IDs only if they appear in the candidate allowlist (allowed_candidate_ids / model_to_classification_suggestions). This is a reasonable default, especially for very large taxonomies.…
No pledges yet. Be the first to back this.
Comments
Similar requests
[Feature Request] Allow overriding the embeddings encoding_format for OpenAI-compatible LLM providers
4 votes · 0 comments
[Feature Request] Different keys for LLM vs. LLM embeddings
6 votes · 0 comments
AI provenance tracking for AI-classified documents
1 vote · 0 comments
[Feature Request] Add grouping options to the document view
1 vote · 0 comments
[Feature Request] AI Suggestions Cache
6 votes · 0 comments
No comments yet.