FeatureFuel
1

Allow AI classification strategy to be selected independently of embeddings

Source: paperless-ngx/paperless-ngx#14358 · opened by @GregDog
Feature request Allow the AI metadata classification strategy to be selected independently of whether embeddings / the LLM index are enabled. This is not a request to disable embeddings, RAG, similar-document retrieval, or document chat. Those remain useful. The request is that turning on the vector index should not force classification to use neighbour-derived taxonomy candidates and neighbour document text. Problem On current paperless_ai, classification is retrieval-assisted: 1. Retrieve similar document chunks (TAXONOMY_CANDIDATE_TOP_K = 15 in ai_classifier.py). 2. Aggregate tag / type / correspondent / storage-path weights from those neighbours (taxonomy.py). 3. Cap tags at MAX_TAG_CANDIDATES = 10. 4. Include RAG text from up to 5 similar documents. 5. Ask the LLM to classify. 6. Accept existing IDs only if they appear in the candidate allowlist (allowed_candidate_ids / model_to_classification_suggestions). This is a reasonable default, especially for very large taxonomies.…

No pledges yet. Be the first to back this.

Make a pledge

Pledge your monetary support if this feature is added.

$

Comments

No comments yet.

Replying to

Add a comment

What do you think about this feature request?


Similar requests