Whether a chatbot answers accurately comes down to the knowledge base about 80% of the time, not the prompt.
The trouble is that the settings all sound abstract — chunk length, indexing method, Top-K, score threshold. This article explains what each one does and when you should touch it.
30-second overview
| Setting | What it decides | What to use first time |
|---|---|---|
| Chunk size | How long each block of text is | The default; adjust when something breaks |
| Indexing method | Vectors or plain keywords | High quality (vectors) |
| Retrieval mode | How relevant chunks are found | Hybrid search |
| Top-K | How many chunks go to the model | 3 to 5 |
| Score threshold | How low a relevance score to discard | Off at first; turn it on once you have data |
What actually happens after you upload a document
Once you understand this pipeline, every setting that follows makes sense:
- Parsing — PDFs, Word files and web pages are converted to plain text
- Chunking — the text is cut into blocks (chunks)
- Indexing — each chunk gets a vector that represents its meaning
- Retrieval — when someone asks a question, the closest chunks by meaning are pulled out
- Generation — those chunks are sent to the model along with the question
Step five is the crux: the model only sees the chunks that were retrieved. Anything that was not retrieved may as well not exist.
Choosing a chunk size
Too small and too large have different symptoms
| Too small | Too large | |
|---|---|---|
| Retrieval precision | High | Low (one chunk mixes several topics) |
| Context completeness | Low | High |
| Symptom | Answers are cut off and lack context | Answers drift off topic and pull in irrelevant material |
How to decide
There is no universal number, but there is a useful principle: one chunk should contain exactly one complete idea.
FAQ-style content suits small chunks, with one question and answer per chunk. Operating manuals suit larger ones, because a complete step should not be split apart.
This is also why documents of different kinds are best kept in separate knowledge bases rather than thrown together under one set of settings.
Separators beat length
If your documents have clear structure (headings, numbering, question-and-answer pairs), using those structural markers as chunk boundaries works far better than a fixed character count. Spending five minutes tidying up the source format usually beats half a day of tuning parameters.
Indexing: high quality or economical
High quality mode calls an embedding model to produce vectors, which captures meaning — a user asking "how do I get my money back" still finds the chunk that says "refund process". The cost is money and time when the index is built.
Economical mode matches keywords. It costs nothing in model fees, but it only sees literal text.
In the vast majority of cases you want high quality. Economical mode fits a narrow set of situations: highly standardized glossary content, or a quick check that the pipeline works at all.
The three retrieval modes
Vector search
Compares semantic similarity. It finds the answer even when the user asks about the same thing in completely different words. Its weakness is proper nouns, part numbers and model numbers — things where the literal string is the point — which it can retrieve poorly.
Full-text search
Compares keywords. Strong on proper nouns, weak on paraphrasing.
Hybrid search
Runs both and merges the results. The best choice in most cases, especially for knowledge bases full of product model numbers and proper nouns.
If your content is dense with part numbers, version numbers or technical jargon, the gap between hybrid and pure vector search is dramatic.
Top-K and the score threshold
Bigger Top-K is not better
Top-K decides how many chunks are retrieved each time. Raising it causes three problems: it is slower, it is more expensive (more content goes to the model), and irrelevant chunks pull the model off course.
Three to five chunks is enough in most situations. If three chunks do not contain the answer, the problem is your chunking or your document quality, and raising Top-K only drags in more noise.
The score threshold cuts both ways
With a threshold set, chunks scoring below it are never retrieved. Less noise is the upside; the downside is that setting it too high makes the bot say "I could not find that" all the time.
Leave it off at first, watch the real distribution of scores for a while, and then pick a threshold. A number you guess at on day one is usually wrong.
Chunking three kinds of document
Abstract principles only go so far, so here is how to cut three common document types:
FAQs and question banks
One question and answer per chunk. This kind of content is naturally well structured, and using the question as the boundary makes retrieval extremely accurate. It is also the document type that gets good results most easily — if you already have a tidy set of support questions and answers, start there.
Operating manuals
One complete step per chunk. Cut them too finely and answers lose context: the user asks "how do I set this up" and gets step three, with the first two steps missing.
If the manual is numbered, using the numbering as the boundary usually beats a fixed character count.
Contracts and terms
One clause per chunk. These documents are dense with proper nouns, so pair them with hybrid search — pure vector search easily confuses similar clauses.
How to split across multiple knowledge bases
Splitting by topic is almost always better than stuffing every document into one knowledge base.
| How you split | When it fits |
|---|---|
| By topic (products / returns / technical specs) | The most common approach, and the most effective |
| By audience (customer-facing / internal) | Some content should never reach customers |
| By currency (current / historical) | Stops old versions competing with new ones |
The third one is the easiest to overlook. Leave documentation for an old product version in the knowledge base and it competes with the new one, and the model cannot tell which is current. Outdated content should be moved out, not kept around for reference.
A debugging order for inaccurate answers
Working through this order is much faster than randomly tweaking parameters:
| Check first | If | The problem is |
|---|---|---|
| Which chunks it cited | The retrieved chunks are irrelevant | Retrieval — change the retrieval mode or the chunking |
| The right chunks came back but the answer is wrong | The prompt — it does not say clearly how to use the material | |
| Nothing was retrieved at all | The threshold is too high, or the documents simply do not cover it | |
| The source document | That passage was cut in half | Chunking — switch to structural separators |
The citations under an answer are the single most useful thing in this whole process. Look at them before anything else.
FAQ
Q: When I update a document, does the knowledge base update itself?
No. After you edit the source document you have to re-upload it and rebuild the index, or the bot keeps answering from the old version. This is the most commonly neglected maintenance task, especially for things that change, like pricing and policies.
Q: How many documents can I put in?
In terms of space, plain text needs very little. But being able to fit them in does not mean you should — as documents pile up, retrieval quality often drops rather than improves, because more similar chunks compete with the correct answer. See what happens as a knowledge base grows.
Q: Can I use scanned PDFs?
Parsing them succeeds unpredictably and eats resources. Run text recognition to turn them into plain text before uploading; both the results and the cost will be much better.
Q: Can one app use several knowledge bases?
Yes. In practice, splitting by topic into several knowledge bases and selecting between them by context works better than piling everything into one.
Sources and further reading
- Dify official documentation—full reference for knowledge bases and retrieval settings
- Dify release notes
Further reading
- Dify Tutorial: Build Your First AI Chatbot from Scratch
- Scaling a Dify Knowledge Base: Retrieval Quality vs. Resources
- Dify Workflows: From Single-Turn Q&A to Multi-Step Processes
- Resource Limits of Managed Dify: When 2 vCPU Is Not Enough
- RoamerHost managed Dify plans
Want someone to do it for you?
If you would rather not build these pipelines yourself, or the scale calls for someone to plan it with you, Roamer Tech (the company behind RoamerHost) takes on business process automation and custom AI agent projects: