The single most common thing new Dify users do is dump every document they have into the knowledge base.
Intuitively that makes sense: more data, more the AI knows. In practice, past a certain size, answer quality starts to drop — and the reason is not easy to see.
The 30-second overview
| Question | Answer |
|---|---|
| Does more data make it more accurate? | Usually the opposite |
| Why | The number of chunks retrieved is fixed, so there is more competition |
| Most effective fix | Split into several knowledge bases by topic |
| Should I raise Top-K? | No — three to five chunks is usually best |
| What breaks first | Indexing time, not queries |
Why more data makes things worse
RAG works like this: the user asks a question → the system finds the most relevant chunks in the knowledge base → those chunks go to the model along with the question → the model answers based on them.
The key is that “the most relevant chunks” has a hard count limit.
Go from 50 chunks to 5,000 and the system still retrieves only three to five. The difference is that there are now many more chunks that look similar but are actually irrelevant, competing with the correct answer.
The result is the phenomenon plenty of people hit without being able to explain it — “questions it used to get right are now answered wrong.”
The chunk size trade-off
Documents are cut into small chunks before indexing. That size directly determines retrieval quality.
| Too small | Too large | |
|---|---|---|
| Retrieval precision | High | Low (one chunk mixes several topics) |
| Context completeness | Low (the answer gets cut off) | High |
| Typical symptom | Truncated answers, missing context | Off-topic answers pulling in unrelated content |
There is no one-size-fits-all number, but there is a rule of thumb: a chunk should contain exactly one complete idea.
FAQ-style content suits small chunks (one question and answer per chunk); operating manuals suit larger ones (a complete step should not be split). That is also why documents of different kinds are better off in separate knowledge bases rather than all mixed together.
Bigger Top-K is not better
Top-K decides how many chunks each retrieval returns. Plenty of people crank it up, figuring “more data can’t hurt”.
It can. Returning a dozen-plus chunks causes three problems: it is slow, it is expensive (more tokens sent to the model), and the model gets pulled off course by irrelevant content.
Three to five chunks is enough in most situations. If three chunks cannot find the right answer, the problem is in your chunking or your document quality, and raising Top-K just drags noise in with it.
How a knowledge base should grow
These approaches work far better than “add more”:
- Split into several knowledge bases by topic. Product docs, return policy and technical specs each stand alone, and the app picks per situation. This is the single most effective step for improving retrieval quality.
- Clear out stale content regularly. An old product description left in the base competes with the new one, and the model cannot tell which is current.
- Clean up the source files first. Splitting a 200-page manual into a few topic files beats dropping the whole book in.
- Test with real questions. Run ten questions customers have actually asked, on a regular cadence — more reliable than any metric.
Where resources give out first
As the knowledge base grows, the first thing you feel is indexing time — every new document has to be parsed and vectorized again, and large files eat memory at this step.
Day-to-day query costs grow more slowly, but they amplify when many people use it at once. The plan’s 2 vCPU / 10 GB is quite generous for a plain-text knowledge base; what actually pushes against the limit is usually a large pile of scanned PDFs.
Before adding resources, though, it is worth confirming whether you have a resource problem or a design problem — the symptoms differ.
Three ways to split a knowledge base
| Split by | When to use it |
|---|---|
| Topic | The most common — product / policy / technical specs |
| Audience | Some content should not be visible to customers |
| Currency | Keep current and historical versions apart |
The third is the easiest to overlook and has a big impact: an old product description left in the same base competes with the new one, and the model cannot tell which one is current.
Once split, the app can pick a different knowledge base per situation, or use a workflow to classify first and then decide which one to query. How to do that is in the workflow orchestration tutorial.
Three stages of a growing knowledge base
| Stage | Symptom | What to do |
|---|---|---|
| Early | Accurate answers, occasionally finds nothing | Just add documents |
| Middle | Starts answering wrong, and very confidently | Split the knowledge base, purge stale content |
| Late | Uploads slow down, indexing takes a long time | Look at resources, or shrink the source files |
The middle stage is the dangerous one. “Questions it used to get right are now answered wrong” — because there are now more chunks that look similar but are irrelevant competing with the correct answer.
The easiest wrong reaction at this stage is “add even more documents”, which makes it worse.
How to know it is time to split
Three concrete signals:
- The same question gets good answers sometimes and bad ones other times — retrieval is returning unstable chunks
- Answers cite chunks from an obviously different topic — you ask about the return policy and it quotes technical specs
- You cannot describe what is in the knowledge base yourself — then the model certainly cannot tell it apart
How to split is covered above. Splitting the knowledge base is the one adjustment that almost always helps, more than tuning any parameter.
FAQ
Q: Is there a limit on the number of entries?
On storage, plain text needs very little. But being able to fit it does not mean you should put it in — the quality ceiling arrives long before the capacity ceiling.
Q: How many should I split it into?
Split by topic, not by count. Product docs, return policy and technical specs as separate bases is a common arrangement. The test is “would these two kinds of content interfere with each other?”
Q: Should I delete old versions of documents?
Yes. An old product description left in the base competes with the new one, and the model cannot tell which is current. Stale content should be moved out, not kept around for reference.
Q: How do I tell whether it got better?
Prepare ten questions that were actually asked and run them after every change. Without a fixed test set, “it feels better” is usually an illusion.
Q: If a document is updated, do I have to re-upload it?
Yes — the knowledge base does not sync automatically. Content that changes, like prices and policies, is especially dangerous: the bot will confidently answer with the old version.
The one thing to do regularly
A knowledge base is not something you configure once and forget; it needs maintenance. The bare minimum is running the same ten questions once a quarter.
Those ten should be questions people actually asked, and they should stay fixed — change the questions and you lose your baseline. Record which ones were answered wrong, and whether the failure was in retrieval or generation.
Two consecutive quarters of declining accuracy usually means one of two things: documents have piled up and it is time to split, or stale content is still in the base competing with the current version.
Sources and further reading
Dify’s features and interface change between versions, so check against the official documentation before working through anything:
- Dify documentation
- Dify release notes — the authority on feature changes
- Dify source code and issue tracker
Further reading
- Resource limits on managed Dify: when 2 vCPU is not enough
- Dify’s technical strengths: RAG knowledge bases, workflow orchestration, and why it is resource-hungry
- Dify tutorial: build your first AI chatbot
- Dify FAQ: API keys, knowledge base tuning and resource allocation
- RoamerHost managed Dify plans
Want someone to build it for you?
If you would rather not assemble these flows yourself, or the scope is large enough that you want someone planning alongside you, Roamer Tech (RoamerHost’s parent company) takes on contract work in business process automation and AI agent development: