Architecture Dify

As Your Dify Knowledge Base Grows: Quality vs Resources

Most people assume more data means better answers. Usually it is the opposite. What happens as a knowledge base grows, and how to tune chunking and Top-K.

E
Eric Founder, Roamer Tech · · 7 min read

Want to start now? Deploy your Dify in 60 seconds

AI app platform — build your AI with no code. From NT$1,599/mo.

Subscribe to Dify

The single most common thing new Dify users do is dump every document they have into the knowledge base.

Intuitively that makes sense: more data, more the AI knows. In practice, past a certain size, answer quality starts to drop — and the reason is not easy to see.

The 30-second overview

QuestionAnswer
Does more data make it more accurate?Usually the opposite
WhyThe number of chunks retrieved is fixed, so there is more competition
Most effective fixSplit into several knowledge bases by topic
Should I raise Top-K?No — three to five chunks is usually best
What breaks firstIndexing time, not queries

Why more data makes things worse

RAG works like this: the user asks a question → the system finds the most relevant chunks in the knowledge base → those chunks go to the model along with the question → the model answers based on them.

The key is that “the most relevant chunks” has a hard count limit.

Go from 50 chunks to 5,000 and the system still retrieves only three to five. The difference is that there are now many more chunks that look similar but are actually irrelevant, competing with the correct answer.

The result is the phenomenon plenty of people hit without being able to explain it — “questions it used to get right are now answered wrong.”

The chunk size trade-off

Documents are cut into small chunks before indexing. That size directly determines retrieval quality.

Too smallToo large
Retrieval precisionHighLow (one chunk mixes several topics)
Context completenessLow (the answer gets cut off)High
Typical symptomTruncated answers, missing contextOff-topic answers pulling in unrelated content

There is no one-size-fits-all number, but there is a rule of thumb: a chunk should contain exactly one complete idea.

FAQ-style content suits small chunks (one question and answer per chunk); operating manuals suit larger ones (a complete step should not be split). That is also why documents of different kinds are better off in separate knowledge bases rather than all mixed together.

Bigger Top-K is not better

Top-K decides how many chunks each retrieval returns. Plenty of people crank it up, figuring “more data can’t hurt”.

It can. Returning a dozen-plus chunks causes three problems: it is slow, it is expensive (more tokens sent to the model), and the model gets pulled off course by irrelevant content.

Three to five chunks is enough in most situations. If three chunks cannot find the right answer, the problem is in your chunking or your document quality, and raising Top-K just drags noise in with it.

How a knowledge base should grow

These approaches work far better than “add more”:

  • Split into several knowledge bases by topic. Product docs, return policy and technical specs each stand alone, and the app picks per situation. This is the single most effective step for improving retrieval quality.
  • Clear out stale content regularly. An old product description left in the base competes with the new one, and the model cannot tell which is current.
  • Clean up the source files first. Splitting a 200-page manual into a few topic files beats dropping the whole book in.
  • Test with real questions. Run ten questions customers have actually asked, on a regular cadence — more reliable than any metric.

Where resources give out first

As the knowledge base grows, the first thing you feel is indexing time — every new document has to be parsed and vectorized again, and large files eat memory at this step.

Day-to-day query costs grow more slowly, but they amplify when many people use it at once. The plan’s 2 vCPU / 10 GB is quite generous for a plain-text knowledge base; what actually pushes against the limit is usually a large pile of scanned PDFs.

Before adding resources, though, it is worth confirming whether you have a resource problem or a design problem — the symptoms differ.

Three ways to split a knowledge base

Split byWhen to use it
TopicThe most common — product / policy / technical specs
AudienceSome content should not be visible to customers
CurrencyKeep current and historical versions apart

The third is the easiest to overlook and has a big impact: an old product description left in the same base competes with the new one, and the model cannot tell which one is current.

Once split, the app can pick a different knowledge base per situation, or use a workflow to classify first and then decide which one to query. How to do that is in the workflow orchestration tutorial.

Three stages of a growing knowledge base

StageSymptomWhat to do
EarlyAccurate answers, occasionally finds nothingJust add documents
MiddleStarts answering wrong, and very confidentlySplit the knowledge base, purge stale content
LateUploads slow down, indexing takes a long timeLook at resources, or shrink the source files

The middle stage is the dangerous one. “Questions it used to get right are now answered wrong” — because there are now more chunks that look similar but are irrelevant competing with the correct answer.

The easiest wrong reaction at this stage is “add even more documents”, which makes it worse.

How to know it is time to split

Three concrete signals:

  • The same question gets good answers sometimes and bad ones other times — retrieval is returning unstable chunks
  • Answers cite chunks from an obviously different topic — you ask about the return policy and it quotes technical specs
  • You cannot describe what is in the knowledge base yourself — then the model certainly cannot tell it apart

How to split is covered above. Splitting the knowledge base is the one adjustment that almost always helps, more than tuning any parameter.

FAQ

Q: Is there a limit on the number of entries?

On storage, plain text needs very little. But being able to fit it does not mean you should put it in — the quality ceiling arrives long before the capacity ceiling.

Q: How many should I split it into?

Split by topic, not by count. Product docs, return policy and technical specs as separate bases is a common arrangement. The test is “would these two kinds of content interfere with each other?”

Q: Should I delete old versions of documents?

Yes. An old product description left in the base competes with the new one, and the model cannot tell which is current. Stale content should be moved out, not kept around for reference.

Q: How do I tell whether it got better?

Prepare ten questions that were actually asked and run them after every change. Without a fixed test set, “it feels better” is usually an illusion.

Q: If a document is updated, do I have to re-upload it?

Yes — the knowledge base does not sync automatically. Content that changes, like prices and policies, is especially dangerous: the bot will confidently answer with the old version.

The one thing to do regularly

A knowledge base is not something you configure once and forget; it needs maintenance. The bare minimum is running the same ten questions once a quarter.

Those ten should be questions people actually asked, and they should stay fixed — change the questions and you lose your baseline. Record which ones were answered wrong, and whether the failure was in retrieval or generation.

Two consecutive quarters of declining accuracy usually means one of two things: documents have piled up and it is time to split, or stale content is still in the base competing with the current version.

Sources and further reading

Dify’s features and interface change between versions, so check against the official documentation before working through anything:

Further reading

Want someone to build it for you?

If you would rather not assemble these flows yourself, or the scope is large enough that you want someone planning alongside you, Roamer Tech (RoamerHost’s parent company) takes on contract work in business process automation and AI agent development:

Ready to get started with Dify?

60 seconds after you subscribe, Dify is installed for you — an isolated container with hard resource limits you never share, and HTTPS out of the box.

Subscribe to Dify

Billed monthly · no contract · cancel anytime

Hi, I'm Roamer! Tap me anytime with a question and I'll help you out.

Roamer

Roamer - AI assistant

Online
Roamer

Ask me anything, anytime — I'll do my best to help!

Powered by RoamerHost AI