Architecture Dify

Dify Hosting Limits: When 2 vCPU Is Not Enough

Dify is resource-hungry, but where is often misunderstood: model inference does not run on your machine. The three real consumers, and when to add resources.

E
Eric Founder, Roamer Tech · · 7 min read

Want to start now? Deploy your Dify in 60 seconds

AI app platform — build your AI with no code. From NT$1,599/mo.

Subscribe to Dify

Dify is one of the most resource-hungry services on this site. The plan gives you 2 vCPU / 10 GB (NT$1,599/mo), a tier above both n8n and NocoDB.

But "resource-hungry" is widely misunderstood, which leads people to spend money in the wrong place.

The 30-second overview

SymptomIs it a resource problem?
Uploading a large PDF times outYes, not enough memory
Slows down when several people use it at onceYes, concurrency is maxed out
Slow replies with a single userNo, that is the model's response time
Inaccurate answersNo, that is a design problem
High API billNo, that has nothing to do with instance specs

First, to be clear: model inference does not run on your machine

The most common misconception is "AI is CPU-heavy, so I should buy something bigger".

In reality you are connecting to external models from OpenAI, Anthropic, and others, and inference runs entirely on their servers. Your Dify instance only assembles the question, sends it out, and receives the answer. That part barely touches your compute resources; what it consumes is your API bill.

So if your pain point is "answers are slow", first work out which part is slow. It is usually the model's own response time, and adding vCPU will not help one bit.

The three things that actually consume resources

1. Document parsing and indexing (the heaviest by far)

When you upload a document into a knowledge base, the system has to parse the file into text, split it into chunks, and send each chunk off to be turned into a vector. This is the heaviest step in the whole pipeline.

Large PDFs hurt in particular, especially scans or files with complex layouts, which consume a lot of memory at the parsing stage alone. Timeouts or failures partway through a large upload usually come from here.

The good news is that it is a one-time cost. Once the index is built, day-to-day queries are much cheaper.

2. Vector retrieval

Every time someone asks a question, the system has to find the most relevant passages in the knowledge base. The larger the knowledge base, the heavier this step.

3. Concurrent conversations

The difference between one user and ten simultaneous users is linear. If you put a Dify app on your website as customer support, peak-hour concurrency is the number you actually need to estimate.

Signs you should add resources

SymptomIs it a resource problem?
Uploading a large PDF times out or failsYes — not enough memory
Noticeably slower when several people use it at onceYes — concurrency is maxed out
Retrieval slows down as the knowledge base growsPossibly — but read the next section first
The answers are inaccurateNo — this is a design problem
Very slow replies with a single userNo — usually the model's response time
High API billNo — that is the model cost, unrelated to instance specs

Three things to check before adding resources

To be honest: for most people the bottleneck is not vCPU, it is design. Check these three before adding resources.

1. Whether your knowledge base contains things that do not belong. Many people throw every document in and end up retrieving nothing but noise. Half the documents at twice the quality usually works better, and saves resources too.

2. Top-K set too high. Pulling back a dozen-plus passages on every retrieval is slow and distracts the model. Three to five is enough in most situations.

3. Process large files outside before uploading. Splitting a 200-page PDF into a few topic-specific files greatly reduces the parsing load and improves retrieval precision as well.

Is 10 GB of storage enough?

For a plain-text knowledge base it is more than enough, since text itself takes very little space. The main consumers are the original files and the vector index.

What usually approaches the limit is storing a lot of images or scanned PDFs. If your source files are scans, consider converting them to plain text before uploading, which brings down both storage and parsing cost.

Two bills to keep separate

Dify spending comes from two sources, and they grow in completely different ways:

Monthly hosting feeModel cost
Paid toThe hosting providerThe model provider
How it variesFixedProportional to usage
What drives it upUpgrading your planLarge Top-K, many conversation turns, a big knowledge base

When the bill goes up, many people go and look at their plan specs, but that money usually grew on the model side. Work out which bill it is first, then decide what to adjust.

Three things to do before adding resources

Honestly: for most people the bottleneck is not vCPU, it is design.

1. Clear out what should not be in the knowledge base

Many people throw every document in and end up retrieving nothing but noise. Half the documents at twice the quality usually works better, and saves resources too.

2. Bring Top-K back down to three to five

Pulling back a dozen-plus passages every time is slow, expensive, and distracting for the model. If three passages cannot find the answer, the problem is in your chunking or document quality, and raising the number just drags the noise in with it.

3. Process large files outside first

Splitting a 200-page PDF into a few topic-specific files greatly reduces the parsing load and improves retrieval precision as well. Convert scanned PDFs to text before uploading.

How to tell which kind of slow it is

This one diagnosis saves a lot of wasted effort:

TestResultConclusion
Ask one question on your ownStill slowModel response time; adding resources will not help
Have three people ask at the same timeNoticeably slowerConcurrency maxed out; add resources
Ask something outside the knowledge baseFastThe slowness is in retrieval; look at knowledge base size first

FAQ

Q: Does the monthly fee include model costs?

No. This is the most commonly misunderstood point: the monthly fee is the platform hosting cost, and model call charges are paid by you to the provider. They are two separate bills.

Q: Is 10 GB of storage enough?

For a plain-text knowledge base it is more than enough, since text takes very little space. What usually approaches the limit is storing a lot of scanned PDFs or images.

Q: What do I do if uploads keep failing?

Start with the file type and size. Scanned PDFs have a low parsing success rate and are resource-hungry, so convert them to text first. Only if it keeps happening should you consider whether your specs are too small.

Q: How do I estimate how much concurrency I need?

It depends where you plan to put the app. Internal team use has very low concurrency; if it is on your website as customer support, what you need to estimate is the number of simultaneous conversations at peak.

Q: Will upgrading my plan make the answers more accurate?

No. Accuracy is a design problem (document quality, chunking, retrieval settings) and has nothing to do with compute resources. These two often get conflated, but the fixes are completely different.

When you really should upgrade

Pulling the above together, there are only three genuine signals to upgrade:

SignalWhy
Large file uploads fail oftenNot enough memory at the parsing stage
Clearly slower when several people use it at peakConcurrency maxed out
A very large knowledge base with noticeably slower retrievalTry splitting it first, upgrade if that does not help

If none of the three applies, adding resources will not solve your problem, and you will only find that out after spending the money.

Sources and further links

Dify's features and interface change between versions, so check against the official documentation before you start:

Further reading

Want someone to build it for you?

If you would rather not assemble these workflows yourself, or the project is large enough that you want someone planning it with you, Roamer Tech (the company behind RoamerHost) takes on contract work for business process automation and AI agents:

Ready to get started with Dify?

60 seconds after you subscribe, Dify is installed for you — an isolated container with hard resource limits you never share, and HTTPS out of the box.

Subscribe to Dify

Billed monthly · no contract · cancel anytime

Hi, I'm Roamer! Tap me anytime with a question and I'll help you out.

Roamer

Roamer - AI assistant

Online
Roamer

Ask me anything, anytime — I'll do my best to help!

Powered by RoamerHost AI