Arama Yap Mesaj Submit
Request a Callback
+90
X
X

Select Your Currency

₺ Turkish Lira $ US Dollar € Euro
X
X

Select Your Currency

₺ Turkish Lira $ US Dollar € Euro

Contact Us

Location Halkali merkez neighborhood fatih st ozgur apt no 46 , Kucukcekmece , Istanbul , 34303 , TR
TECHNICAL GUIDE • TR / EN / DE

Ubuntu 24.04: Production RAG with Qdrant + Ollama + n8n

A production-oriented guide to combining n8n, Ollama and Qdrant on Ubuntu 24.04 LTS with persistent storage, container networking, embedding compatibility, GPU planning, security and recovery.

Important production note

Before running commands in production, validate versions, backups, firewall rules and the rollback plan on your own infrastructure.

Ubuntu 24.04 LTS Qdrant vector store Ollama local LLM n8n RAG workflow
ARCHITECTURE & DIAGNOSTICS
EKA CORE
Ubuntu 24.04 Qdrant + Ollama + n8n RAG Setup

n8n orchestrates the workflows, Ollama serves chat and embedding models, and Qdrant provides vector retrieval. Most failures come from network addressing, embedding dimension mismatch, capacity pressure or exposing services without security rather than from containers simply failing to start.

RAG architecture and data flowProduction-focused technical check
Validated
CPU/RAM/VRAM/NVMe sizingProduction-focused technical check
Validated
Docker and service networkingProduction-focused technical check
Validated
Embedding and Qdrant collection designProduction-focused technical check
Validated
Official sources + measurable tests + rollback plan
What this guide covers

A reliable RAG stack is more than three running containers. Ingestion and retrieval must use compatible embeddings, Qdrant vector dimensions must match the model, n8n must reach services through the Docker network, and all important data must live on persistent storage.

01

What this guide covers

Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.

✓RAG architecture and data flow
✓CPU/RAM/VRAM/NVMe sizing
✓Docker and service networking
✓Embedding and Qdrant collection design
✓n8n ingestion and retrieval
✓Security, TLS and secrets
✓Backup, monitoring and troubleshooting
✓Production benchmarking

Contents

  1. RAG architecture and data flow
  2. Sizing CPU, RAM, VRAM and NVMe
  3. Docker networking and the localhost trap
  4. Embedding, chunks and Qdrant collections
  5. Separate ingestion and question-answer workflows
  6. Production security for Qdrant, n8n and Ollama
  7. Backups, upgrades and rollback
  8. How to benchmark a RAG stack
  9. Component and network matrix
  10. Vector storage pre-calculator
  11. Common RAG failures and misdiagnosis
  12. Commands and deployment blocks
  13. Frequently asked questions
02

RAG architecture and data flow

Documents are cleaned and chunked, each chunk is embedded and stored in Qdrant with metadata. At query time the same embedding family turns the question into a vector, Qdrant retrieves relevant chunks, and n8n sends that context to the Ollama chat model.

Keep generation and retrieval diagnostics separate. A fluent answer does not prove that retrieval is correct, and a successful vector query does not prove the right evidence was selected. Measure retrieval quality, answer quality and latency independently.

03

Sizing CPU, RAM, VRAM and NVMe

Ollama capacity depends on model weights, context length and GPU offload. Qdrant capacity depends on vector count, dimensions, payload indexes, quantization, replication and storage mode.

Context growth increases memory pressure. Test representative prompt sizes and concurrency instead of approving a server just because the model loads once. Fast local SSD or NVMe storage is especially useful for ingestion, snapshots and larger collections.

04

Docker networking and the localhost trap

Inside the n8n container, localhost points to n8n itself. Services in the same Compose network should normally be addressed by service name, such as http://ollama:11434 and http://qdrant:6333.

Keep Qdrant and Ollama private unless external access is genuinely required. Use a reverse proxy, TLS, authentication, network restrictions and rate limits for exposed endpoints.

05

Embedding, chunks and Qdrant collections

The Qdrant vector dimension must match the embedding output. Changing embedding models can require a new collection and re-embedding rather than mixing incompatible vectors.

Chunk size should preserve useful semantic boundaries without flooding the prompt with oversized context. Validate chunking and top-k retrieval against a real question set.

06

Separate ingestion and question-answer workflows

Use one workflow for source loading, cleaning, chunking, embedding and Qdrant upserts; use another for question embedding, vector retrieval, prompt construction and Ollama generation.

Re-embedding all documents on every question wastes resources and makes failures harder to isolate. Correlate n8n executions, Qdrant queries and model latency when possible.

07

Production security for Qdrant, n8n and Ollama

Self-hosted Qdrant requires explicit security configuration. Use API keys, restricted network binding and TLS before production exposure.

Protect the n8n encryption key together with database backups. Keep .env files outside public web paths and restrict Ollama API access to services that actually need it.

08

Backups, upgrades and rollback

Container images are not data backups. Protect Qdrant storage/snapshots, the n8n database and encryption key. Large model caches may be reproducible but still affect recovery time objectives.

Record image versions, test upgrades in staging and treat embedding model changes as data migrations because collections may need to be regenerated.

09

How to benchmark a RAG stack

Track time-to-first-token, end-to-end latency, Qdrant query p95, embedding time, retrieval relevance, context size, GPU/CPU utilization and error rate.

Use the same questions, collection snapshot and model settings when comparing servers. Verify actual model offload and context allocation with `ollama ps`.

MAP

Component and network matrix

Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.

ComponentRoleNetwork / portProduction note
n8nWorkflow orchestration5678 behind proxy/private networkProtect encryption key and database
OllamaChat and embedding inference11434 private networkValidate model, VRAM and context
QdrantVector and payload retrieval6333 HTTP / 6334 gRPCAPI key, binding, TLS and snapshots
PostgreSQLPersistent n8n data5432 internal onlyBackups and storage monitoring
Reverse ProxyHTTPS entry point80/443 publicTLS, rate limits and timeouts
CALC

Vector storage pre-calculator

This result is an estimate; production decisions require real measurements and tests.

ERR

Common RAG failures and misdiagnosis

A reliable RAG stack is more than three running containers. Ingestion and retrieval must use compatible embeddings, Qdrant vector dimensions must match the model, n8n must reach services through the Docker network, and all important data must live on persistent storage.

Symptom / problemLikely layerFirst verification
n8n cannot reach OllamaContainer network / localhost misuseTest http://ollama:11434 from the n8n network.
Qdrant reports a dimension errorEmbedding and collection mismatchCompare embedding output dimension with collection config.
Answers use irrelevant sourcesChunking / embedding / retrievalInspect top-k results on a real test set.
Ollama uses CPU despite a GPUDriver/runtime/offloadVerify `nvidia-smi` and `ollama ps`.
Data disappears after Qdrant restartPersistent volume issueVerify the Qdrant storage mount and snapshots.
n8n credentials fail after migrationEncryption key mismatchRestore the matching encryption key with the database backup.
FLOW

Deployment and validation flow

Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.

1

Validate Ubuntu and Docker

Update the host and verify Docker Compose and optional GPU runtime.

2

Start services on a private network

Confirm service-name connectivity between n8n, Ollama, Qdrant and the database.

3

Freeze the embedding design

Record model, dimension, distance metric and chunking rules.

4

Build the ingestion workflow

Source → clean → chunk → embed → Qdrant upsert.

5

Build retrieval and chat

Question → embed → Qdrant top-k → context → Ollama.

6

Add security and backups

Reduce exposed ports, add TLS/API keys and test recovery.

7

Load test and observe

Measure p95 latency, retrieval quality and resource pressure under concurrency.

CLI

Commands and deployment blocks

Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.

Verify Docker and Compose
docker --version && docker compose version
Official n8n AI Starter Kit
git clone https://github.com/n8n-io/self-hosted-ai-starter-kit.git
cd self-hosted-ai-starter-kit
cp .env.example .env
Start CPU profile
docker compose --profile cpu pull
docker compose --profile cpu up -d
NVIDIA GPU profile
nvidia-smi
docker compose --profile gpu-nvidia pull
docker compose --profile gpu-nvidia up -d
Service checks
docker compose ps
curl -fsS http://127.0.0.1:6333/
curl -fsS http://127.0.0.1:11434/api/tags
Ollama model/context
ollama list
ollama ps
Log diagnostics
docker compose logs --tail=150 n8n qdrant
docker compose logs --tail=150 ollama-cpu ollama-gpu 2>/dev/null || true
TECHNICAL PRE-ASSESSMENT

Plan your RAG server with us

Share model size, dataset size, concurrency, GPU preference and target latency so we can evaluate CPU/RAM/VRAM/NVMe sizing and service separation.

Phone & WhatsApp0850 307 34 58Do not send passwords or API keys initially.
SRC

Official and technical sources

Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.

EKA

Related Eka Sunucu pages

Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.

FAQ

Frequently asked questions

A reliable RAG stack is more than three running containers. Ingestion and retrieval must use compatible embeddings, Qdrant vector dimensions must match the model, n8n must reach services through the Docker network, and all important data must live on persistent storage.

Can n8n, Ollama and Qdrant run on one server?

Yes for many small and medium deployments, but shared CPU, RAM, VRAM and storage contention should be load-tested.

Should Qdrant port 6333 be public?

Usually no. Keep it private when n8n is on the same network. If external clients need access, add authentication, TLS and network restrictions.

Can I keep a collection after changing embedding models?

A different vector space or dimension usually requires a new collection and re-embedding.

Is a GPU mandatory for RAG?

Not for n8n or Qdrant. Ollama can run on CPU, but GPU acceleration may be important for target latency and concurrency.

Is the n8n AI Starter Kit production-ready by itself?

It is an excellent starting point for learning and proof-of-concept work; production still needs explicit HA, security, backups, observability and scaling design.

Can Qdrant use NFS or S3 as its live storage?

Qdrant documents block-level POSIX-compatible storage for persistent data rather than NFS or object storage as live storage.

Why does larger Ollama context use more VRAM?

Longer context increases runtime memory requirements; verify real allocation with `ollama ps`.

How should RAG quality be measured?

Use a curated question set and evaluate retrieved evidence, answer correctness and latency together.

EKA SUNUCU

Plan your RAG server with us

Share model size, dataset size, concurrency, GPU preference and target latency so we can evaluate CPU/RAM/VRAM/NVMe sizing and service separation.

Phone & WhatsApp0850 307 34 58ekasunucu.com
Top