A production-oriented guide to combining n8n, Ollama and Qdrant on Ubuntu 24.04 LTS with persistent storage, container networking, embedding compatibility, GPU planning, security and recovery.
Before running commands in production, validate versions, backups, firewall rules and the rollback plan on your own infrastructure.
n8n orchestrates the workflows, Ollama serves chat and embedding models, and Qdrant provides vector retrieval. Most failures come from network addressing, embedding dimension mismatch, capacity pressure or exposing services without security rather than from containers simply failing to start.
A reliable RAG stack is more than three running containers. Ingestion and retrieval must use compatible embeddings, Qdrant vector dimensions must match the model, n8n must reach services through the Docker network, and all important data must live on persistent storage.
Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.
Documents are cleaned and chunked, each chunk is embedded and stored in Qdrant with metadata. At query time the same embedding family turns the question into a vector, Qdrant retrieves relevant chunks, and n8n sends that context to the Ollama chat model.
Keep generation and retrieval diagnostics separate. A fluent answer does not prove that retrieval is correct, and a successful vector query does not prove the right evidence was selected. Measure retrieval quality, answer quality and latency independently.
Ollama capacity depends on model weights, context length and GPU offload. Qdrant capacity depends on vector count, dimensions, payload indexes, quantization, replication and storage mode.
Context growth increases memory pressure. Test representative prompt sizes and concurrency instead of approving a server just because the model loads once. Fast local SSD or NVMe storage is especially useful for ingestion, snapshots and larger collections.
Inside the n8n container, localhost points to n8n itself. Services in the same Compose network should normally be addressed by service name, such as http://ollama:11434 and http://qdrant:6333.
Keep Qdrant and Ollama private unless external access is genuinely required. Use a reverse proxy, TLS, authentication, network restrictions and rate limits for exposed endpoints.
The Qdrant vector dimension must match the embedding output. Changing embedding models can require a new collection and re-embedding rather than mixing incompatible vectors.
Chunk size should preserve useful semantic boundaries without flooding the prompt with oversized context. Validate chunking and top-k retrieval against a real question set.
Use one workflow for source loading, cleaning, chunking, embedding and Qdrant upserts; use another for question embedding, vector retrieval, prompt construction and Ollama generation.
Re-embedding all documents on every question wastes resources and makes failures harder to isolate. Correlate n8n executions, Qdrant queries and model latency when possible.
Self-hosted Qdrant requires explicit security configuration. Use API keys, restricted network binding and TLS before production exposure.
Protect the n8n encryption key together with database backups. Keep .env files outside public web paths and restrict Ollama API access to services that actually need it.
Container images are not data backups. Protect Qdrant storage/snapshots, the n8n database and encryption key. Large model caches may be reproducible but still affect recovery time objectives.
Record image versions, test upgrades in staging and treat embedding model changes as data migrations because collections may need to be regenerated.
Track time-to-first-token, end-to-end latency, Qdrant query p95, embedding time, retrieval relevance, context size, GPU/CPU utilization and error rate.
Use the same questions, collection snapshot and model settings when comparing servers. Verify actual model offload and context allocation with `ollama ps`.
Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.
| Component | Role | Network / port | Production note |
|---|---|---|---|
| n8n | Workflow orchestration | 5678 behind proxy/private network | Protect encryption key and database |
| Ollama | Chat and embedding inference | 11434 private network | Validate model, VRAM and context |
| Qdrant | Vector and payload retrieval | 6333 HTTP / 6334 gRPC | API key, binding, TLS and snapshots |
| PostgreSQL | Persistent n8n data | 5432 internal only | Backups and storage monitoring |
| Reverse Proxy | HTTPS entry point | 80/443 public | TLS, rate limits and timeouts |
This result is an estimate; production decisions require real measurements and tests.
A reliable RAG stack is more than three running containers. Ingestion and retrieval must use compatible embeddings, Qdrant vector dimensions must match the model, n8n must reach services through the Docker network, and all important data must live on persistent storage.
| Symptom / problem | Likely layer | First verification |
|---|---|---|
| n8n cannot reach Ollama | Container network / localhost misuse | Test http://ollama:11434 from the n8n network. |
| Qdrant reports a dimension error | Embedding and collection mismatch | Compare embedding output dimension with collection config. |
| Answers use irrelevant sources | Chunking / embedding / retrieval | Inspect top-k results on a real test set. |
| Ollama uses CPU despite a GPU | Driver/runtime/offload | Verify `nvidia-smi` and `ollama ps`. |
| Data disappears after Qdrant restart | Persistent volume issue | Verify the Qdrant storage mount and snapshots. |
| n8n credentials fail after migration | Encryption key mismatch | Restore the matching encryption key with the database backup. |
Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.
Update the host and verify Docker Compose and optional GPU runtime.
Confirm service-name connectivity between n8n, Ollama, Qdrant and the database.
Record model, dimension, distance metric and chunking rules.
Source → clean → chunk → embed → Qdrant upsert.
Question → embed → Qdrant top-k → context → Ollama.
Reduce exposed ports, add TLS/API keys and test recovery.
Measure p95 latency, retrieval quality and resource pressure under concurrency.
Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.
docker --version && docker compose versiongit clone https://github.com/n8n-io/self-hosted-ai-starter-kit.git
cd self-hosted-ai-starter-kit
cp .env.example .envdocker compose --profile cpu pull
docker compose --profile cpu up -dnvidia-smi
docker compose --profile gpu-nvidia pull
docker compose --profile gpu-nvidia up -ddocker compose ps
curl -fsS http://127.0.0.1:6333/
curl -fsS http://127.0.0.1:11434/api/tagsollama list
ollama psdocker compose logs --tail=150 n8n qdrant
docker compose logs --tail=150 ollama-cpu ollama-gpu 2>/dev/null || trueShare model size, dataset size, concurrency, GPU preference and target latency so we can evaluate CPU/RAM/VRAM/NVMe sizing and service separation.
Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.
Architecture, sizing, Docker networking, embeddings, security, performance, backups and troubleshooting from installation to production.
A reliable RAG stack is more than three running containers. Ingestion and retrieval must use compatible embeddings, Qdrant vector dimensions must match the model, n8n must reach services through the Docker network, and all important data must live on persistent storage.
Yes for many small and medium deployments, but shared CPU, RAM, VRAM and storage contention should be load-tested.
Usually no. Keep it private when n8n is on the same network. If external clients need access, add authentication, TLS and network restrictions.
A different vector space or dimension usually requires a new collection and re-embedding.
Not for n8n or Qdrant. Ollama can run on CPU, but GPU acceleration may be important for target latency and concurrency.
It is an excellent starting point for learning and proof-of-concept work; production still needs explicit HA, security, backups, observability and scaling design.
Qdrant documents block-level POSIX-compatible storage for persistent data rather than NFS or object storage as live storage.
Longer context increases runtime memory requirements; verify real allocation with `ollama ps`.
Use a curated question set and evaluate retrieved evidence, answer correctness and latency together.
Share model size, dataset size, concurrency, GPU preference and target latency so we can evaluate CPU/RAM/VRAM/NVMe sizing and service separation.