Qwen-Image-2512 ve Z-Image-Turbo özellikle yazı üretimi için incelenmeli. Qwen ailesi karmaşık text rendering ve düzenlemede, Z-Image ise 16 GB tüketici sınıfında fotogerçekçilik ve EN/Çince metinde öne çıkıyor.
Choose an Ollama GPU by model size, VRAM, quantization, context, concurrency, coding, vision and image/video workloads instead of raw GPU speed alone.
The practical limit is determined first by usable accelerator memory, then by quantization, context/KV cache and runtime overhead. This page separates models that fit comfortably from models that require offload or are not a sensible target for this hardware.
Important: model-file size is not the whole memory requirement. KV cache, context, vision encoders, runtime workspace and concurrent requests consume additional memory.
Compatibility is classified as comfortable, borderline/tuning required, or not a natural target. This is more useful than a binary 'runs/does not run' label.
Creator and community benchmarks vary with backend, driver, quantization, context, batch size and power limit. They are shown as real-world evidence, not guaranteed performance.
| Model | Size / memory | Workload | Status | Technical assessment |
|---|---|---|---|---|
| Qwen3.5 9B Q4_K_M | 6,6 GB | LLM · Vision · Kod | Comfortable | 8 GB kartta ağırlıklar sığabilir; KV cache ve vision yükü için bağlamı kontrollü tutmak gerekir. |
| Qwen3.5 27B Q4_K_M | 17 GB | LLM · Vision · Kod | Comfortable | 24 GB sınıfında güçlü genel/kodlama seçeneği; uzun bağlam ayrıca bellek tüketir. |
| Qwen3.5 35B-A3B INT4 | 20 GB | MoE LLM · Vision | Comfortable | 24 GB kartta uygulanabilir; aktif parametre az olsa da toplam ağırlıkların bellekte tutulması gerekir. |
| gpt-oss-20b | 16 GB bellek hedefi | Reasoning · Agent · Kod | Comfortable | OpenAI resmi olarak 16 GB bellekte çalışabildiğini belirtiyor; bağlam ve backend ek yükü ayrıca değerlendirilir. |
| gpt-oss-120b | 80 GB bellek hedefi | Reasoning · Agent | Borderline / tuning required | OpenAI MXFP4 ile 80 GB bellek hedefi veriyor; tipik oyuncu GPU'ları için sistem RAM/offload gerekir. |
For coding and agent workloads, prefer a model that leaves memory headroom for repository context, tool calls and KV cache instead of filling the entire device with weights.
Vision and multimodal inference adds image-encoder and visual-token overhead. A model that fits for text-only use can become borderline with multiple high-resolution images.
For image generation, checkpoint size is only one component: text encoders, VAE, resolution and editing/reference images can raise peak memory. Use the official model-card requirement as the baseline.
Video diffusion is significantly heavier than chat inference. Resolution, frame count, VAE and temporal modules can change memory and generation time substantially.
Q4/INT4 reduces weight memory but does not make KV cache or runtime workspace disappear. System-RAM offload can enable larger models at the cost of throughput and latency.
The local.ai Intelligence score combines τ²-bench, GAIA and GDPval. It is a quality reference, not proof that a model fits your VRAM/RAM.
| Model | Developer | Intelligence | Role | Hardware note |
|---|---|---|---|---|
| Step 3.7 Flash | StepFun | 78,8 | Genel zekâ / reasoning | local.ai Intelligence lideri; skor donanım uyumluluğu anlamına gelmez. |
| DeepSeek V4 Flash 0731 | DeepSeek | 77,5 | Genel / reasoning / kod | Çok büyük model sınıfı; tüketici GPU'suna sırf 'Flash' adına bakarak uygun kabul edilmemeli. |
| Qwen3.6 27B | Qwen / Alibaba | 76,5 | Genel / coding / multimodal | 27B sınıfı güncel kalite referansı; checkpoint ve quantization boyutu ayrıca doğrulanmalı. |
| Qwen3.6 35B-A3B | Qwen / Alibaba | 74,9 | MoE / agent / multimodal | Aktif parametre düşük olsa da model ağırlıklarının bellek yerleşimi ayrıca hesaplanır. |
| Gemma 4 31B IT | 66,2 | Multimodal / coding / reasoning | Yerel kullanımda quantization ve context VRAM ihtiyacını belirler. | |
| Nemotron 3.5 Lightning | NVIDIA | 65,5 | Agent / reasoning | 30B-A3B sınıfı güncel NVIDIA yerel/agent modeli. |
| Qwen3.5 9B Base | Qwen / Alibaba | 59,3 | Küçük-orta local LLM | Ollama Q4 6,6 GB sınıfı nedeniyle 8 GB donanımlarda pratik önemi yüksek. |
| gpt-oss-20b | OpenAI | 43,3 | Reasoning / agent / coding | OpenAI resmi minimumu 16 GB bellek; benchmark skoru tek başına model seçimi değildir. |
local.ai metodolojisi: Intelligence = %10 τ²-bench + %35 GAIA + %55 GDPval. Decode/prefill ölçümleri 32K soğuk context ve tek stream üzerinde raporlanır. Skorlar ve model sürümleri zamanla değişebilir.
| Type | Model | Memory reference | Best fit | Model card |
|---|---|---|---|---|
| Görsel | FLUX.2 Klein 4B | ~13 GB VRAM | Hızlı yerel üretim, düzenleme, multi-reference | Hugging Face ↗ |
| Görsel | Z-Image-Turbo 6B | 16 GB tüketici sınıfı | Fotogerçekçilik, EN/Çince metin, 8 NFE | Hugging Face ↗ |
| Görsel | Qwen-Image-2512 | 20B model ailesi | Metin doğruluğu, layout, multimodal entegrasyon | Hugging Face ↗ |
| Video | Wan2.2 TI2V-5B | 4090 resmi tüketici örneği | T2V + I2V, 720p, 24 fps | Hugging Face ↗ |
| Video | Motif-Video-2B | 2B sınıfı | Üretici değerlendirmesinde VBench toplam 83,76 | Hugging Face ↗ |
| Video + Ses | LTX-2.5 | Checkpoint/precision'a göre | Açık ağırlık, yerel çalıştırma, senkron video + ses | Hugging Face ↗ |
| Video + Ses | MiniMax H3 | 33B sınıfı | 2K'ya kadar, 15 sn'ye kadar, native stereo audio; lisansı ticari kullanım öncesi okunmalı | Hugging Face ↗ |
Qwen-Image-2512 ve Z-Image-Turbo özellikle yazı üretimi için incelenmeli. Qwen ailesi karmaşık text rendering ve düzenlemede, Z-Image ise 16 GB tüketici sınıfında fotogerçekçilik ve EN/Çince metinde öne çıkıyor.
FLUX.2 Klein 4B için resmî yaklaşık 13 GB VRAM hedefi var. 16 GB kartlar burada 8 GB kartlara göre doğrudan farklı bir sınıfa geçiyor.
Wan2.2 5B olgun 720p/24 fps referansı; Motif-Video-2B daha küçük araştırma seçeneği; LTX-2.5 senkron ses-video; MiniMax H3 ise çok daha ağır 33B omni-modal sınıf. “En yeni” ile “donanımınıza en uygun” aynı şey değildir.
| Model | Ölçüm / veri | Source type | What does it show? |
|---|---|---|---|
| 8 GB sınıfı | 9B Q4 | Ollama model boyutları | Giriş seviyesi |
| 16 GB sınıfı | 20B/27B INT4 sınırı | Ollama/OpenAI | Orta sınıf |
| 24 GB sınıfı | 27B/35B quant | Ollama | Güçlü tek GPU |
| 32 GB sınıfı | 35B + headroom | Ollama | Üst tüketici sınıfı |
Creator and community benchmarks vary with backend, driver, quantization, context, batch size and power limit. They are shown as real-world evidence, not guaranteed performance.
Yes, provided the quantized model and context fit the available memory. Start with the models marked comfortable.
No. KV cache, runtime workspace, memory bandwidth, GPU architecture, context and laptop power limits affect speed.
On limited VRAM, Q4 often gives a better overall experience by leaving room for context. Q8 uses substantially more memory.
No. System RAM can be used for offload but does not increase discrete GPU VRAM. Unified-memory systems are different.
Ollama is primarily for LLM/vision inference. Image models such as FLUX, Z-Image and Qwen-Image typically use ComfyUI or Diffusers.
Resolution, frames, precision, VAE and model-card hardware guidance matter in addition to parameter count.
It is useful when capacity matters more than latency. For interactive workloads, larger VRAM is preferable when offload becomes the bottleneck.
No. Use only the context you need; KV cache grows with context and can consume the remaining memory.
Send the exact checkpoint size, target context and concurrent user count and we can size VRAM/RAM correctly.