Recipes
Browse every known recipe for running models on local hardware, filter by machine to find the one that works on yours.
| Hardware | Model | Decode tok/s | Prompt processing | Recipe | Runs |
|---|---|---|---|---|---|
| Hardware 8× RTX Pro 6000 Blackwell server (768 GB) | Model kimi-k2-5-1t-moe @ INT4 (FP8 KV, DCP=8) on vLLM | Decode tok/s batch ~900 @ 40K (100-conc. aggregate) | Prompt processing — | Recipe local-inference-lab/rtx6kpro wiki (Kimi K2.5 high concurrency, Festr) | Runs |
| Hardware Dual RTX Pro 6000 Blackwell build | Model qwen3-6-27b-dense @ FP8 (fp8 KV, MTP spec=3) on vLLM | Decode tok/s batch ~894 @ 175K (32-conc. aggregate) | Prompt processing — | Recipe theogravity/dual-rtx-6000-blackwell-qwen3.6-27b-fp8 (coding sweep, seqs=32) | Runs |
| Hardware 8× H100 80 GB server | Model deepseek-v3-671b-moe @ FP8 (TP=8) on vLLM | Decode tok/s batch ~620 @ 1.024K (100-conc. aggregate) | Prompt processing — | Recipe dzhsurf/deepseek-v3-r1-deploy-and-benchmarks (8xH100 vLLM TP=8, ~100 concurrency) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model gemma4-26b-moe @ native on vLLM cluster TP=2 (triton, ROCm) | Decode tok/s batch ~411 @ 4K (200-conc. aggregate) | Prompt processing — | Recipe kyuz0 amd-strix-halo-vllm-toolboxes (triton cluster tp2 throughput, 200 reqs) | Runs |
| Hardware 8× RTX Pro 6000 Blackwell server (768 GB) | Model qwen3-5-397b-a17b-moe @ NVFP4 on SGLang+MTP | Decode tok/s ~350 @ 4K | Prompt processing — | Recipe local-inference-lab/rtx6kpro wiki (Qwen3.5-397B 8x single-batch) | Runs |
| Hardware 8× H100 80 GB server | Model qwen3-6-35b-a3b-moe @ FP8 on vLLM | Decode tok/s ~320 @ 128K | Prompt processing — | Recipe vLLM Recipes | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model qwen3-6-35b-a3b-moe @ AWQ-4bit / native on vLLM cluster TP=2 (aiter, ROCm) | Decode tok/s batch ~287 @ 4K (200-conc. aggregate) | Prompt processing — | Recipe kyuz0 amd-strix-halo-vllm-toolboxes (aiter cluster tp2 throughput, Dec 2025) | Runs |
| Hardware Single RTX 5090 build | Model qwen3-coder-30b @ Q4_K on llama.cpp (CUDA) | Decode tok/s ~226 @ 4K | Prompt processing ~7093 @ 4K | Recipe hardware-corner.net RTX 5090 LLM benchmarks (GGUF Q4) | Runs |
| Hardware 8× H100 80 GB server | Model mistral-medium-3-5-128b @ FP8 on vLLM | Decode tok/s ~220 @ 128K | Prompt processing — | Recipe vLLM Recipes | Runs |
| Hardware Single RTX 5090 build | Model gemma4-26b-moe @ Q4_K on llama.cpp (CUDA) | Decode tok/s ~180 @ 4K | Prompt processing ~8799 @ 4K | Recipe hardware-corner.net RTX 5090 LLM benchmarks (GGUF Q4) | Runs |
| Hardware Single H100 80 GB workstation | Model qwen3-6-35b-a3b-moe @ FP8 on vLLM | Decode tok/s ~180 @ 128K | Prompt processing — | Recipe vLLM Recipes | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ FP8 (FP8 KV, MTP n=2) on vLLM | Decode tok/s batch ~179.9 @ 393.216K (8-conc. aggregate) | Prompt processing — | Recipe NVIDIA Developer Forum 373808 (jasl vLLM TP=4, n=8 aggregate) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-35b-a3b-moe @ AWQ-4bit / native on vLLM (aiter, ROCm) | Decode tok/s batch ~178 @ 4K (200-conc. aggregate) | Prompt processing — | Recipe kyuz0 amd-strix-halo-vllm-toolboxes (aiter tp1 throughput, Dec 2025) | Runs |
| Hardware Single RTX Pro 6000 Blackwell 96 GB build | Model qwen3-6-35b-a3b-moe @ FP8 on vLLM | Decode tok/s ~170 @ 262K | Prompt processing — | Recipe GitHub lastloop-ai | Runs |
| Hardware Dual RTX Pro 6000 Blackwell build | Model qwen3-6-27b-dense @ NVFP4 on vLLM+MTP | Decode tok/s ~156 @ 262K | Prompt processing ~831 @ 262K | Recipe loFT LLC | Runs |
| Hardware Single RTX 3090 (used) build | Model qwen3-coder-30b @ Q4_K on llama.cpp (CUDA) | Decode tok/s ~153 @ 4K | Prompt processing ~2988 @ 4K | Recipe hardware-corner.net RTX 3090 LLM benchmarks (GGUF Q4) | Runs |
| Hardware Quad RTX Pro 6000 Blackwell build (384 GB) | Model qwen3-5-397b-a17b-moe @ AWQ-INT4 (QuantTrio) on SGLang+MTP | Decode tok/s ~152 @ 4K | Prompt processing — | Recipe local-inference-lab/rtx6kpro wiki (Qwen3.5-397B single-batch decode) | Runs |
| Hardware Dual RTX 3090 (used) build | Model qwen3-6-35b-a3b-moe @ AWQ on vLLM | Decode tok/s ~149 @ 4K | Prompt processing — | Recipe GitHub - tfriedel (RTX 3090 lab) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-35b-a3b-moe @ ROCmFP4 (CHADROCK) on llama-server+ROCmFPX | Decode tok/s ~140 @ 4K | Prompt processing — | Recipe GitHub hogeheer499-commits/strix-halo-guide | Runs |
| Hardware Dual RTX Pro 6000 Blackwell build | Model qwen3-6-27b-dense @ FP8 (native, fp8 KV, MTP spec=3) on vLLM | Decode tok/s ~137 @ 175K | Prompt processing — | Recipe theogravity/dual-rtx-6000-blackwell-qwen3.6-27b-fp8 (benchmark sweep) | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model mistral-small-4-119b-moe @ NVFP4 on vLLM | Decode tok/s batch ~131 @ 262.144K (20-conc. aggregate) | Prompt processing — | Recipe Sebastien67 Medium (DGX Spark vLLM NVFP4, n=20 aggregate) | Runs |
| Hardware Quad RTX Pro 6000 Blackwell build (384 GB) | Model qwen3-5-397b-a17b-moe @ NVFP4 (nvidia checkpoint) on vLLM+MTP | Decode tok/s ~130 @ 4K | Prompt processing — | Recipe local-inference-lab/rtx6kpro wiki (Qwen3.5-397B MTP scaling table, concurrency=1) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model gemma-4-31b @ native (bf16/fp16) on vLLM cluster TP=2 (triton, ROCm) | Decode tok/s batch ~128 @ 4K (200-conc. aggregate) | Prompt processing — | Recipe kyuz0 amd-strix-halo-vllm-toolboxes (triton cluster tp2 throughput, 200 reqs) | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model minimax-m2-7-230b-moe @ MiniMax-M2.5-NVFP4 (modelopt_fp4, fp8 KV) on SGLang+MTP | Decode tok/s batch ~124 @ 196.608K (8-conc. aggregate) | Prompt processing — | Recipe NVIDIA Developer Forum 373676 (SGLang TP=4 EP=4, n=8 aggregate) | Runs |
| Hardware Single RTX Pro 6000 Blackwell 96 GB build | Model qwen3-next-80b-moe @ Q4_K_M on ollama / llama.cpp (CUDA) | Decode tok/s ~124 @ 4K | Prompt processing ~3274 @ 4K | Recipe vaditaslim.com RTX PRO 6000 Blackwell 8-model benchmarks | Runs |
| Hardware Single RTX 3090 (used) build | Model gemma4-26b-moe @ Q4_K on llama.cpp (CUDA) | Decode tok/s ~119 @ 4K | Prompt processing ~3625 @ 4K | Recipe hardware-corner.net RTX 3090 LLM benchmarks (GGUF Q4) | Runs |
| Hardware Single H100 80 GB workstation | Model qwen3-6-27b-dense @ FP16 on vLLM | Decode tok/s ~110 @ 128K | Prompt processing — | Recipe vLLM Recipes | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model qwen3-5-122b-a10b-moe @ cyankiwi AWQ-4bit on vLLM cluster TP=2 (aiter, ROCm) | Decode tok/s batch ~104 @ 4K (200-conc. aggregate) | Prompt processing — | Recipe kyuz0 amd-strix-halo-vllm-toolboxes (aiter cluster tp2 throughput, Dec 2025) | Runs |
| Hardware Single RTX 3090 (used) build | Model qwen3-6-35b-a3b-moe @ Q4_K_XL on llama.cpp | Decode tok/s ~101 @ 65K | Prompt processing ~1171 @ 0.5K | Recipe aminrj.com (Qwen3.6 on 24GB) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-35b-a3b-moe @ IQ4_XS-Q8nextn on llama-server+MTP | Decode tok/s ~101 @ 4K | Prompt processing — | Recipe GitHub hogeheer499-commits/strix-halo-guide | Runs |
| Hardware 8× RTX Pro 6000 Blackwell server (768 GB) | Model kimi-k2-5-1t-moe @ INT4 (BF16 KV, EP=8, overclocked GDDR7) on SGLang+MTP | Decode tok/s ~101 @ 4K | Prompt processing — | Recipe local-inference-lab/rtx6kpro wiki (Kimi K2.5 8x single-batch decode) | Runs |
| Hardware Single RTX Pro 6000 Blackwell 96 GB build | Model qwen3-6-27b-dense @ INT4 (AutoRound) + MTP n=3, FP8 KV on vLLM (flashinfer, MTP) | Decode tok/s ~100 @ 262K | Prompt processing — | Recipe GitHub lastloop-ai | Runs |
| Hardware 8× RTX Pro 6000 Blackwell server (768 GB) | Model glm-51-754b-moe @ NVFP4-MTP (lukealonso/GLM-5.1-NVFP4-MTP, served as GLM-5) on SGLang+MTP | Decode tok/s ~100 @ 4K | Prompt processing — | Recipe local-inference-lab/rtx6kpro wiki (GLM-5 single-batch decode; models/glm5.md = GLM-5.1) | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model qwen3-6-35b-a3b-moe @ NVFP4 on vLLM+DFlash | Decode tok/s ~97 @ 0.5K | Prompt processing ~9090 @ 0.5K (derived) | Recipe GitHub AEON-7 | Runs |
| Hardware MacBook Pro M5 Max 64 GB | Model gemma-4-12b @ MLX NVFP4 on Ollama 0.31 (MLX) + MTP | Decode tok/s ~95 @ 4K | Prompt processing — | Recipe Ollama blog (framework-author first-party; M5 Max, Aider polyglot) | Runs |
| Hardware Single RTX 5090 build | Model qwen3-6-27b-dense @ NVFP4 on vLLM | Decode tok/s ~92 @ 200K | Prompt processing ~5300 @ 47K | Recipe GitHub devnen | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model qwen3-6-35b-a3b-moe @ NVFP4 on vLLM | Decode tok/s ~90 @ 43K | Prompt processing ~2133 @ 32K (derived) | Recipe GitHub technigmaai/dgx-spark | Runs |
| Hardware Dual RTX 3090 (used) build | Model qwen3-6-27b-dense @ AWQ on vLLM | Decode tok/s ~90 @ 100K | Prompt processing — | Recipe GitHub - tfriedel (RTX 3090 lab) | Runs |
| Hardware MacBook Pro M5 Max 64 GB | Model qwen3-6-35b-a3b-moe @ MLX-4bit on MLX-LM | Decode tok/s ~87 @ 4K | Prompt processing ~2447 @ 4K | Recipe oMLX Benchmark | Runs |
| Hardware Dual RTX Pro 6000 Blackwell build | Model minimax-m2-7-230b-moe @ MiniMax-M2.5-NVFP4 on vLLM | Decode tok/s ~85 @ 4K | Prompt processing — | Recipe local-inference-lab/rtx6kpro wiki (MiniMax-M2.5 single-stream table) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-35b-a3b-moe @ Q4_0 on llama.cpp | Decode tok/s ~81 @ 4K | Prompt processing ~1244 @ 0.5K | Recipe GitHub hogeheer499-commits/strix-halo-guide | Runs |
| Hardware Quad RTX Pro 6000 Blackwell build (384 GB) | Model minimax-m2-7-230b-moe @ MiniMax-M2.5-FP8 on vLLM | Decode tok/s ~81 @ 20K | Prompt processing — | Recipe local-inference-lab/rtx6kpro wiki (MiniMax-M2.5 single-stream table) | Runs |
| Hardware DGX B200 — 8× B200 server (1.44 TB HBM3e) | Model nemotron-3-ultra-550b-a55b-moe @ NVFP4 + FP8 KV on Dynamo + vLLM (TP=4, expert-parallel, MTP) | Decode tok/s batch ~80.6 @ 4K (20-conc. aggregate) | Prompt processing — | Recipe NVIDIA ai-dynamo/dynamo recipes (B200 TP4+EP, NVFP4+FP8, MTP) | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model minimax-m3-428b-moe @ MiniMax-M3-AWQ-INT4 (fp8 KV, EAGLE3) on vLLM | Decode tok/s batch ~79 @ 262.144K (4-conc. aggregate) | Prompt processing — | Recipe NVIDIA Developer Forum 375361 (vLLM TP=4, n=4 aggregate) | Runs |
| Hardware Single AMD Radeon AI Pro R9700 32 GB build | Model qwen3-6-35b-a3b-moe @ Q4_K_M on llama.cpp | Decode tok/s ~77 @ 4K | Prompt processing ~1636 @ 32K | Recipe GitHub truelies444 | Runs |
| Hardware 8× H100 80 GB server | Model kimi-k2-6-1t-moe @ FP8 on vLLM | Decode tok/s ~75 @ 256K | Prompt processing — | Recipe HF - RedHatAI (Kimi-K2.6-FP8-BLOCK) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-35b-a3b-moe @ MTP-GGUF UD-Q4_K_XL (draft-mtp n=3) on llama.cpp (Vulkan RADV, MTP) | Decode tok/s ~75 @ 0.5K | Prompt processing — | Recipe kyuz0 amd-strix-halo-toolboxes MTP grid (results-mtp/summary.json, 15 May 2026) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model qwen3-5-122b-a10b-moe @ cyankiwi AWQ-8bit on vLLM cluster TP=2 (aiter, ROCm) | Decode tok/s batch ~74 @ 4K (200-conc. aggregate) | Prompt processing — | Recipe kyuz0 amd-strix-halo-vllm-toolboxes (aiter cluster tp2 throughput, 200 reqs) | Runs |
| Hardware Dual AMD Radeon AI Pro R9700 build (64 GB) | Model qwen3-6-35b-a3b-moe @ Q6_K on llama.cpp | Decode tok/s ~72 @ 4K | Prompt processing ~3038 @ 32K | Recipe GitHub truelies444 | Runs |
| Hardware Single RTX 3090 (used) build | Model qwen3-6-27b-dense @ AWQ/AutoRound-INT4 on vLLM+MTP | Decode tok/s ~72 @ 32K | Prompt processing — | Recipe GitHub devnen | Runs |
| Hardware Single RTX 4090 build | Model qwen3-coder-30b @ Q4_K on llama.cpp (CUDA) | Decode tok/s ~68 @ 64K | Prompt processing ~1502 @ 64K | Recipe hardware-corner.net RTX 4090 LLM benchmarks (GGUF Q4) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ not stated (native precision) on vLLM (Aidendle94/B12X-MoE, TP=2 RoCE) + DSpark spec-decode | Decode tok/s ~65 @ 200K | Prompt processing — | Recipe GitHub 0rand (DeepSeek-V4 DSpark serving stack) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ NVFP4-KV (nvfp4_ds_mla) on vLLM+DSpark | Decode tok/s ~63 @ 200K | Prompt processing — | Recipe NVIDIA Developer Forum (374846) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-35b-a3b-moe @ Q4_K_M (UD) on llama.cpp | Decode tok/s ~62 @ 4K | Prompt processing ~1059 @ 0.5K | Recipe GitHub hogeheer499-commits/strix-halo-guide | Runs |
| Hardware 8× Strix Halo cluster (1024 GB unified) | Model qwen3-6-35b-a3b-moe @ Q8_0 on llama.cpp | Decode tok/s ~62 @ 4K | Prompt processing — | Recipe GitHub - strix-halo-guide | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model qwen3-6-35b-a3b-moe @ FP8 on vLLM | Decode tok/s ~60 @ 32K | Prompt processing ~6520 @ 8K (derived) | Recipe NVIDIA Developer Forum (366822) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-35b-a3b-moe @ UD-Q4_K_XL on llama.cpp (Vulkan RADV) | Decode tok/s ~60 @ 0.5K | Prompt processing ~1114 @ 0.5K | Recipe kyuz0 amd-strix-halo-toolboxes grid (docs/results.json, 16 May 2026) | Runs |
| Hardware DGX H200 — 8× H200 server (1.13 TB HBM3e) | Model nemotron-3-ultra-550b-a55b-moe @ NVFP4 + FP8 KV on Dynamo + vLLM (TP=8, expert-parallel, MTP) | Decode tok/s batch ~58.7 @ 4K (10-conc. aggregate) | Prompt processing — | Recipe NVIDIA ai-dynamo/dynamo recipes (8xH200 TP8+EP, NVFP4+FP8, MTP) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model minimax-m2-7-230b-moe @ cyankiwi AWQ-4bit on vLLM cluster TP=2 (aiter, ROCm) | Decode tok/s batch ~57 @ 4K (200-conc. aggregate) | Prompt processing — | Recipe kyuz0 amd-strix-halo-vllm-toolboxes (aiter cluster tp2 throughput, 200 reqs) | Runs |
| Hardware Single Intel Arc Pro B70 build | Model qwen3-6-35b-a3b-moe @ Q4_K_M (UD) on llama.cpp (SYCL) | Decode tok/s ~55 @ 4K | Prompt processing ~615 @ 0.5K | Recipe GitHub PMZFX | Runs |
| Hardware Tesla V100 32 GB SXM2 mod build | Model qwen3-6-35b-a3b-moe @ Q5_K_M on llama.cpp | Decode tok/s ~55 @ 10K | Prompt processing — | Recipe GitHub - ai-bond (V100 flash-attn) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model mimo-v2-5-310b-a15b-moe @ NVFP4 on vLLM+DFlash | Decode tok/s ~54 @ 0.5K | Prompt processing ~2083 @ 50K (derived) | Recipe GitHub HeNryous (renek) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model gemma4-26b-moe @ UD-Q4_K_XL on llama.cpp (Vulkan RADV) | Decode tok/s ~54 @ 0.5K | Prompt processing ~1324 @ 0.5K | Recipe kyuz0 amd-strix-halo-toolboxes grid (docs/results.json, 16 May 2026) | Runs |
| Hardware MacBook Pro M5 Max 64 GB | Model gemma-4-12b @ MLX NVFP4 on Ollama 0.31 (MLX), no MTP | Decode tok/s ~50.2 @ 4K | Prompt processing — | Recipe Ollama blog (framework-author first-party; M5 Max) | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model qwen3-6-35b-a3b-moe @ FP8 on vLLM+DFlash | Decode tok/s ~50 @ 262K | Prompt processing ~4932 @ 0.5K | Recipe GitHub ZengboJamesWang | Runs |
| Hardware 4× Strix Halo cluster (512 GB unified) | Model qwen3-6-35b-a3b-moe @ Q8_0 on llama.cpp | Decode tok/s ~50 @ 4K | Prompt processing — | Recipe Frame.work Community | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ FP8 (official weights, FP8 KV, MTP n=2) on vLLM | Decode tok/s ~49.4 @ 393.216K | Prompt processing — | Recipe NVIDIA Developer Forum 373808 (jasl vLLM TP=4) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ FP8+FP4 mixed (FP8 KV) on vLLM+MTP | Decode tok/s ~45.5 @ 1000K | Prompt processing ~786 @ 800K | Recipe GitHub tonyd2wild | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model mimo-v2-5-310b-a15b-moe @ NVFP4 on vLLM+DFlash | Decode tok/s ~45 @ 131K | Prompt processing — | Recipe NVIDIA Developer Forum (375607) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model mimo-v2-5-310b-a15b-moe @ NVFP4 on vLLM+DFlash | Decode tok/s ~45 @ 26K | Prompt processing ~5540 @ 0.256K (derived) | Recipe NVIDIA Developer Forum (375923) | Runs |
| Hardware Single H100 80 GB workstation | Model mistral-medium-3-5-128b @ FP8 on vLLM | Decode tok/s ~45 @ 128K | Prompt processing — | Recipe vLLM Recipes | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ FP8+FP4 mixed (FP8 KV) on vLLM+MTP | Decode tok/s ~40 @ 500K | Prompt processing — | Recipe GitHub tonyd2wild | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model qwen3-5-397b-a17b-moe @ NVFP4 on SGLang v0.5.12 + MTP | Decode tok/s ~40 @ 32K | Prompt processing — | Recipe NVIDIA Developer Forum 366325 (ht12) + the dgxarley 4-node test matrix that thread publishes - SGLang v0.5.12 base image, driver 580.159.03, case 18 (flashinfer_cutlass MoE + triton attention + full CUDA graphs + MTP NEXTN s3/d4), Qwen3.5-397B-A17B-NVFP4, 4x GB10 TP=4 EP=1, measured 2026-06-19 | Runs |
| Hardware 8× DGX Spark cluster (1024 GB unified, CUDA) | Model mimo-v2-5-pro-1t-moe @ NVFP4 on vLLM+MTP | Decode tok/s ~40 @ 1K | Prompt processing ~1950 @ 2K | Recipe NVIDIA Developer Forum (370803) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model qwen3-6-35b-a3b-moe @ Q8_0 on llama.cpp | Decode tok/s ~40 @ 4K | Prompt processing — | Recipe Frame.work Community | Runs |
| Hardware 8× DGX Spark cluster (1024 GB unified, CUDA) | Model qwen3-5-397b-a17b-moe @ FP8 (406 GiB) on vLLM | Decode tok/s ~39.5 @ 32K | Prompt processing — | Recipe NVIDIA Developer Forum 369446 (vLLM eugr fork TP=8) | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model qwen3-6-27b-dense @ Q4_K_M on llama.cpp+DFlash | Decode tok/s ~38 @ 256K | Prompt processing — | Recipe GitHub phuongncn | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model mimo-v2-5-310b-a15b-moe @ NVFP4 4-bit weights + NVFP4 4-bit KV cache on vLLM (TP=2 Ray) + DFlash spec-decode | Decode tok/s ~37.8 @ 1000K | Prompt processing — | Recipe GitHub tonyd2wild (MiMo V2.5 DFlash 1M NVFP4-KV) | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model minimax-m3-428b-moe @ MiniMax-M3-MXFP4 (bf16 KV, EAGLE3 k=2) on vLLM | Decode tok/s ~34.8 @ 262.144K | Prompt processing ~2020 @ 262.144K | Recipe NVIDIA Developer Forum 375386 (vLLM TP=4 EAGLE3) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model mimo-v2-5-310b-a15b-moe @ NVFP4 on vLLM+MTP | Decode tok/s ~34 @ 0.5K | Prompt processing ~2609 @ 2K (derived) | Recipe NVIDIA Developer Forum (370459) | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model minimax-m3-428b-moe @ MiniMax-M3-AWQ-INT4 (fp8 KV, EAGLE3) on vLLM | Decode tok/s ~33.7 @ 262.144K | Prompt processing — | Recipe NVIDIA Developer Forum 375361 (vLLM TP=4 EAGLE3) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ FP8+FP4 mixed (FP8 KV) on vLLM | Decode tok/s ~33 @ 32K | Prompt processing ~512 @ 128K (derived) | Recipe NVIDIA Developer Forum (370309) | Runs |
| Hardware 8× H100 80 GB server | Model deepseek-v3-671b-moe @ FP8 (671B, TP=8) on vLLM | Decode tok/s ~33 @ 1.024K | Prompt processing — | Recipe dzhsurf/deepseek-v3-r1-deploy-and-benchmarks (8xH100 vLLM TP=8, concurrency=1) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model minimax-m2-7-230b-moe @ UD-Q3_K_S on llama.cpp (Vulkan RADV) | Decode tok/s ~31 @ 0.5K | Prompt processing ~243 @ 0.5K | Recipe kyuz0 amd-strix-halo-toolboxes grid (docs/results.json, 16 May 2026) | Runs |
| Hardware Dual RTX Pro 6000 Blackwell build | Model mistral-medium-3-5-128b @ FP8 on vLLM | Decode tok/s ~30 @ 98K | Prompt processing — | Recipe HF mistralai discussion #17 | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model qwen3-6-27b-dense @ Q4_K_M on llama.cpp+MTP | Decode tok/s ~28 @ 2K | Prompt processing ~1084 @ 2K | Recipe NVIDIA Developer Forum (370298) | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model mistral-small-4-119b-moe @ NVFP4 on vLLM | Decode tok/s ~27.8 @ 262.144K | Prompt processing — | Recipe Sebastien67 Medium (first-hand DGX Spark vLLM NVFP4 run) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model minimax-m2-7-230b-moe @ MiniMax-M2.5-REAP Q4_K_M (GGUF, pruned REAP variant) on llama.cpp (Vulkan RADV, RPC cluster) | Decode tok/s ~26.7 @ 0.512K | Prompt processing ~272 @ 0.512K | Recipe visorcraft/strix-halo-llm-perf (2-node RPC llama-bench, 2026-02-19) | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model minimax-m2-7-230b-moe @ MiniMax-M2.5-NVFP4 (modelopt_fp4, fp8 KV) on SGLang+MTP | Decode tok/s ~25.5 @ 196.608K | Prompt processing — | Recipe NVIDIA Developer Forum 373676 (SGLang TP=4 EP=4) | Runs |
| Hardware Mac Mini M4 (24 GB) | Model qwen3-6-35b-a3b-moe @ MLX-4bit on MLX-LM | Decode tok/s ~25 est. @ 4K | Prompt processing — | Recipe maloyan.xyz (M4 16GB, scaled) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model minimax-m2-7-230b-moe @ MiniMax-M2.5-REAP MXFP4_MOE (GGUF) on llama.cpp (Vulkan RADV, RPC cluster) | Decode tok/s ~24.5 @ 0.512K | Prompt processing ~299.5 @ 0.512K | Recipe visorcraft/strix-halo-llm-perf (2-node RPC llama-bench, 2026-02-19) | Runs |
| Hardware Single AMD Radeon AI Pro R9700 32 GB build | Model qwen3-6-27b-dense @ Q5_K_M on llama.cpp | Decode tok/s ~24 @ 4K | Prompt processing ~611 @ 32K | Recipe GitHub truelies444 | Runs |
| Hardware Dual AMD Radeon AI Pro R9700 build (64 GB) | Model qwen3-6-35b-a3b-moe @ base weights (fp8 KV) on vLLM | Decode tok/s ~22.8 @ 1.036K | Prompt processing ~4600 @ 1.036K | Recipe mlai.blog (Qwen3.5-35B-A3B on dual R9700, ROCm vLLM) | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model glm-5-2-753b-moe @ AWQ-INT4 (cyankiwi/GLM-5.2-AWQ-INT4, 15% data-free expert pruning) on vLLM+MTP | Decode tok/s ~22 @ 8.192K | Prompt processing ~535 @ 8.192K | Recipe NVIDIA Developer Forum 374125 (CosmicRaisins, AWQ-INT4 TP=4 MTP) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-5-122b-a10b-moe @ UD-Q5_K_XL on llama.cpp (Vulkan RADV) | Decode tok/s ~22 @ 0.5K | Prompt processing ~337 @ 0.5K | Recipe kyuz0 amd-strix-halo-toolboxes grid (docs/results.json, 16 May 2026) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-27b-dense @ UD-Q4_K_M (draft-mtp n=3) on llama.cpp (ROCm, MTP) | Decode tok/s ~21 @ 0.5K | Prompt processing — | Recipe Caleb Coffie - benchmarking llama.cpp MTP on Strix Halo | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model llama-4-scout @ UD-Q4_K_XL on llama.cpp (Vulkan RADV) | Decode tok/s ~20 @ 0.5K | Prompt processing ~103 @ 0.5K | Recipe hardware-corner.net Strix Halo optimization benchmarks | Runs |
| Hardware 8× DGX Spark cluster (1024 GB unified, CUDA) | Model kimi-k2-6-1t-moe @ NVFP4 on vLLM | Decode tok/s ~18 @ 32K | Prompt processing — | Recipe NVIDIA Developer Forum | Runs |
| Hardware Mac Mini M4 (16 GB) | Model qwen3-6-35b-a3b-moe @ MLX-4bit on MLX-LM | Decode tok/s ~17 @ 4K | Prompt processing — | Recipe maloyan.xyz (M4 16GB) | Runs |
| Hardware Tesla V100 32 GB SXM2 mod build | Model qwen3-6-27b-dense @ Q5_K_M on llama.cpp | Decode tok/s ~17 @ 32K | Prompt processing — | Recipe hardware-corner.net (V100 32GB guide) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model qwen3-6-27b-dense @ Q5_K_M on llama.cpp | Decode tok/s ~16 @ 4K | Prompt processing — | Recipe llm-tracker.info (kyuz0) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model deepseek-v4-flash-284b-moe @ IQ2_XXS-w2Q2K imatrix (~80.8 GB) on ds4 (antirez DeepSeek-V4-Flash engine) + MTP, ROCm 7.2.4 gfx1151 | Decode tok/s ~15.25 @ 2K | Prompt processing ~152 @ 2K (derived) | Recipe kyuz0 ds4 Strix Halo toolbox (ds4-bench, single 128GB node); antirez ds4 engine | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model deepseek-v4-flash-284b-moe @ Hybrid Q2/Q4 imatrix (layers 37-42 Q4, ~97 GB) on ds4 (antirez engine) + MTP, ROCm 7.2.4 gfx1151 | Decode tok/s ~15.02 @ 2K | Prompt processing ~138 @ 2K | Recipe kyuz0 ds4 Strix Halo toolbox (ds4-bench, single-node hybrid Q2/Q4) | Runs |
| Hardware MacBook Air M4 (16 GB) | Model qwen3-6-35b-a3b-moe @ MLX-4bit on MLX-LM | Decode tok/s ~15 est. @ 4K | Prompt processing — | Recipe maloyan.xyz (M4 16GB, fanless-adjusted) | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model glm-5-2-753b-moe @ NVFP4 (REAP-less, high-quality 4-bit; cf. nvidia/GLM-5.2-NVFP4 model card) on vLLM | Decode tok/s ~15 @ 131.072K | Prompt processing ~500 @ 131.072K | Recipe NVIDIA Developer Forum 374832 (REAP-less NVFP4, custom vLLM fork TP=4) | Runs |
| Hardware Quad RTX 3090 (used) build | Model devstral-2-123b @ IQ4_KSS (GGUF) on ik_llama.cpp (-sm graph, tensor-parallel) | Decode tok/s ~15 est. @ 4K | Prompt processing ~300 @ 2K | Recipe HF ubergarm Devstral-2-123B-GGUF discussion #2 (phakio, ik_llama.cpp 4-GPU) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model mimo-v2-5-310b-a15b-moe @ UD-Q4/Q5_K_XL GGUF (~180-215 GB, split across 2 nodes) on llama.cpp RPC (2x Strix Halo 128GB, ROCm, USB4net secondary link) | Decode tok/s ~15 @ 10K | Prompt processing ~356 @ 10K (derived) | Recipe r/LocalLLaMA operator report (2x Strix Halo 128GB, llama.cpp RPC over USB4net); AesSedai/unsloth MiMo-V2.5 GGUF | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model nemotron-3-super-120b-a12b-moe @ UD-Q4_K_XL on llama.cpp (ROCm 7.2.3) | Decode tok/s ~14 @ 0.5K | Prompt processing ~276 @ 0.5K | Recipe kyuz0 amd-strix-halo-toolboxes grid (docs/results.json, 16 May 2026) | Runs |
| Hardware 8× DGX Spark cluster (1024 GB unified, CUDA) | Model kimi-k2-6-1t-moe @ NVFP4 (60 shards, ~554 GiB, no spec-decode) on vLLM | Decode tok/s ~13.5 @ 32.768K | Prompt processing — | Recipe NVIDIA Developer Forum 369446 (vLLM eugr fork TP=8) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model deepseek-v4-flash-284b-moe @ Q4 imatrix distributed (~153.3 GB, Q4 experts) on ds4 multi-node (pipeline-parallel, 2x Strix, ROCm 7.2.4 gfx1151) + MTP | Decode tok/s ~13.01 @ 2K | Prompt processing ~62 @ 2K (derived) | Recipe kyuz0 ds4 Strix Halo toolbox (ds4-bench, 2-node distributed Q4) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model glm-5-2-753b-moe @ low-bit (vLLM TP=2) on vLLM (TP over 2 Sparks) | Decode tok/s ~12 @ 40K | Prompt processing — | Recipe NVIDIA Developer Forum 374523 (GLM-5.2 vLLM TP=2 update) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model gemma-4-31b @ UD-Q4_K_XL on llama.cpp (Vulkan RADV) | Decode tok/s ~11 @ 0.5K | Prompt processing ~302 @ 0.5K | Recipe kyuz0 amd-strix-halo-toolboxes grid (docs/results.json, 16 May 2026) | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model glm-5-2-753b-moe @ UD-IQ1_S (~2.3 bpw, 1-bit) on llama.cpp RPC (tensor-split over 2 nodes) | Decode tok/s ~8 @ 2K | Prompt processing ~213 @ 2K | Recipe NVIDIA Developer Forum 374523 (GLM-5.2 on 2x DGX Spark, 1-bit llama.cpp RPC) | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model glm-5-2-753b-moe @ IQ4_XS (GGUF, ~365GB across 4 nodes, DSA sparse attention active) on llama.cpp (RPC multi-node) | Decode tok/s ~6.28 @ 1048.576K | Prompt processing ~222 @ 1048.576K | Recipe NVIDIA Developer Forum 373933 (IQ4_XS llama.cpp RPC, DSA active) | Runs |
| Hardware RTX 3060 12 GB build | Model gemma-4-12b @ Q4_K_M (GGUF) on llama.cpp (Vulkan) | Decode tok/s ~5 @ 4K | Prompt processing — | Recipe Hacker News Gemma 4 12B launch thread (community, Vulkan-throttled) | Runs |
| Hardware 8× Strix Halo cluster (1024 GB unified) | Model kimi-k2-6-1t-moe @ Q5_K_M on llama.cpp | Decode tok/s ~5 @ 4K | Prompt processing — | Recipe Frame.work Community | Runs |
| Hardware Quad Tesla P40 (96 GB) homelab build | Model mistral-medium-3-5-128b @ Q4_K_M on llama.cpp | Decode tok/s ~4 @ 32K | Prompt processing — | Recipe GitHub - llama.cpp #12990 (P40 FA) | Runs |
| Hardware 2× Strix Halo cluster (256 GB unified) | Model mistral-medium-3-5-128b @ Q4_K_M on llama.cpp | Decode tok/s ~3 @ 4K | Prompt processing — | Recipe llm-tracker.info (kyuz0) | Runs |
| Hardware Single AMD Instinct MI50 32 GB (used) build | Model qwen3-6-35b-a3b-moe @ Q4_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe GitHub llama.cpp #19880 (MI50 enablement) | Runs |
| Hardware Quad AMD MI50 32 GB (128 GB) homelab build | Model qwen3-6-35b-a3b-moe @ Q5_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe aibytes.blog (MI50 ROCm vs Vulkan) | Runs |
| Hardware Quad AMD MI50 32 GB (128 GB) homelab build | Model mistral-medium-3-5-128b @ Q4_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe HF - unsloth (GGUF discussion) | Runs |
| Hardware Mac Mini M4 (16 GB) | Model gemma-4-12b @ MLX 4-bit on Ollama (MLX) / mlx-lm | Decode tok/s — | Prompt processing — | Recipe Ollama model library (Apple-Silicon MLX build; runnable recipe). Fit corroborated by Gemma 4 launch HN thread (Q4_K_M ~6.6 GB). | Runs |
| Hardware Mac Mini M4 (24 GB) | Model gemma-4-12b @ MLX 4-bit on Ollama (MLX) / mlx-lm | Decode tok/s — | Prompt processing — | Recipe Ollama model library (Apple-Silicon MLX build). Fit corroborated by HN launch thread. | Runs |
| Hardware MacBook Air M4 (16 GB) | Model gemma-4-12b @ MLX 4-bit on Ollama (MLX) / mlx-lm | Decode tok/s — | Prompt processing — | Recipe Ollama model library (Apple-Silicon MLX build). Fit corroborated by HN launch thread. | Runs |
| Hardware NVIDIA DGX Spark (128 GB) | Model qwen3-6-35b-a3b-moe @ NVFP4 on SGLang+MTP | Decode tok/s — | Prompt processing — | Recipe GitHub r0b0tlab | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ NVFP4-KV (nvfp4_ds_mla) on vLLM+DSpark | Decode tok/s — | Prompt processing — | Recipe HF drowzeys | Runs |
| Hardware 2× DGX Spark cluster (256 GB unified, CUDA) | Model qwen3-6-35b-a3b-moe @ FP8 on vLLM (Ray TP=2) | Decode tok/s — | Prompt processing — | Recipe Medium - Michael Peres | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model deepseek-v4-flash-284b-moe @ FP8 on vLLM | Decode tok/s — | Prompt processing — | Recipe NVIDIA Developer Forum | Runs |
| Hardware 4× DGX Spark cluster (512 GB unified, CUDA) | Model nemotron-3-ultra-550b-a55b-moe @ NVFP4 (FP8 KV, MTP) on vLLM (TP=4) | Decode tok/s — | Prompt processing — | Recipe NVIDIA-NeMo Nemotron Spark Deployment Guide (4x DGX Spark, NVFP4, vLLM TP=4) | Runs |
| Hardware 8× DGX Spark cluster (1024 GB unified, CUDA) | Model deepseek-v4-pro-1t6-moe @ FP8 on vLLM | Decode tok/s — | Prompt processing — | Recipe GitHub - vLLM #43367 | Runs |
| Hardware Single Intel Arc B580 12 GB build | Model qwen3-6-27b-dense @ Q4_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe GitHub - intel/llm-scaler (official) | Runs |
| Hardware Quad RTX 3090 (used) build | Model mistral-medium-3-5-128b @ Q4_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe HF - bartowski (GGUF) | Runs |
| Hardware Quad RTX 3090 (used) build | Model qwen3-6-35b-a3b-moe @ FP16 on vLLM | Decode tok/s — | Prompt processing — | Recipe GitHub - tfriedel (RTX 3090 lab) | Runs |
| Hardware Single RTX 4090 build | Model qwen3-6-27b-dense @ AWQ on vLLM | Decode tok/s — | Prompt processing — | Recipe GitHub - thc1006 (Ampere/Ada spec-decode) | Runs |
| Hardware Single RTX 4090 build | Model qwen3-6-35b-a3b-moe @ Q4_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe GitHub - thc1006 (spec-decode) | Runs |
| Hardware RTX 3060 12 GB build | Model qwen3-6-27b-dense @ Q4_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe GitHub - llama.cpp build docs | Runs |
| Hardware Dual RTX 5090 build | Model qwen3-6-35b-a3b-moe @ NVFP4 on vLLM | Decode tok/s — | Prompt processing — | Recipe HF - RedHatAI (NVFP4) | Runs |
| Hardware Dual RTX 5090 build | Model qwen3-6-27b-dense @ FP8 on vLLM | Decode tok/s — | Prompt processing — | Recipe HF - Qwen (official FP8) | Runs |
| Hardware Single RTX Pro 6000 Blackwell 96 GB build | Model mistral-medium-3-5-128b @ NVFP4 on vLLM | Decode tok/s — | Prompt processing — | Recipe HF nvidia (NVFP4 card) | Runs |
| Hardware Single Tesla P40 24 GB (used) build | Model qwen3-6-27b-dense @ Q4_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe GitHub llama.cpp #19248 (P40) | Runs |
| Hardware Quad Tesla P40 (96 GB) homelab build | Model qwen3-6-35b-a3b-moe @ Q5_K_M on llama.cpp | Decode tok/s — | Prompt processing — | Recipe GitHub - llama.cpp #12990 (P40 FA) | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model mimo-v2-5-310b-a15b-moe @ UD-IQ2_M (~2.7 bpw, ~92.8 GB) on llama.cpp (Vulkan/RADV, kyuz0 container), gfx1151 | Decode tok/s — | Prompt processing ~31 @ 0.5K (derived) | Recipe hogeheer499 strix-halo-guide community evidence map (Corsair AI WS 300, IQ2_M, capacity row); bartowski MiMo GGUF | Runs |
| Hardware AMD Ryzen AI Max+ 395 (128 GB) | Model gemma-4-12b @ UD-Q4_K_XL (GGUF) on llama.cpp (ROCm 7.2.x, gfx1151) | Decode tok/s — | Prompt processing — | Recipe kyuz0 Strix Halo toolboxes (ROCm gfx1151 llama.cpp recipe; 12B not yet in grid) | Runs |
| Hardware 4× Strix Halo cluster (512 GB unified) | Model qwen3-5-397b-a17b-moe @ Q-family GGUF (HIP+RPC, np2, ctx 200k) on llama.cpp | Decode tok/s — | Prompt processing — | Recipe visorcraft/strix-halo-llm-perf (Qwen3.5-397B RPC shape-control, 2026-02-21) | Runs |
| Hardware 4× Strix Halo cluster (512 GB unified) | Model deepseek-v4-flash-284b-moe @ Q4_K-class across 4 nodes on llama.cpp RPC (4x Framework Desktop / Strix mainboards) | Decode tok/s — | Prompt processing — | Recipe frame.work llama.cpp RPC multi-node recipe (extended to 4 nodes) | Runs |
| Hardware 4× Strix Halo cluster (512 GB unified) | Model mimo-v2-5-310b-a15b-moe @ Q4_K_M (~178 GB) / Q5_K_M (~213 GB) on llama.cpp RPC (4x Framework Desktop / Strix mainboards) | Decode tok/s — | Prompt processing — | Recipe bartowski MiMo-V2.5 GGUF (Q4_K_M/Q5) + frame.work llama.cpp RPC 4-node | Runs |
| Hardware MacBook Pro M5 Pro 48 GB | Model qwen3-6-35b-a3b-moe @ MLX-4bit on MLX-LM | Decode tok/s — | Prompt processing — | Recipe GitHub ml-explore/mlx-lm | Runs |
| Hardware MacBook Pro M5 Pro 48 GB | Model gemma-4-12b @ MLX 4-bit on Ollama 0.31 (MLX) + MTP | Decode tok/s — | Prompt processing — | Recipe Ollama blog (framework-author first-party MTP recipe; M5-family). Directional only; M5 Pro not separately measured. | Runs |
| Hardware MacBook Pro M5 Pro 48 GB | Model qwen3-6-27b-dense @ MLX 4-bit on mlx-lm / Ollama (MLX) | Decode tok/s — | Prompt processing — | Recipe Ollama model library (Apple-Silicon MLX build). Fit from db.json sizeQ4=16 vs 40 GB usable. | Runs |
| Hardware DGX H200 — 8× H200 server (1.13 TB HBM3e) | Model kimi-k2-7-code-1t-moe @ native INT4 on vLLM (TP=8, expert-parallel) | Decode tok/s — | Prompt processing — | Recipe vLLM Recipes (official K2.7 Code command, 8xH200 INT4); tok/s = K2.6 same-box SGLang INT4 (paxsaroffcuts) | Runs |
No recipes match the selected hardware.