V100 Multi-GPU Server Options
4×, 6×, 8× PCIe and native SXM2 platforms for V100 inference — plus GPU efficiency analysis. Used market, mid-2025 pricing.
Quick Verdict for Home Rack
Start with the Supermicro 7048GR-TR (4× PCIe, $750 barebone) — it runs on a single 230V/16A circuit and is quiet enough for a home. If you need more VRAM, the 4028GR-TR with 6× V100 hits the sweet spot of 96 GB VRAM while staying within a dedicated 230V/16A circuit at ~2,000 W peak. Full 8× GPU or DGX-1 builds need 208V/30A power and produce extreme noise.
The V100 16GB is the best value GPU for inference: it leads both tok/s per dollar and tok/s per watt among affordable datacenter cards.
Ollama + V100 Compatibility Warning
Ollama v0.6.0+ dropped Volta (V100) CUDA support. Pin Ollama to v0.5.x, or switch to llama-server / vLLM. Your tools (Nomad, Hermes, OpenClaw, Opencode) all use the Ollama API — llama-server with --api-compatibility ollama is a drop-in replacement.
GPU Efficiency Comparison
All benchmarks: Llama 2 7B Q4_0, 128-token generation (tg128). Data from the llama.cpp CUDA benchmark thread.
| GPU | tok/s | TDP (W) | Used Price | tok/s / W | tok/s / $ | W / $ |
|---|---|---|---|---|---|---|
| V100 16GB | 129 | 250 | $155 | 0.52 | 0.83 | 1.61 |
| P100 16GB | 60 | 250 | $70 | 0.24 | 0.86 | 3.57 |
| P40 24GB | 54 | 250 | $85 | 0.22 | 0.64 | 2.94 |
| T4 16GB | 45 | 70 | $215 | 0.64 | 0.21 | 0.33 |
| 1080 Ti 11GB | 62 | 250 | $100 | 0.25 | 0.62 | 2.50 |
| V100 32GB | 129 | 250 | $550 | 0.52 | 0.23 | 0.45 |
| RTX 3090 | 160 | 350 | $700 | 0.46 | 0.23 | 0.50 |
| RTX 4090 | 187 | 450 | $1,500 | 0.42 | 0.12 | 0.30 |
| A100 80GB | 195 | 300 | $6,500 | 0.65 | 0.03 | 0.05 |
What the Efficiency Charts Tell You
V100 16GB wins on combined efficiency. It’s #3 in power efficiency (behind A100 and T4 — both far more expensive), and #2 in value-per-dollar (nearly tied with P100). The P100 is cheaper per token, but at half the speed — the V100 gets you to interactive speed (129 tok/s on 7B) while the P100 doesn’t quite (60 tok/s). The T4 is the power-sipping champion at 70W TDP, but it’s actually slower than a 1080 Ti and costs twice as much used.
The W/$ chart shows ongoing cost structure: cheap GPUs (P100, P40) have high wattage relative to their price — electricity becomes a significant cost over time. More expensive GPUs (T4, RTX 4090, A100) are capital-heavy but power-light. V100 sits in a balanced middle.
Server Platforms at a Glance
| Platform | GPUs | Form | Interconnect | Peak W | Cost | Home? |
|---|---|---|---|---|---|---|
| 7048GR-TR | 4× PCIe | 4U tower | PCIe 3.0 | ~1,200 W | $750 bare | BEST |
| 4028GR-TR (6×) | 6× PCIe | 4U rack | PCIe 3.0 | ~2,000 W | $1,960 bare | DUCT. |
| 4028GR-TR (8×) | 8× PCIe | 4U rack | PCIe 3.0 | ~2,600 W | $1,960 bare | HARD |
| 6049GP-TRT | 10× dbl-w | 4U rack | PCIe 3.0 | ~3,000+ W | $2,500+ bare | HARD |
| DGX-1 V100 | 8× SXM2 | 3U rack | NVLink mesh | ~3,200 W | $5,300+ | HARD |
| SXM2 Adapter 2× | 2× SXM2 | PCIe card | NVLink | ~600 W | $350–500 | YES |
| SXM2 Adapter 4× | 4× SXM2 | baseboard | NVLink | ~1,200 W | $600–1,000 | OK |
Barebones require CPUs, RAM, and drives separately. SXM2 adapter builds need a host system. DGX-1 prices are complete systems. „DUCT.“ = dedicated circuit required.
Option 1: 4× GPU PCIe Server
Supermicro SYS-7048GR-TR 4× PCIe
The practical choice for a home lab. Tower/rackmount convertible with the X10DRG-Q motherboard, dual LGA 2011-3, four direct-attach PCIe 3.0 x16 slots. Single 2000W PSU keeps it on one circuit. Tower mode is far quieter than rackmount servers.
Build Cost
| Configuration | Source | Price |
|---|---|---|
| Barebones (chassis + board + PSU) | eBay | ~$750 |
| + 2× Xeon E5-2680 v4 + 128 GB DDR4 | eBay | ~$1,050 |
| + 4× V100 16 GB PCIe | eBay | ~$1,570 total |
| + 4× V100 32 GB PCIe instead | eBay | ~$3,570 total |
Advantages
- + 64 GB VRAM — 30B+ Q4 models, 70B Q3
- + 3,600 GB/s aggregate bandwidth
- + ~1,200 W peak — single 230V/16A circuit
- + Tower = much quieter than rack servers
- + $750 barebones — cheapest full server
Challenges
- − Limited to 4 GPUs — no room to grow
- − PCIe 3.0 x16 only (~16 GB/s per GPU)
- − Single PSU — no redundancy
- − 4U still takes significant rack height
Option 2: 6× GPU PCIe Server
Supermicro SYS-4028GR-TR (6 of 8 slots) 6× PCIe
There’s no standard 6-GPU server — the cleanest path is the 8-slot 4028GR-TR with 6 slots populated. This drops peak power from ~2,600 W to ~2,000 W, keeping it within a dedicated 230V/16A circuit. The 2 empty slots are a future upgrade path to 8× when you get dedicated power wiring.
Build Cost
| Configuration | Source | Price |
|---|---|---|
| Barebones (chassis + board + 4× PSU) | eBay | ~$1,960 |
| + 2× Xeon E5-2680 v4 + 256 GB DDR4 | eBay | ~$2,400 |
| + 6× V100 16 GB PCIe | eBay | ~$3,190 total |
| Later: + 2× V100 16 GB → full 8× | eBay | ~$3,500 total |
Advantages
- + 96 GB VRAM — 70B Q4 comfortably, 70B Q6
- + 5,400 GB/s aggregate bandwidth
- + ~2,000 W peak — fits dedicated 230V/16A
- + Built-in upgrade path to 8× GPUs
- + 80 PCIe lanes from dual-socket
Challenges
- − Needs a dedicated 230V/16A circuit
- − Extreme noise — datacenter-grade fans
- − 36 kg — needs rack rails, strong rack
- − PCIe 3.0 interconnect only
- − $1,960 just for the barebone
Why 6× Is the Power Sweet Spot
6× V100 at load: 6 × 250 W = 1,500 W GPU + ~500 W system = ~2,000 W peak. A dedicated European 230V/16A circuit delivers 3,680 W — that’s comfortable headroom. Going to 8× pushes to ~2,600 W, which is still under the limit but leaves little margin for load spikes — and you really want 208V/30A or two circuits at that level. The 6-GPU config gives you 50% more VRAM than 4× while staying on a single dedicated breaker.
Option 3: 8× GPU PCIe Server
Supermicro SYS-4028GR-TR (full) 8× PCIe
Same chassis as the 6× option, fully populated. Dual-socket LGA 2011-3, four 2000W PSUs (3+1 redundant), eight PCIe 3.0 x16 slots with passive risers for full-length double-width cards.
Advantages
- + 128 GB VRAM — 70B Q5+ or two concurrent models
- + Maximum aggregate bandwidth for PCIe V100
- + Proven, well-documented platform
Challenges
- − ~2,600 W — needs 208V/30A or dual circuits
- − Extremely loud — datacenter only
- − 8× GPU is only $490 more than 6×, but power jump is real
There’s also the Supermicro 6049GP-TRT — a Xeon Scalable (LGA 3647) platform supporting up to 10× double-width or 20× single-width GPUs. Dual CPU, 4× 2000W PSU, 4U rackmount. Used barebones start at ~$2,500 on eBay. This is overkill for V100 builds and draws even more power, but it’s the ultimate PCIe expansion option if you plan to move to newer GPUs later.
Option 4: Native SXM2 Platforms
SXM2 V100 modules are cheaper than PCIe cards ($100–150 vs $130–180), and they unlock NVLink — 300 GB/s bidirectional GPU-to-GPU, 19× faster than PCIe 3.0 x16. This matters for multi-GPU inference because inter-GPU transfer becomes the bottleneck when splitting large models across many cards.
Nvidia DGX-1 V100 8× SXM2 + NVLink Mesh
Nvidia’s purpose-built 8× V100 server with full NVLink hybrid cube mesh. Dual Xeon E5-2698 v4, 512 GB DDR4, 3U rackmount, four 1600W PSUs.
| Config | Source | Price |
|---|---|---|
| DGX-1 V100 16 GB — complete | eBay | $5,288 – $9,222 |
| DGX-1 V100 32 GB — complete | eBay | ~$8,650 |
| DGX-1 GPU baseboard only (no GPUs) | eBay | ~$1,775 |
Advantages
- + NVLink mesh — 300 GB/s per GPU pair
- + All-in-one: no BOM to assemble
- + 3U — more compact than 4U Supermicro
Challenges
- − ~3,200 W — requires 208V/30A dedicated
- − Extremely loud — datacenter only
- − $5,300+ — most expensive option
- − Proprietary — scarce replacement parts
Custom SXM2 Adapter Boards 2× or 4× SXM2 + NVLink
Third-party adapter boards seat bare V100 SXM2 modules and present them as PCIe devices. Available from eBay vendors like „Angry Sysadmins“ and Chinese board makers. Dual-card variants include NVLink bridges.
| Board | GPUs | NVLink | Board Price | Total Build |
|---|---|---|---|---|
| Single SXM2 Adapter | 1× SXM2 | No | $50–$80 | ~$200 |
| TNS-2SXM2-4P54 | 2× SXM2 | Yes (bridge) | $350–$500 | ~$700 |
| Quad SXM2 Baseboard | 4× SXM2 | Yes | $600–$1,000 | ~$1,400 |
Total = board + SXM2 modules ($100–150 each). Needs host system with PCIe slot(s) and PSU rated for 300 W per GPU. Cooling is DIY — SXM2 modules need active airflow.
Advantages
- + Cheapest NVLink path — 2× V100 for ~$700
- + SXM2 modules cheaper than PCIe cards
- + Drops into any standard PCIe motherboard
- + 300 GB/s GPU-to-GPU interconnect
Challenges
- − Niche product — limited vendor support
- − Cooling is entirely DIY
- − Board availability fluctuates
- − Power delivery can be finicky
Power & Noise Reality
| Configuration | Peak W | Idle W | Circuit | Noise |
|---|---|---|---|---|
| 2× V100 SXM2 (adapter) | ~800 W | ~150 W | Standard outlet | Quiet (DIY cooling) |
| 4× V100 PCIe (7048GR-TR) | ~1,200 W | ~250 W | Shared 230V/16A OK | Moderate (tower fans) |
| 4× V100 SXM2 (quad board) | ~1,200 W | ~250 W | Shared 230V/16A OK | Moderate (DIY fans) |
| 6× V100 PCIe (4028GR-TR) | ~2,000 W | ~400 W | Dedicated 230V/16A | Loud |
| 8× V100 PCIe (4028GR-TR) | ~2,600 W | ~500 W | 208V/30A or dual circuits | Extreme |
| 8× V100 SXM2 (DGX-1) | ~3,200 W | ~700 W | 208V/30A required | Extreme |
European Home Circuit Math
Czech home circuit: 230V × 16A = 3,680 W max per breaker, shared with everything else on that circuit. 4× V100 at ~1,200 W leaves plenty of room. 6× at ~2,000 W works on a dedicated breaker (nothing else on it). 8× at ~2,600 W is marginal — one load spike can trip the breaker, and you really want 208V/30A or two separate circuits.
Noise: 4028GR-TR and DGX-1 reviewers consistently report moving to colocation. The 7048GR-TR tower is livable — comparable to a loud desktop PC.
What Can You Run?
| Model | Quant | VRAM | 2× V100 | 4× V100 | 6× V100 | 8× V100 |
|---|---|---|---|---|---|---|
| Llama 3.1 8B | Q4_K_M | ~5.5 GB | 1 GPU | 1 GPU | 1 GPU | 1 GPU |
| Qwen 2.5 14B | Q4_K_M | ~9.5 GB | 1 GPU | 1 GPU | 1 GPU | 1 GPU |
| Qwen 2.5 32B | Q4_K_M | ~20 GB | 2 GPU | 2 GPU | 2 GPU | 2 GPU |
| Llama 3.1 70B | Q3_K_M | ~35 GB | NO | 3 GPU | 3 GPU | 3 GPU |
| Llama 3.1 70B | Q4_K_M | ~42 GB | NO | 3–4 GPU | 4 GPU | 4 GPU |
| Llama 3.1 70B | Q6_K | ~58 GB | NO | NO (64 GB) | 5 GPU | 5 GPU |
| Llama 3.1 70B | Q8_0 | ~75 GB | NO | NO | 5–6 GPU | 6 GPU |
V100 16 GB each. VRAM estimates include KV-cache overhead for 2K context. llama.cpp --tensor-split distributes layers across GPUs.
The 6× Sweet Spot
The jump from 4× (64 GB) to 6× (96 GB) unlocks 70B models at Q6_K and Q8_0 — noticeably higher quality than the Q3/Q4 that 4× limits you to. The jump from 6× to 8× (128 GB) only adds headroom for context or concurrent models. Unless you plan to run two 70B models simultaneously, 6× covers everything that matters.
Recommended Build Paths
SXM2 Starter
- • 1× V100 SXM2 module + single adapter
- • Drop into existing system
- • 16 GB HBM2 @ 900 GB/s
- • Proves the concept before bigger investment
NVLink Pair
- • TNS-2SXM2-4P54 dual adapter
- • 2× V100 SXM2 with NVLink bridge
- • 32 GB VRAM, 300 GB/s interconnect
- • Runs 32B Q4 models fast
4× V100 Tower
- • Supermicro 7048GR-TR barebone
- • 2× Xeon E5-2680 v4 + 128 GB DDR4
- • 4× V100 16GB PCIe ($520)
- • 64 GB VRAM, single circuit, quiet
- • Runs 70B Q3 comfortably
6× V100 Rack
- • Supermicro 4028GR-TR barebone
- • 2× Xeon E5-2680 v4 + 256 GB DDR4
- • 6× V100 16GB PCIe ($780)
- • 96 GB VRAM, dedicated circuit
- • Runs 70B Q6/Q8, upgradeable to 8×
- • Warning: loud — needs separate room
Software Stack for V100
| Component | Recommended | Notes |
|---|---|---|
| Inference | llama-server | Full Volta support, --tensor-split for multi-GPU, Ollama API compat with --api-compatibility ollama |
| Multi-user | vLLM | Best for concurrent users, tensor-parallel built-in, +20–30% VRAM overhead |
| Avoid | Ollama v0.6.0+ | Dropped Volta CUDA. Use v0.5.x max, or migrate to llama-server |
| CUDA | 12.x | V100 SM 7.0 fully supported |
| Driver | 535+ | Stable on Linux, supports compute mode |
Where to Buy
| Item | Source | Price |
|---|---|---|
| V100 16GB PCIe | eBay.de, eBay.com | $130–180 |
| V100 16GB SXM2 module | eBay, AliExpress | $100–150 |
| V100 32GB PCIe | eBay | $400–700 |
| Single SXM2→PCIe adapter | eBay | $50–80 |
| Dual SXM2 NVLink board | eBay | $350–500 |
| Quad SXM2 baseboard | eBay | $600–1,000 |
| Supermicro 7048GR-TR bare | eBay | $700–850 |
| Supermicro 4028GR-TR bare | eBay | $1,800–2,200 |
| Supermicro 6049GP-TRT bare | eBay | $2,500+ |
| DGX-1 V100 16GB complete | eBay | $5,288–9,222 |
| Xeon E5-2680 v4 pair | eBay, AliExpress | $40–80 |
| 128 GB DDR4 ECC (8×16GB) | eBay, czech-server.cz | $200–300 |
Prices mid-2025 from eBay.com, eBay.de, AliExpress. Czech sources (bazos.cz, aukro.cz, czech-server.cz) may differ. Always check seller ratings and return policies.
SXM2 Adapter Boards & NVLink Options
V100 SXM2 modules use a proprietary mezzanine connector — they don’t plug into standard PCIe slots. You need an adapter board. Three tiers are available, each with different NVLink and PCIe trade-offs.
| Adapter Type | NVLink | PCIe to Host | Price (eBay) | Best For |
|---|---|---|---|---|
| Single SXM2 → PCIe | None | 1× PCIe x16 | $50–80 | Testing, single-GPU inference |
| Dual NVLink Board | 300 GB/s | 2× PCIe x8 (via SlimSAS or direct) | ~$338 | 64 GB unified VRAM, 35B–70B models |
| Quad NVLink Baseboard | 300 GB/s mesh | PLX8749 switch | ~$789 | 128 GB VRAM, 70B+ models |
Why NVLink Matters
NVLink provides 300 GB/s bidirectional bandwidth between GPUs — 19× faster than PCIe 3.0 x16 (15.75 GB/s). This means two V100 32GB cards act as a single 64 GB memory pool with near-linear scaling for large models.
Without NVLink, each GPU is isolated. A 35B parameter model that needs 60+ GB VRAM simply won’t load across two non-NVLink cards efficiently — the PCIe bottleneck destroys performance.
Dual NVLink Board — Recommended Setup
Power Warning
V100 SXM2 modules use 8-pin EPS CPU power connectors, not standard PCIe 6/8-pin. EPS delivers up to 300 W per connector. Most builders use a separate dedicated 750 W PSU just for the GPU board, connected via an ATX jumper or dual-PSU adapter.
Dual V100 NVLink Benchmarks
Source: Angry Sysadmins blog — dual V100 32GB SXM2 with NVLink, llama-server backend.
V100 NVLink Installation Requirements
What your host machine needs to support a dual V100 SXM2 NVLink adapter board.
GTX 1080 Ti Requirements (for comparison)
Workstation PCIe Slot Comparison
Refurbished workstations from gigacomputer.cz evaluated for V100 NVLink and 1080 Ti compatibility. All prices include VAT with Prague pickup.
HP Z6 G4 Workstation
PCIe Slots (single CPU: slots 1, 2, 4, 5 active)
| Slot | Type | Electrical | Source | Notes |
|---|---|---|---|---|
| 1 | PCIe 3.0 | x4 | CPU | — |
| 2 | PCIe 3.0 | x16 | CPU | Primary GPU slot |
| 3 | PCIe 3.0 | x4 | PCH | Needs 2nd CPU |
| 4 | PCIe 3.0 | x8 | CPU | — |
| 5 | PCIe 3.0 | x16 | CPU | Second GPU slot |
| 6 | PCIe 3.0 | x4 | PCH | Needs 2nd CPU |
Advantages
- + 1,000 W PSU with GPU power cables included
- + Two full x16 + one x8 CPU-connected slots
- + Dual Xeon Scalable socket for future expansion
- + 12 DIMM slots, up to 384 GB ECC RAM
- + Confirmed dual Tesla P40 support by HP community
Challenges
- − Xeon Silver 4108 is low clock (1.8 GHz base)
- − Needs 2nd CPU to unlock slots 3 & 6
- − Proprietary HP PSU form factor
Lenovo ThinkStation P920
PCIe Slots (single CPU: slots 1–5)
| Slot | Type | Electrical | Power | Source |
|---|---|---|---|---|
| 1 | PCIe 3.0 | x16 | 75 W | CPU 1 |
| 2 | PCIe 3.0 | x4 | 25 W | CPU 1 |
| 3 | PCIe 3.0 | x16 | 75 W | CPU 1 |
| 4 | PCIe 3.0 | x4 | 25 W | CPU 1 |
| 5 | PCIe 3.0 | x4 | 25 W | PCH |
| 6 | PCIe 3.0 | x16 | 75 W | CPU 2 |
| 7 | PCIe 3.0 | x16 | 75 W | CPU 2 |
| 8 | PCIe 3.0 | x16 | 75 W | CPU 2 |
Advantages
- + 1,400 W PSU — no second PSU needed for dual V100
- + Up to 5× x16 slots with dual CPU
- + 16 DIMM slots, up to 1 TB RAM
- + 12C/24T provides more threads than Z6 G4
Challenges
- − Xeon Silver 4214 has lower single-thread than W-series
- − Only 2× x16 with single CPU (slots 1 & 3)
- − Large and heavy chassis
Lenovo ThinkStation P520
PCIe Slots
| Slot | Type | Electrical | Power | Source |
|---|---|---|---|---|
| 1 | PCIe 3.0 | x8 | 25 W | CPU |
| 2 | PCIe 3.0 | x16 | 75 W | CPU |
| 3 | PCIe 3.0 | x4 | 25 W | PCH |
| 4 | PCIe 3.0 | x16 | 75 W | CPU |
| 5 | PCI (legacy) | — | 25 W | — |
| 6 | PCIe 3.0 | x4 | 25 W | PCH |
Advantages
- + Best single-thread CPU (W-2145 @ 3.7 GHz)
- + Two x16 + one x8 CPU-connected slots
- + Cheapest workstation option
- + Upgradeable PSU (900 W / 1000 W)
Challenges
- − Stock 690 W PSU too small for dual V100
- − Single socket — no expansion path
- − Slot power rated at only 75 W per x16
Dell Precision 5820 Tower
PCIe Slots (Xeon config)
| Slot | Type | Electrical | Power | Source |
|---|---|---|---|---|
| 1 | PCIe 3.0 | x8 | 50 W | CPU |
| 2 | PCIe 3.0 | x16 | 300 W | CPU |
| 3 | PCIe 3.0 | x1 | 25 W | PCH |
| 4 | PCIe 3.0 | x16 | 300 W | CPU |
| 5 | PCIe 3.0 | x4 | 25 W | PCH |
| 6 | PCI 32-bit | — | 25 W | — |
Advantages
- + 300 W per-slot power rating on x16 slots
- + Two x16 + one x8 CPU-connected
- + Cheapest workstation in the list
- + High single-thread clock (3.6 GHz)
Challenges
- − Must verify PSU wattage (425 W variant is useless)
- − Only 4 cores / 8 threads
- − Only 16 GB RAM stock
- − Single socket, no expansion
HP ProDesk 600 G4 MT
PCIe Slots
| Slot | Type | Electrical | Notes |
|---|---|---|---|
| 1 | PCIe 3.0 | x16 | Only real GPU slot |
| 2 | PCIe 3.0 | x16 → x4 wired | Useless for GPU |
| 3 | PCIe 3.0 | x1 | — |
| 4 | PCIe 3.0 | x1 | — |
Not Viable for GPU Work
400 W PSU cannot power any serious GPU. Only one real x16 slot. Consumer board lacks Above 4G Decoding, no ECC. Cannot fit a 1080 Ti (PSU too small) or V100 NVLink (not enough slots, not enough power).
GPU Compatibility Matrix
| Workstation | Price | PSU | 1080 Ti | V100 NVLink Dual | Key Blocker |
|---|---|---|---|---|---|
| HP Z6 G4 | 19,990 Kç | 1,000 W | YES | YES | None — best fit |
| P920 | 19,941 Kç | 1,400 W | YES | YES | Lower single-thread clock |
| P520 | 17,990 Kç | 690 W | YES | PSU SWAP | Stock PSU too small for dual V100 |
| Dell 5820 | 16,141 Kç | 425/950 W | YES | CHECK PSU | Must verify which PSU variant ships |
| ProDesk 600 | 6,612 Kç | 400 W | NO | NO | PSU, slots, BIOS all insufficient |
Recommended Purchase: HP Z6 G4
The HP Z6 G4 at 19,990 Kç is the optimal starting platform for a V100 NVLink build:
- ✓ 1,000 W PSU with GPU power cables already present
- ✓ Two full x16 + one x8 CPU-connected PCIe slots
- ✓ Workstation BIOS with Above 4G Decoding support
- ✓ Dual Xeon socket for future CPU/slot expansion
- ✓ 12 DIMM slots, expandable to 384 GB ECC
- ✓ Confirmed dual Tesla P40 support by HP community
Alternative: The ThinkStation P920 (19,941 Kç) with its 1,400 W PSU eliminates the need for a second PSU entirely, but trades single-thread performance.
Estimated Total Build Cost
| Component | Source | Price |
|---|---|---|
| HP Z6 G4 Workstation | gigacomputer.cz | 19,990 Kç |
| 2× V100 32GB SXM2 modules | eBay | ~$200–300 each (~10,000–15,000 Kç) |
| Dual NVLink adapter board | eBay (olddays168) | ~$338 (~7,800 Kç) |
| 750 W dedicated GPU PSU | Local / Alza.cz | ~2,500 Kç |
| Cooling (4× Arctic P8 80mm) | Alza.cz | ~500 Kç |
| Cables & adapters | Various | ~500 Kç |
| TOTAL | ~35,000–46,000 Kç |
That gives you 64 GB of NVLink-connected VRAM capable of running 35B–70B parameter models at 100–160 tok/s. For comparison, a single RTX 4090 (24 GB VRAM, ~45,000 Kç alone) can’t even load a 35B model.