V100 Multi-GPU Server Guide: 4×, 6×, 8× PCIe and SXM2 Builds for AI Inference

Napsal(a)

v rubrice

,

V100 Multi-GPU Server Options

4×, 6×, 8× PCIe and native SXM2 platforms for V100 inference — plus GPU efficiency analysis. Used market, mid-2025 pricing.

Quick Verdict for Home Rack

Start with the Supermicro 7048GR-TR (4× PCIe, $750 barebone) — it runs on a single 230V/16A circuit and is quiet enough for a home. If you need more VRAM, the 4028GR-TR with 6× V100 hits the sweet spot of 96 GB VRAM while staying within a dedicated 230V/16A circuit at ~2,000 W peak. Full 8× GPU or DGX-1 builds need 208V/30A power and produce extreme noise.

The V100 16GB is the best value GPU for inference: it leads both tok/s per dollar and tok/s per watt among affordable datacenter cards.

Ollama + V100 Compatibility Warning

Ollama v0.6.0+ dropped Volta (V100) CUDA support. Pin Ollama to v0.5.x, or switch to llama-server / vLLM. Your tools (Nomad, Hermes, OpenClaw, Opencode) all use the Ollama API — llama-server with --api-compatibility ollama is a drop-in replacement.

GPU Efficiency Comparison

All benchmarks: Llama 2 7B Q4_0, 128-token generation (tg128). Data from the llama.cpp CUDA benchmark thread.

GPUtok/sTDP (W)Used Pricetok/s / Wtok/s / $W / $
V100 16GB129250$1550.520.831.61
P100 16GB60250$700.240.863.57
P40 24GB54250$850.220.642.94
T4 16GB4570$2150.640.210.33
1080 Ti 11GB62250$1000.250.622.50
V100 32GB129250$5500.520.230.45
RTX 3090160350$7000.460.230.50
RTX 4090187450$1,5000.420.120.30
A100 80GB195300$6,5000.650.030.05
Token Generation per Watt
Higher = more power-efficient. Llama 2 7B Q4_0 tg128.
A100 80GB
0.65 tok/s/W
T4 16GB
0.64 tok/s/W
V100 16GB
0.52 tok/s/W
RTX 3090
0.46 tok/s/W
RTX 4090
0.42 tok/s/W
1080 Ti
0.25
P100 16GB
0.24
P40 24GB
0.22
Token Generation per Dollar
Higher = better value for money. Used market mid-2025 prices.
P100 16GB
0.86 tok/s/$
V100 16GB
0.83 tok/s/$
P40 24GB
0.64 tok/s/$
1080 Ti
0.62 tok/s/$
V100 32GB
0.23
RTX 3090
0.23
T4 16GB
0.21
RTX 4090
0.12
A100 80GB
0.03
Power Draw per Dollar Invested
Higher = more power-hungry per dollar of GPU cost. Low values mean capital cost dominates over electricity.
P100 16GB
3.57 W/$
P40 24GB
2.94 W/$
1080 Ti
2.50 W/$
V100 16GB
1.61 W/$
RTX 3090
0.50 W/$
V100 32GB
0.45 W/$
T4 16GB
0.33 W/$
RTX 4090
0.30 W/$
A100 80GB
0.05 W/$

What the Efficiency Charts Tell You

V100 16GB wins on combined efficiency. It’s #3 in power efficiency (behind A100 and T4 — both far more expensive), and #2 in value-per-dollar (nearly tied with P100). The P100 is cheaper per token, but at half the speed — the V100 gets you to interactive speed (129 tok/s on 7B) while the P100 doesn’t quite (60 tok/s). The T4 is the power-sipping champion at 70W TDP, but it’s actually slower than a 1080 Ti and costs twice as much used.

The W/$ chart shows ongoing cost structure: cheap GPUs (P100, P40) have high wattage relative to their price — electricity becomes a significant cost over time. More expensive GPUs (T4, RTX 4090, A100) are capital-heavy but power-light. V100 sits in a balanced middle.

Server Platforms at a Glance

PlatformGPUsFormInterconnectPeak WCostHome?
7048GR-TR4× PCIe4U towerPCIe 3.0~1,200 W$750 bareBEST
4028GR-TR (6×)6× PCIe4U rackPCIe 3.0~2,000 W$1,960 bareDUCT.
4028GR-TR (8×)8× PCIe4U rackPCIe 3.0~2,600 W$1,960 bareHARD
6049GP-TRT10× dbl-w4U rackPCIe 3.0~3,000+ W$2,500+ bareHARD
DGX-1 V1008× SXM23U rackNVLink mesh~3,200 W$5,300+HARD
SXM2 Adapter 2×2× SXM2PCIe cardNVLink~600 W$350–500YES
SXM2 Adapter 4×4× SXM2baseboardNVLink~1,200 W$600–1,000OK

Barebones require CPUs, RAM, and drives separately. SXM2 adapter builds need a host system. DGX-1 prices are complete systems. „DUCT.“ = dedicated circuit required.

Total Build Cost (server + V100 16GB at $130–155 each)
Barebones + CPUs + RAM + GPUs. DGX-1 is complete system price.
SXM2 2× + board
$700
7048GR-TR + 4×
$1,570
SXM2 4× + board
$1,400
4028GR-TR + 6×
$3,190
4028GR-TR + 8×
$3,680
DGX-1 16GB
$5,300–$9,200

Option 1: 4× GPU PCIe Server

Recommended

Supermicro SYS-7048GR-TR 4× PCIe

The practical choice for a home lab. Tower/rackmount convertible with the X10DRG-Q motherboard, dual LGA 2011-3, four direct-attach PCIe 3.0 x16 slots. Single 2000W PSU keeps it on one circuit. Tower mode is far quieter than rackmount servers.

Form Factor
4U Tower / Rack
GPU Slots
4× PCIe 3.0 x16
CPU Platform
Dual LGA 2011-3
Max RAM
2 TB DDR4 ECC
PSU
1× 2000 W
Noise
Moderate (tower)

Build Cost

ConfigurationSourcePrice
Barebones (chassis + board + PSU)eBay~$750
+ 2× Xeon E5-2680 v4 + 128 GB DDR4eBay~$1,050
+ 4× V100 16 GB PCIeeBay~$1,570 total
+ 4× V100 32 GB PCIe insteadeBay~$3,570 total

Advantages

  • + 64 GB VRAM — 30B+ Q4 models, 70B Q3
  • + 3,600 GB/s aggregate bandwidth
  • + ~1,200 W peak — single 230V/16A circuit
  • + Tower = much quieter than rack servers
  • + $750 barebones — cheapest full server

Challenges

  • Limited to 4 GPUs — no room to grow
  • PCIe 3.0 x16 only (~16 GB/s per GPU)
  • Single PSU — no redundancy
  • 4U still takes significant rack height

Option 2: 6× GPU PCIe Server

Supermicro SYS-4028GR-TR (6 of 8 slots) 6× PCIe

There’s no standard 6-GPU server — the cleanest path is the 8-slot 4028GR-TR with 6 slots populated. This drops peak power from ~2,600 W to ~2,000 W, keeping it within a dedicated 230V/16A circuit. The 2 empty slots are a future upgrade path to 8× when you get dedicated power wiring.

Form Factor
4U Rackmount
GPU Slots
8× PCIe 3.0 x16
Using
6 of 8 slots
CPU Platform
Dual LGA 2011-3
PSU
4× 2000 W (3+1)
Weight
~36 kg (80 lbs)

Build Cost

ConfigurationSourcePrice
Barebones (chassis + board + 4× PSU)eBay~$1,960
+ 2× Xeon E5-2680 v4 + 256 GB DDR4eBay~$2,400
+ 6× V100 16 GB PCIeeBay~$3,190 total
Later: + 2× V100 16 GB → full 8×eBay~$3,500 total

Advantages

  • + 96 GB VRAM — 70B Q4 comfortably, 70B Q6
  • + 5,400 GB/s aggregate bandwidth
  • + ~2,000 W peak — fits dedicated 230V/16A
  • + Built-in upgrade path to 8× GPUs
  • + 80 PCIe lanes from dual-socket

Challenges

  • Needs a dedicated 230V/16A circuit
  • Extreme noise — datacenter-grade fans
  • 36 kg — needs rack rails, strong rack
  • PCIe 3.0 interconnect only
  • $1,960 just for the barebone

Why 6× Is the Power Sweet Spot

6× V100 at load: 6 × 250 W = 1,500 W GPU + ~500 W system = ~2,000 W peak. A dedicated European 230V/16A circuit delivers 3,680 W — that’s comfortable headroom. Going to 8× pushes to ~2,600 W, which is still under the limit but leaves little margin for load spikes — and you really want 208V/30A or two circuits at that level. The 6-GPU config gives you 50% more VRAM than 4× while staying on a single dedicated breaker.

Option 3: 8× GPU PCIe Server

Supermicro SYS-4028GR-TR (full) 8× PCIe

Same chassis as the 6× option, fully populated. Dual-socket LGA 2011-3, four 2000W PSUs (3+1 redundant), eight PCIe 3.0 x16 slots with passive risers for full-length double-width cards.

Total VRAM
128 GB (8×16)
Aggregate BW
7,200 GB/s
Peak Power
~2,600 W
Total Build
~$3,680

Advantages

  • + 128 GB VRAM — 70B Q5+ or two concurrent models
  • + Maximum aggregate bandwidth for PCIe V100
  • + Proven, well-documented platform

Challenges

  • ~2,600 W — needs 208V/30A or dual circuits
  • Extremely loud — datacenter only
  • 8× GPU is only $490 more than 6×, but power jump is real

There’s also the Supermicro 6049GP-TRT — a Xeon Scalable (LGA 3647) platform supporting up to 10× double-width or 20× single-width GPUs. Dual CPU, 4× 2000W PSU, 4U rackmount. Used barebones start at ~$2,500 on eBay. This is overkill for V100 builds and draws even more power, but it’s the ultimate PCIe expansion option if you plan to move to newer GPUs later.

Option 4: Native SXM2 Platforms

SXM2 V100 modules are cheaper than PCIe cards ($100–150 vs $130–180), and they unlock NVLink — 300 GB/s bidirectional GPU-to-GPU, 19× faster than PCIe 3.0 x16. This matters for multi-GPU inference because inter-GPU transfer becomes the bottleneck when splitting large models across many cards.

Nvidia DGX-1 V100 8× SXM2 + NVLink Mesh

Nvidia’s purpose-built 8× V100 server with full NVLink hybrid cube mesh. Dual Xeon E5-2698 v4, 512 GB DDR4, 3U rackmount, four 1600W PSUs.

Form Factor
3U Rackmount
NVLink BW
300 GB/s per link
Total VRAM
128 GB / 256 GB
Peak Draw
~3,200 W
ConfigSourcePrice
DGX-1 V100 16 GB — completeeBay$5,288 – $9,222
DGX-1 V100 32 GB — completeeBay~$8,650
DGX-1 GPU baseboard only (no GPUs)eBay~$1,775

Advantages

  • + NVLink mesh — 300 GB/s per GPU pair
  • + All-in-one: no BOM to assemble
  • + 3U — more compact than 4U Supermicro

Challenges

  • ~3,200 W — requires 208V/30A dedicated
  • Extremely loud — datacenter only
  • $5,300+ — most expensive option
  • Proprietary — scarce replacement parts

Custom SXM2 Adapter Boards 2× or 4× SXM2 + NVLink

Third-party adapter boards seat bare V100 SXM2 modules and present them as PCIe devices. Available from eBay vendors like „Angry Sysadmins“ and Chinese board makers. Dual-card variants include NVLink bridges.

BoardGPUsNVLinkBoard PriceTotal Build
Single SXM2 Adapter1× SXM2No$50–$80~$200
TNS-2SXM2-4P542× SXM2Yes (bridge)$350–$500~$700
Quad SXM2 Baseboard4× SXM2Yes$600–$1,000~$1,400

Total = board + SXM2 modules ($100–150 each). Needs host system with PCIe slot(s) and PSU rated for 300 W per GPU. Cooling is DIY — SXM2 modules need active airflow.

Advantages

  • + Cheapest NVLink path — 2× V100 for ~$700
  • + SXM2 modules cheaper than PCIe cards
  • + Drops into any standard PCIe motherboard
  • + 300 GB/s GPU-to-GPU interconnect

Challenges

  • Niche product — limited vendor support
  • Cooling is entirely DIY
  • Board availability fluctuates
  • Power delivery can be finicky

Power & Noise Reality

ConfigurationPeak WIdle WCircuitNoise
2× V100 SXM2 (adapter)~800 W~150 WStandard outletQuiet (DIY cooling)
4× V100 PCIe (7048GR-TR)~1,200 W~250 WShared 230V/16A OKModerate (tower fans)
4× V100 SXM2 (quad board)~1,200 W~250 WShared 230V/16A OKModerate (DIY fans)
6× V100 PCIe (4028GR-TR)~2,000 W~400 WDedicated 230V/16ALoud
8× V100 PCIe (4028GR-TR)~2,600 W~500 W208V/30A or dual circuitsExtreme
8× V100 SXM2 (DGX-1)~3,200 W~700 W208V/30A requiredExtreme

European Home Circuit Math

Czech home circuit: 230V × 16A = 3,680 W max per breaker, shared with everything else on that circuit. 4× V100 at ~1,200 W leaves plenty of room. 6× at ~2,000 W works on a dedicated breaker (nothing else on it). 8× at ~2,600 W is marginal — one load spike can trip the breaker, and you really want 208V/30A or two separate circuits.

Noise: 4028GR-TR and DGX-1 reviewers consistently report moving to colocation. The 7048GR-TR tower is livable — comparable to a loud desktop PC.

What Can You Run?

ModelQuantVRAM2× V1004× V1006× V1008× V100
Llama 3.1 8BQ4_K_M~5.5 GB1 GPU1 GPU1 GPU1 GPU
Qwen 2.5 14BQ4_K_M~9.5 GB1 GPU1 GPU1 GPU1 GPU
Qwen 2.5 32BQ4_K_M~20 GB2 GPU2 GPU2 GPU2 GPU
Llama 3.1 70BQ3_K_M~35 GBNO3 GPU3 GPU3 GPU
Llama 3.1 70BQ4_K_M~42 GBNO3–4 GPU4 GPU4 GPU
Llama 3.1 70BQ6_K~58 GBNONO (64 GB)5 GPU5 GPU
Llama 3.1 70BQ8_0~75 GBNONO5–6 GPU6 GPU

V100 16 GB each. VRAM estimates include KV-cache overhead for 2K context. llama.cpp --tensor-split distributes layers across GPUs.

The 6× Sweet Spot

The jump from 4× (64 GB) to 6× (96 GB) unlocks 70B models at Q6_K and Q8_0 — noticeably higher quality than the Q3/Q4 that 4× limits you to. The jump from 6× to 8× (128 GB) only adds headroom for context or concurrent models. Unless you plan to run two 70B models simultaneously, 6× covers everything that matters.

Recommended Build Paths

SXM2 Starter

~$200
  • • 1× V100 SXM2 module + single adapter
  • • Drop into existing system
  • • 16 GB HBM2 @ 900 GB/s
  • • Proves the concept before bigger investment

NVLink Pair

~$700
  • • TNS-2SXM2-4P54 dual adapter
  • • 2× V100 SXM2 with NVLink bridge
  • • 32 GB VRAM, 300 GB/s interconnect
  • • Runs 32B Q4 models fast
Best Start

4× V100 Tower

~$1,570
  • • Supermicro 7048GR-TR barebone
  • • 2× Xeon E5-2680 v4 + 128 GB DDR4
  • • 4× V100 16GB PCIe ($520)
  • • 64 GB VRAM, single circuit, quiet
  • • Runs 70B Q3 comfortably

6× V100 Rack

~$3,190
  • • Supermicro 4028GR-TR barebone
  • • 2× Xeon E5-2680 v4 + 256 GB DDR4
  • • 6× V100 16GB PCIe ($780)
  • • 96 GB VRAM, dedicated circuit
  • • Runs 70B Q6/Q8, upgradeable to 8×
  • • Warning: loud — needs separate room

Software Stack for V100

ComponentRecommendedNotes
Inferencellama-serverFull Volta support, --tensor-split for multi-GPU, Ollama API compat with --api-compatibility ollama
Multi-uservLLMBest for concurrent users, tensor-parallel built-in, +20–30% VRAM overhead
AvoidOllama v0.6.0+Dropped Volta CUDA. Use v0.5.x max, or migrate to llama-server
CUDA12.xV100 SM 7.0 fully supported
Driver535+Stable on Linux, supports compute mode

Where to Buy

ItemSourcePrice
V100 16GB PCIeeBay.de, eBay.com$130–180
V100 16GB SXM2 moduleeBay, AliExpress$100–150
V100 32GB PCIeeBay$400–700
Single SXM2→PCIe adaptereBay$50–80
Dual SXM2 NVLink boardeBay$350–500
Quad SXM2 baseboardeBay$600–1,000
Supermicro 7048GR-TR bareeBay$700–850
Supermicro 4028GR-TR bareeBay$1,800–2,200
Supermicro 6049GP-TRT bareeBay$2,500+
DGX-1 V100 16GB completeeBay$5,288–9,222
Xeon E5-2680 v4 paireBay, AliExpress$40–80
128 GB DDR4 ECC (8×16GB)eBay, czech-server.cz$200–300

Prices mid-2025 from eBay.com, eBay.de, AliExpress. Czech sources (bazos.cz, aukro.cz, czech-server.cz) may differ. Always check seller ratings and return policies.


SXM2 Adapter Boards & NVLink Options

V100 SXM2 modules use a proprietary mezzanine connector — they don’t plug into standard PCIe slots. You need an adapter board. Three tiers are available, each with different NVLink and PCIe trade-offs.

Adapter TypeNVLinkPCIe to HostPrice (eBay)Best For
Single SXM2 → PCIeNone1× PCIe x16$50–80Testing, single-GPU inference
Dual NVLink Board300 GB/s2× PCIe x8 (via SlimSAS or direct)~$33864 GB unified VRAM, 35B–70B models
Quad NVLink Baseboard300 GB/s meshPLX8749 switch~$789128 GB VRAM, 70B+ models

Why NVLink Matters

NVLink provides 300 GB/s bidirectional bandwidth between GPUs — 19× faster than PCIe 3.0 x16 (15.75 GB/s). This means two V100 32GB cards act as a single 64 GB memory pool with near-linear scaling for large models.

Without NVLink, each GPU is isolated. A 35B parameter model that needs 60+ GB VRAM simply won’t load across two non-NVLink cards efficiently — the PCIe bottleneck destroys performance.

Dual NVLink Board — Recommended Setup

Board Size
170 × 236 mm (mATX)
GPU Interface
2× SXM2 sockets
Host Interface
2× PCIe x8 via SlimSAS
NVLink Speed
300 GB/s bidirectional
Power Input
2× 8-pin EPS per GPU
Total GPU Power
~600 W (300 W each)
Cooling
6× fan headers, needs active cooling
eBay Seller
olddays168 (Shenzhen, 3-mo warranty)

Power Warning

V100 SXM2 modules use 8-pin EPS CPU power connectors, not standard PCIe 6/8-pin. EPS delivers up to 300 W per connector. Most builders use a separate dedicated 750 W PSU just for the GPU board, connected via an ATX jumper or dual-PSU adapter.

Dual V100 NVLink Benchmarks

Source: Angry Sysadmins blog — dual V100 32GB SXM2 with NVLink, llama-server backend.

Gemma-4-26B
164 tok/s
Gemma-4-26B (min)
141 tok/s
Qwen3.6-35B
112 tok/s

V100 NVLink Installation Requirements

What your host machine needs to support a dual V100 SXM2 NVLink adapter board.

PCIe Slots
2× PCIe x8 minimum (x16 preferred)
BIOS: Above 4G Decoding
Must be ENABLED
BIOS: Resizable BAR
Must be ENABLED
BIOS: Boot Mode
UEFI only (disable CSM/Legacy)
GPU Power
2× 8-pin EPS per GPU (600 W total)
PSU Headroom
750 W dedicated GPU PSU recommended
Cooling
4× 80mm fans @ 3000 RPM minimum
Physical Space
Board sits externally or open-frame

GTX 1080 Ti Requirements (for comparison)

PCIe Slot
1× PCIe 3.0 x16
TDP
250 W
Power Connectors
1× 6-pin + 1× 8-pin PCIe
Card Length
267 mm (dual-slot width)
Min PSU
600 W recommended

Workstation PCIe Slot Comparison

Refurbished workstations from gigacomputer.cz evaluated for V100 NVLink and 1080 Ti compatibility. All prices include VAT with Prague pickup.

HP Z6 G4 Workstation

BEST PICK 19,990 Kç
CPU
Xeon Silver 4108 (8C/16T)
RAM
32 GB DDR4 ECC (12 slots, max 384 GB)
Storage
512 GB SSD
GPU (stock)
Quadro P2000
PSU
1,000 W (90%, 2× 6/8-pin GPU cables)
Dual CPU
Yes — Intel Xeon Scalable

PCIe Slots (single CPU: slots 1, 2, 4, 5 active)

SlotTypeElectricalSourceNotes
1PCIe 3.0x4CPU
2PCIe 3.0x16CPUPrimary GPU slot
3PCIe 3.0x4PCHNeeds 2nd CPU
4PCIe 3.0x8CPU
5PCIe 3.0x16CPUSecond GPU slot
6PCIe 3.0x4PCHNeeds 2nd CPU

Advantages

  • + 1,000 W PSU with GPU power cables included
  • + Two full x16 + one x8 CPU-connected slots
  • + Dual Xeon Scalable socket for future expansion
  • + 12 DIMM slots, up to 384 GB ECC RAM
  • + Confirmed dual Tesla P40 support by HP community

Challenges

  • Xeon Silver 4108 is low clock (1.8 GHz base)
  • Needs 2nd CPU to unlock slots 3 & 6
  • Proprietary HP PSU form factor

Lenovo ThinkStation P920

RUNNER-UP 19,941 Kç
CPU
Xeon Silver 4214 (12C/24T)
RAM
32 GB DDR4 ECC (16 slots, max 1 TB)
Storage
512 GB SSD
GPU (stock)
Quadro P1000
PSU
1,400 W (!!) 80+ Platinum
Dual CPU
Yes — unlocks slots 6–8

PCIe Slots (single CPU: slots 1–5)

SlotTypeElectricalPowerSource
1PCIe 3.0x1675 WCPU 1
2PCIe 3.0x425 WCPU 1
3PCIe 3.0x1675 WCPU 1
4PCIe 3.0x425 WCPU 1
5PCIe 3.0x425 WPCH
6PCIe 3.0x1675 WCPU 2
7PCIe 3.0x1675 WCPU 2
8PCIe 3.0x1675 WCPU 2

Advantages

  • + 1,400 W PSU — no second PSU needed for dual V100
  • + Up to 5× x16 slots with dual CPU
  • + 16 DIMM slots, up to 1 TB RAM
  • + 12C/24T provides more threads than Z6 G4

Challenges

  • Xeon Silver 4214 has lower single-thread than W-series
  • Only 2× x16 with single CPU (slots 1 & 3)
  • Large and heavy chassis

Lenovo ThinkStation P520

BUDGET PICK 17,990 Kç
CPU
Xeon W-2145 (8C/16T, 3.7 GHz)
RAM
32 GB DDR4 ECC
Storage
512 GB NVMe SSD
GPU (stock)
Quadro P2000
PSU
690 W (options: 900 W, 1000 W)
Dual CPU
No — single socket

PCIe Slots

SlotTypeElectricalPowerSource
1PCIe 3.0x825 WCPU
2PCIe 3.0x1675 WCPU
3PCIe 3.0x425 WPCH
4PCIe 3.0x1675 WCPU
5PCI (legacy)25 W
6PCIe 3.0x425 WPCH

Advantages

  • + Best single-thread CPU (W-2145 @ 3.7 GHz)
  • + Two x16 + one x8 CPU-connected slots
  • + Cheapest workstation option
  • + Upgradeable PSU (900 W / 1000 W)

Challenges

  • Stock 690 W PSU too small for dual V100
  • Single socket — no expansion path
  • Slot power rated at only 75 W per x16

Dell Precision 5820 Tower

CHECK PSU 16,141 Kç
CPU
Xeon W-2123 (4C/8T, 3.6 GHz)
RAM
16 GB DDR4 ECC
Storage
512 GB SSD
GPU (stock)
Quadro P2000
PSU
425 W or 950 W (check before buying!)
Dual CPU
No — single socket

PCIe Slots (Xeon config)

SlotTypeElectricalPowerSource
1PCIe 3.0x850 WCPU
2PCIe 3.0x16300 WCPU
3PCIe 3.0x125 WPCH
4PCIe 3.0x16300 WCPU
5PCIe 3.0x425 WPCH
6PCI 32-bit25 W

Advantages

  • + 300 W per-slot power rating on x16 slots
  • + Two x16 + one x8 CPU-connected
  • + Cheapest workstation in the list
  • + High single-thread clock (3.6 GHz)

Challenges

  • Must verify PSU wattage (425 W variant is useless)
  • Only 4 cores / 8 threads
  • Only 16 GB RAM stock
  • Single socket, no expansion

HP ProDesk 600 G4 MT

NOT SUITABLE 6,612 Kç
CPU
i5-8400 (6C/6T)
RAM
16 GB DDR4
Storage
256 GB SSD
GPU (stock)
Intel UHD 630
PSU
400 W
Dual CPU
No

PCIe Slots

SlotTypeElectricalNotes
1PCIe 3.0x16Only real GPU slot
2PCIe 3.0x16 → x4 wiredUseless for GPU
3PCIe 3.0x1
4PCIe 3.0x1

Not Viable for GPU Work

400 W PSU cannot power any serious GPU. Only one real x16 slot. Consumer board lacks Above 4G Decoding, no ECC. Cannot fit a 1080 Ti (PSU too small) or V100 NVLink (not enough slots, not enough power).


GPU Compatibility Matrix

WorkstationPricePSU1080 TiV100 NVLink DualKey Blocker
HP Z6 G419,990 Kç1,000 WYESYESNone — best fit
P92019,941 Kç1,400 WYESYESLower single-thread clock
P52017,990 Kç690 WYESPSU SWAPStock PSU too small for dual V100
Dell 582016,141 Kç425/950 WYESCHECK PSUMust verify which PSU variant ships
ProDesk 6006,612 Kç400 WNONOPSU, slots, BIOS all insufficient

Recommended Purchase: HP Z6 G4

The HP Z6 G4 at 19,990 Kç is the optimal starting platform for a V100 NVLink build:

  • 1,000 W PSU with GPU power cables already present
  • Two full x16 + one x8 CPU-connected PCIe slots
  • Workstation BIOS with Above 4G Decoding support
  • Dual Xeon socket for future CPU/slot expansion
  • 12 DIMM slots, expandable to 384 GB ECC
  • Confirmed dual Tesla P40 support by HP community

Alternative: The ThinkStation P920 (19,941 Kç) with its 1,400 W PSU eliminates the need for a second PSU entirely, but trades single-thread performance.

Estimated Total Build Cost

ComponentSourcePrice
HP Z6 G4 Workstationgigacomputer.cz19,990 Kç
2× V100 32GB SXM2 moduleseBay~$200–300 each (~10,000–15,000 Kç)
Dual NVLink adapter boardeBay (olddays168)~$338 (~7,800 Kç)
750 W dedicated GPU PSULocal / Alza.cz~2,500 Kç
Cooling (4× Arctic P8 80mm)Alza.cz~500 Kç
Cables & adaptersVarious~500 Kç
TOTAL~35,000–46,000 Kç

That gives you 64 GB of NVLink-connected VRAM capable of running 35B–70B parameter models at 100–160 tok/s. For comparison, a single RTX 4090 (24 GB VRAM, ~45,000 Kç alone) can’t even load a 35B model.

Komentáře

Napsat komentář

Vaše e-mailová adresa nebude zveřejněna. Vyžadované informace jsou označeny *

opentree.cz
Přehled ochrany osobních údajů

Tyto webové stránky používají soubory cookies, abychom vám mohli poskytnout co nejlepší uživatelský zážitek. Informace o souborech cookie se ukládají ve vašem prohlížeči a plní funkce, jako je rozpoznání, když se na naše webové stránky vrátíte, a pomáhají našemu týmu pochopit, které části webových stránek považujete za nejzajímavější a nejužitečnější.