BenchmarksNewsPC Components

Loud AI rig: Nvidia V100 added to gaming PC for $266

Featured image Loud AI rig Nvidia V100 added to gaming PC for 266

In a world obsessed with massive cloud computing, some enthusiasts are taking a radical detour: turning industrial surplus into cutting-edge AI power. One dedicated PC builder recently achieved this feat by repurposing an old, noisy enterprise GPU to fuel powerful local Large Language Model (LLM) inference.

The result? A system that doubled its total available VRAM to 32GB for a remarkably low cost of just $266 (£200). This ingenious hack perfectly exemplifies the spirit of resourcefulness thriving in the current RAMpocalypse.

The core piece of hardware was a Tesla V100 SXM2, typically relegated to data centers. While these cards are largely obsolete, they packed substantial VRAM and provided the raw power needed for serious AI work.

To make this relic fit into a consumer setup, the enthusiast didn’t stop there. They utilized an SXM2-to-PCIe adapter and a specialized PWM modifier to tackle the most irritating obstacle: the notorious noise of the cooler. The original V100 fan was described as loud enough to rival a lawnmower, generating 82dB of noise.

The solution required some surgical tinkering. By carefully rerouting the existing fan wires and connecting them directly to the motherboard’s PWM header, the system could be dramatically quieted. Furthermore, by adjusting the fan speed to just 10% under full load, the temperatures were kept safely below 50C.

This hardware setup became a unique pairing: a powerful RTX 4080 (with 16GB VRAM) paired with the Tesla V100 (also 16GB VRAM), creating a combined 32GB of unified VRAM capacity.

The system was successfully configured using NixOS and a legacy Nvidia driver, which cleverly managed support across both the Volta and Ada architectures. This setup proved highly effective for running demanding local LLMs.

When testing a large 27 billion parameter model, the enthusiast achieved an impressive inference speed of 32 tokens per second—a rate deemed fast enough for genuinely interactive use and often quicker than relying on cloud API services.

This project demonstrates that sometimes, the most powerful AI solutions aren’t found in the newest components, but in the creative repurposing of existing, underutilized technology. It’s a fantastic example of how innovation can turn industrial leftovers into bleeding-edge local computing power.

Image credit: Nvidia