YOU CAN BUILD A 48 GB AI SERVER FROM SCRAPPED DATACENTER GPUS
Two scrapped datacenter GPUs in a used Dell server give you 48 GB of VRAM - enough to run 70B-parameter models quantized at home.
by magiccodingman (GitHub username)
AIOpen-hardware
- difficulty
- ●●●●○
- time
- a weekend-plus
- license
- license not specified
- repo
- repo ACTIVE59 stars
●●●●○ · a weekend-plus · license not specified · 59 stars · repo ACTIVE
WHAT YOU’LL NEED
Jump to section
COMPAREE VERDICT
This is for someone who already knows how to deploy LLMs and wants to escape cloud bills by owning the metal. The guide pairs two Tesla P40s (48 GB of VRAM) with a used Dell PowerEdge R730, for 1,092 dollars at the creator's late-2023 prices, with the P40s listed at 175 dollars each; used prices have moved since, so check current listings. There is nothing to print: the server's own fans cool the passive cards, and the creator shares a Python fan-control script because the R730 otherwise runs its fans at full blast. The hard part is not the hardware - the guide is honest that it is not a step-by-step software walkthrough, so if you have never installed CUDA drivers or run a model locally, plan on external guides. Speeds are modest: the creator measured around 1.4 tokens per second on a 70B model at 4-bit in his first benchmarks. Also note the creator's warning that NVIDIA's maintenance support for the P40 ran until July 2026, so newer drivers may drop it; he now suggests looking at other GPUs.
IN THE REPO
GOOD TO KNOW
- —GitHub guide has a complete parts list with eBay and Amazon links, current as of the last commit.
- —No printed parts: the P40s are cooled by the Dell R730's own fans, and the repo includes a Python fan-control script so the server is not deafening.
- —No firmware or software setup guide in the repo; the guide assumes you know how to install Linux and run inference tooling.
- —AI HOMELAB video walks the physical build but does not cover model deployment.
- —No license file in the repo; the guide is public but its terms are unstated.
- —This is a 2016 Pascal card; the guide does not claim it matches modern consumer GPUs in training speed or power efficiency.
Parts to buy
6 itemsFrom our check of the build. Exact quantities and part numbers are in the creator’s BOM.
Can I build this?
Build at your own risk. Projects involve tools, electronics and sometimes mains voltage — follow the creator’s safety notes.
Videos
DIY 4x Nvidia P40 Homeserver for AI with 96gb VRAM!
Community video documenting a 4x P40, 448 GB RAM build including the cooling solution. Covers physical assembly; does not walk software setup.
More builds like this
All projectsGallery
Start here
Navigation into the creator’s own docs — we don’t rewrite the guide, we route you to the source.
- 1.Read the full parts list and verify your chosen server has two x16 slots with physical clearance for full-length cards (The guide links current eBay and Amazon listings. Prices drift; check before buying.)
- 2.Set up fan control: iDRAC minimum fan speed plus the creator's Python fan-control script(Without it the R730 does not recognise the P40s and runs its fans at full speed.)
- 3.Watch the AI HOMELAB build video for physical assembly reference (Shows the cooling install and slot spacing. Software setup is not covered.)
Resources
Documentation, files and community threads for this build — we link straight to the original sources and never rehost the creator’s files.
KNOWN ISSUES
- Most common mistake: buying the cards before confirming your server has two full-length x16 slots with physical clearance for them (the R730 needs the Riser 3 GPU addition the guide lists) and PSUs that can feed them.
- P40s have no fans of their own. In the R730 the chassis fans cool them, but the server does not recognise the cards and spins fans to full speed until you set up the fan-control script.
- The repo does not cover CUDA driver install, Docker setup, or model deployment. If you have never run a local LLM, you will need external guides.
- Two P40s add about 500 W under load on top of a dual-Xeon server. Use the 1100 W Dell PSUs the guide specifies (the creator found a 1600 W unit was not compatible) and check what else shares that circuit.
- Used enterprise hardware often ships with no documentation. Know how to check BIOS, verify RAM slots, and flash firmware if needed.
- Used enterprise cards come with no warranty. Buy from sellers with returns and test each card as soon as it arrives.
Can I run this on a normal desktop motherboard?
Not easily. Consumer boards rarely have the slot spacing, power and airflow these full-length passive cards need. The guide uses a used Dell PowerEdge R730 with an extra GPU riser for a reason.
How does a P40 compare to a modern RTX card?
The P40 is a 2016 Pascal datacenter card with 24 GB of VRAM but no Tensor cores. For inference, VRAM decides which models fit, and the P40 gives a lot of it cheaply. The author's comparison: an RTX 3090 was up to about 2.4 times faster, for about four times the money. For training or fine-tuning, he suggests 3090s instead.
Do I need water cooling?
No. The P40s are passive cards, and in the Dell R730 the server's own fans cool them. The guide's fan-control script keeps the noise down; expect it to be loud under load.
What models can I actually run with 48 GB?
The guide's target is a 70B model quantized to 4-bit, which uses about 35 GB of the 48 GB; the author measured about 1.4 tokens per second on Llama 2 70B 4-bit, and about 5 to 6 on 13B and 7B models. Full-precision 70B models do not fit. There is no model compatibility chart, so check each model's memory requirements yourself (most are listed on Hugging Face).
Community builds
No community builds yet — be the first, we feature the best ones.
Discussion1
FROM THE COMPAREE TEAM
Two used Tesla P40s in a retired Dell server give you 48 GB of VRAM, and the guide is honest that the server fans will sound like a jet until you tame them with a script. What stopped you the first time you tried to run a big model locally?
magiccodingman (GitHub username)
The creator maintains the Magic-AI-Wiki, a collection of budget AI build guides. The P40 guide is one of several cost-optimized paths to local inference hardware. No personal site or bio is in the repo.
DISCLAIMER
- Comparee is not the author of the projects featured here. All rights to each project belong to its creator — every page links to the original source, and we never host creators’ files.
- Information is provided without warranty and may become outdated as projects evolve. Prices are indicative bands only — always check the creator’s parts list for current costs.
- Building and operating any project is at your own responsibility. Protective equipment, safe workshop practice and compliance with local regulations are the builder’s responsibility.
CompareeTEAM1mo agoedited
Practical notes from our verification: the guide is one Markdown page in a larger personal wiki, not a standalone project, and it builds a two-card machine: two used Tesla P40s (48 GB of VRAM in total) in a retired Dell PowerEdge R730, which the author priced at about 1,100 dollars at late-2023 prices. The linked AI HOMELAB video is a separate community build with four P40s, not the guide author's machine. The single biggest trap is noise: the R730 does not recognise the P40s, so its fans run at full speed until you set up iDRAC fan control and the author's Python fan script, which the guide walks through. Use the 1100 W Dell PSUs it lists — the author withdrew an earlier 1600 W PSU recommendation because it did not work in the R730. Software setup is only sketched (Ubuntu recommended), and the author says you will need to look up many commands yourself. Photo note: the photographs on this page come from other P40 builds, each credited to its creator, not from the guide author's server. Correction (4 October 2026): we re-checked this page line by line against the project's own repository, documentation and videos, and fixed errors in earlier versions.