Quick answer

Buy the RTX 5080 at $1,399.99 if your model fits 16GB; our own bench has it decoding a 20B model at 178 to 190 tokens per second. Only a 48GB card holds the 42.52GB Llama 3.3 70B file. Skip the RTX 5090 for 70B work, because its 32GB misses that line and still costs $4,829.99.

On this page

Forty-two and a half gigabytes. That single number settles most of this purchase, and almost nothing else about either card does.

The exact Llama 3.3 70B Q4_K_M file is 42.52GB of weights. It does not fit on a 32GB RTX 5090. It does not fit across two 16GB cards, because those still add up to 32GB. A 48GB RTX A6000 is the only card in this comparison that holds it — and the A6000 is a 2020 board.

Everything that fits 16GB belongs to the RTX 5080.

Some retailer links below are affiliate links; if you buy through them TechFuel HQ may earn a small commission at no extra cost to you. As an Amazon Associate I earn from qualifying purchases. Commissions never influence what gets recommended — see our disclosure.

The line, and which cards cross it

CardMemoryBoard powerHolds 42.52GB of weights?Lowest August price
RTX A600048GB GDDR6 ECC300WYes — 5.48GB left over$5,490.00
RTX 509032GB GDDR7575WNo — short by 10.52GB$4,829.99
RTX 508016GB GDDR7360WNo$1,399.99

Memory and power figures come from the manufacturer spec tables on the NVIDIA RTX A6000, RTX 5090 and RTX 5080 pages; PNY publishes 768GB/s of bandwidth and a dual-slot active fansink for its A6000 board.

The fit is binary and the price is not. That is the whole shape of this decision: one card clears a hard capacity threshold, two do not, and the cheapest of the three is the one you probably want.

16GB is already fast, which is the awkward part

Our RTX 5080 throughput dataset recorded gpt-oss 20B at MXFP4 decoding 178 to 190 tokens per second across two dated captures, with the whole model sitting in 13,381 to 14,067 MiB of the card’s 16GB. Qwen 2.5 14B at Q4_K_M decoded at 94 to 97 tokens per second. Every cell is the median of three runs, and the published CSV carries the driver version, the Ollama build, the model digests and the resident memory readings.

That result is the reason the 48GB card is not an automatic upgrade. A 20B model and a 14B model are already resident and already quick on a $1,399.99 card. Paying nearly four times that for 48GB buys nothing at all unless the extra capacity changes which model can stay loaded. The VRAM-tier guide maps the full ladder, and the VRAM calculator works out the context-sensitive version of the same question.

Skip the RTX 5090 for 70B work

This is the finding that surprised us. If you are buying a card to reach a 70B model, the RTX 5090 is the worst of the three, and it is not close.

At $4,829.99 it costs $3,430 more than the RTX 5080 — and it stops 10.52GB short of the file that would justify the upgrade. It buys a bigger pool that lands on the wrong side of the threshold this page is about, then asks 575W to do it. Meanwhile the A6000 sits $660 above it and actually clears the line at 300W.

One band does belong to it, and the RTX 5080 cannot take it: a working set between 16GB and 32GB — a 32B model at Q4, say — is too big for 16GB and nowhere near needing 48GB. If that is your workload, the 5090 is the only card in this table that fits it. Nothing here recommends it above that band, and no matched throughput number for it was captured on this page.

Availability made the point again. All three RTX 5090 card rows in the August feed were checked on September 7 and every one had gone out of stock. If you are shopping for a 32GB pool you are paying a scarcity price for capacity that does not reach.

Buy the RTX 5080

Start at $1,399.99 and stop there unless your model genuinely will not fit 16GB.

That is the recommendation for nearly everyone reading this. A 20B model runs at 178 to 190 tokens per second on it by our own measurement, it draws 360W, it is a current-generation card with a current warranty, and it costs less than a third of either 48GB or 32GB alternative.

Check NeweggLink checked 2026-09-07Search AmazonLink checked 2026-09-07

Thirteen RTX 5080 cards were in stock in the August feed between $1,399.99 and $1,999.99, so there is real room to shop the exact board. Prices move — check before you order.

The one case for going after a used A6000

Buy 48GB when your resident working set lands above 32GB and comfortably below 48GB. That is a narrow window and it is the only window. A dense 70B at Q4 fits it. A 20B does not, and a 120B does not either.

Watch the headroom while you are in there. The weights take 42.52GB and leave 5.48GB for the KV cache and runtime buffers. Ordinary context lengths live in that space; long ones do not, and the failure shows up as an out-of-memory error partway through a session rather than at load time. Work out your context budget before you commit to the card, not after.

Three things the A6000 gives you beyond raw capacity: ECC memory, a 300W ceiling that fits a normal power supply, and a two-slot board in a market where the fast consumer cards have grown to three and four. NVIDIA also supports it on the RTX Enterprise driver branch rather than the Game Ready branch, which is the right trade for a machine that works and the wrong one for a machine that games.

What to make the seller put in writing

A used A6000 is a private-party or reseller transaction, and the coverage rides entirely on that paperwork. PNY publishes a three-year limited warranty on a new card, and NVIDIA’s warranty policy sends most claims to the reseller or board partner first. Neither of those facts follows a card to its second owner on its own. So before you pay, get:

  1. the exact part number and a photo of the serial;
  2. whether the card is used, refurbished, open-box, or pulled from a workstation;
  3. the return window, and who pays return shipping;
  4. the warranty provider and the remaining duration;
  5. enough of that return window left to run a full 48GB memory test and an hour of sustained load.

The fifth one is the one people skip. Run it the day the card arrives.

And know the ceiling. A new PNY board lists at $5,490.00 — $660 above the RTX 5090 and roughly four times the RTX 5080. A used card asking within a few hundred dollars of that is not a deal; the discount is the entire reason to accept a five-year-old GPU with someone else’s hours on it. If you would rather have the new-card warranty than the savings, this is the exact board:

Check NeweggLink checked 2026-09-07Search AmazonLink checked 2026-09-07

How this was checked

The capacity argument uses one exact model artifact rather than a rounded parameter count: the Hugging Face repository reports 42,520,398,816 bytes for Llama-3.3-70B-Instruct-Q4_K_M.gguf, which is 42.52GB. Memory, power and cooling figures were read off the current NVIDIA and PNY specification tables on September 7. Throughput comes from our own RTX 5080 dataset, captured on an RTX 5080 under Ollama with the method and raw CSV published alongside it — no A6000 was benchmarked here, which is why this page recommends the 48GB card on capacity and never on speed.

Prices are the lowest in-stock USD offer per exact model from the August 17 Newegg feed, restricted to actual video cards. No used price appears anywhere on this page: every used-A6000 listing checked on September 7 was expired, dead, or unreachable. That is itself the finding. Used A6000 supply is thin and it moves, so treat any single listing you find as a sample of one and price it against the $5,490.00 new-card ceiling above.

This page is re-checked every six months, and sooner if a new consumer card moves the 32GB ceiling or a major 70B release changes file size.

Frequently asked questions

What can a 48GB RTX A6000 run that a 32GB RTX 5090 cannot?
A dense 70B model at Q4. The exact Llama 3.3 70B Q4_K_M file is 42.52GB of weights, which overshoots 32GB by 10.52GB before the KV cache and runtime buffers are allocated at all. A 48GB card takes the weights and leaves 5.48GB for everything else. That is enough for ordinary context lengths and not enough for very long ones, so size your context before you buy the card.
Do two 16GB GPUs equal one 48GB card for local LLMs?
No. Two 16GB cards give you 32GB total when the runtime can split a model across them, and 32GB is exactly the ceiling the 42.52GB file already misses. Two cards solve a budget problem, not this capacity problem. If your model fits 16GB, one card is simpler and quicker than a split. Between 16GB and 32GB a pair can work, and no measured multi-GPU result here ranks it against one 32GB card.
Is a used RTX A6000 faster than an RTX 5080?
Buy it for capacity, not speed. The A6000 is a 2020 Ampere board with 768GB/s of GDDR6 and 48GB of it; the RTX 5080 is a current card with 16GB of much faster GDDR7. On models that fit 16GB, our own dataset has the 5080 decoding gpt-oss 20B at 178 to 190 tokens per second across two dated captures. The A6000 earns its price only when 48GB changes which model can stay resident.
Does a used RTX A6000 still have a warranty?
Assume nothing either way. PNY lists a three-year limited warranty on a new card and NVIDIA routes most warranty claims through the reseller or board partner, which means a used card’s coverage lives entirely in the paperwork between you and that seller. Get the return window, the warranty provider and the remaining duration in writing before you pay, and test the full 48GB under sustained load while the return window is still open.

Evidence ledger

Last updated
Methodology
This guide was written and edited by Lowell K. Wood IV in St. Louis County, MO. Specs and prices verified against vendor and project documentation current on the date above. Full editorial standard: methodology.
Update log
  • 2026-09-07 — Last reviewed and updated.
Corrections
Spotted an error or a stale number? Email hello@techfuelhq.com. Confirmed corrections are added to the update log above.

About the author

Written by Lowell K. Wood IV, who builds and runs TechFuelHQ from St. Louis, Missouri.