Post
The repaste worked. The VRAM got 12 °C hotter.
Follows on from Putting an RTX 3090 in an HP ProLiant ML350 Gen9. The card is a second-hand blower-style 3090 in an enterprise tower, power-capped to 275 W.
I repasted the die and replaced every thermal pad on it. The paste job was an unambiguous success: the die hotspot dropped 26 °C and 38 minutes of accumulated thermal throttling went to zero. The GDDR6X memory, which I had also just repadded, came back 12 °C hotter.
That combination looks exactly like a botched repad. It isn’t one. The explanation is annoying enough, and non-obvious enough, that it seemed worth writing down.
§Why I opened the card
Not for the memory. An earlier benchmark had memory junction peaking at 90 °C under sustained memory-bound load. My own notes from that run said, in as many words, no repad indicated today.
That deserves a digression, because “90 °C is fine” sounds absurd if you’ve come from CPUs, and there are two different bars that get quoted interchangeably:
| Micron’s own rating for GDDR6X | 95 °C case (later revised to 95/105) |
| NVIDIA’s throttle point on the 3090 | 110 °C junction |
| What a stock 3080/3090 typically does at full load | 96–108 °C |
| What people who care about longevity aim for | under ~95–100 °C |
GDDR6X really does run this hot — PAM4 signalling is power-hungry, and on a 3090 half the 24 modules sit on the back of the board with nothing but the backplate to cool them. So 96–108 °C is genuinely normal and the card will not complain until 110.
But normal is not the same as harmless. Sustained operation above 100 °C has been associated with accelerated memory degradation in long-term community testing, and Micron’s own number is 95. “Within spec” here means “NVIDIA won’t throttle it”, not “this is a good place to live.” That distinction matters later in this post, so hold onto it.
So much for the memory. The die was the actual problem. Under a compute load at the 275 W cap, the card read 78 °C at the edge sensor while the hotspot sat at 102–103 °C. That ~25 °C gap is the signature of dried-out die paste: the edge sensor is nowhere near the hot part, so the card looks fine while the silicon cooks. It had also quietly accumulated about 38 minutes of cumulative thermal slowdown.
The delta is the diagnostic, not the absolute temperature. A hotspot-to-edge gap of 10–20 °C is normal. Around 25 °C means the interface has degraded. The delta is also the number to trust across seasons, because it cancels ambient — a hotspot reading on its own tells you as much about the room as the card.
§The guide
I followed a card-specific pad guide written by Pratheep Mahesan, posted to Reddit as u/Parking_Movie_2569 in February 2022, originally shared as a Google Drive PDF and since mirrored on the Internet Archive by someone who wanted it to survive the Drive link going away. Good instinct. (The document is signed Pratheep; the archive metadata spells it Pradeep.)
It’s a photographed, per-zone teardown for exactly this board, which is rarer than it should be — most repad guides are written for Founders Edition cards. His specification: GELID GP-Ultimate 12 W/mK, 1 mm front and 2 mm back, Thermal Grizzly Kryonaut on the die, and no pad behind the die.
It has no stated licence, so I’m linking and crediting rather than reproducing it. If you’re doing this job, go and read his document — the pad-placement photographs are the valuable part and I’m not going to re-host them.
Two details of his are easy to skip and worth calling out. He dry-fits the cooler with a single dot of paste and re-opens it to confirm transfer before peeling the pad films — a two-minute check that proves the cold plate actually meets the die. And he deliberately puts no pad behind the die, because bridging it to the backplate moves heat into a surface far worse at shedding it than the heatsink you just mounted.
§What I actually used, and where I deviated
Two deviations, both deliberate, and one of them matters more than anything else here.
Pads: a generic 12.8 W/mK silicone sheet, not the GELID he specifies. I bought 1 mm and 2 mm and ended up doing the entire card with the 1 mm, double-layering where 2 mm was called for. That is not thermally equivalent to a single 2 mm pad — you’ve introduced an extra pad-to-pad interface, and interfaces are where thermal resistance lives. I’ll come back to this. (Incidentally his document specifies “GELID GP-ULTIMATE 12 W/mK”, which conflates two products: GP-Extreme is 12 W/mK, GP-Ultimate is 15.)
Die: PTM7950 phase-change material, not Kryonaut.
This is the deviation I’d defend hardest, because it treats the cause rather than the symptom. My die interface hadn’t failed from bad luck — it had pumped out. Thermal paste under a bare GPU die lives with large temperature swings and high mounting pressure, and over time it migrates out of the gap. Kryonaut, which the guide specifies, is a strong performer with exactly this as its known weakness: it’s a relatively dry compound with a reputation for drying and pumping out on bare-die mounts. Re-pasting with Kryonaut would have bought me a couple of good years and then the identical 25 °C delta.
PTM7950 is a phase-change material: solid at room temperature, softening around 45 °C into a very thin, well-wetted bond line that doesn’t migrate. On paper it’s worse — 8.5 W/mK against Kryonaut’s ~12.5 — but bulk conductivity is a poor predictor of an interface, where bond-line thickness and contact resistance dominate. The measured delta below is the answer to that argument.
Two honest caveats. PTM7950 counterfeits are widespread on general marketplaces, and I can’t verify provenance on mine; what I can say is what it measured. And PTM beds in — it improves over dozens of heat cycles as the material thins and expels trapped air, so my numbers were taken on a fresh mount with very few cycles on it. Which makes a prediction, below.
§What nvidia-smi won’t tell you
On this card nvidia-smi reports the edge temperature and nothing else. Memory Current Temp is N/A, and hotspot isn’t exposed at all. Both of the numbers that actually matter for this job are invisible to the standard tool.
They’re readable from BAR0 registers. I used gddr6 for memory junction and ThomasBaruzier’s fork, which reports core, hotspot and memory in one shot. Worth knowing: on GDDR6 the memory figure comes from one register — the hottest of the 24 modules, not a per-module readout. Per-channel readout only exists on the GDDR7 path.
§The paste result
Compute-bound load, 275 W, steady state:
| before | after | Δ | |
|---|---|---|---|
| edge | 76–80 °C | 66 °C | −10…−14 |
| die hotspot | 102–103 °C | 76 °C | −26 |
| hotspot − edge | ~25 °C | 10 °C | back in the healthy band |
| cumulative thermal slowdown | ~38 min | 0 µs | gone |
A colleague independently reproduced this on the real production workload rather than my synthetic one and landed on hotspot 75.9 °C, delta 10.8 °C — agreement within noise from a completely different kind of load. Decode throughput stayed flat across a 15-minute soak with no droop, which is the behaviour that actually demonstrates throttling is gone; a card still throttling droops as it heat-soaks.
A prediction I’m willing to be wrong about in public: that 10 °C delta was measured on a PTM7950 mount a few hours old. The material is known to improve over dozens of thermal cycles as it thins and expels trapped air — reports of hotspots settling several degrees over the first weeks are common. So this number should get better, not worse, and 10 °C is probably a ceiling on the delta rather than a floor. I’ll re-run the same benchmark in a month and update this post. If it hasn’t moved, that’s worth knowing too — and would raise a fair question about whether my sheet was genuine.
§The memory result, which looked like a failure
Same memory-bound benchmark as the pre-repaste baseline:
| before | after | |
|---|---|---|
| memory junction peak | 90 °C | 102 °C |
| core peak | 72 °C | 65 °C |
| memtest errors | 0 | 0 |
| sustained bandwidth | 662 GB/s | 678 GB/s |
Hotter memory, right after replacing every pad on the card. The obvious reading is that I’d done the pads badly — a gap, a wrong thickness, a module not making contact.
Then I pinned the fan and ran it again:
fan on auto (61–64%): memory junction 96–102 °C core 64–65 °C
fan forced to 100%: memory junction 90 °C core 55 °C
At matched airflow the memory returns to exactly its pre-repaste number.
§The fan curve keys off the edge sensor, and is blind to GDDR6X
That’s the whole mechanism. The GeForce fan curve responds to edge temperature. It cannot see memory junction at all.
Before the repaste, the die was hot, so the edge sensor was hot, so the fan ran hard — and the memory got a lot of airflow as a free side-effect of the die being in trouble. Fixing the paste dropped the edge temperature by about 14 °C, the fan curve responded by backing off from near-100% to 61–64%, and the memory lost airflow it had been getting for free.
The memory didn’t get worse. Its cooling budget did.
So a successful repaste can raise your VRAM temperature, and the symptom is indistinguishable from the pad job you performed in the same sitting. If I’d only run the memory-bound benchmark — which was my existing baseline, and the obvious thing to re-run — I’d have concluded I’d botched the pads and taken the card apart a second time.
Never compare junction temperatures across runs at an uncontrolled fan speed. Pin the fan, or you’re measuring the fan curve.
§What I can and can’t say about the pads
I can’t check 24 modules individually; the sensor reports one number. But that number is a maximum, which is the right statistic for is any single module bad — a module with a bad pad becomes the hottest one by definition and can’t hide behind the others.
Three things say the pads are fine:
- at matched airflow the maximum returns to precisely the pre-repaste value
- sustained bandwidth is the highest I’ve recorded on this card, with zero memtest errors, so nothing is mechanically wrong
- the cooldown curve. A bad pad is a high-resistance path, so it runs hotter and cools slower. With the fan pinned and the load cut, memory fell 18 °C in the first 12 seconds, with a time constant of ~37 s against the die’s ~25 s. A module sitting on an air gap cannot shed heat that fast.
What that doesn’t rule out: a module that isn’t currently the hottest could have mediocre contact and sit invisibly below the maximum. Only a thermal camera on the backplate under load, or pad witness-marks on disassembly, would settle that.
§My numbers don’t match the guide’s, and that’s the useful part
Pratheep reports memory junction going 102 °C → 76–78 °C. I got no memory improvement at all. Five reasons, none of which means his guide is wrong:
- His card was actually broken. Mine wasn’t. His 102 °C was measured with the fan already at 100%. Mine was 90 °C on auto, which I’d already assessed as healthy. He recovered genuinely cooked pads; I replaced pads that weren’t the problem, so there was nothing to recover.
-
He added a backplate cooler — four heat pipes and its own fan, at step 17. Half of a 3090’s 24 memory modules are on the back, cooled only through the backplate. That’s a large share of his gain, and I didn’t do it.
It has also, as far as I can tell, never been separated out. A commenter on his thread asked exactly the right question — “do you recall what were your memory VRAM temps without the additional cooling heatsink on the back?” — and I can’t find an answer to it. So his headline 76–78 °C has always been a pads plus backplate-cooler figure, and the isolated contribution of the pads alone was never published, by him or anyone else I could find. Which means my failure to reproduce his number isn’t really a failed replication: there was no pads-only number to replicate. 3. His after-configuration is a mining tune that deliberately favours memory: core clock −300 MHz, so far less die heat into the shared heatsink, with the fan pinned at 75–80%. 4. Different thermal environment — and mine is probably the better one. It is tempting to assume a blower in a server chassis is at a disadvantage against an open desktop case. The measurements say otherwise, and I’ll come back to it below: a tower server moves a great deal of air through a long, unobstructed chassis, and this card was benefiting from it. 5. I didn’t follow his pad spec. Generic 12.8 W/mK sheet rather than GELID, and 1 mm double-layered where he calls for a single 2 mm. The extra pad-to-pad interface is real thermal resistance — though I want to be fair to it, because stacking is established practice: one 3090 FTW3 owner’s best result came from deliberately stacking 2.0 + 0.5 mm pads on the backplate side. So it’s a cost, not a blunder. Either way: if you want his result, use his pads at his thicknesses. I can’t claim to have tested his procedure, only to have been informed by it.
The lesson isn’t repads don’t work. It’s measure your starting point before you open the card. A repad on memory already sitting at 90 °C has almost nothing to give. The same job on a card at 102 °C with the fan pegged is transformative. Same procedure, same pads, completely different value — and the only way to know which one you have is to measure first.
§Where that leaves my numbers against everyone else’s
“Unchanged” is only a good outcome if you know what good looks like, so I went looking for other people’s post-repad results. The important thing I got wrong at first was comparing against the wrong class of card: most published repad results are Founders Edition or triple-fan coolers in open cases. Mine is a blower, which has one radial fan and one exhaust path for the whole board, and a small backplate doing all the work for the twelve rear memory modules.
Blower-specific results, under sustained load:
| card | before | after | notes |
|---|---|---|---|
| Gigabyte 3090 Turbo (mining farm) | ~105 °C | 90 °C | backplate pads only |
| Gigabyte 3090 Turbo (teardown) | — | 102–104 °C | 1 mm front / 2 mm back, 80% fans |
| ASUS 3090 Turbo | 98–100 °C | 94 °C | copper shims, plus an extra backplate blower |
| ASUS 3090 Turbo | 98 °C | 80 °C | pads plus backplate heatsinks and two extra fans |
| 3090 FE | 110 °C | 84–86 °C | open-air, dual axial |
| triple-fan AIB | 102–110 °C | 80–92 °C | open-air |
Mine: 90 °C on a memory-bandwidth torture test with the fan pinned, 94.5 °C on the real workload at auto fan, 89.7 °C on the real workload with a fan floor.
Two things fall out of that, and the second one is the useful one.
First: for a blower, this is a fine result. It matches the better of the two published Gigabyte Turbo outcomes and beats the other by more than 12 °C. The blower band post-repad is roughly 80–94 °C, and I’m sitting in it.
Second, and more interesting: every blower result below about 88 °C in that data was achieved with supplementary backplate cooling. The ASUS owner who got to 80 °C did it with backplate heatsinks and two added fans. The one who got to 94 °C already had an extra backplate blower fitted. Nobody reached the low 80s on pads alone.
That is exactly step 17 of the guide I followed — the four-heat-pipe backplate cooler with its own fan — and it’s the step I skipped. So the data doesn’t just say my result is normal; it says the remaining headroom on this card isn’t in the pads at all, it’s on the back of the board. If I want the low 80s, that’s where it is.
And it reframes the whole job retrospectively. My card was at 90 °C before I touched it — the same number a Gigabyte Turbo owner reported after a successful repad. The pads weren’t merely “not the problem”. They were already performing as well as a good repad would have made them.
§The fan floor, and what it costs
If a cooler die means a lazier fan, the fix is a fan floor. On the real workload, pinning 80% pulled memory from 94.5 °C to 89.7 °C.
It isn’t free, which surprised me. The same arm decoded slightly slower — and at first it looked like a 2.9% penalty, which would have changed the decision. It wasn’t: that figure came from a single warmup request. Across 52 requests per arm the difference was 0.8%, matching an 0.66% difference in sustained clock.
The mechanism is real though. power.draw is total board power, and the blower is on that budget. At a hard 275 W cap, fan watts come out of core watts. The tell was the ordering: the pinned-fan arm ran second and cooler, and thermal leakage would make a cooler arm clock higher. It clocked lower. Cooler-but-slower at identical board power can only be the fan spending the budget.
So: about 5 °C of memory junction for under 1% of throughput.
And this is where the two bars matter. On auto, under a memory-heavy load, the card now sits at 96–102 °C — normal by NVIDIA’s standard, well short of the 110 °C throttle, and squarely in the band long-term testing associates with accelerated memory degradation. With a floor applied it sits at ~90 °C, under Micron’s own 95 °C figure.
That reframes the decision. It isn’t “5 °C cooler for a rounding error of throughput” — it’s moving the memory out of the degradation-associated band for under 1% of throughput, on a card whose memory I’d like to still be working in three years. Put that way it stops being a tuning preference and starts being maintenance.
Worth taking. But it’s a trade, not a free win, and I’d rather say so.
§If you’re about to do this
- Measure first, with the right instrument.
nvidia-smicannot see either number that matters here. - Judge the paste by the hotspot−edge delta, not absolute temperature. ~10–20 °C healthy, ~25 °C degraded, and it cancels ambient.
- Use two loads. A memory-bound test barely heats the die and cannot evaluate a repaste. A compute-bound test doesn’t stress memory the way real work does.
- Pin the fan before comparing anything to anything.
- Trust the throttle counters. Cumulative thermal slowdown going to zero is a hard, unambiguous signal in a field full of soft ones.
- Don’t mistake “it isn’t throttling” for “it’s fine.” 110 °C is where the card protects itself, not where the memory is happy. If you want it to last, aim under 95–100 °C and treat the gap between those two numbers as yours to manage.
- On a blower, the last 10 °C is on the back of the board, not in the pads. Every published blower result below ~88 °C used added backplate cooling. If you’ve fitted good pads and you’re still in the 90s, stop buying better pads and put a heatsink and a fan on the backplate.
A note on evidence. The load-bearing comparison here is the within-session fan-matched A/B, which was controlled. The comparison to the historical baseline is supporting: I did not record fan speed on the original run, because it didn’t occur to me that it mattered — which is, of course, the entire point of this post. All post-repaste numbers are from a fresh PTM7950 mount and should be treated as a first data point in a curve, not a settled value.