CPU Overheating or High Temperature (Gaming/Videos)

I hope everybody on the forum is doing well today. Recently, I suspect I have been having legitimate overheating issues on my system, so I will give some context. While playing games a couple weeks ago, I saw major framerate stutters, higher temperatures than I had ever noticed (qualitatively), and after another few minutes if I kept playing, my system would shut down. In addition, I would see very high CPU usage (multiple cores) on my system monitor. This was evidently a problem, and my first suspicion was that my CPU was failing and/or my primary CPU cooler was insufficient. I decided to attempt solving both issues simultaneously: replacing my AMD Ryzen 9 5900X (in use for approximately 3.5-4 years) with a very similar AMD Ryzen 5900XT, and my Scythe Mugen 5 with a Noctua NH-D15 (dual-fan setup). While I did not get quantitative temperature measurements before the hardware swaps, my temperature measurements now still seem concerning when playing more graphically intensive games, up to around 92°C sustained, with high heat tangibly radiating from the side of the case nearest the CPU heat sink. I will post screenshots below of temperature sensor readings from bottom while playing different games. Here are my current specs, please let me know if you need more info, e.g. from inxi:

Current Specs

Desktop PC

CPU: AMD Ryzen 9 5900XT (16-core)

Mobo: ASUS TUF Gamning X570-PRO (WI-FI)

Primary CPU Cooler: Noctua NH-D15 (dual-fan setup)

SSD (with OS partition): Western Digital WD Black SN770 1TB

RAM: 2x16GB Corsair Vengeance DDR4

GPU: NVIDIA GA102 [GeForce RTX 3080 12GB]

What I have tried troubleshooting so far:

HARDWARE:

-I opened up my case, cleaned everything thoroughly with compressed air and microfiber alcohol wipes, unmounted and cleaned the GPU, unmounted the new CPU cooler, cleaned the CPU and heat sink contact plate thoroughly, reapplied new thermal paste, and re-mounted the new cooler.

-In the process, I ensured that my heat sink contact and thermal paste were done well and cleanly, and feel confident they are, but the temperature problems (as I see them) persist. I also tested to make sure the fans were set up and running properly, and they certainly appear to be.

-I have not tried much with power profiles or performance throttling to limit temperatures (e.g. https://discovery.endeavouros.com/hardware/tools-to-lower-power-consumption-and-device-cooling/2021/03/), in part because I would like advice on the best approach for my system, and in part because I did not, qualitatively, see such poor thermal performance on my system until a few weeks ago with an older CPU and a, theoretically, worse CPU cooler, so I wonder if there is another root cause. I am happy to take any advice here.

SOFTWARE:

-Multiple system updates

-Made sure voltage supply to the CPU was adequate so that it is stable when running at peak frequencies (see: https://wiki.archlinux.org/title/Ryzen#Troubleshooting)

-A few other minor tests following previous forum posts, but no major changes (happy to elaborate)

Outstanding questions

Does this seem like it could still be a problem with my hardware, somewhere else in the system? A failing SSD (I have tested this with smarctl, for instance, and it appears okay)? Or does this seem more likely to be an issue with control of the cooling system from something in the software? Is this something simply resolved by throttling performance and maximizing the cooling output of the hardware I do have? As I mention above, I am happy to do more work to troubleshoot, and try stop-gap solutions, but would also like to identify likely root causes. Please also forgive my inexperience, and the length of this post, I am trying to be comprehensive, and show that I have done some research before just blabbing a forum post. I appreciate any help that anyone is willing to offer, thank you!

Screenshots of bottom outputs after playing two different games for approximately ten minutes: Blasphemous (less intense), and Elden Ring (more intense), respectively (from Steam via Proton)

From what I can gather so far, you’re doing everything right. The Noctua NH-D15 is the ants pants.

Under load, that’s what you would expect when things are performing as they should. If you didn’t feel heat being expelled under load, that might be indicative of a CPU cooling issue. Hot air being expelled is a cooling system performing it’s job.

I’d suggest turning your attention to your BIOS/UEFI, and first, carefully updating your BIOS, before then calibrating your cooling with the ASUS tools, and reviewing your performance settings.

It’d also be good to get an idea of your case’s cooling too. A CPU fan that’s the bees knees, is only really able to perform properly if the case is able to bring in cool air and remove hot air.

I’d only really flag the heat as an issue, if your system is having to throttle down in order to maintain safe temperatures.

Bink,

Thank you for the prompt reply. It does make sense to check how my case cooling setup/fans are performing. For reference, I have a Fractal Pop Air, with two case fans on the front bringing air toward the internal hardware, and a single case fan exhausting out the back (i.e. no fan exhausting out the top). Along with updating the BIOS and reviewing the CPU cooling settings, I can review the fan settings for the case fans. Please let me know if you have any specific recommendations for this, including ways to keep the optimized fan settings consistent with each boot.

Regarding the heat, don’t those values seem high to you? The Noctua NH-D15 certaintly has larger heat sinks, which bring it closer to the side of the case, so perhaps I’m just noticing the heat radiating out more distinctly than I did before, but the sustained high temperatures of over 90°C definitely seem like a problem, and I do not think I was witnessing quite these high of temperatures before. What’s more, I still see high idle temperatures on a lot of the sensors even after shutting games down for over an hour. Were there perhaps some adjustments I should have made after installing the new CPU? It’s hard to know for sure, but, for instance, when I start up now, I get warnings on my system monitor regarding high CPU_IOWAIT times, and I think I may not have the new CPU configured correctly.

Finally, after sorting out any potential software or fan control issues, if I do not want to invest signficantly in new hardware, my best bet to limit temperatures seems to be throttling performance (framerate, output resolution), and while that is less than ideal, I may be able to find an acceptable performance level before considering new purchases in thi$ in$ane consumer hardware environment. I’ve seen a few forum posts recommending some performance monitors/adjustments for gaming, but do you have any specific suggestions, too? Again, thank you for your time and input.

I’m running a Ryzen 9 5900X with a BeQuiet Dark Rock CPU cooler (I can’t recall the exact cooler model, a friend actually gifted it to me when he upgraded).

It’s currently idling at around 50°C, but if I put it under load, like optimising a directory full of PNG’s with oxipng, it’ll pretty quickly jump to 90°C. Sounds pretty similar to your situation actually.

btop

Given you have a Noctua NH-D15, which I would argue is superior to the BeQuiet, I am somewhat surprised your results are similar to mine. I’ve often wondered if my cooling was insufficient too, so I wouldn’t regard mine as the baseline. The 5900X has a base clock of 3.7GHz, so pushing all cores to 100% sustaining around 4.0GHz seems ok.

Ultimately though, you’re within acceptable tolerances if the CPU isn’t needing to throttle itself down, in order to maintain operational temperatures at full load.

With respect to tuning your cooling and CPU/memory, you’d do this primarily from your BIOS/UEFI, and I would certainly recommend that. I touched on that with this remark:

Thank you, Bink. Your performance information is a helpful benchmark that helps corroborate I’m not totally out of a reasonable range. Regardless, I still think I need to optimize the fan setup, so I will do that tomorrow morning when I have a bit more time. I am also open to further suggestions if anybody else on the forum might have input, or suggestions for further testing.

Based on your sensor readings, the actual CPUINT looks fine, only SMBUSMASTER 1, TSI0_TEMP, TSI1_TEMP could be addressed / investigated. Only in case of the nVME temperatures, I’ld definitely recommend to fix the issue.

If you don’t have an heatsink on your nVME, I’ld address that at first. Also, you may want to revisit the actual airflow within your case and consider adding case fans to prevent heat build-ups. Why ? The SMBUSMASTER 1 as well as the TSI*_TEMP sensors aren’t located in the CPU package, but on the motherboard itself, monitoring different thermal zones on the motherboard. They’re part of the Super I/O chip of the motherboard. But I can’t tell which physical location they’re actually monitoring on that specific motherboard.

Additionally, I doesn’t seems that you’re using one of the DKMS packages (e.g. zenpower3-dkms) to read out Ryzen specific temperatures. It’s only CPUINT which is accessible via bottom, but I don’t see Tccd1, Tctl or Tdie temperature data points from the CPU package.

I definitely would address the nVME temperatures. The other high readings might be not that critical.

In case you decide to install zenpower3-dkms, which is the one I’m using on my Ryzen 5 5600, don’t forget to run sensors-detect afterwards to register the additional sensor readings.

Also, there might be some weirdness due to the use of zenpower3-dkms. In short, on my system I’ve got some pretty high readings for some sensors (3892314°C). Which is most likely due to the fact that these sensors aren’t populated at all on the motherboard that I’ve got, which is an ASRock X300-ITX. But the most important ones are consistent and not skewed in their values.

There is also zenpower5-dkms within the AUR which is relatively new. But I haven’t tested that one, it may eventually fix the issue with some sensor readings. And isn’t exclusive to Zen 5, but does also support the earlier architecture versions of the Zen family of processors (and their respective motherboard chipsets).

Thank you for your input. I am actually able to see temperature readings from the CPU package for Tccd1, Tccd2, and Tctl; they were just higher up in the list of temperature sensors that I wasn’t able to display with a single screenshot. I have attached some new screenshots below of the output of bottom: showing temperature (and other) readings for, respectively, the system idling, with the CPU sensor readings between 43°C-57°C, and after ten minutes of playing Elden Ring, with the CPU temperature readings between 75°C-81°C.
IDLING:

AFTER TEN MINUTES OF ELDEN RING:

At that same time as the ten minutes of gaming screenshot above (lower down in the sensors list, so not visible in the screenshot), the SMBUSMASTER 1 sensor and nvme Sensor 1 show the much higher temps we were discussing before, so it looks like overheating is impacting those systems in a more pronounced way, i.e. that poor cooling/airflow in the case is the likelier culprit. This has been very helpful, thank you. I had read yesterday that SMBUSMASTER 1 was a sensor on the motherboard itself, and inferred that the other sensors listed under “nct6775.656” likely were as well; thank you for pointing out that Tccd1, Tccd2, and Tctl are CPU-specific sensors. I will look into using zenpower3-dkms to monitor these sensors more specifically.

My SSD does actually have a heatsink (see the “Cooling” tab on the manufacturer spec website here), but it may not be performing all that well. It is located toward the bottom of the motherboard, which is situated a little ways above the shroud for the power supply. As Gamer’s Nexus does a good job discussing in their review of the Fractal Pop Air case, the large size of the shroud all across the bottom of the case can impede airflow leading to worse thermal performance. At this point it seems like I may need to consider changing out my case, honestly. I can get better fans than the stock case fans, and add an exhaust fan to the top of the case behind the CPU heat sink to achieve better airflow, but I think the case itself might be quite flawed for air cooling, and at least on this system, a case with stronger fans and more space for good airflow seems a better bet. Do you happen to have any recommendations for quality case fans? Thank you again for your time.

I also have that cooler, the Noctua NH-D15, but as a processor I have a Ryzen 7 5700X, which probably heats up less.

However, due to the height of the RAM I have (they are an RGB model) and the shape of the case, I removed a 140mm fan and in its place I put 2 120mm ones, leaving a 140mm fan in the central part of the heatsink. So in total I have three fans on the heatsink.

But, aside from the fact that I think my case has decent airflow, I also added three 120mm (and 15mm thick) fans to the top of the case (where you put the fan tray for the liquid cooler, you either put that one there or you can optionally put three 120mm fans or two 140mm fans as far as my case is concerned).

However, since they are all Noctua PWM fans, I also “manually” set the speed based on the temperature from the bios, having more stringent parameters than the default ones.

By doing this I rarely exceed 70 degrees, even in summer and with long workloads.

But in any case your processor will heat up more than mine, because it has a higher thermal design power (TDP). An air heat sink therefore must also be chosen on the basis of that.

For your processor the TDP should be 105W (higher than mine, which is around 65W), so that means it will generate more heat under load than mine. As mentioned above, that should also be considered.

I don’t think there is necessarily a need for liquid cooling, but a good air heatsink suitable for that TDP is necessary. But the Noctua NH-D15 is fine for this purpose.

At most, if space allows, I would add another 140mm fan.

Your motherboard has two M.2 slots, both are capable of delivering PCIe 4.0x4. So if it’s the M.2_2 socket that you are using currently, consider switching it over to the upper slot.

That slot won’t be restricted to PCIe 3.0x4 as the Ryzen 9 5900XT doesn’t have integrated graphics.

Additionally, as bottom doesn’t has the capability to log the temperatures, I’ld recommend to use LACT which would not only offer temperature logging during a game session. But also would allow modifying the fan curves. And it provides overclocking / undervolting functionality as well as automatic power profile switching based upon launching specific binaries, such as steam.

Setting the CPU voltage a bit lower could have an significant impact on the temperatures of the CPU. With essentially no performance penalties. You’ll just dig around a little to find some suitable values for your specific CPU.

Also, as a frequency scaling governor, amd_pstate_epp set to active does work quite well.

Additionally, there is also the option to enable eco mode of the CPU via BIOS, limiting the TDP of the CPU which should be 105 Watt out of the box down to 65 Watt could also bring down CPU temperatures. This won’t have a huge impact on single threaded performances in the use case of gaming.

Just based on the k10temp tctl temperature of 82°C: I don’t see it that critical. But there isn’t much headroom until the system may throttle thermally.

The temperatures don’t seem to be out of range, it could be called over-, if you’re trending over 100ºC, if I remember correctly, the maximum allowable temperature for today’s CPUs is 110ºC, and 90ºC to 100ºC is acceptable, as seen in CPUs with hardware-vendor’s own stock coolers. There also is a system auto-shutdown mechanism.

[The TDP talk on the net and in the forum, is often misleading, because many confuse TDP with electricity consumption under load. TDP is the amount of energy that a cooling system has to transform into ambient heat, in order to reduce the CPU heat to the target heat of 99ºC or less. It’s different for winters and summers.]

@Sermor , @1093i3511 , and @cc_spicuous_2

Thank you very much for your input. @Sermor , regarding your fan setup, I can see the rationale for the changes to the fans on the heat sink, but I think I may not have enough space to add more than one fan to the top of the case if I want to maintain air flow through the heat sink which is then exhausted out of the case. This may not be accurate, but on that Gamers Nexus case review I linked above, the editor Steve was discussing that if your general airflow route is intake at the front of the case, through heat sinks/hot components on the mobo, and exhaust out the back/top, if you have an exhaust fan near the front of the case at the top, it will exhaust a fair portion of the front panel intake before it has a chance to hit the heat sink. Do you use an intake fan orientation at that location? Regardless, the Pop Air has a solid portion of metal near the front top of the case, so I could not place a fan there completely in front of the heat sink, anyway.

I appreciate the suggestion to try out lact, especially since it has the capability to log performance during a gaming session. I will do that after I have made the adjustments to my actual case fans and see how I like it, as well as post some results for reference.

Today I picked up some new fans: three Noctua NF-A14x25 G2s (140 mm square), and one Noctua NF-A12x25 G2 (120 mm square) to replace my three current case fans as well as add a new one to the back/top of the case for more exhaust. I will have time to install those tomorrow, and can report back on the new performance. Thank you so much again for the perspective, everybody! The old splitscreen meme of “Gaming during summer” versus “Gaming during winter” is really hitting home today, haha.

Okay, I think this post constitutes a solution. I installed two Noctua NF-A14x25 G2s (140 mm. square) fans to replace the front-of-case stock fans, and one Noctua NF-A12x25 G2 (120 mm. square) fan to replace the back-of-case stock fan. Unfortunately, with the configuration that my case forces for the motherboard/heat sink, the third 140 mm. fan is too large to fit the top of the case for an additional exhaust fan, so I will have to install another 120 mm. fan once I get it. Nevertheless…the thermal performance is already WAY better, even without messing very much with fan settings in the BIOS. After sustained workload from graphically-intensive games running for thirty minutes, sustained CPU temperatures are around 61ºC. GPU temperatures get up to about 80ºC at worst, and the nvme sensor for the SSD is at around 76-78ºC. See bottom outputs below. Granted, I finished the case strip/clean/fan installation late tonight, with the windows long having been open to the cool air outside, so conditions are currently ideal for summertime, but I don’t think performance would be that much worse in a room 20ºF or so warmer. Thank you very much for all of the help, everybody!

Bottom output after thirty minutes of gaming, showing GPU and CPU temperature sensor readings

Bottom output after thirty minutes of gaming, showing motherboard and SSD temperature sensor readings

Definitely learned a lot in this process!

…also, uh, I may have initially installed one of the two fans on my new heatsink (the one between the two towers of the heatsink, and thus less visible) backwards…so, yeh…I might have been screwing myself a bit on airflow. :man_facepalming: