Frequent crashes

Hi, I’m a very long-time Linux noob using Linux as my sole OS. I’ve been sick for a while now, and it really limits my ability to think or analyse, so I’d appreciate if you could help me.

I ran memtest and it had 0 errors, full pass. I have 32GB RAM. I had to reinstall my EOS recently because I couldn’t re-hook partition after a faulty update. I always update with yay. I pretty much have nothing foreign installed - I have steam and even firefox and dragon installed through flatpak. I reinstalled firefox as flatpak due to the crashes, but that didn’t help.

I don’t remember if this was the case with the previous install, but I have no swap. I have a lot of RAM and a good CPU, I think, so I didn’t think I’d need swap.

Operating System: EndeavourOS 
KDE Plasma Version: 6.7.1
KDE Frameworks Version: 6.27.0
Qt Version: 6.11.1
Kernel Version: 7.0.14-arch1-1 (64-bit)
Graphics Platform: Wayland
Processors: 24 × AMD Ryzen 9 5900X 12-Core Processor
Memory: 32 GiB of RAM (31.2 GiB usable)
Graphics Processor: AMD Radeon RX 6600
Manufacturer: ASUS

are my specs. All packages are fully updated.

journalctl only shows me the logs since the current session, so I can’t copy whatever error I get whenever the system crashes. It seems to crash at times I wouldn’t expect it to be overwhelmed - it runs modern games grand, but it crashes when I interact with UIs, be it in-game, in my browser, or I think even when interacting with the KDE menus. I think I even had it crash just after start-up a couple of times.

My bios is fully updated.

I believe my SSD is fully updated.

fwupdmgr update shows:

Devices with the latest available firmware version:
 • UEFI CA
 • UEFI dbx
Devices with no available firmware updates:
 • ASUSTeK KEK Certificate
 • Windows Production PCA
 • KEK CA
 • Master Certificate Authority
 • Master Certificate Authority
 • Option ROM UEFI CA
 • SSD 970 EVO Plus 500GB
 • SW Key CA
 • System Firmware

I tried looking around in UEFI to see if I can change anything. All the things which can be set to auto are set to auto.

I have almost 300GB of free space on my SSD.

I don’t use sleep nor hibernate - both of these I had fully disabled.

inxi -Fxxc0z --no-host

System:
  Kernel: 7.0.14-arch1-1 arch: x86_64 bits: 64 compiler: gcc v: 16.1.1
  Desktop: KDE Plasma v: 6.7.1 tk: Qt v: N/A wm: kwin_wayland dm: N/A
    Distro: EndeavourOS base: Arch Linux
Machine:
  Type: Desktop System: ASUS product: N/A v: N/A serial: <superuser required>
  Mobo: ASUSTeK model: PRIME B550M-K v: Rev X.0x
    serial: <superuser required> part-nu: SKU Firmware: UEFI
    vendor: American Megatrends v: 4101 date: 10/16/2025
CPU:
  Info: 12-core model: AMD Ryzen 9 5900X bits: 64 type: MT MCP arch: Zen 3+
    rev: 2 cache: L1: 768 KiB L2: 6 MiB L3: 64 MiB
  Speed (MHz): avg: 3593 min/max: 567/4955 boost: enabled cores: 1: 3593
    2: 3593 3: 3593 4: 3593 5: 3593 6: 3593 7: 3593 8: 3593 9: 3593 10: 3593
    11: 3593 12: 3593 13: 3593 14: 3593 15: 3593 16: 3593 17: 3593 18: 3593
    19: 3593 20: 3593 21: 3593 22: 3593 23: 3593 24: 3593 bogomips: 177252
  Flags-basic: avx avx2 ht lm nx pae sse sse2 sse3 sse4_1 sse4_2 sse4a ssse3
Graphics:
  Device-1: Advanced Micro Devices [AMD/ATI] Navi 23 [Radeon RX 6600/6600
    XT/6600M] vendor: Sapphire driver: amdgpu v: kernel arch: RDNA-2 pcie:
    speed: 16 GT/s lanes: 16 ports: active: HDMI-A-1 empty: DP-1, DP-2, DP-3,
    Writeback-1 bus-ID: 0d:00.0 chip-ID: 1002:73ff
  Display: wayland server: Xwayland v: 24.1.12 compositor: kwin_wayland
    driver: gpu: amdgpu display-ID: 0
  Monitor-1: HDMI-A-1 model: Idek Iiyama PL2492H res: 1920x1080 hz: 100
    dpi: 93 diag: 604mm (23.8")
  API: EGL v: 1.5 platforms: device: 0 drv: radeonsi device: 1 drv: swrast
    gbm: drv: radeonsi surfaceless: drv: radeonsi wayland: drv: radeonsi x11:
    drv: radeonsi
  API: OpenGL v: 4.6 vendor: amd mesa v: 26.1.3-arch1.2 glx-v: 1.4
    direct-render: yes renderer: AMD Radeon RX 6600 (radeonsi navi23 ACO DRM
    3.64 7.0.14-arch1-1) device-ID: 1002:73ff display-ID: :0.0
  API: Vulkan v: 1.4.350 surfaces: N/A device: 0 type: discrete-gpu
    driver: mesa radv device-ID: 1002:73ff
  Info: Tools: api: clinfo, eglinfo, glxinfo, vulkaninfo
    de: kscreen-console,kscreen-doctor wl: wayland-info x11: xdpyinfo,xprop
Audio:
  Device-1: Advanced Micro Devices [AMD/ATI] Navi 21/23 HDMI/DP Audio
    driver: snd_hda_intel v: kernel pcie: speed: 16 GT/s lanes: 16
    bus-ID: 0d:00.1 chip-ID: 1002:ab28
  Device-2: Advanced Micro Devices [AMD] Starship/Matisse HD Audio
    vendor: ASUSTeK driver: snd_hda_intel v: kernel pcie: speed: 16 GT/s
    lanes: 16 bus-ID: 0f:00.4 chip-ID: 1022:1487
  API: ALSA v: k7.0.14-arch1-1 status: kernel-api
  Server-1: sndiod v: N/A status: off
  Server-2: PipeWire v: 1.6.7 status: active with: 1: pipewire-pulse
    status: active 2: wireplumber status: active 3: pipewire-alsa type: plugin
    4: pw-jack type: plugin
Network:
  Device-1: Realtek RTL8111/8168/8211/8411 PCI Express Gigabit Ethernet
    vendor: ASUSTeK RTL8111H driver: r8169 v: kernel pcie: speed: 2.5 GT/s
    lanes: 1 port: f000 bus-ID: 0a:00.0 chip-ID: 10ec:8168
  IF: enp10s0 state: up speed: 1000 Mbps duplex: full mac: <filter>
Bluetooth:
  Device-1: TP-Link Bluetooth USB Adapter driver: btusb v: 0.8 type: USB
    rev: 1.1 speed: 12 Mb/s lanes: 1 bus-ID: 3-3:3 chip-ID: 2357:0604
  Report: rfkill ID: hci0 rfk-id: 0 state: down bt-service: disabled
    rfk-block: hardware: no software: no address: see --recommends
Drives:
  Local Storage: total: 465.76 GiB used: 136.15 GiB (29.2%)
  ID-1: /dev/nvme0n1 vendor: Samsung model: SSD 970 EVO Plus 500GB
    size: 465.76 GiB speed: 31.6 Gb/s lanes: 4 serial: <filter> temp: 48.9 C
Partition:
  ID-1: / size: 455.41 GiB used: 135.99 GiB (29.9%) fs: ext4
    dev: /dev/nvme0n1p2
Swap:
  Alert: No swap data was found.
Sensors:
  System Temperatures: cpu: 54.4 C mobo: N/A gpu: amdgpu temp: 53.0 C
    mem: 54.0 C
  Fan Speeds (rpm): N/A gpu: amdgpu fan: 0
Info:
  Memory: total: 32 GiB available: 31.24 GiB used: 3.92 GiB (12.5%)
  Processes: 443 Power: uptime: 15m wakeups: 0 Init: systemd v: 261
    default: graphical
  Packages: 959 pm: pacman pkgs: 946 pm: flatpak pkgs: 13 Compilers:
    gcc: 16.1.1 Shell: Bash v: 5.3.15 running-in: konsole inxi: 3.3.40

Can you help me?

Welcome to the community @pallid :waving_hand::smiley: :enos_flag:

You’ve shared good information so far, and ruled out a number of possibilities, so you’ve already made some good diagnostic headway.

You can retrieve previous session logs like this:

Current:

journalctl -b -0

Previous:

journalctl -b -1

One prior to that (and so on…):

journalctl -b -2

You might pipe that into eos-sendlog and share the link, rather than pasting your system log here. So for example:

journalctl -b -1 | eos-sendlog

Always a good idea to have swap setup. While I don’t fully understand it from what I have learnt it is different to RAM.

Thanks for the explanation. eos-sendlog only allowed me to check one -[x] (x=number) and now gives an error. Checking in the terminal is difficult. I can’t remember what the error looks like, even though I remembered it by heart before. I’ll try to do this when I get another crash, and then I’ll respond here again.

Ok, rather than eos-sendlog for a moment, can you try sharing the last 100 entries of a system log you know had the issue?

Replace 1 with the log entry you want:

journalctl -b -1 | tail -n 100

welcome to the forum.

you had a bad update that caused “I couldn’t re-hook partition”. boot or /root or the whole thing? Not sure what that means. the bad update kept you from booting eos?

instead of chrooting (don’t blame you chrooting is a pita) you re-installed EOS. I would have too.

now you crash all the time and hard reboot I assume.

can you blame SWAP alone? probably not. you can shrink /root in gparted and add some swap just to rule it out.

memtest verifies the sticks themselves are fine.

but what does s.m.a.r.t tests say about disk? because it very much sounds like hardware to me: new install, constant crashes

fwupd manager just might be a red herring. firmware was massively updated on all distros recently. my own output looked identical to yours two days ago.

Do everything @Bink says because that may or may not rule out software as an issue. at this point in the troubleshooting you have to reach the point where you can determine if it’s one or the other then act

there’s other factors you did not mention (eos gets the whole SSD? do you have more than one SSD that boots?) etc.

just throwing out Random Thoughts. good luck to you

edit/typo

Often what leads to a thought that has a goal, they’re like gatcha games some are awesome others are why, just why lol

Your system as you say is currently up to date. Even the UEFI Bios is current. What exactly are you experiencing when you say you have frequent crashes. Is the system crashing or is it a browser crashing? As @Bink has suggested you could provide some journal logs. The hardware output looks fine.

How to do the S.M.A.R.T. tests?

_________________________________

By a crash I mean my system reboots itself.

_________________________________

I haven’t been able to get it to crash today. The only thing I found is when I was trying to clear my cache with -Scc, it would pop errors speaking of directories, so I went to the cache location manually, and there was multiple folders with names which were seemingly random strings of numbers and letters. I forcefully deleted them by opening the directory as an admin. I’m not sure if this is related, although I haven’t had the chance to use my computer much today at all. If I end up getting the crash, I’ll get the journalctl report.

From other oddities I’ve spotted, in system info, next to my GPU it says (discrete). I don’t have a discrete GPU. My CPU was bought specifically so it didn’t have integrated GPU. Does discrete mean something else in this case? The GPU is a Sapphire AMD Radeon RX 6600.

Also, I’ve noticed my custom sensors watching over my GPU fan and temp disappeared from the System Monitor, which now also says “This page is missing some sensors and will not display correctly.”. Is it possible the system somehow wrongly detects my GPU, causing problems?

I’m very sorry for no journalctl, although I’m happy over no crash. I was recalled to work in spite of my leave due to staff shortages, so I’m worse off than normally, so sorry for anything which doesn’t make sense. :frowning:

Don’t think these folders would have caused problems, as far as I know that is a long standing bug.
This link can tell you more.

And I also noticed that it has to do with packages that have been installed (and updated) from the AUR

I’ve just crashed. :frowning:

I did the log thingy:

https://dpaste.com/76M7B66LW

It said ‘system fatal error’

In this topic I can see just about the same error messages as you get.

This seems to be related to the Baloo file indexer.
Disabling Baloo might fix the problem for now.

I don’t see that thread mentioning any crashing…?

Could it be a Steam or Flatpak issue?

I believe I’d get crashes even after turning off Steam, but my brain is very hazy nowadays.

The smartctl command lets you interact with the disks, get various health info, launch tests, etc.
e.g.
smartctl -a /dev/sda
to get most info

smartctl -t short /dev/sda
to launch short test, etc.

I hope I’m not responding too fast.

I get the result: Smartctl open device: /dev/sda failed: No such device

so I did lsblk from which I got

NAME        MAJ:MIN RM   SIZE RO TYPE MOUNTPOINTS
nvme0n1     259:0    0 465.8G  0 disk 
├─nvme0n1p1 259:1    0     2G  0 part /efi
└─nvme0n1p2 259:2    0 463.8G  0 part /

I only have 1 SSD and it is all for my EOS.

I tried smartctl -a nvme0n1p2 but that didn’t work, so I tried nvme list and got

Node                  Generic               SN                   Model                                    Namespace  Usage                      Format           FW Rev  
--------------------- --------------------- -------------------- ---------------------------------------- ---------- -------------------------- ---------------- --------
/dev/nvme0n1          /dev/ng0n1            S4EVNX0RB13312H      Samsung SSD 970 EVO Plus 500GB           0x1        166.04  GB / 500.11  GB    512   B +  0 B   2B2QEXM7

and so I tried smartctl -a /dev/nvme0n1, got permission denied, then sudo-d it and got:

smartctl 7.5 2025-04-30 r5714 [x86_64-linux-7.0.14-arch1-1] (local build)
Copyright (C) 2002-25, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Number:                       Samsung SSD 970 EVO Plus 500GB
Serial Number:                      S4EVNX0RB13312H
Firmware Version:                   2B2QEXM7
PCI Vendor/Subsystem ID:            0x144d
IEEE OUI Identifier:                0x002538
Total NVM Capacity:                 500,107,862,016 [500 GB]
Unallocated NVM Capacity:           0
Controller ID:                      4
NVMe Version:                       1.3
Number of Namespaces:               1
Namespace 1 Size/Capacity:          500,107,862,016 [500 GB]
Namespace 1 Utilization:            166,036,140,032 [166 GB]
Namespace 1 Formatted LBA Size:     512
Namespace 1 IEEE EUI-64:            002538 5b11b084f1
Local Time is:                      Tue Jun 30 21:34:50 2026 IST
Firmware Updates (0x16):            3 Slots, no Reset required
Optional Admin Commands (0x0017):   Security Format Frmw_DL Self_Test
Optional NVM Commands (0x005f):     Comp Wr_Unc DS_Mngmt Wr_Zero Sav/Sel_Feat Timestmp
Log Page Attributes (0x03):         S/H_per_NS Cmd_Eff_Lg
Maximum Data Transfer Size:         512 Pages
Warning  Comp. Temp. Threshold:     85 Celsius
Critical Comp. Temp. Threshold:     85 Celsius

Supported Power States
St Op     Max   Active     Idle   RL RT WL WT  Ent_Lat  Ex_Lat
 0 +     7.80W       -        -    0  0  0  0        0       0
 1 +     6.00W       -        -    1  1  1  1        0       0
 2 +     3.40W       -        -    2  2  2  2        0       0
 3 -   0.0700W       -        -    3  3  3  3      210    1200
 4 -   0.0100W       -        -    4  4  4  4     2000    8000

Supported LBA Sizes (NSID 0x1)
Id Fmt  Data  Metadt  Rel_Perf
 0 +     512       0         0

=== START OF SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

SMART/Health Information (NVMe Log 0x02, NSID 0x1)
Critical Warning:                   0x00
Temperature:                        41 Celsius
Available Spare:                    100%
Available Spare Threshold:          10%
Percentage Used:                    3%
Data Units Read:                    54,629,966 [27.9 TB]
Data Units Written:                 74,191,859 [37.9 TB]
Host Read Commands:                 407,988,596
Host Write Commands:                894,701,063
Controller Busy Time:               1,937
Power Cycles:                       2,789
Power On Hours:                     1,685
Unsafe Shutdowns:                   146
Media and Data Integrity Errors:    0
Error Information Log Entries:      5,519
Warning  Comp. Temperature Time:    0
Critical Comp. Temperature Time:    0
Temperature Sensor 1:               41 Celsius
Temperature Sensor 2:               45 Celsius

Error Information (NVMe Log 0x01, 16 of 64 entries)
Num   ErrCount  SQId   CmdId  Status  PELoc          LBA  NSID    VS  Message
  0       5519     0  0x0014  0x4004      -            0     0     -  Invalid Field in Command

Self-test Log (NVMe Log 0x06, NSID 0xffffffff)
Self-test status: No self-test in progress
Num  Test_Description  Status                       Power_on_Hours  Failing_LBA  NSID Seg SCT Code
 0   Extended          Completed without error                1601            -     -   -   -    -
 1   Extended          Completed without error                1601            -     -   -   -    -
 2   Short             Completed without error                1601            -     -   -   -    -

Since you’ve ruled out RAM issues, I would stress-test the GPU to check for possible hardware faults. A failing PSU can also cause seemingly random, unexplained crashes, especially under load.

Note: A discrete GPU is a graphics card that’s separate from the processor. It has its own dedicated video memory (VRAM), which is not shared with the CPU.

@pallid
Can you post the following?

journalctl -b -1 -p 3 -xb

Edit: Run this after a reboot.

Edit: Trying to determine if it’s cpu or gpu related.

[nero@NerosLinux ~]$ journalctl -b -1 -p 3 -xb
Jun 30 21:28:23 NerosLinux org_kde_powerdevil[1143]: [  1143][  1.743591] Time since library i>
Jun 30 21:28:23 NerosLinux org_kde_powerdevil[1143]: [  1143][  1.743596] Extra delay starting>
lines 1-2/2 (END)

I had an obscure memory of the GPU fan not being reactive out-of-the-box and needing extra stuff just to make sure it runs, so I’ve installed LACT. I’m not sure how to read the logs to tell if GPU was overheating, or something, but maybe it was that.