[Sad News] My SSD broke...
Hello, this is Rcat.
This time it's a very trivial matter, but I'm so sad that I have to write about it. My SSD broke...
Next time: Automating monitoring
I was able to replace it
Introduction
Terms of Service
Please check the terms of service in advance when using information or works.
https://note.com/rcat999/n/nb6a601a36ef5
About comments
Please check the terms of service guidelines before commenting.
Situation
Target
This story is about the following PC.
The additional data SSD I installed has broken.
Situation
While I was leaving a backup (writing to Linux) running, various bots crashed or kept restarting, and the backup server also went down.
It was a hassle, so Irestarted it, but it wouldn't boot up at all...
When I connected a display, the following message appeared.
Welcome to emergency mode! After logging in, type "journalctl -xb" to view
system logs, "systemctl reboot" to reboot, "systemctl default" to try aga into boot into default mode.
Give root password for maintenance
(or type Control-D to continue):Seeing 'emergency' is way too scarythough.
Anyway, it seems it wouldn't boot because it was stuck on this screen.
Analysis
This is the first time this has happened, so I asked Chappy various things, and it seems the SSD is causing the trouble.
It's really helpful that I can just run the commands I'm told to and throw the results into image recognition text for analysis.

In other words,Linux is fine, but it couldn't boot because the filesystem that needed to be mounted was missing. That seems to be the situation.
Countermeasures
Since it just couldn't boot because it couldn't mount, I deleted the mount settings.
#/etc/fstabGoodbye to the bottom line
# /etc/fstab: static file system information.
#
# Use 'blkid' to print the universally unique identifier for a
# device; this may be used with UUID= as a more robust way to name devices
# that works even if disks are added and removed. See fstab(5).
#
# <file system> <mount point> <type> <options> <dump> <pass>
# / was on /dev/nvme0n1p5 during installation
UUID=4344a6a1-ca75-4b5a-8a53-fb0962fae053 / ext4 errors=remount-ro 0 1
# /boot/efi was on /dev/nvme0n1p1 during installation
UUID=D288-7E29 /boot/efi vfat umask=0077 0 1
/swapfile none swap sw 0 0
#UUID=6842d5c4-598b-445b-b172-6f6e223679fa /SSD ext4 defaults 0 0It booted successfully
Impact
It was a separate drive for data and was disconnected from the OS, so there is almost no impact on the server itself, but all external dependencies are dead.
Dify
Because I was basically using external storage for my work folders, Dify was stored here.
Therefore, the LLM chatbots I had been creating were completely wiped out. Several AI systems went down one after another.
The saving grace was that I had exported some data for distribution. Thanks to this, I was able to restore some things immediately, but
the prompt information for the ones I use privately was lost.
Future countermeasures
I rarely relied on Dify workflows anyway and only used it as a stepping stone for AI APIs, so from now on, I will write all prompts in the program itself and keep Dify as an empty bot.
Even so, in the rapidly changing field of AI, it is difficult to switch models, so considering that Dify makes it easy to change models, I will continue to use it.

Empty. I will use it only to change the model in the top right.
Image Generation AI
The image generation data, including the generated files, on a separate account was completely lost.
The ones I posted to note were downloaded, so they are safe, but the thousands of images I had been slacking off on recently are definitely lost.
Stable Diffusion itself and the bot programs for the API were all stored there, so those are lost too.
I have the bots on hand, but I remember rewriting the ones on the server from time to time, so I'm not sure if they are truly the latest versions.
I'll stop for now as a clean break, and maybe renew it if I feel like doing it again...
Backup
The backup destination for the backup app I run here was this SSD.
It's not a problem that the backups are gone, but if a failure occurs on the backup source at this timing, it would be seriously bad.
Future countermeasures
I will change to mutual backups.
Since it's based on running with a GUI, I'll have to change the client a bit, but since the server also runs on Windows, I'll run it on Windows as well and transfer data from Linux to the Windows PC as needed.
Isn't the impact bigger than I thought?
Well, that's true. My Sunday was completely ruined, though.
In the first place, the Linux server isn't my main PC, so even if the data is gone, it's not a loss that would cause me mental distress.
Of course, since there are various things running on the assumption of Linux, it's troublesome if it stops.
Rcat's data loss countermeasures
This time, a disk that didn't have much in the way of countermeasures broke, but I protect my main data quite robustly.
-
Main PC
RAID1
This writes to two disks simultaneously, so it's fine even if one of them breaks.
Since it's not the C drive, I don't put important things on C.Google Drive
I also keep a very small portion of important data in the cloud.
Introduction of the broken SSD
Finally, here is an introduction to the product that broke.
This is the SSD that broke.

I bought it two years ago. Since it's a 24-hour server, I guess it's about 18,000 hours.
*Since it's broken, I can't confirm. The number of power-on cycles might be low.
Was it a mistake to jump at a 1TB drive in the 5,000 yen range? It didn't last as long as I expected. Looking at the product name, it says 3-year warranty, so maybe it was a mistake to assume they were confident in it.
For now, I think I'll just assert that it's within the warranty period.

I was supposed to talk about stock price analysis, but that's gone to waste... See you next time.
いいなと思ったら応援しよう!
情報が役に立ったと思えば、僅かでも投げ銭していただけるとありがたいです。