this post was submitted on 19 Jul 2024

633 points (98.6% liked)

Technology

58073 readers

3072 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

founded 1 year ago

MODERATORS

[email protected]

633

An angry admin shares the CrowdStrike outage experience (www.theregister.com)

submitted 1 month ago by [email protected] to c/[email protected]

108 comments fedilink hide all child comments

IT administrators are struggling to deal with the ongoing fallout from the faulty CrowdStrike update. One spoke to The Register to share what it is like at the coalface.

Speaking on condition of anonymity, the administrator, who is responsible for a fleet of devices, many of which are used within warehouses, told us: "It is very disturbing that a single AV update can take down more machines than a global denial of service attack. I know some businesses that have hundreds of machines down. For me, it was about 25 percent of our PCs and 10 percent of servers."

He isn't alone. An administrator on Reddit said 40 percent of servers were affected, along with 70 percent of client computers stuck in a bootloop, or approximately 1,000 endpoints.

Sadly, for our administrator, things are less than ideal.

Another Redditor posted: "They sent us a patch but it required we boot into safe mode.

"We can't boot into safe mode because our BitLocker keys are stored inside of a service that we can't login to because our AD is down.

top 50 comments

sorted by: hot top controversial new old

[–] [email protected] 201 points 1 month ago (2 children)

Pity the administrators who dutifully kept a list of those keys on a secure server share, only to find that the server is also now showing a screen of baleful blue.

Lol, can you imagine? It empathetically hurts me even thinking of this situation. Enter that brave hero who kept the fileshare decryption key in a local keepass :D

[–] [email protected] 128 points 1 month ago* (last edited 1 month ago) (1 children)

That's why the 3-2-1 rule exists:

3 copies of everything on
2 different forms of media with
1 copy off site

For something like keys, that means:

secure server share
server share backup at a different site
physical copy (either USB, printed in a safe, etc)

Any IT pro should be aware of this "rule." Oh, and periodically test restoring from a backup to make sure the backup actually works.

[–] [email protected] 38 points 1 month ago (1 children)

We have a cron job that once a quarter files a ticket with whoever is on-call that week to test all our documented emergency access procedures to ensure they’re all working, accessible, up-to-date etc.

[–] [email protected] 8 points 1 month ago

Are you hiring!?

[–] [email protected] 64 points 1 month ago (1 children)

Seems like an argument for a heterogeneous environment, perhaps a solid and secure Linux server to host important keys like that.

[–] [email protected] 54 points 1 month ago (6 children)

Linux can shit the bed too. You need to maintain a physical copy.

[–] [email protected] 52 points 1 month ago

Their point is not that linux can't fail, it's that a mix of windows and linux is better than just one. That's what "heterogeneous environment" means.

You should think of your network environment like an ecosystem; monocultures are vulnerable to systemic failure. Diverse ecosystems are more resilient.

[–] gnutrino 28 points 1 month ago (1 children)

Sure but the chances of your Windows and Linux machines shitting the bed at the same time is less than if everything is running Windows. It's exactly the same reason you keep a physical copy (which after all can break/burn down etc.) - more baskets to spread your eggs across.

[–] [email protected] 9 points 1 month ago (1 children)

Very few businesses are going to spend the money running redundant infrastructure on two different operating systems. Most of them won't even spend the money on a proper DR plan.

[–] [email protected] 26 points 1 month ago (1 children)

Then they get to suffer the consequences when shit like this happens

load more comments (1 replies)

[–] [email protected] 12 points 1 month ago

Hey Ralph can you get that post-it from the bottom of your keyboard?

load more comments (3 replies)

[–] [email protected] 117 points 1 month ago (3 children)

Lmao this is incredible

Another Redditor posted: "They sent us a patch but it required we boot into safe mode.

"We can't boot into safe mode because our BitLocker keys are stored inside of a service that we can't login to because our AD is down.

"Most of our comms are down, most execs' laptops are in infinite bsod boot loops, engineers can't get access to credentials to servers."

N.B.: Reddit link is from the source

I hope a lot of c-suites get fired for this. But I’m pretty sure they won’t be.

[–] MagicShel 86 points 1 month ago

C-suites fired? That's the funniest thing I've heard yet today. They aren't getting fired - they are their own ass-coverage. How can they be to blame when all these other companies were hit as well?

I guess this is a good week for me to still be laid off.

[–] [email protected] 78 points 1 month ago (1 children)

Our administrator is understandably a little bitter about the whole experience as it has unfolded, saying, "We were forced to switch from the perfectly good ESET solution which we have used for years by our central IT team last year.

Sounds like a lot of architects and admins are going to get thrown under the bus for this one.

"Yes, we ordered you to cut costs in impossible ways, but we never told you specifically to centralize everything with a third party, that was just the only financially acceptable solution that we would approve. This is still your fault, so we're firing the entire IT department and replacing them with an AI managed by a company in Sri Lanka."

load more comments (1 replies)

[–] [email protected] 29 points 1 month ago

Fired? I hope they get class-actioned out of existence as a warning to anyone who skimps on QA

[–] [email protected] 110 points 1 month ago (4 children)

Lemmy appears to be weathering the storm quite well.....

..probably runs on linux

[–] [email protected] 95 points 1 month ago* (last edited 1 month ago)

The overwhelming majority of webservers run Linux ~~(it's not even close, like high 90 percent range)~~ Edit: Upon double-checking it's more like mid-80s, but the point stands

[–] [email protected] 68 points 1 month ago (1 children)

It runs on hundreds of servers. If any of them ran windows they might be out but unless you got an account on them you'd be fine with the rest. That's the whole point of federation.

load more comments (1 replies)

[–] [email protected] 12 points 1 month ago

I doubt many Lemmy servers are running enterprise level antivirus.

load more comments (1 replies)

[–] [email protected] 82 points 1 month ago (3 children)

This is why every machine I manage has a second boot option to download a small recovery image off the Internet and phone home with a shell. And a copy of it on a cheap USB stick.

Worst case I can boot the Windows install in a VM with the real disk, do the maintenance remotely. I can reinstall the whole thing remotely. Just need the user to mash F12 during boot and select the recovery environment, possibly input WiFi credentials if not wired.

I feel like this should be standard if you have a lot of remote machines in the field.

[–] [email protected] 20 points 1 month ago (2 children)

Just need the user to mash F12 during boot and select the recovery environment, possibly input WiFi credentials if not wired

In theory that sounds great, now just do it 1000+ times while your phone is ringing off the hook and you're working with some of the most tech illiterate people in your org.

[–] [email protected] 16 points 1 month ago

I'm pressing F and 1 and 2, but nothings happening!

load more comments (1 replies)

[–] [email protected] 9 points 1 month ago* (last edited 1 month ago) (1 children)

Sounds like a nightmare for security, and a dream for attackers.

More companies need to do this, solid job security.

load more comments (1 replies)

[–] [email protected] 79 points 1 month ago* (last edited 1 month ago) (1 children)

If you have EC2 instances running Windows on AWS, here is a trick that works in many (not all) cases. It has recovered a few instances for us:

Shut down the affected instance.
Detach the boot volume.
Move the boot volume (attach) to a working instance in the same region (us-east-1a or whatever).
Remove the file(s) recommended by Crowdstrike:
Navigate to the C:\Windows\System32\drivers\CrowdStrike directory
Locate the file(s) matching “C-00000291*.sys”, and delete them (unless they have already been fixed by Crowdstrike).
Detach and move the volume back over to original instance (attach)
Boot original instance

Alternatively, you can restore from a snapshot prior to when the bad update went out from Crowdstrike. But that is not always ideal.

[–] [email protected] 23 points 1 month ago (1 children)

A word of caution, I've done this over a dozen times today and I did have one server where the bootloader was wiped after I attached it to another EC2. Always make a snapshot before doing the work just in case.

load more comments (1 replies)

[–] [email protected] 70 points 1 month ago (6 children)

I didnt know so many servers still run windows.

[–] [email protected] 39 points 1 month ago (2 children)

I'm the corporate world, very much Windows gets used. I know Lemmy likes a circle jerk around Linux. But in the corporate world you find various OS's for both desktop and servers. I had to support several different OS's and developed only for two. They all suck in different ways there are no clear winners.

[–] [email protected] 9 points 1 month ago* (last edited 1 month ago) (1 children)

It's not just a circle jerk in this case. Windows is dominant for desktop usage but Linux has like 90% of the server market and is used for basically all new server projects.

Paying for Windows licensing when it doesn't benefit you, it's silly, and that's been realized for years.

load more comments (1 replies)

[–] [email protected] 32 points 1 month ago (2 children)

Issue is not just on servers, but endpoints also. Servers are something that you can relatively easily fix, because they are either virtualized or physically in same location.

But endpoints you might have thousand physical locations, and IT need to visit all of them (POS, info/commercial displays, IoT sensors etc.).

load more comments (2 replies)

[–] [email protected] 14 points 1 month ago (2 children)

My former employer had a bunch of windows servers providing remote desktops for us to access some proprietary (and often legacy) mission critical software.

Part of the security policy was that any machines in the possession of end users were assumed to be untrustworthy, so they kept the applications locked down on the servers.

load more comments (2 replies)

[–] [email protected] 13 points 1 month ago

On prem AD. At least for my MSP's clients. Have been pushing hard last few years to migrate to azure.

[–] [email protected] 8 points 1 month ago (1 children)

I can’t imagine how much work it would be to migrate all your services onto Linux. The problem was people adopting windows in the first place.

[–] [email protected] 10 points 1 month ago (12 children)

I love the Linux bros coming out of the woodwork on this one when this could have very well have been Linux on the receiving end of this shit show. Given that it's a kernal level software issue, and not necessarily an OS one.

It's largely infeasible to use Linux for many, most, of these endpoints. But facts are hard.

[–] [email protected] 14 points 1 month ago* (last edited 1 month ago) (3 children)

Hey man, let us have this one. Any immutable/atomic distribution could have either prevented this or easily rolled back the update. Not to mention a Linux offering by something like Red Hat, for example, wouldnt recommend installing closed source third party kernel modules for exactly this reason. Not sure about the feasibility of these endpoints, but the way things are generally done on, and the philosophy of, Linux could very well have avoided this catastrophe.

load more comments (3 replies)

[–] [email protected] 9 points 1 month ago

The is no single Linux. It's not a monoculture like that. There are many distros with different build options, different configurations and different components.

Also culture is different. Very few Linux admins would be happy putting in a closed blob kernel driver for anything. In Windows world that's the norm, but not Linux.

What's just happened to Windows world would be harder in Linux world. At worse, one distros rolls out a killer update. Some distros would just reboot to the previous kernel.

load more comments (10 replies)

load more comments (1 replies)

[–] [email protected] 64 points 1 month ago* (last edited 1 month ago) (1 children)

At least no mission critical services were hit, because nobody would run mission critical services in Windows, right?
..
RIGHT??

[–] [email protected] 48 points 1 month ago

[–] [email protected] 31 points 1 month ago (6 children)

Sounds like the best time to unionize

[–] [email protected] 16 points 1 month ago (6 children)

I'm in. This world desperately needs an information workers union. Someone to cover those poor fuckers in the help desk and desktop support as well as the engineers and architects that keep all of this shit running.

Those of us that aren't underpaid are treated poorly. Today is what it looks like if everybody strikes at once.

load more comments (6 replies)

[–] [email protected] 8 points 1 month ago (1 children)

Any time is a good time to unionize

load more comments (1 replies)

load more comments (4 replies)

[–] [email protected] 28 points 1 month ago

lmao

[–] [email protected] 18 points 1 month ago (3 children)

It might be CrowdStrike's fault, but maybe this will motivate companies to adopt better workflows and adopt actual preproduction deployment to test these sort of updates before they go live in the rest of the systems.

[–] [email protected] 19 points 1 month ago* (last edited 1 month ago) (6 children)

I know people at big tech companies that work on client engineering, where this downtime has huge implications. Naturally, they've called a sev1, but instead of dedicating resources to fixing these issues the teams are basically bullied into working insane hours to manually patch while clients scream at them. One dude worked 36 hours straight because his manager outright told him "you can sleep when this is fixed", as if he's responsible for CloudStrike...

Companies won't learn. It's always a calculated risk, and much of the fallout of that risk lies with the workers.

load more comments (6 replies)

[–] [email protected] 9 points 1 month ago

Oh sweet summer child.

[–] [email protected] 8 points 1 month ago

Might be hard to do. Crowdstrike release several updates per day to the channel files to match changes in adversarial behaviour. In this case, BCP and backup are what need to be done.

[–] [email protected] 13 points 1 month ago

80% of our machines were hit. We were working through 9pm on Friday night running around putting in bitlocker keys and running the fix. Our organization made it worse by hiding the bitlocker keys from local administrators.

Also gotta say... way the boot sequence works, combined with the nonsense with raid/nvme drivers on some machines really made it painful.

[–] [email protected] 8 points 1 month ago

I got super lucky. got paid for my car just before the dealership systems went down, got my return flight 2 days before this shit started.

load more comments