Skip to content
TORNLIFE More

Down time to replace a failed component!

Started by ellarae [1001245] on in General Discussion.

24 replies · 545 views · thread synced · 5 days ago · View on torn.com
About this thread

Posts archived: 25 / 25 posts (100%) · the total is Torn's reply count + the opening post at the last fetch

Counted by TornLife from the archived posts.

Archived posts
25
Discussion span
→
Authority score
57 / 100
Historical score
24 / 100
Story score
57 / 100
Engagement score
68 / 100

People posting, likes and official posts are not counted for this thread yet: on threads longer than one page they come from a periodic pass over the archive, which has not covered it.

Most-liked replies

Mentioned in this thread

JustJanet [911066] ×1

Dbeer [2043567]

Or because too much money is going to be added due to the real madrid vs bayern **** up and now they gotta catch and fed a few multies to remove the cash from the system.
Swatsicle [1593313]

Wow! You too have read the message at the top of the screen! Amazing!

Thanks for this awesome thread. I never know whether a message from the admin team is real or completely imagined unless someone makes a thread about it.

From the bottom of my heart, I thank you. The General Discussion forum thrives thanks to the incredibly sharp posters such as yourself.
Orionis [1924258]

I know you're just trying to be funny to earn some Karma, but probably not. It specifically says a database server, and they did have a mongo DB error recently that resulted in chat ceasing to function for a time, so if I was a betting man, I'd say it is related to that.
Zanoab [1877054]

I remember last time this happened, I claimed they were stupid for not having sufficient backup/redundant hardware. I got chewed then right before another failure took the site completely offline... Safe to say they haven't learned?
sharkeyfive [1151690]

Is no one happy with the free stuff they get these days ?

Ched may i make a suggestion, stop giving free stuff away. If people are going to moan let them moan about getting nothing rather than them moan about getting something for free.
JustJanet [911066]

I hate liking your comment since you just hospitalized me! hmmph! Rudeness! but... you make a really great point! I agree.


Multipass [1955315]



Replied to wrong post. Cant remember what this one said now. Something about redundancy not necessarily equaling high availability.
Zanoab [1877054]

The setups are designed specifically to maximize uptime with minimal risk of data loss. When done right, you can mitigate risks to the point they only happen in edge cases. With a proper RAID setup (which is safe to assume is in place), the system functions without data loss depending on how many drives are in play.

Let's say their RAID setup allows them to only be missing one drive at any point in time. A drive failing isn't unheard of and they usually don't happen often. They have a drive failing and needs to be replaced. If a drive was already available, they only need to ask a tech to replace the (correct) drive. Once completed, they don't have to worry until another drive fails. However, every minute that dead drive remains in the array is another chance another drive could fail. Not only would you have to consider the time it takes to have a tech replace the dead drive, you also need time to rebuild the replacement drive. Best case scenario, that is the minimum amount of time you would be at risk of data loss. The time it takes to ship a new drive (that could be dead on arrival which is common) is unnecessary risk that you don't want when data is at stake.

The real fun part are extra factors. If two or more drives in the array are from the same production batch, they are likely to fail near the same time. This could easily be mitigated by collecting drives from different manufacturers so you reduce the chances of multiple drives failing at the same time. If they are buying a batch of drives once their supply of fresh drives is exhausted, it is safe to assume the entire batch is from the same manufacturer and same production batch. Because last year multiple drives had to be replaced, it is a safe bet that additional drives would fail once one does.

If they only had to replace a single failed drive, we more than likely won't notice a reduction in performance. If additional drives fail while replacing a failed drive, the server would go down due to data loss. Backups are for worst case scenarios while redundancies keep people going like nothing happened. If you have to resort to backups due to the same problem at two different periods, then there is something seriously wrong that needs to be fixed.

EDIT:

Source from last year's data loss: https://www.torn.com/forums.php?p=threads&f=65&t=15985133&b=0&a=0

"The secondary drive on our mongoDB server unexpectedly failed yesterday, while a replacement was on order, the primary drive then failed earlier this morning."

Source on migration of data and replacing multiple drives at the same time (potentially from the same batch): https://www.torn.com/forums.php?p=threads&f=65&t=15985457&b=0&a=0

"We will be migrating our DB3 server to new hardware and installing new drives in the mongoDB server that went down on Monday morning."

Mentions: Downtime : Friday 7th 09:30 TCT [Complete] · Unexpected downtime earlier today

Multipass [1955315]

Yes I'm familiar with raid topologies ; ) I'm not overly familiar with mongo but, depending on system design, I can easily imagine that taking down a shard would require downtime. Working with databases while people continue to write to them is a non trivial exercise and when possible it is nice to avoid it.

Ched said they replaced the drive in the MongoDB server AND replaced the hardware on DB3. DB3 could (potentially) be an entirely different database or just a shard of the primary mongo. Either way thats a lot of dicking around with live, clustered services to want to engage in when up-time is not absolutely critical.

Also when Ched said "backups" he may have just been trying to use terminology his users might understand. If he said they had to "replicate the shard" that might lead to more questions but would be a lot less dramatic than it (or backup) sounds.

In any case my point is simply that this its easy to assume people are being stupid when in fact we just don't have even half the information needed to make a valid judgement.