It is much better now. Was totally Ched OK for a few hours though.
About that server upgrade...
Started by CRLF [2095076] on in General Discussion.
Posts archived: 174 / 174 posts (100%) · the total is Torn's reply count + the opening post at the last fetch
How do you think making something take 5 more e was going to fix anything? common sense should of made you think that was bs when it was said.
To be fair how was something to do with items going to fix people attacking?
To be fair how was something to do with items going to fix people attacking?
I honestly didn't expect for so many people to absolutely lose their shit over 3 hours of area specific lag :D It really shows how far we've come along over the past few years for even that to be totally subscription-cancellingly unacceptable.
As most people know, I wanted to entirely cancel this event this year, especially so because it falls on a weekend where we have less resources available. It's a terrible amount of anxiety, my heart rate and blood pressure seem to be directly connected to the Torn status monitor.
However I was talked into it again - for one reason alone - it's a really good stress test for Torn. The fixes that result from the problems that this event uncovers in Torn make the rest of the year faster and more reliable - it shows us what our next bottleneck is and points out the things we need to improve on. Stuff that there's literally no possible way to find otherwise. It gives us a roadmap of improvements for us to follow over the next 6 months.
For this reason, I think it's a great idea to keep it. It will always be high-risk, and it'll likely be uncomfortable for at least a short amount of time. But it's worth it.
The event started at 12:00, and everything seemed pretty smooth. Load times did increase, as is expected, but there were no serious problems. Then at just before 19:00 some critical limit was reached. Our DB3 server once again was throwing a tantrum, specifically affecting the inventory pages and attacking. This caused load times to spike in these areas.
As most of you know, we've been searching for a server admin for months now and been unsuccessful, an ideal candidate and player of Torn joined in September but had to leave due to health issues immediately after entirely wiping our admin server & monitoring infrastructure. Since then we've had adverts on Indeed and Linkedin, and even been searching Upwork for worldwide candidates - due to the niche and expertise we're looking for, it's been a real challenge. We've now enlisted a consultancy firm to find us someone and we've had some good progress already.
Anyway, I digress.. since we don't have a dedicated server admin currently, we suck at finding any server related issues in really quick time. I eventually interrupted the evenings of a couple of developers and Alex, and we were able to track down the issue. It seems players had been gradually opening more and more tabs over the evening, tabs which poll a specific query every few seconds, this resulted in a huge amount of queries being processed, more than the database can handle. Amusingly, this database is already on pretty much the best hardware that currently exists. We can perhaps buy a 5th database server and split it in two, but really this isn't a hardware or networking issue, it's a development issue. Today we were able to get rid of the additional load by simply reducing the poll frequency, this brought it down from the cap that was causing all the issues. We now have a plan to rebuild this particular system using websockets - to skip the DB entirely.
Unfortunately as this stage we can't really just buy our way out of trouble with new servers, we have to sensibly rebuild and refactor old code when its issues become apparent - hence making them apparent via Valentines Day's Love Juice event (Aka, Torn's stress test).
So I apologise for the issues, I know they're frustrating. But it is kind of a necessity for us, this is how we grow. I'll be sure to perhaps add clearer service warnings next time if we think that might help. I'm very much open to feedback on how we could handle things differently, however I'm pretty confident that we're doing alright.
Thanks.
As most people know, I wanted to entirely cancel this event this year, especially so because it falls on a weekend where we have less resources available. It's a terrible amount of anxiety, my heart rate and blood pressure seem to be directly connected to the Torn status monitor.
However I was talked into it again - for one reason alone - it's a really good stress test for Torn. The fixes that result from the problems that this event uncovers in Torn make the rest of the year faster and more reliable - it shows us what our next bottleneck is and points out the things we need to improve on. Stuff that there's literally no possible way to find otherwise. It gives us a roadmap of improvements for us to follow over the next 6 months.
For this reason, I think it's a great idea to keep it. It will always be high-risk, and it'll likely be uncomfortable for at least a short amount of time. But it's worth it.
The event started at 12:00, and everything seemed pretty smooth. Load times did increase, as is expected, but there were no serious problems. Then at just before 19:00 some critical limit was reached. Our DB3 server once again was throwing a tantrum, specifically affecting the inventory pages and attacking. This caused load times to spike in these areas.
As most of you know, we've been searching for a server admin for months now and been unsuccessful, an ideal candidate and player of Torn joined in September but had to leave due to health issues immediately after entirely wiping our admin server & monitoring infrastructure. Since then we've had adverts on Indeed and Linkedin, and even been searching Upwork for worldwide candidates - due to the niche and expertise we're looking for, it's been a real challenge. We've now enlisted a consultancy firm to find us someone and we've had some good progress already.
Anyway, I digress.. since we don't have a dedicated server admin currently, we suck at finding any server related issues in really quick time. I eventually interrupted the evenings of a couple of developers and Alex, and we were able to track down the issue. It seems players had been gradually opening more and more tabs over the evening, tabs which poll a specific query every few seconds, this resulted in a huge amount of queries being processed, more than the database can handle. Amusingly, this database is already on pretty much the best hardware that currently exists. We can perhaps buy a 5th database server and split it in two, but really this isn't a hardware or networking issue, it's a development issue. Today we were able to get rid of the additional load by simply reducing the poll frequency, this brought it down from the cap that was causing all the issues. We now have a plan to rebuild this particular system using websockets - to skip the DB entirely.
Unfortunately as this stage we can't really just buy our way out of trouble with new servers, we have to sensibly rebuild and refactor old code when its issues become apparent - hence making them apparent via Valentines Day's Love Juice event (Aka, Torn's stress test).
So I apologise for the issues, I know they're frustrating. But it is kind of a necessity for us, this is how we grow. I'll be sure to perhaps add clearer service warnings next time if we think that might help. I'm very much open to feedback on how we could handle things differently, however I'm pretty confident that we're doing alright.
Thanks.
Aight bruv
But half the fun is crying on the forums!
Seriously though, good luck finding a sys admin.
Seriously though, good luck finding a sys admin.
It's ok ched, will you be my valentine?
This was an excellent response and I'm happy you're keeping the event. Personally I think these events add something to the game - whether they lag or not - and there should be more, not fewer.
Thank you for responding and not just downvoting the thread.
Thank you for responding and not just downvoting the thread.
'Thank you for responding and not just downvoting the thread.'
SHOTS FIRED AT BOGIE
[image: media0.giphy.com]
SHOTS FIRED AT BOGIE
[image: media0.giphy.com]
Thank you for your response and thank you for being open about the purpose of the event, however, it would probably to have a couple devs on stand by all throughout a stress test in case something does go wrong so you can mitigate the effects immediately and prevent such a large amount of displeasure.
Also, it might have been smart to cancel or delay the event until you had at least found a server admin to monitor things properly.
Thank you for taking care of everything, take care of yourself and have a great weekend.
Also, it might have been smart to cancel or delay the event until you had at least found a server admin to monitor things properly.
Thank you for taking care of everything, take care of yourself and have a great weekend.
I have messaged bogie and ched about the horrible lag but they are useless unless you make a forum post for everyone to see I guess.....
I just dont understand how they say everything is fine when a sloth is faster than i can attack. Anyone else noticing that!?
Maybe ee should stress test the servers for a month straight so we can fix all the problems right away?
Do you want them to work on tuning the server and restarting broken chains, or answering two hundred messages complaining about the lag?
The thread is one-stop shopping. Everyone can post their complaints and bogie or chedburn can answer them collectively with a single reply.
The thread is one-stop shopping. Everyone can post their complaints and bogie or chedburn can answer them collectively with a single reply.
I propose next year we have a half-way to valentines day, stress test, that way the staff and Ched can have 6 months to get their shit together BEFORE the event, not after.
quick, someone call Ched. We have a dozen sys admins right here in this thread from the sounds of it.
(that'd be sarcasm)
(that'd be sarcasm)
So you understand what a stress test is, right? It's when you have a massive number of concurrent tasks hitting the server. If we said, "everybody show up for a server stress test stacked and burn a lot of E hitting each other" we'd get 50 people.
If we say "whoo hoo! special event! free love juice! attacks only 15e!" then everyone, their brother, their grandma and both of their dogs show up stacked and ready to go.
The event must occur to cause the massive number of concurrent tasks. It's the only way to have a proper test. A simulation will only take you so far.
If we say "whoo hoo! special event! free love juice! attacks only 15e!" then everyone, their brother, their grandma and both of their dogs show up stacked and ready to go.
The event must occur to cause the massive number of concurrent tasks. It's the only way to have a proper test. A simulation will only take you so far.
Noo my karma