Showing posts with label sorry. Show all posts
Showing posts with label sorry. Show all posts

Thursday, October 22, 2020

Linden Lab Asks For Patience For The Next Few Weeks With Cloud Migration

 

It's no secret things have been more glitchy in Second Life than usual. Last Friday, Linden Lab through Oz Linden came out and apologized for the various glitches, crashes, and bugs recently, saying they were from it's moving things to cloud servers. Yesterday, Wednesday October 21, in the "Tools and Technology" blog, it was April Linden's turn. She gave an update on how much has been done so far.

We’re in a really exciting time in the history of Second Life. We’re in the home stretch on moving the grid to the cloud. We hit a fun milestone a few days ago, and now there’s over 1,000 regions running in the cloud!

Everyone in the Lab is working hard on this project, and we’re moving very quickly. I just got out of a leadership meeting where we went over what’s currently in flight, and there’s so many things moving that I lost track of them all. It’s amazing!


However, not everything was good news. She asked residents to be patient.

I’ve come to ask for a favor.  ... I’ve come to ask you for is your patience. 

We’re doing our very best to fix things that come up as we go. This means that we might need to restart regions more often than you’re used to, and things may break just a little more often than we’ve all been accustomed to.

In order to get this project done as fast as possible and minimize the time (and resulting bugs) we have to spend with one foot in our datacenter and the other in the cloud, we don’t want to limit ourselves to restarting regions just once a week. We’re ready to get this project done! We’ve seen how much better Second Life runs in the cloud, and we’re ready to have everyone on the grid experience it.

I’m sorry that things might be a little rough over the next few weeks. It’s our goal to finish the cloud migration by the holidays, so that everyone, Resident and Linden alike, can have a nice quiet holiday with our friends and families.

 We can’t promise we’ll make it by then, but we’re sure giving it all we’ve got. The mood around the Lab is really positive right now, and we’re all working hard together to make it happen. I’m really proud to be a part of the team that’s transforming Second Life as we know it.

Thanks so much for hanging in there with us. We know it’s frustrating at times, but it won’t last for too long, and there’s a better future on the other side of this. We truly appreciate your understanding and patience as we finish up this project.


So it appears in addition to the glitches we've been seeing, we're very likely to have to deal with sim restarts as well. Perhaps the Lab may slow things down a bit next week interfere with Halloween less. If not, residents may have to contend with having to suddenly leave sims in a hurry in addition to the ghosts and goblins.

To read the Linden blog post in full, Click Here.

Bixyl Shuftan

Monday, October 19, 2020

Linden Lab: "We Apologize For Recent Service Disruptions"


Glitches are more or less a part of the Second Life experience. But if it seems they've been getting worse lately, it's not your imagination. And even Linden Lab is admitting things are buggier than usual. In the Tools and Technology blog on Friday October 16, Oz Linden stated in the "Uplift Update" that while there's been overall improvement in the Second Life infrastructure, "it’s the bumps in the road that are most noticeable to our residents. We apologize for recent service disruptions ..." He would point out several "rough spots."

Region Crossings
One of the first troubles we found was that region crossings were significantly worse between a cloud region and a datacenter region. We did a deep dive into the code for objects (boats, cars, planes, etc) and produced an improvement that made them significantly faster and more reliable even within the datacenter. This has been applied to all regions already and was a good step forward.

Group Chat stalls
Many users have reported that they are not able to get messages in some of their groups; we're very much aware of the problem. The start of those problems does coincide with when the chat service was uplifted; unfortunately the problems did not become clear until moving that service back to the datacenter was not an option. We haven’t been able to get that fixed as quickly as we would like, but the good news is that we have some changes nearly ready that we think may improve the service and will certainly provide us with better information to diagnose it if it isn't fixed. Those changes are live on the Beta grid now and should move to the main grid very soon.

Bake Failures
Wednesday and especially Thursday of this past week were bad days for avatar appearance, and we're very much aware of how important that is. The avatar bake service has actually been uplifted for some time - it wasn't moving it that caused the problem, but another change to a related service. The good news is that thanks to a great cross-team effort during those two days we were able to determine why an apparently unrelated simulator update triggered the problem and got a fix deployed Thursday night.

Increased Teleport Failures
We have seen a slight increase in the frequency of teleport failures. I know that if it's happened to you it probably doesn't feel like a "slight" problem, especially since it appears to be true that if it's happened to someone once, it tends to keep happening for a while. Measured over the entire grid, it's just under two percentage points, but even that is unacceptable. We're less sure of the specific causes for this (including whether or not it's Uplift related), but are improving our ability to collect data on it and are very much focused on finding and fixing the problem whatever it is.

Marketplace & Stipend Glitches

We've had some challenges related to uplift for both the Marketplace and the service that pays Premium Stipends. Marketplace had to be returned to the datacenter yesterday, but we'll correct the problems that required the rollback and get it done soon. The Stipends issues were both good and bad for users; there were some delays, but on the other hand we sent some users extra stipends (our fault, you win - we aren't taking them back); those problems are, we believe, solved now.


So why all of the glitches and bugs? Oz called the Uplift to move Second Life's data to cloud servers "a massive, complicated project that I've previously compared to converting a steam-driven railroad to a maglev monorail -- without ever stopping the train." He would later try to reasure the residents, "While this week in particular has seen some bumps in the road, it's actually going well overall. Lots of the infrastructure you don't interact with directly, and some you do, has been uplifted and has worked smoothly." He would go on to say almost all of the Beta Grid's sims were in the Cloud, and "we've uplifted around a hundred regions on the main grid. Performance of those regions has been very very good, and stability has been excellent." He would go on to say more sims would be uplifted over the weekend and that "Uplift of the Release Candidate regions, which will bring the count into the thousands, will begin soon."

Oz would say the Uplift to the Cloud should be completed, or at least be close to it, by the end of the year.

To read the blog entry in full, Click Here.

Bixyl Shuftan

Friday, November 1, 2019

April Linden Explains Last Sunday's Outage


If you tried to get on Second Life on Sunday October 27th, then you may have been blocked by issues that persisted all afternoon. On Wednesday, April Linden, the Operations Manager of the Grid, posted in the Tools and Technology blog just what happened in "The Sunday That Was Not a Fun Day."

The trouble started back on Thursday. We had some pretty bad problems talking to key services (packet loss) on one of our Internet links. It didn’t impact everything, but the stuff it hit was pretty important. It started for a few hours on Thursday, but just magically went away on its own. It’s never good when a problem just magically fixes itself because it’s pretty likely it’s gonna happen again. 

The same thing happened again on Friday, but once again it went away on its own before we were able to debug what actually happened. 

We were nervous going into the weekend, and sure enough, it bit us again on Sunday. We started having more Internet communication problems, but this time, it didn’t just go away. 

Now that we had it in a bad state we started to troubleshoot and figured out really fast that it wasn’t our equipment. Our stuff was (and still is) working just fine, but we were getting intermittent errors and delays on traffic that was routed through one of our providers. We quickly opened a ticket with the network provider and started engaging with them. That’s never a fun thing to do because these are times when we’re waiting on hold on the phone with a vendor while Second Life isn’t running as well as it usually does. 

After several hours trying to troubleshoot with the vendor, we decided to swing a bigger hammer and adjust our Internet routing. It took a few attempts, but we finally got it, and we were able to route around the problematic network. We’re still trying to troubleshoot with the vendor, but Second Life is back to normal again.

April went on to say that the power outages affecting much of California had nothing to do with the outage, although one of the people working on the issue was, "One of our engineers was working in the dark in a house without power, with his laptop being powered by a long string of extension cords that ran to a generator outside."

We’re really sorry that this past Sunday wasn’t very fun. The weekend before Halloween is a really fun time to be Inworld, and it was a frustrating day all the way around. (I personally love the way our Residents really get into Halloween in a way that’s only possible in Second Life!) Knowing that it wasn't as awesome as it could have been makes me sad, and we’re working to make it better in the future. 

April requested that if anyone suspects an issue that may have been caused by the outage, please let the Support team know.

Source: BBC, Tools and Technology

Bixyl Shuftan

Thursday, August 25, 2016

April Linden Explains Tuesday's Grid Problems, "We're Sorry About This Outage"


As some of you already know, there were some problems with the Second Life Grid on Tuesday. For a while around Noon and the early afternoon, residents were unable to log in and some were being logged out. At the time, Linden Lab posted an "Unscheduled Maintenance" notice at 11:03 AM stating this, and asking residents to temporarily stop making purchases inworld and buying Lindens.

We are aware that some users are currently being logged out, and all users are currently unable to log in to Second Life. Additionally, please refrain from transacting on the LindeX on in-world as purchases are likely to fail. Some regions are also unavailable at this time, and teleports to other regions may fail currently. We have identified the issue causing this situation and are currently working to resolve it. Please monitor this blog for additional updates.

It was 2:49 PM when they finally gave the all-clear. A few residents I talked to later mentioned the downtime. One would later tell me adjustments she had made at the time while she was going through a number of sets of clothes had been undone, and wondered if what was going on with the Grid had something to do with it.

Several years ago, downtimes happened regularly. With the sometimes self-depreciating humor of the days of Philip Linden, they'd post a picture parodying the monolith and apes scene in 2001: A Space Odyssey, "The Grid Is Down While We Bang On Things." Today, downtimes are much less frequent. And this time, April Linden decided to make a statement the following day explaining what happened.

Shortly after 10:30am, the master node of one of the central databases crashed. This is the same type of crash we’ve experienced before, and we handled it in the same way. We shut down a lot of services (including logins) so we could bring services back up in an orderly manner, and then promptly selected a new master and promoted it up the chain. This took roughly an hour, as it usually does.

A few minutes before 11:30am we started the process of restoring all services to the Grid. When we enabled logins, we did it in our usual method - turning on about half of the servers at once. Normally this works out as a throttle pretty well, but in this case, we were well into a very busy part of the day. Demand to login was very high, and the number of Residents trying to log in at once was more than the new master database node could handle.

Around noon we made the call to close off logins again and allow the system to cool off. While we were waiting for things to settle down we did some digging to try to figure out what was unique about this failure, and what we’ll need to do to prevent it next time.

We tried again at roughly 12:30pm, doing a third of the login hosts at a time, but this too was too much. We had to stop on that attempt and shut down all logins again around 1:00pm.

On our third attempt, which started once the system cooled down again, we took it really slowly, and brought up each login host one at a time. This worked, and everything was back to normal around 2:30pm.

My team is trying to figure out why we had to turn the login servers back on much more slowly than in the past. We’re still not sure. It’s a pretty interesting challenge, and solving hard problems is part of the fun of running Second Life.
Voice services also went down around this time, but for a completely unrelated reason. It was just bad luck and timing.

We did have one bright spot! Our status blog handled the load of thousands of Residents checking it all at once much better. We know it wasn’t perfect, but it showed much improvement over the last central database failure, and we’ll keep getting better 

My team takes the stability of Second Life very seriously, and we’re sorry about this outage. We now have a new challenging problem to solve, and we’re on it.

"We're sorry." Not words that we hear from Linden Lab every day. But this time, it's on the record that they spoke them.

Hat Tip: Inara Pey

Bixyl Shuftan

Friday, June 10, 2016

Newser Beach Party This July


About twice a year, sometimes more, the Newser crew holds a beach party, to both have a little fun and get to know our readers. Usually we have one close to our anniversary on June 6. But this time, there have been a lot of things going on. Second Life is a busy time anyway in June with the Second Life Birthday celebrations, of which the Newser has been a participant as well as covering it, and the Relay events, which include our neighbors the Sunweavers / Sunbeamer Team. And lately there's been some other things going on, including in real life for some of the staff. So we will be postponing the party until a weekend between the SLB and the Relay Weekend, which means sometime in July. No one event caused this decision. We've simply had to deal with a lot on top of what we usually do, and something had to give.

Sorry to those whom were looking for a little fun in the sun with us. If anyone wants to invite the team to their own beach party, by all means those of us who can make it will drop by. In the meantime, keep the sunscreen and swimwear handy for next month.

Bixyl Shuftan