Velocirooster adminensis 
In the middle like a bird without a beak 🐓 Admin of this here instance 🦖 I have no idea what I'm doing 🦣 he/him 🙆♂️
Alt Text: Avi is the head of a rooster with a velociraptor's face instead of a beak, seen in profile and looking majestic af. The header is a hilariously inaccurate 19th-century woodcut depicting an iguanodon and a megalosaurus as big lumbering quadrapedal lizards biting each other. Neither of them seems bothered by this, in fact they both are sporting big goofy toothy grins.
Attention Beige Party-goers!
Looks like we had another unplanned outage. This time the culprit was the hot standby db server. At some point it stopped responding to the production server and when that happens, the prod server starts caching all changes locally until the standby comes back up. If it doesn't come back up, that cache grows until it fills the entire hard drive is filled up, which is what happened this morning.
We are back up and running now. I need to check out the standby server to see what went wrong, but that's something I can do in the background and it won't require any more downtime today.
Beige-bless ![]()
Attention Beige Party-goers!
Sorry for that unplanned outage. We were getting overwhelmed with Sidekiq jobs again, the result being the Beige Bubble effect where posts on the timeline are delayed by several hours.
In trying to add some more DB connections for SideKiq to use, I found out that pgbouncer likes to switch its IP around when you restart its docker container. Postgres was not configured to accept that new IP address, so the whole site went down. Oops.
Anyways, we're back up now. I've added more DB connections which allowed me to add some more SideKiq workers. Queues appear to be trending in the right direction though it might take a bit for them to clear out entirely. If they start going up again, I will add some more Sidekiq workers as needed.
Thanks as always for bearing with me as I continue to figure this stuff out, and Beige-Bless ![]()
Attention Beige Party-goers! ![]()
GREAT NEWS! After much struggle, I have finally settled on a Postgres migration strategy that I'm comfortable with. It's going to mean 2-3 hours of downtime but it has the advantage of leaving the old database intact so I can roll back in case something goes horribly wrong, which is nice. I did a dry run last night and everything went smoothly, so I feel good about moving forward with the actual database migration now.
So, what this means is that we now have a date for when we will upgrade to Mastodon 4.5. I am going to do this in 3 stages. The first, and most important is the migration to Postgres 17. That will happen this Sunday night, May 31st/June 1st. The second stage will happen a week later, on June 7th/8th. The third and final stage will occur on June 14/15. This last stage will be the actual upgrade to Mastodon 4.5.
The dry run took about 2 and 1/2 hours, so I am going to plan for a 4 hour maintenance window to account for any hiccups that might occur. The maintenance window will start at 3AM UTC on Monday, June 1st and continue until 7AM. In the Western Hemisphere, that's 11PM EDT on Sunday the 31st until 3AM on the 1st. I don't anticipate the site being down for the full 4 hours, but we will definitely experience at least 2 hours of downtime while I am performing the database migration. As always, be prepared for the site to go down at any time during the maintenance window.
If everything goes well with the database migration I will post a new announcement sometime next week confirming the next maintenance window scheduled for the following Sunday, June 7th. Thanks for bearing with me as I figure this out. The good news is that after this migration everything will be containerized, so upgrades should be a lot simpler in the future. I really do appreciate you all hanging on with this (at times) rickety trolley.
Thanks, and Beige-bless ![]()
Attention Beige Party-goers ![]()
I have completed the upgrade and we are now on version 4.6.2. There should be no additional downtime tonight.
Beige-bless ![]()
Attention Beige Party-goers! ![]()
We are currently having issues with the full text search. This is because migrating the database requires reindexing elasticsearch. We've got a lot of toots in this baby so a full reindex can take a day or more. I just checked and the process is still running and the progress bar still has a ways to go. I'll post another update once the reindex process is complete.
Beige-bless ![]()
Attention Beige Party-goers! ![]()
I need to make a couple more tweaks to the web server, which will require taking it offline briefly. Any downtime should be less than a minute.
Beige-bless ![]()
Sorry folks, was updating docker in preparation for the 4.6 tomorrow and it ended up taking the database offline for several minutes. We should be good now.
Attention Beige Party-goers! ![]()
Welp, looks like the days of me being able to update docker and have the site go down for just a few seconds are over. Since I've already started, I'm going to go ahead and update the other servers. You may see the site go down for a couple of minutes. I will post another announcement when I'm done.
Beige-bless ![]()
Just a reminder: No matter how much you wish violence upon someone, no matter how evil and vile and reprehensible they are, no matter how sure you are that they deserve to be removed from this earth, either by intentional action or happenstance, Beige Party is not the platform for expressing this opinion.
I can make a lot of compelling moral arguments for this position, but the most immediate and salient point is that advocating for violence in general, and political violence in particular, puts you, me, and the entire instance in legal jeopardy. Don't do it.
Going offline for Mastodon upgrade. See you in a couple hours...
Ok, going offline for few hours for server maintenance. See you on the other side.
Attention Beige Party-goers! ![]()
I'm still getting reports of media loading slowly. It seems to happen more around the evenings (in the US) and on weekends, which is when traffic is tends to be higher. The weird thing is that I have not been able to catch the problem in the wild. When I check a sample of images both local and remote, everything seems to be loading fine. What's been reported is images loading fine on the TL but when clicking to enlarge, the images load at dial-up speeds—20 seconds+. Is anyone else having this issue?
Thanks, and Beige-bless ![]()
The ceremony of innocence is drowned;
The best lack all conviction, while the worst
Are full of passionate intensity.
Surely some revelation is at hand
@woof@aria.dog @christymarx@beige.party This issue has been resolved. The images in question have been taken down. Stop spamming this thread or I will be defederating you for harassment.
@OohOkayKay@beige.party @BaronessWinter@beige.party We have been having issues at least since the 11th. I think what happened was I was running too many threads for the web server so it was eating up a lot of memory, then I tried running the reindex for elasticsearch and the combined processes were occasionally using 100% of available memory and grinding everything to a halt. That should hopefully have been fixed last Sunday when I did the 4.6 upgrade. The only time it should have gone down this week was tonight for the upgrade to 4.6.2.
