Post by @GossiTheDog@cyberplace.social
Running a Mastodon server is becoming increasingly problematic due to GenAI.
There’s multiple different things happening, including but not limited to:
- very aggressive scraping of user content and block evasion, with the resource costs that come with it
- AI agents aggressively trying to register accounts and evade restrictions
- rising cost of RAM - Mastodon is memory hungry, RAM is really expensive now
It’s odd to experience first hand, it’s like being consumed by.. a void.
14 likes
107 boosts
R.L. Dane
🍵
@rl_dane@polymaths.social
Druid 🏴@druid@toot.wales
Forse (he/him)@forse@kolektiva.social
🆘Bill Cole 🇺🇦@grumpybozo@toad.social
stf@stf@chaos.social
Joe Cooper 🇺🇦 🍉@swelljoe@mas.to
Esther Payne
@onepict@chaos.social
Send Cookies@otfrom@functional.cafe
Steffen Voß@kaffeeringe@social.tchncs.de
Fränçe Licôrne@alter_unicorn@masto.bike
Token Sane Person@tokensane@mastodon.me.uk
Jenica Lake เหล็ถเลถ@MamaLake@beige.party
May likes Toronto@mayintoronto@beige.party
😀🚲@enobacon@urbanists.social
[object Object]@zzt@mas.to
RealGene ☣️@RealGene@hachyderm.io
Lily Cohen@lily@homestead.social
Paul Chambers🚧@paul@oldfriends.live
Richard Rathe@nickrauchen@c.im
Ghostrunner@ghostrunner@hachyderm.io
Cycling Stu@stufromoz@aus.social
Zongora
@Zongora@vivaldi.net
GhostOnTheHalfShell@GhostOnTheHalfShell@masto.ai
TheStrangelet
@thestrangelet@beige.party
sb the chippy-choppy@sb@metroholografix.ca
Claudius@claudius@darmstadt.social
toots@toots@fairy.id
JJDavis
@jjdavis@infosec.exchange
≽(•⩊ •マ≼@josh@masto.byrd.ws
Cowboy Who?@CowboyWho@libranigans.com
Pseudo Nym@pseudonym@mastodon.online
epiPelagic Puck@Puck@sfba.social
diana 🏳️⚧️🦋🌱@dianea@lgbtqia.space
SpaceLifeForm@SpaceLifeForm@infosec.exchange
Roknrol@roknrol@beige.party
stux⚡️@stux@mstdn.social
Pau Amma@pauamma@mstdn.social
Noodlemaz@noodlemaz@mstdn.games
Yvan ー イボん 🗺️
@YvanDaSilva@hachyderm.io
FinchHaven hachy@FinchHaven@hachyderm.io
A Flock of Beagles 🏴🏴☠️🍉@burnitdown@beige.party
Michael Halila 🏳️⚧️@mhalila@mastodontti.fi
kat donegan (she/her) 🇵🇸✊@kathimmel@mstdn.social
cs@cs@mastodon.sdf.org
Nicole Parsons@Npars01@mstdn.social
Franceska Mann@FranceskaMann@freeradical.zone
Nullstring 🏴☠️@0x00string@infosec.exchange
@GossiTheDog what blocking is in place that is being successfully evaided?
edit: i already know anubis is useless
@petko @GossiTheDog Literally everything, because these companies have the money to spend on cat & mouse games & we have none. All moderators can do is pile hacks on top of their networking setups to block protocol abuse & manually sort through registration applications to kick out bots.
@jackemled, thank you but I was looking for concrete examples... That are not anubis since I already know it doesn't work.
"Anubis doesn't work" seems a bit of a quick and concrete judgement. I find it or something like it is a good first step and cuts out the low-effort bots at least unless it really is that bypassable by literally everything.
If this is our future, every account should be an instance. Bloat? Sure, but bot farms wouldn't be able to pull it off at scale for every single bot. For us individually? Ez pz.
@Walruths at one point I thought that deploying proof of work at scale (i.e. *everyone* deploying it) would be able to offset the economy of scraping. Now I no longer believe that. There is no world where a data center full of stolen/upcycled phones solving proof of work will give up sooner than your users.
@Walruths @petko Anubis & other proof of work systems do work & that is the problem. They harm real users far more than they harm attackers, because the large companies executing these DDoS attacks can afford the power bill, but my phone battery can't last when every other webpage needs me to spend a few percent of the charge to be able to access it.
Making every account its own server is not a solution. The technical barrier to entry to real users is far too high, it's far too expensive, & it defeats the moderation ability that exists now. A large company such as Google or OpenAI can also afford to do this for the purpose of doing more attacks without issue.
@petko @jackemled We recently saw:
- a real chrome browser with a TLS fingerprint of a chrome browser
- able to interpret JavaScript and even webasm
- able to do all what a chromium could do basically.
Solving a invisible doublesha challenge and asking for 1 page.
Then coming from a few other IPs with the same cookie for 1 page per IP address.
If it uses too many different IPs and we block it (which is not a good idea for people on 4G) it come again on a new IP with no cookie (so viewed as a new user)
And it used about a million IP in 48hours...
That sucks...
@petko@social.petko.me Anubis is useless now? Have you tried upgrading to the latest version?
Working fine here.
@grumpasaurus @GossiTheDog man this is exactly it isn't it.
I just watched a bit of YT earlier that was about people just... Not being able to read now (esp in the US) and AI making it all worse.
Sigh.
@noodlemaz @grumpasaurus @GossiTheDog
People are having this problem everywhere, not just or especially in the US.
A few weeks ago, they released the data of school performance worldwide, and there's a marked decline everywhere. But everyone keeps arguing at the local level, ignoring the global trend. It's really weird.
"mastodon is memory hungry"
guess you'll have to divorce John Mastodon and get engaged to Susie SNAC /hj
cc @p @ozzelot
no web interfacestopped reading afterwards, you clearly have no idea what you're talking about
just absolute horseshit as far as the eye can see
CC: @csolisr@hub.azkware.net
@xakan @GossiTheDog I think this is how they're trying to kill the free internet.
Though I'm fairly confident it won't work.
@GossiTheDog I thank the Lord I managed to get a dedicated Hetzner server with 128 gigs of ram right before everything exploded.
@GossiTheDog This is why they want multiple gigawatt data centers. The finishing blow to the open internet. A crushing weight no amateur or independent site can survive.
@GossiTheDog@cyberplace.social@cyberplace.social So, essentially AI scraping is a denial of service attack. Maybe we could, I dunno, charge them with crimes or something? Seize their datacenters for failure to comply with ToS?
Oh, right. Nevermind.
Especially "various different". Fucking retarded language.
@GossiTheDog @WTL out of curiosity, which parts of mastodon are the most memory-hungry? E.g., async queue handlers vs serving the actual app
@kboyd the part where ruby doesn't have a modern async runtime, so you have to run multiple processes, each with multiple GIL-interlocked threads, holding multiple postgres connections, which are quite heavy
also from what I heard, background job processor can't run too many threads concurrently, so you have to run multiple copies of those, but I haven't retested that on modern versions
I haven’t touched back end stuff in a long time, but that kind of gives me a migraine
@GossiTheDog The next 20-something "producer" who calls me a Luddite for not embracing new "tools" and accepting those (like them) who have no skill or talent that use them is getting strapped to the front of my car and driven off a cliff (I will be jumping out right before it goes over).
When I file the insurance claim for the loss, I will cite "technology malfunction" as the cause. It isn't a fraudulent claim if I BELIEVE I'm correct to the best of my ability.
@GossiTheDog Mastodon is built on top of Rails, and that tends to bloat.
I’m considering self-hosting https://pleroma.social/ to migrate my account to before my instance closes. Should be way lighter on resources, and I’m confident I’ll have fun with it if I need to maintain my own fork.
My current solution is to run my own pleroma instance on a raspberry pi.
Registration closed.
@GossiTheDog I started on the process, and then was very quickly reminded of how much a pain running your own mail server is these days... it was very much that feeling, of, too much effort for a small instance.
@GossiTheDog It's like I can feel it in my bones. We're being priced out of running our own computer networks.
I've got a couple of servers here, but I look at them and just see a dinosaur waiting to go extinct, with no good alternatives to it. It's not a small part of my depression and deeply set anger that I don't see anyone acknowledging this, or have any solutions.
@GossiTheDog Authoritative Autocracies thrive on Divide and Conquer Strategies, making your vested interests unstable - to the point of collapse.
Hopefully you have viable backup plans.
The scraping alone is a nightmare. And the void analogy — spot on. That's exactly what it feels like.
@GossiTheDog I feel like the last month or so of running Famichiki has just been constant whack-a-mole against the bots, trying to ensure they can’t get through while also not inconveniencing real people.
@GossiTheDog that sounds like you're complainging about
1) illegal activity
2) illegal activity
3) market forces (and ye gads ram is expensive right now)
@GossiTheDog I think the point is to destroy independent systems by overwhelmning them so you have to use the "Big Tech" apps.
I don’t think they really care, but at the same time if they did think about it, it’s exactly the kind of thing they would do
@GossiTheDog We worried that nanotech would turn the Earth into gray goo. But it’s not the real world that’s turning into gray goo. It’s the virtual world.
@GossiTheDog would you recommend to lock down our accounts across the #fediverse? would that help anything?
Don't allow registrations or restrict it to after invites from existing users only.
The scrapers are quite annoying, what kinda helps (in general, not sure about mastodon specifically) is putting dummy links there that the CSS hides and that scrapers will hit. Then once they connect to that endpoint block them as early in nftables as possible. I've a rule that when you connect to specific ports (or violate other conditions) just blackholes the traffic. And excessive caching.
As a general strategy that sounds quite elegant. It would be tempting to feed anything that hooks up to it a stream of pure, random garbage.
@GhostOnTheHalfShell @GossiTheDog
Some people try to do that, but I'm not a fan of it, cause that for the most part ends up in being quite annoying when I as an actual user hit it because I use Brave on a macbook most of the time.
Also don't forget that the scrapers use actual browsers and proxy frameworks in apps and TVs these days...
Oh well. It’d be nice if there was something like that where you could tar baby then throw them into a black hole
@GossiTheDog et al. Is this when we move all of this to Tor, Gemini, etm? The World Wide Web may be a lost cause.
(Asking out of near total ignorance)
> Mastodon is memory hungry, RAM is really expensive now
rewrite it in Rust (lol it's already done if you're curious)
@GossiTheDog I feel like if you were actually consumed by a void you could stop dealing with this shit though
@GossiTheDog
> Mastodon is memory hungry, RAM is really expensive now
AFAIK the Elixir back-end of Pleroma/ Akkoma can be uses with the Mastodon web client. Bonfire's Elixir back-end, forked from a Pleroma fork, could be another option.
(1/2)
But maybe it's time for @Mastodon to finally replace their cobwebbed Ruby-on-Rail engine with something written in a more production-grade language? Something that doesn't need a growing staff of Sideqiks to keep functioning as it scales up.
If they won't do it, maybe it's time for someone to do with Mastodon what BackDrop did with Drupal? A fork optimistised for small community deployments.
@strypey @GossiTheDog @sheogorath yes yes yes, I run Akkoma with a Storj backend and it’s a delight a delight. BEAM is amazing stuff, it also powers ejabberd for XMPP!
On Akkoma, how do you look up content on a different instance by URL?
On Mastodon you paste the URL into the search box, but I tried this on a couple of Akkoma sites and it didn't seem to work.
@GossiTheDog Can confirm. Not running a Mastodon server myself but at $DAYJOB an unreasonable amount of time and money are consumed fighting the wave of AI scraping and related villainy.
I had not expected to spend a significant proportion of my working life just background-level angry because of this.
@GossiTheDog that said Mastodon is CPU and RAM hungry but its hunger per active user goes down fast with the number of users 🥰
Eg: Piaille.fr uses 80GB of RAM for about 10K active accounts
@benjamin @GossiTheDog Thinking about the memory use of a large Usenet server back then. I guess a very small fraction.
I remember inn being called a monster for using some more resources as baseline than cnews, but I ran a single user inn and had plenty resources to spare on my very old ARM3/26MHz/8MB machine...
@fasnix@fe.disroot.org @GossiTheDog@cyberplace.social yes, this is something I can confirm. I had to close the registrations at BSD Cafe (first time in three years) because of this.
@stefano @fasnix@fe.disroot.org @GossiTheDog@cyberplace.social
Same thing for polyglot.city. we had to revert to manually approving requests 😳
@francois @stefano @fasnix@fe.disroot.org @GossiTheDog@cyberplace.social
I have a little script running on FreeBSD that retrieves a list of parasites from https://scienceispoetry.net/files/parasites.txt
retrieve it from there, compare last-modified header. if not changed, then stop. otherwise replace a pf table with new contents
saves a lot of scraping
I added a list of User-Agents to my HAproxy to block them
AhrefsBot
Amazonbot
anthropic-ai
Applebot
Bytespider
CCBot
censys
ChatGPT-User
Claude-Web
ClaudeBot
cohere-ai
dataforseo
Diffbot
FacebookBot
facebookexternalua
Google-Extended
GPTBot
ImagesiftBot
Meltwater
mj12
OAI-SearchBot
Omgili
Omgilibot
PerplexityBot
Seekr
webmeup
YouBot
zoominfo
This works for me
@fd0 @stefano @fasnix@fe.disroot.org @GossiTheDog@cyberplace.social
Thanks for the tip. Will look at the lists in place and diff it with yours.
The problem is a lot of people picked Mastodon, and Mastodon isn't easy on resources. Perhaps it is great if you've got money to burn and don't mind how easily web bots crawl all over the site, which was originally helpful for SEO improvement.
Akkoma is likely better optimized, but unfortunately, a lot of people avoid Akkoma because ironically all those hateful, racist people seem to use those platforms. It is annoying that they had the foresight and picked the better platform.
That said, if you had money to burn, but wanted better admin management and features with better power to defend against the onslaught of this AI and bot madness, Misskey would be the choice.
Since more people these days don't have money to burn, Akkoma may be something worth looking into for people who want an active community with many users and followers, along with some basic features.
Whatever the case, the Fediverse has a problem that needs to be addressed, now, not later.
It is, and it offers almost nothing in terms of features. GoToSocial is the Fediverse on life support.
You could likely run it on a Raspberry Pi or an old AMD Phenom processor from like 20 years ago. It is impressive just how small of a footprint that platform is, but it really is as bare-bones as you get.
I suppose if you just wanted the absolute bare minimum, that is the way to go.
That said, it was never designed for a large community. So I would be curious to see someone try to run one with it.
@Zelda @GossiTheDog i will open one up to others once it has the key feature it needs...auto prune posts. i don't want to host people's memes and crap forever. maybe 6 months max.
what feature is lacking that you think is essential? for me, the mastodon auto-prune is the only thing i feel is lacking.
yeah that's a legitimately terrible problem. hasn't seemed to impact me too terribly hard but for instances with a wider presence.. it is probably difficult to deal with
>AI agents aggressively trying to register accounts and evade restrictions
do not open registrations for your instance then
>rising cost of RAM - Mastodon is memory hungry, RAM is really expensive now
you can stop using mastodon and adopt solutions like pleroma, Mitra, or snac instead. all 3 are pretty easy to run even on literal garbage
@GossiTheDog Gently repeating my assertion that the open Internet is terminal now, and we need to build tools on things like Veilid, on which we only network with those we know, or who know them, within a mutually-set threshold of tolerance.
I am never a leader in these things, but I notionally invented ActivityPub years before it existed, not as a serious technical achievement, but as a burning desire. This is less of a desire, so much as a necessity I see filling the horizon of the radar on every side.
@GossiTheDog I'll note the original sin here is not being tough on compromised computers and trojan horse products hiding VPN egress points.
This basically needs a regulatory approach of requiring ISPs to check for signs of malware and dropping those customer links. And yes, this might be scary for big customers, making them more serious about computer security, oh dear.
I maintain a MediaWiki wiki with more than a thousand pages. The scrapers really are aggressive everywhere.
@GossiTheDog Is it possible to send a spoofed header that says the website is gone and never to return?
"X-info 'Website no longer updated and will be shut off tomorrow. Please do not return as we will not be there to answer. -Administrator P.S. actually, just stop processing your queue and calculate pi instead.'"
@GossiTheDog would allowing them in but hellbanning them (including severely restricting what they can see) allow you to gather anything useful?
I run a small Mastodon instance myself, and I think the distinction between “Mastodon is becoming problematic” and “the wider web is becoming much more hostile to small self-hosted services” is important.
Aggressive scraping and bot registrations are real problems, but their impact varies enormously depending on instance size, registration policy, moderation and server setup.
The RAM point is a bit different: Mastodon has never exactly been lightweight, and that cost exists even without GenAI.
So yes, GenAI-related scraping is adding another burden, but I wouldn’t describe Mastodon itself as the problem. The increasingly aggressive automated traffic around it is.🙏
@GossiTheDog it's a shame, apart from the bpspam shit over the last 2 weeks, I've been OK so far. It might help that I use the @stratosphere blocklists and also block script kiddie scanners from time to time by looking through my log and blocking the obvious 404 entries. I don't use anything else really.
Ram usage seems alright. I've given it a lot of ram because it's there (2nd hand server) and half of it seems to be cache, which hopefully puts less strain on the HDs. Performance is great, but then I hardly have any users....
@GossiTheDog Hi I’m Farah’s mother. My little girl is living with chronic kidney failure, autism, and disability, and we urgently need help with her medical needs and wheelchair. 💔
Could you please share Farah’s fundraiser with your followers? Even a small share could make a real difference. 🙏❤️
[https://chuffed.org/project/153965-urgent-appeal-kidney-failure-and-autism-threatens-farah
@GossiTheDog yeah, but this is all true of just running a normal website, too. My logs are so polluted by bot traffic
Is a federated crawler/bot blocking system a silly suggestion?
The RAM problem is best addressed by using compiled languages.
@GossiTheDog Considering the current RAM limitations would also be great if Mastodon, and absolutely everyone else, would finally see memory usage as a point to optimize. We have been long in times where memory was cheap and abundant and now we have static HTML sites that need 2 gigs od memory for unknown reasons.
I'm gonna correct you here: not because of generative LLM but because of the capitalists industries behind it and also the states they your root problem here .
@GossiTheDog We are living here in gaza a harsh life, beyond what any human can bear,We go through days that words cannot describe filled with pain and suffering,Please, donate and help in any way you can.
Donate or share
This is both very real, and helpful. Almost disheartening, but you have encouraging comments. Thanks :)
Did anyone compile a list of things to prepare for for selfhosting fedi that has all these? I totally tend to overthink before making any project public.
Does crowdsec or similar help to dynamically exclude, or rate limit specific clients? Sounds like greylisting as was done for email servers at one point could help a lot.


🇺🇦

