Dumb Stormtroopers...
Replaced my front pages with anti-AI articles and quotes a while back.
With AI's rise, AI web crawlers are strip-mining the web in their perpetual hunt for ever more content to feed into their Large Language Model (LLM) mills. How much traffic do they account for? According to Cloudflare, a major content delivery network (CDN) force, 30% of global web traffic now comes from bots. Leading the way …
"It's not like most sites change that much day to day,"
I read rumors that they don't store the results. They rather retrieve data on demand. That's supposed to be cheaper than setting up a copy of the web.
And then there is AI that actually goes out onto the web to look for the answer to your question.
Like any public space, anyone can come in. If they misbehave, they'll be asked to leave. If they persist, police and judges can get involved. Banning from all public spaces can be ordered by a court.
Visiting a public website should be treated in the same way
"Visiting a public website should be treated in the same way"
I don't see that the metaphor can be extended that easily. A scraper bot uses up resources and you don't have much of an opportunity to keep them beyond that gates as it's a game of new ICBM v. new countermeasure. It's not just one bot either. There are loads of companies starting up every week and doing their own reaping. Not just in one country or region but worldwide.
A person asks how often a site would be scraped which sounds like their intention is to infer "not often". If it's cheap and there is value in doing it frequently, it will be done so the master of the iron eeks out a few more groats over costs multiplied by billions. I get almost no junk mail anymore since it's likely more than $1 per piece of mail in total v. email spam (past it's best-by date) that's almost free. There is a benefit that I do look more closely at the mail I do get so if somebody sends me marketing stuff, I might actually look at longer than 1s. In years gone by, I learned to walk past the rubbish bins on the way back to the house with the contents of the mail box. A mentor taught me to sort my business mail the same way. If it is marketing and doesn't apply today, bin it. Leave nothing for later.
I remember seeing suggestions many many years ago that we could eliminate spam if it cost 5 cents to send an email - and you were paid that same 5 cents for receiving an email! Since most of us receive more emails than we send, we wouldn't complain about such an arrangement. Companies that send out marketing stuff that isn't considered traditional spam would have to think about how often they send, just like they had to think about how often they sent stuff through snail mail. But spammers could never afford those charges, it would totally break their economic model. Thus, end of spam. Yes we'd need to figure out how to handle stuff like mailing lists who are sending us stuff we want - maybe create some whitelists for senders we don't charge. I always thought that would work, providing we had a payment infrastructure set up to handle it - sort of a TCP/IP protocol for payments. Too bad blockchain wasn't invented in 1995, that might work well for this.
I wonder if we could do the same for the web? I wouldn't care if I had to pay five cents per page I visited, if there was a way to recoup that money similarly to how I'd recoup the cost of sending emails. But it would have to be a way that you and I could take advantage of to recoup, but AI could not. I can't think of a way to do that off the top of my head. It would also provide the infrastructure for micropayments for people generating useful content to have a way to get paid that isn't filling the page with ads that most of us block, or charging a subscription which isn't really tenable for following one-off links.
It's infeasible. The email idea was infeasible because the obvious thing to do if you're paid five cents per email you receive is to create email accounts and subscribe to every mailing list in existence. Spambots would be trying to eliminate you from their lists, but normal marketing earns you five cents per message. Subscribe to a thousand marketing lists that send a message every week and that's 10 currency units (depending on whose cents these are) per week. You don't even have to read them.
But web micropayments work even worse. How are you going to get that back, as you've stated you intend to? You could host your own website and try to get people to read it. Other than that, it looks like you're out of luck. I don't think we should be paid to comment here. Also, the same gaming that would happen to email would happen online instead. From the page makers point of view, that means the more clicks the better. The Register used to have long articles split into multiple parts. Now, they just put the article, however long it is, on one page. I prefer that, but that's half the page click revenue right there. Less scrupulous sites could easily turn a single page view into ten or twenty clicks, and you're going to stop being happy to pay those bills when they stop being small and you still can't make them back. That's to say nothing of the privacy and security aspects, though there are a lot of both.
Whether or not we could make AI do it, that idea can't happen and any attempt at it would not be accepted. Voluntary micropayments have been tried a few times. Even those have proven unpopular failures.
Two things. Firstly what counts as a visit? Since I use NoScript I'm quite aware of the often shocking number of different sites that the place you *wanted* to go tries to pull stuff from. Are those visits?
Secondly, I would not trust my ISP to manage this correctly, nor would I trust the government not to do stupid things like mandate differential pricing for national/foreign sites.
And from having that infrastructure, it's a short slippery slope to having "pay per megabyte" in order to support the physical wires and machines that make the national telephony system work, as if my fifty f***ing five euros a month isn't quite enough...
The email idea was infeasible because the obvious thing to do if you're paid five cents per email you receive is to create email accounts and subscribe to every mailing list in existence
One objection with a number of simple ways to prevent it doesn't render the whole idea infeasible. There may be some good reasons why it is infeasible, but any "people would game it to make money" is easy to prevent: You can't cash out the money you make. It probably wouldn't be "money", as such. It would be some type of email tokens, which when received are only good for sending emails. And are non-transferable and you can only collect so many before your account "maxes out".
Think of it like if in the days when snail mail was all there was (and before metered postage and "no stamp necessary" type stuff so everything had a stamp) you were allowed to re-use stamps. So you could peel the stamp off letters you received and use them to send out letters of your own, but it was illegal to sell them or give them to another person.
Like I said there are good reasons why this wasn't feasible, but not because of what you said.
you were allowed to re-use stamps.
Is that some US or non-UK thing? Here in the UK stamps have always been cancelled by applying another stamp on top of them (of the ink transferred by rubber hand stamp or later machine variety) so that the stamp cannot be re-used. Modern UK stamps have perforated sections inside them which would tear away if you tried to remove a stamp to re-use it to counteract this, plus (probably) unique 3D barcodes to counter counterfeit stamps.
Surely the problem is going to limit itself? We of this website are the technerati, we're very much agreed that other than a few niche applications and some low value tasks there's f-all value in AI, but for the past couple of years we've been voices in the wilderness. However, eventually the twerps bankrolling the steamroller of AI will work out that they're never going to make a return, and stop funding the idiocy. Tech companies whose real skills are in ad-placement, search, running servers, or selling tat will likewise find that AI is dilutive to those revenue streams. All the standalone LLM that have sprung up like mushrooms will vanish when there's no income stream to pay their bills.
It's always risky to call out when a bubble will burst, could be next week, could still be 18 months away, but absent real value generation then it will burst, and we'll see a re-run for the dot com bubble. So the question for content creators and site operators is not what to do today, but what will tomorrow look like?
From various reports on ElReg it would seem the bubble will have burst within the next 2 years., although it won’t stop those who were foolish enough to make massive investments in AI data centres trying to talk up AI, in much the same way as people will still go on about blockchain, NFT etc.
Though like cold fusion and quantum computing the 2 year window is always 2 years away....
Until the next big fad comes along where it will be forgotten about for 10/20/30 years and then "rediscovered" in a slightly less extreme form (just like fashion)
This post has been deleted by its author
"Surely the problem is going to limit itself? "
The issue is that AI might buy a longer moment of your attention which is one more tap on the wedge They are trying to drive in. If there's enough value there for some company, the attempts won't go away in a self-limiting fashion as the costs get ever smaller so the rewards don't need to be all that large to make it up in volume.
Actual payment is too complex, but you want a way to make the user or sender incur a cost.
There was a proposal for email called HashCash or something I think that was a "proof of work" which required the sender to provably execute a computation which would cost them effort. Bitcoin used a similar idea.
That might be more practical to achieve, the browser has to calculate the next hash in the series or something before you can scroll further. But it would slow the browser down, and like bitcoin, be a massive waste of CPU and hence energy. On a laptop it would rinse the battery.
"Actual payment is too complex, but you want a way to make the user or sender incur a cost."
You want the spammers to pay but not companies providing a useful service. There was a neat Twitter account that would send you a tweet when ISS was going to be overhead. It was a lot of fun and free but Twitter shut it down since it operated much like spam and took up too much resources. Maybe there was a company that started a coupon club where you sign up for a few bob each month and they email you offers that are within a certain radius of where you are that can be customized. That could appear as spam from a certain vantage point.
"I've thought of a modest surcharge per Internet connection, say £1/ month, which is then distributed proportionally to websites visited, thus eliminating the need of obnoxious adverts."
One "anything" per month is nowhere near enough to sustain the internet. In the UK that would raise about £800m a year, which is laughable - the Reg alone needs around £15m a year to stay in business. An internet surcharge would need to be around five times that. Good luck selling the idea of a £600 a year connection tax, and that's before we get to the nightmare of "fair distribution" of the money.
Good luck selling the idea of a £600 a year connection tax
This idea has been sold long ago, it's the daily standing charge that every household in the UK pays for the privilege of being connected to gas or electricity.
It is several hundred £ per year, too.
"This idea has been sold long ago, it's the daily standing charge that every household in the UK pays for the privilege of being connected to gas or electricity."
It's not payment for a privilege, it's the cost to process the billing, customer support and keep the physical infrastructure in good repair. They could simply start charging a fee for billing, every phone call and only cover repair costs to the pole outside (or underground junction). The last thing you'd want after a storm is a bill for restoring the downed power line to your home. I remember when the phone company had gone from doing everything from end to end to only covering repairs to a MPOE (minimum point of entry). One could pay a small fee each month to cover inside wiring if they liked. With my manufacturing company, we were fine with doing all of our inside phone work ourselves, but my mom was better served by having complete coverage where they'd come out and fix it if it stopped working for just about any reason.
If I were to be totally on solar, I'd still keep an electricity account. The standing charge is quite reasonable and I don't have a choice anyway. The city requires electrical and water for the house to be considerable suitable for occupancy. It's hard to evict a non-paying tenant, but if the house is condemned, the city can come over and nail up the doors and windows (with notice). Breaking back in would be a criminal offense.
The initial problem we were trying to solve was bot access, and this would not work. Whether it would work for helping more sites exist without ads is a different problem, and I don't think it would, but we can get to that. In the area of bots, though, a per-connection charge isn't going to work. At £1 per connection per month, they could easily farm their connections to a couple thousand bots and run their scans for a couple months. They're going to spend at least a hundred and probably thousands as much on GPUs. They can easily afford that and the sites that get a small fraction of that revenue will not care.
For funding sites, it would provide some funding, but consider how much it would really do. How many individual sites do you think you've visited all August, counting sites where you viewed one page and may not even have read the whole thing? I would have trouble calculating even a mildly close estimate to that. How much of that revenue would go to search engines? I perform a lot of searches, but that's a small number of page views compared to sites where I browse to a lot of pages, so if we paid website owners by number of pages viewed, the search engines would get little. How much would that compare to the per-search ad revenue they get now?
Let's consider DuckDuckGo. It recently gave us some revenue and market statistics which make these calculations work well. It does not give us a number for UK market share, but it claims a 2.11% market share in the US, higher than its global one. That would mean 7.176 million active users, using a basic calculation of everyone in the United States, although since infants or prisoners don't do many web searches, that's probably an overestimate. £1 = $1.35, so these users will be contributing $116M per year. Let's assume that web searches account for one in fifty page views, again probably a significant overestimate. This gives DuckDuckGo annual revenue of $2.3M. Their global revenue now, using advertising, is $100M. I don't think they'll be cutting the ads any time soon.
I'm one of those DDG users in the US. I have it as my default on both Linux/Firefox on iOS/Safari. If I can't find what I'm looking for on DDG I'll try Google, but that rarely helps.
One of the reasons I started using DDG was because Google had seriously fallen in quality (partially SEOs, partially Google's profit motives about what results they show) so I figured it wouldn't make much difference, and I liked not contributing to Google's data collection or ad revenue. And I was right, it made no appreciable difference. I think since then DDG has improved slightly, while Google has continued to get worse.
I've been trying to use DDG as my preferred search engine for a few years, but the absolute worst thing it does is to return "hits" which don't include all my search terms. I wasted time reading irrelevant web pages until I was sure it was happening. Now I can recognize unlikely hits and try Google instead.
"but the absolute worst thing it does is to return "hits" which don't include all my search terms."
So a lot like Google, then.
I do real estate photography for agents as one of my gigs. If I search for "real estate photographer" in my area, I get almost none and I show up 6-20 pages in as I don't pay Google for adwords, placement, etc. What I get back in the results are the franchise real estate offices, their parent franchisers, ads for cameras, etc. Just not much in the way of photographers. I can write a long regular expression for better results, but agents looking for a photographer can't tie their shoes, much less write a regular expression to get better search results. The want to (I was going to say "type in") use voice search with "real estate photographer in "city"" as their query. The results are no help to them and zero to me, but they might fat-finger a result that earns Google a cent or two. Or the searcher might spot a photographer only to find they are 1,000 miles away in another state. For some strange reason, if I search a phone number I'll often get listing for hospitals or certain specialist doctors where the web page matches maybe 3 of the numbers in the string. Nobody is gong to convince me there isn't some paid reason that happens.
This post has been deleted by its author
> Needs to be something that costs the bot at least an order of magnitude more than the site.
I wonder if we couldn’t borrow from the malware community and write misleading URLs that seem to reference something from your legitimate website, but actually result in the AI bot doing the ChatGPT request…
The attack is once ou’ve identified the AI bot is to issue a DMCA takedown..
I've had to automate a way to replace all pages on my forum with a short message when AI bots hit the site, to reduce the bandwidth costs and keep the server response acceptable for everything else. This happens several times a day, sometimes going on for hours. Even though the bots are getting a plain text response with no links, they won't take no for an answer. I wish there was one!
...of BBS (Bulletin Boards), I remember comments then about how things would eventually degenerate once businesses/enterprise got involved. Then the Internet/web burst on the scene. Finally it has all come to fruition - the total enshittification of the internet/web and its destruction by the greed of businesses, who are too avaricious to see the collapse coming. Bring it on.
I miss the web of the late 90s where you could talk to "real" people without bots (spam or otherwise), hand holding "ranking systems" Lycos did this to their chat service and kiddified it with cute cartoon graphics out of the blue for whatever marketing bs reason.
Shit take me back and let me stay there....please?
"I miss the web of the late 90s where you could talk to "real" people without bots "
All of the internet BBS software seems to have gone away. I was on a few BBS's and made some great friends both online and IRL from them. My buddy Dave and I would let our cats chat as well. We both hit the floor one night when Panther and Ogre started meowing back and forth. The audio chat was a bit buggy and there was no video. I think we were using Hotline. I have some old Macs and it might be fun to see if the software might work although there would be no tracker to learn of new BBSs.
Skype killed the last of the BBS software firms that were working towards business oriented functions to help make it a paying venture. I never cared for Skype as it required a man in the middle. Too often our discussions and other activities were NSFW.
My own webserver got hit yesterday causing it to crash with 304 errors on anything that involved the MySQL server. Looking at the logs found all sorts of bots were hammering my phpBB instance.
I looked into installing Anubis but that would require reconfiguting my webserver as a reverse proxy.
As an alternative, I ended up using ngx_http_limit_req_module to limit access of .php files on the phpBB instance to 1 per second (burst 5) for any particular IP address, and it immediately made my server responsive again. Didn't bother doing the same to my WordPress server (yet) as the caching plugin is helping with the responsiveness, and the site is small enough a scraper can come and go in short order.
Count yourself lucky. :/
My personal site is of no appreciable consequence, hosting mainly dribbles of code I've open-sourced, along with the VCS and ticket system I stuff it all into. Actual real human traffic typically amounts to, well, me committing things to version control, and a couple of requests to remember which ticket some commit needed to be tagged with, and possibly as many as five actual third party visitors per month, most of whom almost certainly (I haven't waded through the log swamp to find them) only request static files from my ~user subsite rather than the "main" site, such as it is.
I have the good fortune to have two moderately beefy couple-of-generations-old enterprise servers colocated at work for more or less free, with my main site recently converted to a VM scaled to about an eighth of one machine's CPU/RAM. I have plenty of capacity to throw hardware at the problem if I figured it would even be worth the bother of trying.
I have recently noticed instances of 5+ hours (verified at least as far as the Proxmox CPU usage graph) of being hammered with absolute gibberish requests that are clearly straight up hallucinations and have no just cause to have been requested from me at all... on top of the way-faster-than-justifiable requests for paths in the VCS infrastructure from at least six different identifiable bots, and an unknown horde that don't have the courtesy to identify themselves. I noticed because my source code commits and other access to my own resources were failing because the bots were hitting everything so hard the DB was hitting its connection limit.
My Apache log file for one month should be on the order of maybe a couple of megabytes. It's been >1GB for close to a year. For July it hit 2.5GB. Last month the useless garbage bots pushed it up to 2.8GB.
AI bots can just go fold themselves until they're all corners, stuff themselves up their own backsides, and launch themselves into the nearest supernova.
AV-type detector for AI bot DoS attacks http/s request would seem a way forward. Maybe Apache and the like could build all that right in?
The tech bros would soon learn to cut access deals, just as advertisers do with browsers and adblockers.
No doubt some of the big site providers would scrape client site updates for themselves, and license the anonymised stream to AI Daddies - thus monetising the solution. And, of course, demand big dollar from any client who wishes to opt out.
Hey-ho, on we go.
Sounds like we need a website filter that blocks the bots.
Trouble is this is relatively simple to implement for a self hosted/on-prem website, but for those using hosting providers things might not be so simple.
Publish the IP address ranges that these attacks come from and encourage whoever is upstream from you to block traffic from them at the edge router/firewall. Brute force, sure, but what's good for the goose is good for the gander.
Sounds like we need a website filter that blocks the bots.
Where's the fun in that?
I'm thinking we have robots.txt, which gets ignored. So maybe a foad.txt which bots ignore, but then if you can characterise a bot request, serve it with poison.txt. Direct the bots into a smorgasbord of random or pseudo-random content and let the LLMs eat garbage. No idea how easy that would be to do in a way that would be less resource intensive than a regular bot hit & run, but with enough adoption, it could be entertaining.
> I'm thinking we have robots.txt, which gets ignored.
I wonder if it gets totally ignored, ie. Not even fetched, or if it gets fetched.
My simplistic assumption, is that only bots (and the curious) actually look for and fetch robots.txt et al., so that in itself should enable some bots to be automatically identified and messed with…
That would help with a few of them, but the problem is that many bots try to bypass them. Some bots read robots.txt and find that they're blocked. Others don't, but you can implement a filter on similar means which detects when a bot has announced itself and blocks it then. The problem is that quite a few, possibly all, of the bots then come back but identify themselves as normal browsers instead. Now that they're no longer announcing their presence, simple filters aren't as effective. Now you need to profile their activity to try to guess whether this is a bot flood or not and only ban those bots that are engaging in it, and that's not so easy to do on your own systems either. There are methods you can go to, but it's no longer as simple as inserting a basic config into your webserver.
If your website is being attacked by bots which are ignoring the rules in robots.txt then start feeding them garbage. Feed content which is itself AI generated, possibly containing malicious, actively wrong information. Generate images which are not what they say, generate text which is gobbledegook, fake news, libellous content, fake sports / places / people & facts, medical misinfo. Try and make it something likely to stand out from the homogeneous slop and generates new paths in their models. Just delberate garbage. Either they don't detect the garbage and end up poisoning themselves, or they do and will probably leave you alone. Either way, it's their fault for being dicks.
This post has been deleted by its author
"Feed content which is itself AI generated, possibly containing malicious, actively wrong information. Generate images which are not what they say, generate text which is gobbledegook, fake news, libellous content, fake sports / places / people & facts, medical misinfo."
Just create a normal social media website then?
For each page that you post, request from chatgpt to scrabble it.
On the following text, replace each name with another name, each verb with another verb at the same tense and each adjective by another adjective...
And post this page and add a hidden link to it in several of your existing pages.
The AI, cannot determine that this text is cheat, because it is syntactically correct, even if there is no meaning for any human... And the AI will get poisoned.
If enough people do it, AI companies will have to negotiate so that you don't serve them this cheat... Negotiate meaning pay...
Web hosting is cheap and convenient, but it will not handle these loads (as was mentioned above).
Part of the issue here is that most apache websites run with the default configuration of:
KeepAlive On
ServerLimit 16
StartServers 2
MaxClients 200
MinSpareThreads 25
MaxSpareThreads 75
ThreadsPerChild 25
Which can only handle 50 requests before having to start another server and the overhead it requires.
And then maxes out at 400 concurrent users.
This is usually ok for most sites as they are not really that busy and are serving static with some dynamic pages.
AI clobbers this.
The best cure for this is to run your own servers.
We run HP Z840's each with 2 Xeon processors 28 total threads, 56 vcpu's and 128 GB of ram, nvme SSD's on a 2GB network connection.
Cost of the HP Z840's $700.00 CAD each.
Internet connection $145.00 CAD month. - Remember, you already pay for an internet connection, now you have a bigger one!
Our Apache2 configuration:
KeepAlive Off
StartServers 50
MaxRequestWorkers 20000
ServerLimit 200
ThreadsPerChild 100
ThreadLimit 200
MinSpareThreads 3000
MaxSpareThreads 4000
PerlInterpStart 50
PerlInterpMax 100
PerlInterpMaxSpare 50
This gives us 5,000 apache concurrent users available and 2,500 mod_perl interpreters available.
Only takes a couple of GB for space as everything is multi-threaded.
Scaling to 10,000 and 5,000 respectively.
We get hundreds of requests per second just from hackers looking for wordpress, drupal, php, javascript and other security vulnerabilities.
AI not so much as you have to log in to use your free account, and there are not many webpages available.
Not an unusual amount of Firewall activity - Does spike up sometimes.
Site never goes down.
We use CloudFlare (trying it out) to cache images but since they open your https packets to change it to their SSL we are not really sure of the benefits.
Also, they only work on your domain name and not your IP address (why the Ai can get around it).
It's always tempting to take a look, and they are american based i believe?
Don't want to rant, but running your own web servers can be very advantageous to your business and prices for really good quality servers is/are really good.
Remember, that is the way it all started out.
If we define AI crawlers in violation of copyright and sovereignty of content, we can declare them to be hacking agents and treat the companies who utilize them accordingly..
Which means in practice they will be dealt with as once was the practice in rhe wild west west with horse thieves. By using an AI sheriff app to document the ransacking of the sites visited we can provide a record of wrong doing and use class action lawsuits in European Union competent courts to slap fines running into tens of billions of dollars, effectively choking these companies to death.
The biggest peeve I have is them totally disregarding the effort I've gone to to highlight the useless content. They grab it anyway.
We generate lots of filter links on our sites and there a a LOT of permutations of filters. The content looks largely the same just, well, filtered down a bit. Simply loading no filters and following paginations links will get you all the content. So I've diligently explained when there is no point in indexing this page or there is no point in following the links (as they can be followed from other, indexed pages). AI knows betters so lets just hammer my sites and cost me bandwidth, money, slow down the service to the detriment of paying customers etc. Such a pain. Wish they could play nice.
My blog has the fake path to the content being /blog/entry/date-goes-here. For historical reasons it uses the <base> tag in order to set the base URL as the "..." part.
Every browser handles this properly. Yes, even that crappy ancient one that wouldn't know what a PNG was or what to do with it.
Bots, not so much. It's mildly amusing seeing in the log that something has tried to fetch /blog/entry/entry/entry(repeat about twenty times)/some_image.jpeg (that isn't even at that path, images are in /blog/images !).
So I don't think the problem here is that their directive from on high is "scrape everything" and they put together the quickest, cheapest, shittiest bit of code that does that without crashing (too often). That's why it screws up on my <base> and why it ignores your noindex/nofollow. These are both things just a nudge out of the ordinary that would require somebody to actually think about what they're doing...
I think that although we may be in a period when AI results are useful sometimes, this is because they have been trained on human-generated content. As a larger proportion of the internet becomes filled with AI slop, the bots will ingest more of it and the "answers" will become progressively worse. At some point, the bubble will burst.
"they also tend to disregard crawl delays or bandwidth-saving guidelines"
Well, fair is fair so one must point out that the Googlebot also ignores crawl delays, this is specifically mentioned in the bot info page, in order to use some opaque algorithm that...
...let's just say if it wasn't a known search engine, I'd have blocked it for its misbehaviour.
"where you must pay cash money to access almost anything"
It's already fast heading that way due to corrupt data agencies and the inability of the EU to do anything useful - we have far too many "consent to invasive tracking by our 726 partners or pay us".
How much of the traffic is crawling for training materials, and how much is AIs searching on the user's behalf?
For example, I might wonder before a purchase, whether the tape deck on the Aiwa BBTC-550MG boombox detects different IEC tape types. Before AI I'd spend an afternoon searching through hundreds of web shops who copy paste the same useless marketing material, hundreds of review sites that don't actually review and don't actually know anything about the subject. With AI, the LLM churns through 150 web search results a minute, much much much faster than I could search.
Is that the traffic that I causing problems?
On a side note it would be nice if there was some way of organising the web so one could reliable bundle all the copypasta together.
I run an internet forum on a niche subject. There used to be thousands of these, and they were wonderful, and you could use Google to query all of them. They were all owned by different people and not an American multinational. As the article says, it was the real, open web.
Then along came Facebook and its free groups. In one swoop, forums were drowned out by these groups which are actually inferior in every way, but are far easier to setup and require no technical skill. There are about 30 poor quality groups covering the same subject as my forum - all badly organised, hard to search, and full of misinformation.
My forum is hanging on, mostly because of my technical skills and determination, but the fight against the AI bots might just be the final straw. I'm using every tool that the CloudFlare Free product offers to provide defence, but it's a constant battle. Their AI Labyrinth tool is superb, as are their managed challenges. Without Cloudflare my site would collapse (it was collapsing before I set it up).
I don't like being reliant on cloudflare, but i dont see any alternatives that i can afford.
I dont see how any independent website can survive without IT expertise being involved, and that will remove most of them. Facebook it is then...