Sometimes you have to fight fire with fire. But when it comes to AI-produced slop and hateful content threatening the safety and value of social media platforms, adding more fire—in this case, more AI—can make the problem worse.
At its best, social media can be a haven for people who want to share their experiences and knowledge. It gets closest to this ideal when users contribute authentic, valuable content, whether that’s a uniquely thoughtful blog post or a helpful video on how to build a PC. Relying primarily on AI tools to preserve that authenticity misses what makes social media worthwhile in the first place: the people behind it.
Key facts
- In April, Reddit’s AI moderation tools removed dozens of old posts and comments from r/AskHistorians, including content by experts.
- Reddit says AI has increased enforcement actions on hate and violent content by more than 200% and reduced exposure to harmful content by more than 40%, but false positives remain a huge problem.
- Discord admitted its AI moderation system wrongfully banned about 8,400 accounts between May and early July due to falsely labeling images with square grids as CSAM.
- Facebook, Instagram, and Tumblr users have reported wrongful automated bans and content flags.
- Research indicates marginalized groups are disproportionately affected by AI moderation false positives.
Erroneous erasures
In April, a Slack channel for moderators of the r/AskHistorians Reddit community was unusually busy. The channel, which automatically receives links to modmail messages, was flooded with alerts after dozens of comments and posts dating back 10 years were automatically removed from the subreddit.
“And there was nothing we or the experts [who posted the deleted content] could do about it,” Dr. Sarah Gilbert, one of the moderators, said.
This was particularly damaging to the subreddit because its users view the community as an archive of detailed responses that continue to educate people long after content is posted. The moderators believe Reddit’s recently revamped AI moderation tools were responsible for the removals. After recovering the text of some posts, one of AskHistorians’ moderators noticed that all the deleted content linked to a historical image-sharing website. The moderators think Reddit might have designated that website—and thus any post using its content for explanatory illustrations—as spam.
Reddit has not responded to a request for comment. The deletion erased valuable information that had taken time to aggregate. Some people spend hours, sometimes days, researching and writing responses to questions submitted to the subreddit. Yet it’s possible that those erroneous removals, and others like them, have contributed to metrics intended to demonstrate how effective AI moderation is on Reddit.
Reddit says that thanks to AI, it has increased enforcement actions on hate and violent content by more than 200 percent and that AI drives faster, higher-volume enforcement. AI has helped reduce exposure to potentially harmful content by more than 40 percent, Reddit said this month. It also said that it uses large language models (LLMs) to catch highly subtle, coordinated patterns of fake behavior and artificial hype.
But as the AskHistorians ordeal illustrates, more enforcement doesn’t necessarily mean better enforcement.
The false positives problem
The growth of generative AI has created new obstacles for social media moderation. Gilbert noted that large language models have made spam detection a lot harder, as they seek to mimic real human voices. “Over the last two to three months, we’ve been absolutely flooded by LLM-powered spambots,” she said.
Marketing agencies are creating social media content designed to get brands cited by generative AI chatbots. Marketers have long used inauthentic social media posts to boost visibility, but the rise of chatbots has opened a new front. Startup ReachLLM, for example, focuses specifically on marketing through chatbots. As part of that effort, company representatives have created and moderate subreddits on Reddit.
These challenges have led some social media companies to explore new AI-based moderation techniques. Reddit says its AI tools have revoked nearly 2 million fake votes daily and that it uses LLMs to catch coordinated patterns of fake behavior and artificial hype that older systems once missed.
But many social media platforms have become overly reliant on AI moderation tools that have been quick to penalize users for innocuous content.
Recently, Discord admitted that its AI moderation system wrongfully banned about 8,400 accounts from May to early July. The AI mistakenly labeled images containing square grids, such as chessboards or spreadsheets, as CSAM and subsequently issued a permanent ban to the uploaders. Discord says all affected accounts have since been reinstated.
The company said its AI moderation was not intended for use without human supervision. It claimed that a human employee is supposed to review AI-flagged content before Discord takes action, but a bug caused the AI to bypass the human step and ban accounts.
The mishap highlights why human guardrails remain essential in content moderation. Without meaningful oversight, an AI-based moderation system can make thousands of mistakes in a matter of weeks, with lasting consequences.
Since 2025, many Facebook and Instagram users have complained about mass bans they blame on AI moderation. The lack of human moderation has only fueled frustration among users who say they did not violate any rules, especially since there has been no way to speak with a Meta employee about what caused the ban or how to get an account reinstated. Meta has not said whether AI is behind the bans, but the company has increasingly relied on generative-AI-based moderation rather than humans in recent years—a shift that some people, including Meta employees, say is happening too quickly.
Tumblr is another social community where automated moderation systems have failed. In March, a Tumblr spokesperson said that Tumblr’s automated systems wrongfully banned fewer than 200 accounts in one afternoon. In 2025, Tumblr users complained after the platform’s automatic content moderation systems inaccurately flagged content as mature, reducing its visibility. In both cases, users blamed AI. Tumblr never confirmed that AI caused these problems, but the company has said it uses a mix of machine-learning classification and human moderation.
AI moderation can save social media companies money and help remove harmful content faster. But until these systems can eliminate basic mistakes—like labeling a checkerboard picture as CSAM—they need human oversight.
“Back when there was more transparency in the system, we would routinely report hate and get an automated response that it wasn’t actually in violation of Reddit’s rules, prompting us to start an appeals process,” Gilbert said. “So it’s hard to trust the numbers because it’s hard to trust the ‘judgment’ of Reddit’s systems.” False positives are a “huge problem” on Reddit, she said.
AI’s biases
Typical social media AI-based moderating systems use machine learning classifiers to analyze posts and identify and flag content that breaks platform rules. But it’s difficult for a machine to understand the nuances of sarcasm, satire, and slang. A machine may mistake a quote of a hateful phrase for the real thing, or it may miss a slur disguised with deliberate misspellings.
Further, research suggests that marginalized groups can be disproportionately affected by AI moderation. Without human oversight, AI can end up penalizing the very communities most vulnerable to the hateful content the systems are designed to combat.
Gilbert, who is also the research director of Cornell’s Citizens and Technology Lab, says that marginalized and vulnerable populations are among those who experience the highest rates of moderation, and that typically this is a result of false positives, often driven by instances of counter-speech, language reclamation, and responses to hateful content. “False positives are an equity issue. They mean that groups that are already marginalized are further silenced and censored,” she added.
AI moderators can also make communities less effective at moderating themselves. On Reddit, for example, some subreddit moderators would prefer to ban users who use hateful or violent rhetoric. But if Reddit’s AI removes such content before a human moderator sees it, those moderators lose the ability to assess whether a ban is warranted.
In terms of giving human moderators more control, Reddit this week announced expanding testing for Rules Hub, a suite of tools that lets human moderators choose which rules should be automatically enforced, decide what happens when a rule is triggered (send to queue, filter, or remove), preview the experience before enabling it, and review logs and insights. Reddit expects Rules Hub to eventually replace the Automod tool, which relies primarily on exact keywords.
AI is a tool, not the solution
Moderators have repeatedly blamed the generative AI boom for a spike in content that breaks community-specific or broader platform rules. That’s a serious problem for social media sites that rely on user contributions. As generative AI becomes more accessible, bad actors can produce spam, harassment, and disinformation at near-zero cost. This arms race forces moderators to consider new tools for detection, but the answer is not to eliminate human judgment.
Automated systems can help triage the overwhelming volume of content, identify obvious violations, and surface suspicious patterns. But a machine cannot fully understand context, intent, or the social norms that distinguish meaningful discussion from harmful behavior. Human moderators bring cultural awareness, empathy, and the ability to make nuanced decisions. They also can be held accountable for their decisions, while AI systems often remain opaque.
Social media companies also need to design better appeals processes. Users who are wrongfully banned should not be left with no way to contact a real person. Transparency is essential: if a user knows why content was removed and can appeal, trust in the system improves. Some platforms have started to provide more detailed explanations, but many still rely on vague warnings and automated emails.
Another key layer is community involvement. Volunteer moderators, like those on Reddit, have deep knowledge of their communities’ history and rules. Platforms should give them tools to configure automated moderation to match their needs, while still allowing them to review decisions and intervene. Reddit’s Rules Hub is one attempt to do this, but similar approaches should become more widespread.
There is also a broader responsibility for AI developers. Companies training moderation models must pay attention to biases in training data and test for edge cases. A system that labels chessboards as CSAM is not merely a bug; it reflects a failure of design and testing. These systems must be validated in real-world conditions before being deployed at scale.
Regulators and policymakers are beginning to look at algorithmic decision-making, including content moderation. There are proposals to require transparency, accountability, and human review. While such regulation could help, platforms should not wait for legal mandates. Protecting users and preserving free expression should be the guiding principles.
Generative AI is not going away. It will continue to create new challenges for content moderation, from spam and coordinated inauthentic behavior to more subtle manipulations. But the solution is not to let flawed AI systems act without supervision.
Companies will continue to try new methods of moderating more reliably and effectively, but reducing human input is a step backward. Low-effort AI-generated content is changing the challenges moderation teams face, but that makes stronger approaches more necessary, where machine-scale detection can be combined with human judgment and expertise. Just as social media has no value without people, content moderation can’t succeed without human judgment at the forefront.
Source: Ars Technica News