AI Moderation Faces Scrutiny Amidst Errors and Evolving Online Threats
Social media platforms are increasingly leveraging AI for content moderation, but recent incidents highlight its limitations and the critical need for human oversight.

Social media environments, ideally serving as platforms for genuine knowledge exchange and shared experiences, are most effective when users contribute authentic and valuable content. However, an over-reliance on artificial intelligence (AI) for content moderation risks undermining the human element that forms the core value of these communities.
For instance, in April, the r/AskHistorians subreddit experienced the automatic deletion of numerous posts and comments, some dating back a decade. Moderators of the community, including Dr. Sarah Gilbert, noted that neither they nor the experts who had originally contributed the content could reverse these removals. This event was particularly detrimental to AskHistorians, which functions as a historical archive for its users. The moderators suspect that Reddit's updated AI moderation tools were responsible, possibly misidentifying content linked to the historical image-sharing site Rare Historical Photos as spam. Such erroneous deletions eradicated valuable information that had often required extensive research and time to compile.
Despite these issues, Reddit has reported increased enforcement actions and reduced exposure to potentially harmful content through its AI tools, including large language models (LLMs) designed to detect subtle patterns of fake behavior. The company states that AI has led to a more than 200 percent increase in enforcement against hate and violent content and a 40 percent reduction in exposure to harmful material. However, the AskHistorians situation suggests that a higher volume of enforcement does not always equate to more effective or accurate moderation.
The proliferation of generative AI has introduced new complexities for social media moderation. Dr. Gilbert observed that LLMs have made spam detection considerably more difficult, as they can mimic human communication convincingly. Over recent months, communities have reported being inundated with LLM-powered spambots. Marketing firms are also now creating social media content specifically to influence generative AI chatbots, with some, like startup ReachLLM, even establishing and moderating subreddits to promote brands through chatbots.
In response, some social media platforms are exploring advanced AI moderation techniques. Reddit, for example, claims its AI tools revoke nearly two million fake votes daily and utilize LLMs to identify sophisticated patterns of artificial hype missed by older systems. Yet, several platforms have been criticized for an over-reliance on AI that penalizes innocuous content. Discord recently acknowledged that its AI moderation system mistakenly banned approximately 8,400 accounts between May and early July. The AI incorrectly flagged images containing square grids, such as chessboards, as child sexual abuse material (CSAM), leading to permanent bans. While Discord stated that all affected accounts have since been reinstated and that a bug bypassed human review, this incident highlights the indispensable role of human oversight.
Similarly, Meta has faced user complaints since 2025 regarding mass bans on Facebook and Instagram, which users attribute to AI moderation. While Meta has not confirmed AI as the cause, the company has increased its reliance on generative AI for moderation, a shift some employees believe is occurring too rapidly. Tumblr has also experienced similar issues, with its automated systems mistakenly banning hundreds of accounts and incorrectly flagging content as 'mature,' reducing its visibility. Although Tumblr has not definitively attributed these problems to AI, it uses a combination of machine learning and human moderation.
While AI moderation offers advantages in cost-efficiency and rapid removal of harmful content, its effectiveness is limited by its inability to grasp nuances like sarcasm or satire. Research indicates that AI moderation can disproportionately affect marginalized groups, often resulting in 'false positives' that silence vulnerable communities under the guise of counter-speech or language reclamation. Dr. Gilbert, also the research director at Cornell's Citizens and Technology Lab, emphasizes that false positives represent an equity issue, further marginalizing already vulnerable populations.
Furthermore, AI's preemptive removal of content can hinder community-based moderation. If AI removes hateful content before human moderators can assess it, they lose the ability to determine if a user ban is warranted. Reddit is addressing this by expanding testing for Rules Hub, a new suite of tools empowering human moderators to customize rule enforcement, preview outcomes, and review logs. This initiative aims to replace the keyword-dependent Automod tool. The surge in AI-generated content poses a significant challenge for social media platforms that depend on user contributions. While companies will continue to refine moderation methods, reducing human input is a retrograde step. Effective content moderation requires a synergy of scalable machine detection and human judgment, ensuring that the valuable human element remains central.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



