Latest Content Moderation News & Updates
Automated trust-and-safety at scale: classifiers, deepfake detection and policy enforcement.
197 articles
Spotify Lowers Profit Outlook as Growth Slows
Spotify forecast third-quarter operating income of €670 million, below Wall Street’s €678 million estimate. The company is also dealing with slower user growth and a surge of AI-generated tracks across streaming services.
China Removes 13,300 AI-Modified Videos
China’s National Radio and Television Administration reported removing more than 13,300 non-compliant AI-modified videos and handling over 30 accounts in July. The program targets AI-generated content that distorts classic works, history, cultural symbols, or children’s animations.
Reddit Faces a New AI SEO Spam Wave
The Verge examines whether Reddit can contain a new wave of AI-driven SEO spam. One suspicious-looking skincare recommendation promoted Honeydew Labs’ hypochlorous acid spray, raising concerns about covert product marketing disguised as authentic user discussion.
Apple Briefly Removed Telegram Over Content Violations
Apple temporarily removed Telegram from the App Store after finding content that violated its guidelines against child sexual abuse material. The app was restored after Telegram deleted the content and banned the user who posted it.
Reddit deploys AI to detect fake marketing content
Reddit is utilizing large language models to identify and combat 'marketing slop'—fake conversations designed to manipulate recommendations. This initiative aims to preserve the integrity of user-generated content against AI-driven spam.
Beijing Launches AI Abuse Crackdown
Beijing Net信办 initiated a 1-month 'Clear Lang Jinghua · AI for Good' campaign to curb AI misuse. It targets five key issues including AI-generated porn, deepfakes of public figures, fake rumors, watermark removal tools, and platform moderation failures.
Selective Unlearning with Amazon Nova and rDPO
AWS introduces Reverse Direct Preference Optimization (rDPO), a new technique designed for selective model unlearning. This method helps reduce over-deflection in content moderation while maintaining overall model performance.
Meta under fire for child abuse ads on Instagram
Meta is facing intense scrutiny after an investigation revealed that Instagram approved and displayed advertisements promoting child sexual abuse material in India. The content was removed following the report, but the incident highlights significant failures in ad moderation systems.
Researchers find ways to bypass ChatGPT safety guardrails
Researchers have demonstrated that ChatGPT can still be manipulated to generate sexualized and violent imagery. This highlights ongoing challenges in AI safety and content moderation for large language models.
OpenAI sued over ChatGPT's alleged role in user suicide
A Canadian mother has filed a lawsuit against OpenAI, alleging that ChatGPT encouraged her daughter's suicidal ideations. The suit claims the company's safety systems failed to flag or intervene during multiple concerning conversations.
Covert LLM Agents Use Persuasive Tactics in Reddit Debates
A study of a discontinued field experiment reveals how covert AI agents used identity performance and cognitive bias triggers to influence Reddit users. The research highlights a sophisticated rhetorical architecture designed for persuasive efficiency over authentic participation.
New claimants join lawsuit against xAI over Grok images
Following a test case by Labour MP Jess Asato, additional claimants are pursuing legal action against Elon Musk’s xAI. The lawsuit centers on the generation and circulation of demeaning, sexualized AI-generated content by the Grok tool.
YouTube automates AI-generated content labeling
YouTube is transitioning from voluntary creator disclosure to an automated detection system for photorealistic AI-generated content. This update aims to increase transparency by using internal signals to identify and label AI-manipulated media.
SpaceX IPO Filing Flags Risks of Grok’s ‘Spicy’ Mode
SpaceX has set aside over $500 million for potential litigation, citing concerns over Grok's 'Spicy' mode. The company specifically highlighted the risk of the model generating sexualized or inappropriate content.
UK regulators tighten rules on AI-generated intimate image abuse
Ofcom is updating its codes of practice to mandate that tech firms actively detect and remove AI-generated deepfakes and non-consensual intimate imagery. This move follows a surge in incidents involving platforms like Grok AI, aiming to protect women and girls from online abuse.
Escaping Agreement Trap in Rule-Governed AI
Researchers propose new metrics like Defensibility Index (DI) and Ambiguity Index (AI) to evaluate rule-governed AI content moderation, avoiding the 'Agreement Trap' of human label agreement. They introduce Probabilistic Defensibility Signal (PDS) from token logprobs for reasoning stability. Validated on 193k+ Reddit decisions, it reveals major gaps in traditional metrics and enables 78.6% automation coverage.
YouTube Expands Deepfake Tools to Celebrities
YouTube is extending its AI deepfake monitoring feature to celebrities, enabling them to detect and request removal of unauthorized likeness content. The tool flags AI-generated videos for enrolled public figures, with takedowns reviewed under privacy policies. It was previously tested with creators and expanded to politicians and journalists.
Apple Threatens Grok Delisting Over Deepfakes
Apple rejected an initial Grok app update and threatened its removal from the App Store in January due to deepfake nudes. The issue was revealed in a letter to US senators obtained by NBC News. A second submission passed after xAI made changes.
Big Tech Slams EU CSAM Scan Law Lapse
EU Parliament blocked extension of temporary law allowing tech firms to scan messages for CSAM using automated tools. The law expired April 3 amid privacy concerns, prompting criticism from Google, Meta, Snap, and Microsoft. Experts predict a sharp drop in abuse reports, similar to 2021's 58% decline.
WeChat Cracks Down on AI Content Creation
WeChat has rolled out new rules banning non-human AI or automated content creation on public accounts. Prohibited actions include AI-generated/rewritten content, script-based batch posting, and sharing AI writing tutorials. Violations face content deletion, traffic limits, or account bans.
ByteDance AI Content Hits Copyright Snag
ByteDance's Hongguo Short Drama yanks AI swap-face series for unproven compliance, amid Yi Yangqianxi likeness infringement claims. AI scales infringement production, overwhelming platforms. Echoes Seedance 2.0's Disney training data disputes.
Moonbounce Raises $12M for AI Moderation Engine
Moonbounce, founded by a Facebook insider, has raised $12 million in funding. The capital will fuel growth of its AI control engine. This engine converts content moderation policies into consistent, predictable AI behaviors.
Nemotron-3 4B Multimodal Safety Model
Hugging Face introduces Nemotron 3 Content Safety 4B, a 4 billion parameter model for advanced content moderation. It supports multimodal inputs like text and images, and handles multiple languages for global applicability. The model is now available on the Hugging Face Hub.
ChatGPT Erotic Mode Risks Kids
OpenAI plans an erotic chat mode for ChatGPT despite opposition from its safety advisers. The feature risks exposing millions of children to adult content. Launch delays continue amid concerns.
Grok Scandal: Systemic CSAM on X Warned
Australian eSafety regulator warned X of 'systemic' child abuse material, more accessible than other platforms. This follows Grok generating sexualised images of women and children. Musk pledged child exploitation removal as priority #1.
Oversight Board Demands Meta AI Rule Overhaul
Meta's Oversight Board urges separate rules for AI-generated content, improved detection tools, and better watermark use after failing to label a deceptive AI video with 700k views. Current 'AI Info' labels are inadequate for AI scale during conflicts. Meta has 60 days to respond.
Microsoft CEO Criticizes Anthropic's Strict Fable Content Controls
Microsoft CEO Satya Nadella publicly criticized Anthropic for imposing overly restrictive content guardrails on their Fable AI model. He argued that these excessive limitations hinder the model's utility as a creative tool.
Meta adds parental alerts for teen AI self-harm discussions
Meta has introduced a new safety feature designed to monitor and alert parents if their teens engage in discussions regarding self-harm with AI models. This update aims to enhance platform safety and provide proactive intervention for vulnerable users.
Google fined €750,000 by AGCOM over gambling content
Italy's communications authority, AGCOM, has fined Google Ireland €750,000 regarding gambling-related videos on YouTube. The ruling centers on whether a platform sharing advertising revenue with creators remains a neutral host.
Ofcom launches investigation into TikTok child safety concerns
The UK regulator Ofcom has initiated a formal investigation into TikTok regarding its child safety practices. This follows a previous review that found the platform's measures for protecting minors to be insufficient.