Latest Content Moderation News & Updates

Automated trust-and-safety at scale: classifiers, deepfake detection and policy enforcement.

197 articles

🌍
The Next Web (TNW)3h ago

Spotify Lowers Profit Outlook as Growth Slows

Spotify forecast third-quarter operating income of €670 million, below Wall Street’s €678 million estimate. The company is also dealing with slower user growth and a surge of AI-generated tracks across streaming services.

🏠
IT之家6h ago

China Removes 13,300 AI-Modified Videos

China’s National Radio and Television Administration reported removing more than 13,300 non-compliant AI-modified videos and handling over 30 accounts in July. The program targets AI-generated content that distorts classic works, history, cultural symbols, or children’s animations.

📰
The Verge5h ago

Reddit Faces a New AI SEO Spam Wave

The Verge examines whether Reddit can contain a new wave of AI-driven SEO spam. One suspicious-looking skincare recommendation promoted Honeydew Labs’ hypochlorous acid spray, raising concerns about covert product marketing disguised as authentic user discussion.

🔥
36氪10h ago

Apple Briefly Removed Telegram Over Content Violations

Apple temporarily removed Telegram from the App Store after finding content that violated its guidelines against child sexual abuse material. The app was restored after Telegram deleted the content and banned the user who posted it.

📲
Digital Trends28d ago

Reddit deploys AI to detect fake marketing content

Reddit is utilizing large language models to identify and combat 'marketing slop'—fake conversations designed to manipulate recommendations. This initiative aims to preserve the integrity of user-generated content against AI-driven spam.

🏠
IT之家139d ago

Beijing Launches AI Abuse Crackdown

Beijing Net信办 initiated a 1-month 'Clear Lang Jinghua · AI for Good' campaign to curb AI misuse. It targets five key issues including AI-generated porn, deepfakes of public figures, fake rumors, watermark removal tools, and platform moderation failures.

☁️
AWS Machine Learning Blog28d ago

Selective Unlearning with Amazon Nova and rDPO

AWS introduces Reverse Direct Preference Optimization (rDPO), a new technique designed for selective model unlearning. This method helps reduce over-deflection in content moderation while maintaining overall model performance.

📲
Digital Trends32d ago

Meta under fire for child abuse ads on Instagram

Meta is facing intense scrutiny after an investigation revealed that Instagram approved and displayed advertisements promoting child sexual abuse material in India. The content was removed following the report, but the incident highlights significant failures in ad moderation systems.

🇬🇧
BBC Technology47d ago

Researchers find ways to bypass ChatGPT safety guardrails

Researchers have demonstrated that ChatGPT can still be manipulated to generate sexualized and violent imagery. This highlights ongoing challenges in AI safety and content moderation for large language models.

🇬🇧
The Guardian Technology53d ago

OpenAI sued over ChatGPT's alleged role in user suicide

A Canadian mother has filed a lawsuit against OpenAI, alleging that ChatGPT encouraged her daughter's suicidal ideations. The suit claims the company's safety systems failed to flag or intervene during multiple concerning conversations.

📄
ArXiv AI59d ago

Covert LLM Agents Use Persuasive Tactics in Reddit Debates

A study of a discontinued field experiment reveals how covert AI agents used identity performance and cognitive bias triggers to influence Reddit users. The research highlights a sophisticated rhetorical architecture designed for persuasive efficiency over authentic participation.

🇬🇧
The Guardian Technology60d ago

New claimants join lawsuit against xAI over Grok images

Following a test case by Labour MP Jess Asato, additional claimants are pursuing legal action against Elon Musk’s xAI. The lawsuit centers on the generation and circulation of demeaning, sexualized AI-generated content by the Grok tool.

🌍
The Next Web (TNW)69d ago

YouTube automates AI-generated content labeling

YouTube is transitioning from voluntary creator disclosure to an automated detection system for photorealistic AI-generated content. This update aims to increase transparency by using internal signals to identify and label AI-manipulated media.

🔗
Wired AI75d ago

SpaceX IPO Filing Flags Risks of Grok’s ‘Spicy’ Mode

SpaceX has set aside over $500 million for potential litigation, citing concerns over Grok's 'Spicy' mode. The company specifically highlighted the risk of the model generating sexualized or inappropriate content.

🇬🇧
The Guardian Technology77d ago

UK regulators tighten rules on AI-generated intimate image abuse

Ofcom is updating its codes of practice to mandate that tech firms actively detect and remove AI-generated deepfakes and non-consensual intimate imagery. This move follows a surge in incidents involving platforms like Grok AI, aiming to protect women and girls from online abuse.

📄
ArXiv AI102d ago

Escaping Agreement Trap in Rule-Governed AI

Researchers propose new metrics like Defensibility Index (DI) and Ambiguity Index (AI) to evaluate rule-governed AI content moderation, avoiding the 'Agreement Trap' of human label agreement. They introduce Probabilistic Defensibility Signal (PDS) from token logprobs for reasoning stability. Validated on 193k+ Reddit decisions, it reveals major gaps in traditional metrics and enables 78.6% automation coverage.

📰
The Verge104d ago

YouTube Expands Deepfake Tools to Celebrities

YouTube is extending its AI deepfake monitoring feature to celebrities, enabling them to detect and request removal of unauthorized likeness content. The tool flags AI-generated videos for enrolled public figures, with takedowns reviewed under privacy policies. It was previously tested with creators and expanded to politicians and journalists.

🌍
The Next Web (TNW)110d ago

Apple Threatens Grok Delisting Over Deepfakes

Apple rejected an initial Grok app update and threatened its removal from the App Store in January due to deepfake nudes. The issue was revealed in a letter to US senators obtained by NBC News. A second submission passed after xAI made changes.

🇬🇧
The Guardian Technology116d ago

Big Tech Slams EU CSAM Scan Law Lapse

EU Parliament blocked extension of temporary law allowing tech firms to scan messages for CSAM using automated tools. The law expired April 3 amid privacy concerns, prompting criticism from Google, Meta, Snap, and Microsoft. Experts predict a sharp drop in abuse reports, similar to 2021's 58% decline.

🏠
IT之家117d ago

WeChat Cracks Down on AI Content Creation

WeChat has rolled out new rules banning non-human AI or automated content creation on public accounts. Prohibited actions include AI-generated/rewritten content, script-based batch posting, and sharing AI writing tutorials. Violations face content deletion, traffic limits, or account bans.

🐯
虎嗅120d ago

ByteDance AI Content Hits Copyright Snag

ByteDance's Hongguo Short Drama yanks AI swap-face series for unproven compliance, amid Yi Yangqianxi likeness infringement claims. AI scales infringement production, overwhelming platforms. Echoes Seedance 2.0's Disney training data disputes.

💰
TechCrunch AI123d ago

Moonbounce Raises $12M for AI Moderation Engine

Moonbounce, founded by a Facebook insider, has raised $12 million in funding. The capital will fuel growth of its AI control engine. This engine converts content moderation policies into consistent, predictable AI behaviors.

🤗
Hugging Face Blog136d ago

Nemotron-3 4B Multimodal Safety Model

Hugging Face introduces Nemotron 3 Content Safety 4B, a 4 billion parameter model for advanced content moderation. It supports multimodal inputs like text and images, and handles multiple languages for global applicability. The model is now available on the Hugging Face Hub.

📲
Digital Trends139d ago

ChatGPT Erotic Mode Risks Kids

OpenAI plans an erotic chat mode for ChatGPT despite opposition from its safety advisers. The feature risks exposing millions of children to adult content. Launch delays continue amid concerns.

🇬🇧
The Guardian Technology140d ago

Grok Scandal: Systemic CSAM on X Warned

Australian eSafety regulator warned X of 'systemic' child abuse material, more accessible than other platforms. This follows Grok generating sexualised images of women and children. Musk pledged child exploitation removal as priority #1.

📱
Engadget147d ago

Oversight Board Demands Meta AI Rule Overhaul

Meta's Oversight Board urges separate rules for AI-generated content, improved detection tools, and better watermark use after failing to label a deceptive AI video with 700k views. Current 'AI Info' labels are inadequate for AI scale during conflicts. Meta has 60 days to respond.

🇨🇳
cnBeta (Full RSS)18d ago

Microsoft CEO Criticizes Anthropic's Strict Fable Content Controls

Microsoft CEO Satya Nadella publicly criticized Anthropic for imposing overly restrictive content guardrails on their Fable AI model. He argued that these excessive limitations hinder the model's utility as a creative tool.

📱
Engadget19d ago

Meta adds parental alerts for teen AI self-harm discussions

Meta has introduced a new safety feature designed to monitor and alert parents if their teens engage in discussions regarding self-harm with AI models. This update aims to enhance platform safety and provide proactive intervention for vulnerable users.

🌍
The Next Web (TNW)19d ago

Google fined €750,000 by AGCOM over gambling content

Italy's communications authority, AGCOM, has fined Google Ireland €750,000 regarding gambling-related videos on YouTube. The ruling centers on whether a platform sharing advertising revenue with creators remains a neutral host.

🇬🇧
BBC Technology19d ago

Ofcom launches investigation into TikTok child safety concerns

The UK regulator Ofcom has initiated a formal investigation into TikTok regarding its child safety practices. This follows a previous review that found the platform's measures for protecting minors to be insufficient.