SevenTnewS

content moderation

4 published articles

Mistral AIFeatured4 min read

AI Safety: 3B Classifier, Apache 2.0 Weights

Mistral's Shieldstral puts your moderation policy in the prompt, not the weights

Shieldstral frames moderation as a binary question: an instruction, a yes/no query, and the content to judge. Mistral says the 3B model matches open guardrails up to seven times its size on text safety, with Apache 2.0 weights that run on one 16GB GPU.

2026-08-05

LLMs & ModelsFeatured4 min read

AI Safety

Shieldstral, the 3B classifier that outguns models seven times its size

A 3B safety classifier matches text models nearly seven times its size and sets a new multimodal moderation state of the art, per a July 2026 arXiv paper. The trick: moderation reframed as binary question answering, trained on roughly 54.1 million samples.

2026-08-05

AI3 min read

Integrity Report

Stability AI's first transparency report counted 13 CSAM cases. The real story is what it doesn't say.

Stability AI's first transparency report counts 13 CSAM reports, details its safety stack, and reveals gaps in content provenance for open models. The story is in the limitations.

2026-07-28

AIFeatured2 min read

AI governance

China's tech giants just deleted millions of AI relationships

ByteDance, Tencent, and Alibaba are simultaneously shutting down AI companion features that allowed millions of users to form emotional bonds with customizable virtual characters. The move, widely seen as a response to China's new AI content regulations taking effect in April 2026, has triggered user outrage and complaints about losing irreplaceable emotional investments.

2026-07-08