content moderation
4 published articles
AI Safety: 3B Classifier, Apache 2.0 Weights
Mistral's Shieldstral puts your moderation policy in the prompt, not the weights
Shieldstral frames moderation as a binary question: an instruction, a yes/no query, and the content to judge. Mistral says the 3B model matches open guardrails up to seven times its size on text safety, with Apache 2.0 weights that run on one 16GB GPU.
2026-08-05
AI Safety
Shieldstral, the 3B classifier that outguns models seven times its size
A 3B safety classifier matches text models nearly seven times its size and sets a new multimodal moderation state of the art, per a July 2026 arXiv paper. The trick: moderation reframed as binary question answering, trained on roughly 54.1 million samples.
2026-08-05
Integrity Report
Stability AI's first transparency report counted 13 CSAM cases. The real story is what it doesn't say.
Stability AI's first transparency report counts 13 CSAM reports, details its safety stack, and reveals gaps in content provenance for open models. The story is in the limitations.
2026-07-28
AI governance
China's tech giants just deleted millions of AI relationships
ByteDance, Tencent, and Alibaba are simultaneously shutting down AI companion features that allowed millions of users to form emotional bonds with customizable virtual characters. The move, widely seen as a response to China's new AI content regulations taking effect in April 2026, has triggered user outrage and complaints about losing irreplaceable emotional investments.
2026-07-08