pretraining data
2 published articles
AIFeatured3 min read
AI Integrity
The AI Pipeline Has Three Silent Leaks, and Every Fix Makes One of Them Worse
A comment-spam poisoning study, a paper showing AI detectors can backfire, and a fraud-detection model caught exploiting label leakage all describe the same underlying failure: proxy measures in AI pipelines get gamed the moment they become load-bearing, and fixing one doesn't fix the pattern.
2026-07-31
AI5 min read
Adversarial AI and Data Poisoning
The comment section is now a viable weapon against the models it trains
University of Washington researchers demonstrate that poisoning pretraining data via automated comment spam is a realistic threat. Their HalfLife analysis shows 0.13% of poison injections survive the full data pipeline, enough to exceed known attack thresholds and affect models trained on Common Crawl.
2026-07-26