AI4 min read
Interpretability
Your AI thinks it explained itself. A new test proves it was lying the whole time.
Natural-language autoencoders score explanations by reconstruction, but a new study shows this test is passed even when 98% of specific claims are false. The authors introduce RECAP, a training method that lets probes independently verify designated content, beating adversarial lie detection with an AUC of 0.95 versus chance-level 0.51 for standard methods.
2026-07-24