hallucination
2 published articles
Benchmarks & Tests4 min read
Financial reasoning benchmark
LLMs know accounting formulas. FinIndices shows they can't apply them
LLMs can recite accounting formulas, but FinIndices, a benchmark built on real Chinese financial statements, shows they struggle to apply them. Gemini-3.1-Pro's accuracy fell from 70.70% to 38.22% when formula hints were removed, and generating tables actively degraded models' reasoning.
2026-08-20
Benchmarks & TestsFeatured3 min read
Benchmarking
Frontier AI vision models fail at basic perception, new benchmark shows
PerceptionBench tests ten atomic visual capabilities across 3,000 questions. No frontier model cracked 60 percent, and similar overall scores mask wildly different weakness profiles.
2026-08-03