LLM benchmark
4 published articles
Financial reasoning benchmark
LLMs know accounting formulas. FinIndices shows they can't apply them
LLMs can recite accounting formulas, but FinIndices, a benchmark built on real Chinese financial statements, shows they struggle to apply them. Gemini-3.1-Pro's accuracy fell from 70.70% to 38.22% when formula hints were removed, and generating tables actively degraded models' reasoning.
2026-08-20
K12-Bench: New Test of AI Curriculum Understanding
Why your AI tutor can't see how math builds on itself
Peking University's K12-Bench reveals that even the best language models barely understand how school concepts connect. Scores of 57% and 46% show a blind spot in AI's ability to handle prerequisite chains, concept taxonomies, and visual grounding, skills that real tutors use every day.
2026-08-02
Benchmark
LLMs can describe data. They cannot reason through it. A new benchmark proves the gap is real.
SDABench, a new capability-oriented benchmark spanning six core scientific reasoning skills and five domains, tests 15 LLMs and finds that models are strong on descriptive analysis but collapse on inferential and causal tasks. The paper provides a five-stage error analysis framework to localize failures.
2026-07-25
LLM Performance
A Chinese video-generation startup just quietly beat Claude Opus at coding
MiniMax's M2.7 scores 56.22% on SWE-Pro, matching near-Claude Opus performance, while touting 97% skill adherence on complex tasks and superior office productivity editing. The model signals a shift from benchmark chasing to real-world agent deployment.
2026-07-14