SevenTnewS

LLM benchmark

4 published articles

Benchmarks & Tests4 min read

Financial reasoning benchmark

LLMs know accounting formulas. FinIndices shows they can't apply them

LLMs can recite accounting formulas, but FinIndices, a benchmark built on real Chinese financial statements, shows they struggle to apply them. Gemini-3.1-Pro's accuracy fell from 70.70% to 38.22% when formula hints were removed, and generating tables actively degraded models' reasoning.

2026-08-20

LLMs & Models2 min read

K12-Bench: New Test of AI Curriculum Understanding

Why your AI tutor can't see how math builds on itself

Peking University's K12-Bench reveals that even the best language models barely understand how school concepts connect. Scores of 57% and 46% show a blind spot in AI's ability to handle prerequisite chains, concept taxonomies, and visual grounding, skills that real tutors use every day.

2026-08-02

Benchmarks & Tests4 min read

Benchmark

LLMs can describe data. They cannot reason through it. A new benchmark proves the gap is real.

SDABench, a new capability-oriented benchmark spanning six core scientific reasoning skills and five domains, tests 15 LLMs and finds that models are strong on descriptive analysis but collapse on inferential and causal tasks. The paper provides a five-stage error analysis framework to localize failures.

2026-07-25

LLMs & Models4 min read

LLM Performance

A Chinese video-generation startup just quietly beat Claude Opus at coding

MiniMax's M2.7 scores 56.22% on SWE-Pro, matching near-Claude Opus performance, while touting 97% skill adherence on complex tasks and superior office productivity editing. The model signals a shift from benchmark chasing to real-world agent deployment.

2026-07-14