Benchmarks & Tests4 min read
Benchmark
LLMs can describe data. They cannot reason through it. A new benchmark proves the gap is real.
SDABench, a new capability-oriented benchmark spanning six core scientific reasoning skills and five domains, tests 15 LLMs and finds that models are strong on descriptive analysis but collapse on inferential and causal tasks. The paper provides a five-stage error analysis framework to localize failures.
2026-07-25