SevenTnewS

arXiv

18 published articles

Vision & Diffusion4 min read

Computer Vision, World Models & AI Research

PhiZero: teaching video AI to think in physics before it renders

PhiZero, a CASIA world model, learns a compact discrete "physical language" from raw video and uses it to reason about how a scene will evolve before rendering frames. The authors argue this reason-then-render design produces more physically coherent video than direct pixel prediction.

2026-07-31

Vision & Diffusion4 min read

AI Research

VideoCoCo fixes AI video's broken physics by thinking in Blender code

VideoCoCo treats executable Blender code as a chain of thought: a coding agent scripts a scene, a simulator plays it out, and a video engine makes the result photorealistic. The split targets text-to-video's physics problem and posts best average scores on PhyGenBench and VBench-2.0.

2026-07-30

AI3 min read

Artificial Intelligence

Better AI detectors might make people use AI more, not less

Imperfect LLM detectors can distort user incentives, leading to more AI use and lower quality outputs. The paper's findings challenge the naive assumption that detection tools cleanly reduce machine-generated content.

2026-07-26

Benchmarks & Tests4 min read

Benchmark

LLMs can describe data. They cannot reason through it. A new benchmark proves the gap is real.

SDABench, a new capability-oriented benchmark spanning six core scientific reasoning skills and five domains, tests 15 LLMs and finds that models are strong on descriptive analysis but collapse on inferential and causal tasks. The paper provides a five-stage error analysis framework to localize failures.

2026-07-25

LLMs & Models4 min read

AI Research

Four minds, one answer: why AI that thinks differently beat the biggest models at humanity's hardest test

PoTRE breaks inference into four agents working in parallel, then reconciles their answers dynamically. It hit 49.92% on Humanity's Last Exam, beating the previous official best. The approach works with fewer tokens than scaled-up baselines.

2026-07-24

NLP & MLFeatured4 min read

AI Research

GPT-5.5's big win reveals something missing from every agent benchmark

EvoPolicyGym isolates a critical but understudied capability: an agent's ability to refine an executable policy through repeated feedback-constrained edits. The benchmark reveals GPT-5.5 as the strongest performer across 16 environments, and provides trajectory-level diagnostics that expose how different agents allocate budget and convert feedback into tuned parameters.

2026-07-11

← PreviousPage 2 / 2 · 18 articlesNext →