arXiv
18 published articles
Computer Vision, World Models & AI Research
PhiZero: teaching video AI to think in physics before it renders
PhiZero, a CASIA world model, learns a compact discrete "physical language" from raw video and uses it to reason about how a scene will evolve before rendering frames. The authors argue this reason-then-render design produces more physically coherent video than direct pixel prediction.
2026-07-31
AI Research
VideoCoCo fixes AI video's broken physics by thinking in Blender code
VideoCoCo treats executable Blender code as a chain of thought: a coding agent scripts a scene, a simulator plays it out, and a video engine makes the result photorealistic. The split targets text-to-video's physics problem and posts best average scores on PhyGenBench and VBench-2.0.
2026-07-30
Artificial Intelligence
Better AI detectors might make people use AI more, not less
Imperfect LLM detectors can distort user incentives, leading to more AI use and lower quality outputs. The paper's findings challenge the naive assumption that detection tools cleanly reduce machine-generated content.
2026-07-26
Benchmark
LLMs can describe data. They cannot reason through it. A new benchmark proves the gap is real.
SDABench, a new capability-oriented benchmark spanning six core scientific reasoning skills and five domains, tests 15 LLMs and finds that models are strong on descriptive analysis but collapse on inferential and causal tasks. The paper provides a five-stage error analysis framework to localize failures.
2026-07-25
AI Research
Four minds, one answer: why AI that thinks differently beat the biggest models at humanity's hardest test
PoTRE breaks inference into four agents working in parallel, then reconciles their answers dynamically. It hit 49.92% on Humanity's Last Exam, beating the previous official best. The approach works with fewer tokens than scaled-up baselines.
2026-07-24
AI Research
GPT-5.5's big win reveals something missing from every agent benchmark
EvoPolicyGym isolates a critical but understudied capability: an agent's ability to refine an executable policy through repeated feedback-constrained edits. The benchmark reveals GPT-5.5 as the strongest performer across 16 environments, and provides trajectory-level diagnostics that expose how different agents allocate budget and convert feedback into tuned parameters.
2026-07-11