Lab
Research, experimentation and open source: hardware, IoT, robotics, biotech and edge AI.
47 published articles
Benchmark
LLMs can describe data. They cannot reason through it. A new benchmark proves the gap is real.
SDABench, a new capability-oriented benchmark spanning six core scientific reasoning skills and five domains, tests 15 LLMs and finds that models are strong on descriptive analysis but collapse on inferential and causal tasks. The paper provides a five-stage error analysis framework to localize failures.
2026-07-25
Benchmark & Tests
Your graph model breaks on dirty data. A new benchmark shows exactly where.
OpenRTAG, a benchmark from an academic team, organizes TAG quality issues into a 3×3 taxonomy covering text, structure, and label degradation. It tests traditional GNNs, large language model-enhanced GNNs, and graph foundation models across nine datasets, revealing different sensitivity patterns that prior fragmented studies missed.
2026-07-24
Agent evaluation
Your AI agent keeps failing? It might be the harness, not the brain
PawBench, an open-source benchmark from the AgentScope team, systematically evaluates models and agent harnesses together. Results show that harness design can swing scores by over 11 points for smaller models, exposing a blind spot in how AI agents are currently judged.
2026-07-22
Sakana AI's collective intelligence finally goes physical
These identical bricks taught themselves to recognize their own shape. No central brain needed.
A stack of identical cubes, each running the same tiny neural network and talking only to its immediate neighbors, can figure out their global shape, detect damage, and guide regrowth. No single brick knows where it is. The system hit 100 percent accuracy in hardware, published in Nature Communications.
2026-07-21
Agentic coding
Grok 4.5 just broke the coding agent leaderboard: the lead is real, the margins are tiny
Grok 4.5 now leads the SWE Marathon leaderboard, beating Claude 4 Opus and GPT-5. The benchmark tests real software engineering skills: bug fixes, feature additions, and code understanding across real repositories. The margin is slim, but the trend lines point toward a shrinking gap between what agents can do and what they need to do.
2026-07-20
AI Research
BMW's 16 GB GPU just did what needs an A100
BMW researchers show that Hierarchical Global Attention, paired with truncated backprop and external KV storage, lets a 16 GB GPU train on 16K tokens, four times the limit of dense attention, with no measurable loss in adapter quality.
2026-07-20
NVIDIA
Nvidia and Hugging Face just made distributed diffusion training boring (that's the point)
Nvidia's NeMo Automodel now integrates directly with Hugging Face Diffusers, enabling production-grade distributed training for models like FLUX, Wan 2.1, and HunyuanVideo. The Apache 2.0 library handles parallelism as a config toggle and lets fine-tuned checkpoints load straight back into inference pipelines.
2026-07-19
AI & Robotics
Nvidia's robot model just unlocked a three-order-of-magnitude advantage: time
NVIDIA Research's RoboTTT scales robot model context to 8K timesteps, three orders of magnitude beyond current policies, unlocking one-shot imitation from human video and 87% performance gains, suggesting context length as a new scaling axis for robot foundation models.
2026-07-17
Hardware & Electronics
A 40nm chip just made Nvidia's A100 look 478 times slower
A joint team from Peking University and the Chinese Academy of Sciences has built a neuromorphic chip that processes neural dynamics 50 to 478 times faster than an NVIDIA A100 GPU while using a fraction of the power. The 40nm phase-change memristor chip achieves millisecond-level real-time neural dynamics, opening the door to surgical navigation and brain digital twins.
2026-07-16
Market momentum
Nvidia just crossed $3 trillion again. The AI spending machine isn't slowing down.
Nvidia's market cap crossed $3 trillion as AI infrastructure demand keeps growing. The milestone underscores that enterprise GPU purchases remain strong despite growing competition from AMD and startups offering cheaper inference chips.
2026-07-16
Machine Learning Theory
Google just proved why diffusion models invent, not just copy
Google researchers reveal that the creativity of diffusion models stems from a 'score smoothing' effect caused by neural network regularization. This theoretical framework explains why models interpolate between training data points rather than merely memorizing them, opening the path for controlled novelty in generative AI.
2026-07-16
Humanoid robotics
Xiaomi's robot did what Musk said Optimus couldn't: real factory work in six months
Xiaomi's humanoid robot achieved 90%+ success on complex flexible workpiece tasks after six months of factory deployment, challenging Elon Musk's January claim that Optimus isn't ready for useful factory work. The leap from single-station screw insertion to bimanual panel sorting and box folding sets a new pace for humanoid industrialization.
2026-07-15