SevenTnewS

Lab

Research, experimentation and open source: hardware, IoT, robotics, biotech and edge AI.

47 published articles

4 min read

Benchmark

LLMs can describe data. They cannot reason through it. A new benchmark proves the gap is real.

SDABench, a new capability-oriented benchmark spanning six core scientific reasoning skills and five domains, tests 15 LLMs and finds that models are strong on descriptive analysis but collapse on inferential and causal tasks. The paper provides a five-stage error analysis framework to localize failures.

2026-07-25

3 min read

Benchmark & Tests

Your graph model breaks on dirty data. A new benchmark shows exactly where.

OpenRTAG, a benchmark from an academic team, organizes TAG quality issues into a 3×3 taxonomy covering text, structure, and label degradation. It tests traditional GNNs, large language model-enhanced GNNs, and graph foundation models across nine datasets, revealing different sensitivity patterns that prior fragmented studies missed.

2026-07-24

2 min read

Agent evaluation

Your AI agent keeps failing? It might be the harness, not the brain

PawBench, an open-source benchmark from the AgentScope team, systematically evaluates models and agent harnesses together. Results show that harness design can swing scores by over 11 points for smaller models, exposing a blind spot in how AI agents are currently judged.

2026-07-22

Featured3 min read

Sakana AI's collective intelligence finally goes physical

These identical bricks taught themselves to recognize their own shape. No central brain needed.

A stack of identical cubes, each running the same tiny neural network and talking only to its immediate neighbors, can figure out their global shape, detect damage, and guide regrowth. No single brick knows where it is. The system hit 100 percent accuracy in hardware, published in Nature Communications.

2026-07-21

Featured3 min read

Agentic coding

Grok 4.5 just broke the coding agent leaderboard: the lead is real, the margins are tiny

Grok 4.5 now leads the SWE Marathon leaderboard, beating Claude 4 Opus and GPT-5. The benchmark tests real software engineering skills: bug fixes, feature additions, and code understanding across real repositories. The margin is slim, but the trend lines point toward a shrinking gap between what agents can do and what they need to do.

2026-07-20

Featured6 min read

AI Research

BMW's 16 GB GPU just did what needs an A100

BMW researchers show that Hierarchical Global Attention, paired with truncated backprop and external KV storage, lets a 16 GB GPU train on 16K tokens, four times the limit of dense attention, with no measurable loss in adapter quality.

2026-07-20

Featured3 min read

NVIDIA

Nvidia and Hugging Face just made distributed diffusion training boring (that's the point)

Nvidia's NeMo Automodel now integrates directly with Hugging Face Diffusers, enabling production-grade distributed training for models like FLUX, Wan 2.1, and HunyuanVideo. The Apache 2.0 library handles parallelism as a config toggle and lets fine-tuned checkpoints load straight back into inference pipelines.

2026-07-19

6 min read

AI & Robotics

Nvidia's robot model just unlocked a three-order-of-magnitude advantage: time

NVIDIA Research's RoboTTT scales robot model context to 8K timesteps, three orders of magnitude beyond current policies, unlocking one-shot imitation from human video and 87% performance gains, suggesting context length as a new scaling axis for robot foundation models.

2026-07-17

3 min read

Hardware & Electronics

A 40nm chip just made Nvidia's A100 look 478 times slower

A joint team from Peking University and the Chinese Academy of Sciences has built a neuromorphic chip that processes neural dynamics 50 to 478 times faster than an NVIDIA A100 GPU while using a fraction of the power. The 40nm phase-change memristor chip achieves millisecond-level real-time neural dynamics, opening the door to surgical navigation and brain digital twins.

2026-07-16

Featured3 min read

Market momentum

Nvidia just crossed $3 trillion again. The AI spending machine isn't slowing down.

Nvidia's market cap crossed $3 trillion as AI infrastructure demand keeps growing. The milestone underscores that enterprise GPU purchases remain strong despite growing competition from AMD and startups offering cheaper inference chips.

2026-07-16

5 min read

Machine Learning Theory

Google just proved why diffusion models invent, not just copy

Google researchers reveal that the creativity of diffusion models stems from a 'score smoothing' effect caused by neural network regularization. This theoretical framework explains why models interpolate between training data points rather than merely memorizing them, opening the path for controlled novelty in generative AI.

2026-07-16

5 min read

Humanoid robotics

Xiaomi's robot did what Musk said Optimus couldn't: real factory work in six months

Xiaomi's humanoid robot achieved 90%+ success on complex flexible workpiece tasks after six months of factory deployment, challenging Elon Musk's January claim that Optimus isn't ready for useful factory work. The leap from single-station screw insertion to bimanual panel sorting and box folding sets a new pace for humanoid industrialization.

2026-07-15

← PreviousPage 3 / 4 · 47 articlesNext →