SevenTnewS

AI

Artificial intelligence: LLMs, agents, diffusion, vision, NLP and the latest from top labs.

565 published articles

4 min read

AI Infrastructure & Policy

Groq bets its future on the American AI Stack

Groq endorsed the White House's AI Action Plan the day it launched, naming itself the compute layer of the American AI Stack. The release reveals a company betting that national policy and cheap domestic inference will carry it into the global market.

2026-08-08

3 min read

Inference Infrastructure

AI inference is becoming a location problem: Groq opens a UK data center

Groq opened a UK data center with Equinix as GroqCloud tops 3.5 million developers and production AI traffic keeps climbing. The move is a bet that inference is now a geography problem, with latency and locality deciding where real-time AI can run.

2026-08-08

5 min read

Video generation

MiniMax scrapped its proven architecture to make H3 do everything

MiniMax says H3 unifies text, image, video and audio generation in one model, prices 2K output below a third of mainstream models, and plans to open the weights within days. The small print is the real story: the company abandoned the architecture that gave it an edge to get there.

2026-08-07

4 min read

Computer Vision

CABiNet stays within 2 points of YOLO26x at an eighth of the compute

The VDD Semantic Segmentation Model Zoo brings YOLO26 and CABiNet models trained on varied drone footage to Hugging Face. YOLO26x-sem leads at 78.83% mIoU, but CABiNet-Large's 77.76% at 54.8 GFLOPs makes the efficiency case.

2026-08-07

Featured3 min read

AI Safety

Why Microsoft is opening AI safety testing to the world

Microsoft has funded 18 university labs across six continents to independently red team AI systems, acknowledging that internal teams cannot catch every risk.

2026-08-07

3 min read

RAG research

BM25 beats agentic RAG when corpora pass 10 million tokens

Agentic search wins on small corpora, but a muset-ai study across 28 nested tiers shows BM25 overtaking it near 10 million tokens. The agent burns 39x more query tokens, and graph RAG stalls in construction.

2026-08-07

4 min read

Generative AI

Pippa pays artists $0.005 per image. Its models still train on scraped art

Pippa promises artists a cut of every generation made in their style: $0.005 per image, $0.003 per video second, plus a share of a 5 percent royalty pool. With around 800 subscribers and four signed artists, and base models still trained on scraped content, the ethical pitch has a math problem.

2026-08-07

5 min read

Agentic UI

Qoder Canvas: a design system built for agents, not humans

Alibaba's Qoder team argues the chat window is the wrong container for complex agent output. Qoder Canvas applies design-system thinking to agent interfaces, teaching agents to build interactive, codebase-aware artifacts instead of walls of Markdown.

2026-08-07

4 min read

Qwen / Alibaba Cloud

Qwen2.5-Omni outperforms Gemini-1.5-Pro on OmniBench, fits under 12GB

Alibaba's open-source Qwen2.5-Omni outscored Gemini-1.5-Pro on OmniBench and topped the MMAU audio reasoning leaderboard. Quantized builds cut VRAM below 12GB and MNN support brings real-time voice chat to phones.

2026-08-07

Featured1 min read

AI Coding

OpenCode bets token efficiency wins, adding Ling 3.0 Flash for free

OpenCode now offers inclusionAI's Ling 3.0 Flash for free, days after release. The move extends a cost-conscious strategy betting that token efficiency wins the AI coding race.

2026-08-07

Featured3 min read

AI Video Generation

Wan3.0-Video charges by the second: a 30-second clip runs to $6

Alibaba's Wan3.0-Video bills per second on DashScope, with a 30-second 1080P clip costing $6 per generation. We break down the pricing tiers, the multi-input workflow, and the questions the listing leaves unanswered.

2026-08-07

4 min read

AI Research

Voice Memory: a 776-byte file that tells speech recognition when to do nothing

A new inference-only scheme for speech recognition learns restraint: a frozen corrector reads a per-domain memory file and decides when to abstain. Unconstrained correction breaks correct tokens on up to 64% of edits; Voice Memory cuts that to 35% and lowers weighted WER from 8.36% to 7.52%.

2026-08-06

← PreviousPage 6 / 48 · 565 articlesNext →