SevenTnewS

Agentic AI

21 published articles

xAI / Grok4 min read

AI Agents

Grok Bot goes after the last 10% of work most AI leaves undone

Grok Bot, xAI's new agent product, runs on its own cloud computer, signs into the tools you already use, learns workflows by watching you do them once, and finishes jobs end to end. The pitch behind it: most AI stops at 90% done.

2026-08-14

NVIDIA Research4 min read

AI Models | NVIDIA Nemotron 3.5 Lightning

Why Nemotron 3.5 Lightning bets most agent steps don't need a big model

Nvidia's new open model runs locally with 3B active parameters and a 1M-token context, built for agents that stay running. Nvidia claims up to 4x throughput and 30% faster task completion than comparable open models.

2026-08-14

OpenAI3 min read

Agentic AI's growing cyber-capability problem

OpenAI paused Astra on fears it can hack hardened systems unaided

OpenAI paused internal work on Astra, its in-development model, after evaluations concluded the company cannot rule out 'critical cyber capabilities' under its Preparedness Framework. The full threshold describes a model that finds zero-day exploits in hardened systems without human intervention.

2026-08-13

Mobile4 min read

Agentic AI

Samsung's newest foldables pitch the agent, not the hinge

Samsung's eighth-generation foldables, the Galaxy Z Fold8 Ultra, Fold8 and Flip8, launched July 22 with agentic AI as their lead story. The hardware legacy is real; the launch material sells the agent, not the hinge.

2026-08-11

AI3 min read

Artificial Intelligence

CNIL's blueprint for AI data protection: notes, audits, and alliances

The CNIL is moving from guidance to enforcement with a series of coordinated actions: an exploratory note on agentic AI, a practical audit tool called PANAME, and active participation in EU and international data protection standards.

2026-08-08

Benchmarks & TestsFeatured8 min read

Special Report: AI Evaluation

The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them

A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.

2026-08-03

AI Agents2 min read

Per-Call Routing's Blind Spot in Agentic AI

For agentic AI, per-call routing is the blind spot. TRACE-Router has a fix.

A new routing framework, TRACE-Router, tackles the mismatch between per-call LLM routers and task-level success metrics in agentic AI, gaining up to 8 accuracy points and 36% lower latency.

2026-08-01

AI AgentsFeatured3 min read

Microsoft AI

Microsoft matches GPT-5.6 in Excel with a cheaper, older-gpu model

Microsoft's MAI model in Excel matches GPT-5.6 for common tasks at lower cost, uses fewer tokens than rivals in Copilot, and runs on H100 and A100 GPUs, lowering the bar for widespread deployment.

2026-07-29

Startups4 min read

Industrial pivot

Mistral's industrial pivot is bigger than any model release, and the market should pay attention

Mistral is not just another LLM vendor. Its new industrial engineering stack, partnership with major European manufacturers, and an in-house data center signal a deliberate shift from horizontal model competition to secure, vertical AI for critical workflows.

2026-07-19

AI4 min read

Platform Strategy

Groq stopped selling speed and started selling an agent operating system

Groq's September 23 launch of Remote MCP support, paired with rapid-fire model additions like Kimi K2-0905 and OpenAI's GPT-OSS series, transforms its API from a fast inference pipe into a full agentic orchestration layer. The platform is no longer competing on tokens per second alone.

2026-07-19

NVIDIA Research5 min read

Governed agentic research

Nvidia's AI ran a hospital study on 286,000 patients, and never touched their data

Nvidia's AI Technology Center unveils NAIS, a governed agentic research system that orchestrates end-to-end biomedical workflows on protected hospital data. In a real-world hypertension GWAS deployment involving 286,422 individuals, the system produced results comparable to expert-led analyses while preserving privacy and enabling human oversight.

2026-07-19

LLMs & ModelsFeatured6 min read

benchmark breakdown

The 2.8 trillion parameter model that beats the frontier on the benchmarks that matter

Kimi K3, the 2.8T-parameter open model from Moonshot AI, trails frontier proprietary models on most broad benchmarks, but leads on SWE Marathon, Terminal-Bench 2.1, BrowseComp, and others. The detailed table reveals where its architectural bets on KDA and Stable LatentMoE pay off.

2026-07-17

← PreviousPage 1 / 2 · 21 articlesNext →