Agentic AI
21 published articles
AI Agents
Grok Bot goes after the last 10% of work most AI leaves undone
Grok Bot, xAI's new agent product, runs on its own cloud computer, signs into the tools you already use, learns workflows by watching you do them once, and finishes jobs end to end. The pitch behind it: most AI stops at 90% done.
2026-08-14
AI Models | NVIDIA Nemotron 3.5 Lightning
Why Nemotron 3.5 Lightning bets most agent steps don't need a big model
Nvidia's new open model runs locally with 3B active parameters and a 1M-token context, built for agents that stay running. Nvidia claims up to 4x throughput and 30% faster task completion than comparable open models.
2026-08-14
Agentic AI's growing cyber-capability problem
OpenAI paused Astra on fears it can hack hardened systems unaided
OpenAI paused internal work on Astra, its in-development model, after evaluations concluded the company cannot rule out 'critical cyber capabilities' under its Preparedness Framework. The full threshold describes a model that finds zero-day exploits in hardened systems without human intervention.
2026-08-13
Agentic AI
Samsung's newest foldables pitch the agent, not the hinge
Samsung's eighth-generation foldables, the Galaxy Z Fold8 Ultra, Fold8 and Flip8, launched July 22 with agentic AI as their lead story. The hardware legacy is real; the launch material sells the agent, not the hinge.
2026-08-11
Artificial Intelligence
CNIL's blueprint for AI data protection: notes, audits, and alliances
The CNIL is moving from guidance to enforcement with a series of coordinated actions: an exploratory note on agentic AI, a practical audit tool called PANAME, and active participation in EU and international data protection standards.
2026-08-08
Special Report: AI Evaluation
The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them
A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.
2026-08-03
Per-Call Routing's Blind Spot in Agentic AI
For agentic AI, per-call routing is the blind spot. TRACE-Router has a fix.
A new routing framework, TRACE-Router, tackles the mismatch between per-call LLM routers and task-level success metrics in agentic AI, gaining up to 8 accuracy points and 36% lower latency.
2026-08-01
Microsoft AI
Microsoft matches GPT-5.6 in Excel with a cheaper, older-gpu model
Microsoft's MAI model in Excel matches GPT-5.6 for common tasks at lower cost, uses fewer tokens than rivals in Copilot, and runs on H100 and A100 GPUs, lowering the bar for widespread deployment.
2026-07-29
Industrial pivot
Mistral's industrial pivot is bigger than any model release, and the market should pay attention
Mistral is not just another LLM vendor. Its new industrial engineering stack, partnership with major European manufacturers, and an in-house data center signal a deliberate shift from horizontal model competition to secure, vertical AI for critical workflows.
2026-07-19
Platform Strategy
Groq stopped selling speed and started selling an agent operating system
Groq's September 23 launch of Remote MCP support, paired with rapid-fire model additions like Kimi K2-0905 and OpenAI's GPT-OSS series, transforms its API from a fast inference pipe into a full agentic orchestration layer. The platform is no longer competing on tokens per second alone.
2026-07-19
Governed agentic research
Nvidia's AI ran a hospital study on 286,000 patients, and never touched their data
Nvidia's AI Technology Center unveils NAIS, a governed agentic research system that orchestrates end-to-end biomedical workflows on protected hospital data. In a real-world hypertension GWAS deployment involving 286,422 individuals, the system produced results comparable to expert-led analyses while preserving privacy and enabling human oversight.
2026-07-19
benchmark breakdown
The 2.8 trillion parameter model that beats the frontier on the benchmarks that matter
Kimi K3, the 2.8T-parameter open model from Moonshot AI, trails frontier proprietary models on most broad benchmarks, but leads on SWE Marathon, Terminal-Bench 2.1, BrowseComp, and others. The detailed table reveals where its architectural bets on KDA and Stable LatentMoE pay off.
2026-07-17