NVIDIA
20 published articles
AI Models | NVIDIA Nemotron 3.5 Lightning
Why Nemotron 3.5 Lightning bets most agent steps don't need a big model
Nvidia's new open model runs locally with 3B active parameters and a 1M-token context, built for agents that stay running. Nvidia claims up to 4x throughput and 30% faster task completion than comparable open models.
2026-08-14
AI & Robotics
Training surgical robots just got faster: Nvidia's Cosmos-H-Dreams runs at 160 fps
Nvidia's Cosmos-H-Dreams turns surgical simulation into an interactive, real-time experience. Running at 160 fps on a single GPU, it lets surgeons, trainees, and AI policies practice inside a closed loop of generated video without needing a physical robot.
2026-08-01
AI Research
Video generation gets 120x faster on one GPU, inside Nvidia's hybrid attention breakthrough
Nvidia Research's SANA-Video 2.0 combines linear and periodic softmax attention in a 3:1 ratio to run 120x faster than Wan 2.2 on one H100. The hybrid design recovers full-rank expressiveness while keeping inference at 13 seconds for a 720p clip. The caveats: the benchmarks stop at 720p and bundle the attention change with proprietary optimizations.
2026-07-31
Open Source AI
Nvidia's AV-Flamingo solves the one thing every other video AI gets wrong: time
Nvidia releases AV-Flamingo, an open audio-visual large language model designed for long, complex video understanding. It uses a three-stage curriculum and a timestamped chain-of-thought framework to handle temporal and cross-modal reasoning that trips up most models.
2026-07-28
Physical AI
Nvidia's 4B model does what most small robots can't: act, not just talk
Nvidia's Cosmos 3 Edge packs world modeling into 4B parameters, ranking first on VANTAGE-Bench among its size class. Its dual-transformer design blends reasoning with real-time action prediction for robots, all deployable on Jetson and RTX hardware.
2026-07-28
Infrastructure & Funding
Together AI's $800 million bet on open inference math
Together AI raised $800 million at an $8.3 billion valuation, backed by Aramco Ventures, NVIDIA, and others. The company claims annual bookings topped $1.15 billion last quarter, fueled by customers like Decagon, Eleven Labs, and Cursor who report 6x to 20x cost reductions switching to open inference.
2026-07-27
Open-weight reasoning
Nvidia's Nemotron-4 runs 30% cheaper than GPT-4. The catch is buried in the benchmark.
Nvidia's Nemotron-4 line challenges Llama 3.3 and Mistral Large with competitive MMLU scores and a claimed 30% inference savings. The open license and dual-size strategy position it as a practical alternative for teams running agentic workloads at scale, pending third-party verification.
2026-07-25
Multi-agent orchestration
Nemotron meets Fugu: Sakana AI's bet that open models win as a swarm, not alone
Sakana AI integrates NVIDIA's Nemotron open model family into the Fugu multi-agent orchestration system. Fugu dynamically selects and combines specialized models. The collaboration aims to show that coordinated open models can match monolithic frontier systems.
2026-07-23
Governed agentic research
Nvidia's AI ran a hospital study on 286,000 patients, and never touched their data
Nvidia's AI Technology Center unveils NAIS, a governed agentic research system that orchestrates end-to-end biomedical workflows on protected hospital data. In a real-world hypertension GWAS deployment involving 286,422 individuals, the system produced results comparable to expert-led analyses while preserving privacy and enabling human oversight.
2026-07-19
AI Models & Infrastructure
Nvidia just proved that better embeddings pay for themselves in agent runtime
Nvidia's Nemotron 3 Embed collection claims the #1 spot on the RTEB benchmark and introduces 1B variants that retain 99% of the 8B model's accuracy. Company data shows stronger retrieval reduces downstream token costs in agentic systems.
2026-07-18
Market momentum
Nvidia just crossed $3 trillion again. The AI spending machine isn't slowing down.
Nvidia's market cap crossed $3 trillion as AI infrastructure demand keeps growing. The milestone underscores that enterprise GPU purchases remain strong despite growing competition from AMD and startups offering cheaper inference chips.
2026-07-16
special report / edge ai
Nvidia just cracked open DeepStream. Your edge AI project will never be the same.
The full source code of Nvidia's DeepStream video analytics SDK is now on GitHub under Apache 2.0 and CC-BY-4.0, opening edge AI development to a wider audience. Version 9.1 brings LLM-based coding agents, Triton Inference Server integration, and consolidated repositories for end-to-end pipelines.
2026-07-16