distributed training
2 published articles
LabFeatured5 min read
NVIDIA
Nvidia and Hugging Face just made distributed diffusion training boring (that's the point)
Nvidia's NeMo Automodel now integrates directly with Hugging Face Diffusers, enabling production-grade distributed training for models like FLUX, Wan 2.1, and HunyuanVideo. The Apache 2.0 library handles parallelism as a config toggle and lets fine-tuned checkpoints load straight back into inference pipelines.
2026-07-19
Tools & Frameworks2 min read
Infrastructure deep-dive
Nous Research's MoE field notes: what actually happens when 1 trillion parameters hit 1024 GPUs
Nous Research publishes field notes on scaling MoE expert parallelism with DeepEP. The report details throughput, configuration trade-offs, and bottlenecks from pretraining a 1T-parameter MoE model, offering practical deployment insights for large-scale distributed training.
2026-07-19