Introducing NVIDIA Nemotron 3.5 Lightning⚡
An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster.
It delivers up to 4x the output speed of similar-sized models.
Already running an inference engine? So where does NVIDIA Dynamo fit in?
In five minutes, we break down how Dynamo sits around engines like @sgl_project, @vllm_project and TensorRT-LLM to scale inference across GPUs and nodes.
Full video in the comments 🔽
10 million downloads for NVIDIA Warp 🎉
Warp started with a simple idea: you shouldn’t have to leave Python to get real GPU performance for physics and simulation.
Since then, developers have used it to accelerate work across physics simulation, computational engineering,
Congrats to @Alibaba_Qwen on releasing Qwen3.8-Flash-Next, an experimental open-weight model that previews the Qwen4 architecture.
We’ve got Day 0 support to fine-tune with NVIDIA NeMo AutoModel and NeMo RL, plus recipes to run it with @sgl_project, @vllm_project and
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram