09/04/2026
Local agents are getting a speed boost. ⚡
New optimizations deliver up to 1.9x higher llama.cpp throughput on GeForce RTX 5090, 1.2x vLLM performance on RTX PRO 6000 Blackwell and up to 1.4x on a two-system DGX Spark cluster.
Available now: https://nvda.ws/3TeVBOj