Back to News Feed
NVIDIA Blog23h agoNVIDIA Writers

NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents

The landscape of artificial intelligence is shifting toward the edge, as the open-source community empowers developers to build, customize, and deploy sophisticated AI agents directly on local hardware. Throughout August, NVIDIA is spearheading a celebration of this movement, highlighting the collaborative efforts, innovative models, and developer tools that are defining the next generation of local AI.

This month-long initiative serves as a hub for the latest advancements in agentic AI. NVIDIA is showcasing a suite of new open models, software frameworks, and educational resources designed to lower the barrier to entry for developers. To stay updated on these developments, enthusiasts can follow the NVIDIA RTX Spark channels on social media or subscribe to the RTX AI PC newsletter.

Introducing Nemotron 3.5 Lightning: Speed Meets Specialization

A major highlight of this month’s announcements is the expansion of the Nemotron 3 model family with the release of Nemotron 3.5 Lightning. This 30B mixture-of-experts (MoE) model is engineered specifically for "always-on" agentic workflows, prioritizing both efficiency and customization.

"Nemotron 3.5 Lightning delivers up to 4x faster token generation and 30% faster time to completion compared to open models in its class."

Because the model features open weights, developers have the freedom to fine-tune it to meet specific requirements. Whether it is mastering a unique writing style, learning the technical nuances of 3D design, or adhering to specific coding conventions and frameworks, Nemotron 3.5 Lightning provides the flexibility needed for personalized AI experiences. When integrated with local files and tools, these models can evolve into highly capable assistants for email management, smart-home automation, or collaborative coding.

Optimized Deployment Ecosystem

NVIDIA has partnered with key industry players to ensure seamless deployment. Developers can leverage:

  • Deployment Platforms: Collaboration with vLLM, Ollama, llama.cpp, and LM Studio ensures support for NVFP4 and GGUF formats.
  • Efficiency Tools: Unsloth provides day-one support, offering quantized models for streamlined local execution via Unsloth Studio.
  • Hardware Scalability: The model is optimized to run on everything from NVIDIA RTX PCs and Jetson modules to high-end RTX PRO workstations and enterprise-grade DGX systems.

Managing Costs with NeMo Switchyard

As enterprises scale their generative AI operations, balancing performance with infrastructure costs has become a primary challenge. To address this, NVIDIA has introduced NeMo Switchyard, an open-source routing library designed to optimize agent workflows.

NeMo Switchyard intelligently directs each step of a task to the most appropriate model based on a balance of speed, accuracy, and cost. By dynamically routing tasks, the library allows developers to maintain frontier-level performance while significantly reducing operational expenses. Internal benchmarks indicate that this approach can reduce task completion costs to roughly one-third of what would be required by relying solely on a single, high-end model like Opus 4.8.

Getting Started with Local AI

The tools and resources announced this month are designed to be accessible across the entire NVIDIA ecosystem.

  • For Developers: NeMo Switchyard is currently available on GitHub, and technical documentation for Nemotron 3.5 Lightning can be found on the NVIDIA technical blog.
  • For Edge Computing: Those interested in building at the edge should explore the Jetson AI Lab for tutorials and real-world project inspiration.
  • Cloud Accessibility: For those requiring cloud-based deployment, Nemotron 3.5 Lightning is accessible via OpenRouter, as an NVIDIA NIM microservice on build.nvidia.com, and through a wide network of cloud service providers and inference platforms.

With a robust hardware lineup—including new Blackwell systems from partners like Dell, HP, Lenovo, and Supermicro—the infrastructure for local, intelligent agents has never been more accessible or powerful.

#agents#open source#developer#modelnvidia