While NVIDIA is renowned for its powerful GPUs, it has quietly become a leading force in the world of open AI models, releasing a diverse range of models that are highly downloaded on platforms like Hugging Face. These models span various domains, from complex reasoning and world simulation to applications in robotics, autonomous vehicles, drug discovery, and even quantum computing. This article delves into NVIDIA's strategy behind developing these advanced, open-source models and the compelling reasons for making them freely available.

NVIDIA's open model ecosystem is broadly categorised. On one end are 'reasoning models' like their Nemotron family, which are large language models designed to "think" through tasks by generating intermediate steps, crucial for complex problem-solving in areas like coding and mathematics. At the other end are 'physical AI' models, such as Cosmos, which are world models capable of understanding and predicting physical interactions. These are vital for robotics and autonomous systems, enabling them to perceive and act within the real world. Cosmos 3, an advanced version, even integrates 'world action models' to directly generate robot control data. Bridging these are 'vision-language-action' (VLA) models, like the Isaac GR00T, which translate visual input and instructions into physical actions for humanoid robots, allowing for human-like decision-making and manipulation.
The article highlights NVIDIA's focus on creating models that are both cutting-edge and efficient. This is achieved through a hybrid architecture that combines Mamba layers for efficient processing of long sequences with attention layers for precise recall, offering the best of both worlds. Additionally, the use of Mixture-of-Experts (MoE) layers activates only a subset of parameters per token, maintaining speed and low cost. A key innovation is the co-design of their GPUs and models, exemplified by training in 4-bit precision (NVFP4) enabled by their Blackwell GPU architecture, leading to faster training and reduced power consumption without sacrificing accuracy. Post-training, involving supervised fine-tuning and extensive reinforcement learning, further enhances model capability. A unified foundation, akin to their CUDA software, allows for efficient development and reuse of components across different model families, fostering a culture of collaboration and "laziness" that max imises resource utilisation.
NVIDIA's commitment to "open" extends beyond releasing model weights; it includes publishing training data, post-training datasets, and the methodologies used, enabling others to reproduce and build upon their work. This transparency is crucial for advancing AI safely and effectively. The company's motivation for this open approach is twofold: firstly, it provides invaluable deep insights into AI trends, essential for guiding future hardware development and ensuring honest progress. Secondly, by supporting the AI ecosystem rather than competing with it, NVIDIA aims to foster growth that ultimately drives demand for their compute hardware. Lessons learned from this open-source journey include the importance of a sustained program over single releases, prioritising developer experience, and adopting trusted community licenses.
Fuente Original: https://blog.bytebytego.com/p/how-nvidia-builds-open-models-for
Artículos relacionados de LaRebelión:
- Kimi K3 AI Open Weights But Understand the Enterprise Caveats
- Chinese AI Models Cheaper Smarter Taking On The US
- Bots Rule the Web Humans Now Outnumbered Online
- AI Models Crumble Multi-Turn Attacks Expose Major Flaws
- Microsofts AI Models Slash Costs 89 Versus OpenAI
Artículo generado mediante LaRebelionBOT
No hay comentarios:
Publicar un comentario