Skip to main content

AI assistant

Sign in to chat with this filing

The assistant answers questions, extracts KPIs, and summarises risk factors directly from the filing text.

NVIDIA CORP Call Transcript 2026

Jun 1, 2026

Call Transcript

NVIDIA CORP

Download source file

Welcome to the stage, NVIDIA Founder and CEO, Jensen Huang. Welcome to GTC Taiwan. Great to see all of you. Very good to be home. I brought my parents home. Where are my parents? Everybody give a round of applause to my mum and dad. A round of applause for our pre-game show superstars, ladies and gentlemen. Look how adorable they are. The superstars of Taiwan. There are so many of you here today. We are broadcasting this right now to 70 other watch parties across Taiwan. 70 different conferences are going at the same time. Everybody is watching this keynote. We have so much to tell you, and I have so many partners to thank. It is incredible how large our ecosystem in Taiwan has become. Most of the time, when people think about ecosystem, they think about our software stack. They think about the developer ecosystem above the computing systems that NVIDIA builds. NVIDIA's ecosystem spans all the way upstream to all of our supply chain here in Taiwan, where it all begins, and downstream all the way to data centers, and eventually to end users. Today, we're going to talk about almost all of the ecosystem. There are so many people to thank. I love my ecosystem here. There are so many companies here, and some of my favorite ecosystem partners. So many. Taiwan's rich ecosystem, the richest ecosystem, the world's best supply chain ecosystem. Unbelievable. Well, thank you all for being here. This year, our businesses together are growing incredibly. In fact, somebody told me last night that the annual GDP of Taiwan is going to grow almost 10%. Unbelievable. Well, we have a lot to talk about. Let's get going. Two years ago, when I was here, I started to talk to you about how AI has moved from generative AI and the other waves of AIs that are coming. The next wave of AI was agentic AI. Today, we can say that agentic AI has arrived, that useful AI has arrived. Now, what does this mean? This is GitHub. Of course, one of the first applications of agentic AI is software coding, one of the most valuable professions. Incredibly large ecosystem, 30 million, 40 million professional software developers, probably another couple of hundred who are students and enthusiasts and so on and so forth. Say 30 million, 40 million software developers in the world, code for a living. This represents most of them. This is GitHub. The pull request is when they download software, they modify it, and commit is when they push it back up. Okay? If you could look at this, in 2023, the number of commits was 300 million, 2024, 400 million, 2025, 500 million commits. In the first few months of 2026, it has nearly tripled. Now, what does that mean? 30 million software developers representing about $3 trillion worth of GDP. That's what they're paid. $3 trillion worth of salaries per year, which is generating economic growth for the rest of the industries. Say $100 trillion of the world's industries is generated by $3 trillion worth of salary. That $3 trillion worth of salary is now producing nearly 3x as much output. It's effectively a $9 trillion productivity from $3 trillion of salaries. Does that make any sense? The difference is absolutely extraordinary. This is the potential. This is the promise of AI. The number of software engineers is actually increasing. People talk about AI reducing jobs. Complete nonsense. It's causing more software engineers to be hired. The reason for that is very simple. If you can hire a software engineer and you could generate $9 trillion worth of productive work, why wouldn't you want to hire more software engineers? If that line was flat, obviously people will hire fewer software engineers. Because the output is so incredible, people want to hire more software engineers. This is going to show up in our economy somehow, soon. The first thing is useful AI has arrived. Now, what does that mean from the industry's perspective? From the industry's perspective, that means that tokens are now in extraordinary demand, because if you could do this, you're going to want to produce more of it. Because tokens are now profitable units. Tokens are now profitable units of revenues. Because it is now profitable, the AI companies want to build a lot more tokens, generate a lot more tokens, build more AI factories, which is the reason why compute demand here in Taiwan has skyrocketed. It is precisely the reason why all of you are so busy and your businesses are doing so well. In fact, that looks like some of your stock price. The compute pattern has changed. Everything has changed. The first idea is that useful AI has arrived. AI is now a profit generator. AI is now a GDP generator. Behind it is a whole new kind of computing pattern, not just a large language model, but an agent. Today, almost everything we are going to talk about is going to be based on this. Let me take a quick moment and show you what I am talking about. Inside, this is an agent. It's an agent application. In the old days, this would be application, this would be code, and this would be operating system. Application code running inside an operating system. Today it is agent, which consists of a large language model or many, sitting inside a harness, and that harness helps it, orchestrates it to do productive work. This is the input. When that input comes, it has to understand, observe, reason, act, use tools. That tool could be a spreadsheet, web browser, a data processing engine, database engine, for example. This is orchestrated, this harness orchestrates this routing of information every single time it touches, either processing the context, understanding what is happening, reasoning about what to do, coming up with a plan that it acts on. That orchestration path is orchestrated by some software. This is fundamentally an agent. It deals with short-term memory, called working memory, long-term memory, just like we do, we have long-term memory. The memory management system is incredibly important. This entire system is called an agent. The large language model is used to do the thinking, and the harness connects everything together, just like an operating system. Okay? This is the new computing model, and this is what an agent. It could do incredible things. This is the big breakthrough. The simultaneous convergence of large language models that are now able to do a really good job thinking, reasoning, planning, using tools, and the fact that we have now these harnesses that manages memory, the orchestration, uses tools, we can now do amazing things. Let me give you some example. This is the prompt. This is the code that is generated, and this comes out. This is the input. This is the input, and that's the output. What do you guys think? It's pretty amazing, right? We use Claude Code here, but Codex does an incredible job as well. Here's another example. This is the input, "Create a GIF, NVIDIA green dots on black scatter, form Taiwan 101 building, morph to GTC Taipei 2026, morph to NVIDIA AI logo, then scatter and repeat." Right? You saw that. That was the prompt. Here's the next one, "I lost my remote control battery clip. It looks like this. Create a CAD file. It uses a tool. Create a CAD file ready for 3D printing to create a new one." Make sense? This is now the new computing pattern. Whereas we used to launch an application, click and type, we now replace that with explaining to the AI what we want, our intent, and the AI generates the code or uses tools and produce the necessary output. This is how computers are going to work in the future. This is agentic AI. For two years, we've been building towards this, and now it has arrived. One of the big breakthroughs, of course, is tool use. A lot of people have said, "Jensen, AI is coming, agentic AI is coming, therefore, all of the software companies are going to go out of business." I said, it's exactly the opposite, because there are going to be so many agents, the world is no longer limited by the number of people. Therefore, those agents are going to use more tools than ever. This is actually an incredible time to be a software company. The software has to be presented to the agent in a way that the agent can use it. This is a big breakthrough. In fact, what we have done, as you know, what NVIDIA's treasure is all of our CUDA libraries. I call them CUDA-X libraries. This is NVIDIA's treasure. Today, we're able to now present these CUDA-X libraries to agents who can use it much more effectively than even humans. So this is a wonderful time for CUDA-X libraries. Let's take a look. 20 years ago, we built CUDA, a single architecture for accelerated computing. We reinvented computing. 1,000 CUDA-X libraries help developers make breakthroughs in every field of science and engineering. CUDA-X libraries are tools for agents; cuLitho for computational lithography, cuOpt for decision optimization, cuDSS for direct sparse solvers. AI-Q for deep research across structured and unstructured documents. Aerial for AI-RAN. Warp for differentiable physics. Parabricks for genomics. At their foundation are algorithms, and they are beautiful. A round of applause for math. Math is beautiful. The computing pattern of software is going to change. Let's come back to this. This is the agent. It is the ultimate disaggregated and distributed computing model. Many different computers are going to be activated in order to process this agent. The agent consists of model, harness, tools and skills, and a runtime. All of that is running at different places in a data center. You can think of the model as the brain, the harness as the body. The tools that it uses, working in a runtime, think of it as a workshop. This is a person, a worker, working with tools in a workshop. Of course, this is being done at extraordinarily large scales, and each one of those steps are running in a different part of the computer. You could see the large language model is thinking, context processing, observing, understanding the environment, reasoning, coming up with a plan, and acting on the plan. Every single time that happens, an entire rack of Grace Blackwell NVL72 is activated. It's thinking with a large language model. Whenever it uses a tool, a CPU is used. That tool could be a C compiler, it could be Python, it could be JavaScript, or it could be accelerated computing. Today's agents are relatively simple users of tools. Tomorrow they're going to be very sophisticated users of tools, which is the reason why the CUDA-X libraries that I showed you are going to be incredibly popular with agents. They solve some of the most important problems the world knows, and all of our CUDA-X libraries are now going to come with skills that the AI could learn how to use. The CUDA-X library, some skills, basically a manual, the AI reads it and go, "Uh-huh, that's how you use it." The ability to use these libraries by agents are going to be incredible. The tools run on CPUs and GPUs and large language models. The security harness runs on CPUs and a security processor called a DPU, NVIDIA's BlueField. The orchestration of all this runs on a CPU. This is the entire harness, and the CPU is orchestrating all of the work. One of the hardest parts is memory. You could just imagine the working memory is called KV Caching. What to remember, compaction, not just compression, but how to retrieve. Do you retrieve structured data? Do you retrieve unstructured data? What is the ontology, the relationship of all of these different data to itself? That entire processing is incredibly complicated. The memory system of AIs is going to cause the storage system to be completely revolutionized. As you could see, every aspect of this computing model, this computing pattern, this new application called an agent, is fundamentally different than the way that applications used to run. A whole bunch of software sitting inside a binary, sitting inside an operating system. This is the reason, this disaggregated, this distributed, this heterogeneous computing problem is precisely the reason we built our next generation, Vera Rubin. Vera Rubin is not one chip. Vera Rubin is not a GPU only, it's starts with a GPU. Vera Rubin is incredible. This entire thing is Vera Rubin from end to end. It has GPUs, Vera Rubin NVLink 72. It is orchestrated by Vera CPUs that I'm going to tell you more about. The storage systems, revolutionary. Vera, along with CX-9, our software stack called DOCA, the security processor that's inside so that everything is encrypted at rest, in motion, as well as in use. Everything across this is secure because the AI model is so precious. This is the reason why this entire system obeys confidential computing. Each one of these systems would be a complete revolution in itself. Vera Rubin is the most ambitious endeavor in the history of our company. The whole company worked on Vera Rubin across all 40,000 engineers. Not to mention all of you. All of you participated in the creation of this entire system. Vera Rubin is really a miracle, and it's not just one chip, it is so many. Well, it's even beyond that. A long time ago, NVIDIA used to be a GPU company. Over the years, we've evolved to become a systems company. You're looking here now for the most complex system, most complex and ground-up system ever designed. Ultimately, our customers, our partners, don't want to buy computers, they want to build AI factories. The reason why NVIDIA has really started to transform ourself yet again. You could see so much of our technology is now at the entire infrastructure scale. Our partners are at infrastructure scale. Power generators, cooling systems, the grid providers. Many industrial companies are now part of our ecosystem because ultimately, we're trying to build an entire stack just like GPUs, just like when we were building Grace Blackwell NVL72, just like now, we are building a full stack system so that our customers could build amazing AI infrastructure. Let's take a look. The world is racing to build AI factories, the largest infrastructure build-out in human history. AI factories are incredibly complex. Every layer, chip, rack, network, power, cooling, grid, must be designed together from end to end, because compute is revenues. NVIDIA DSX is the blueprint, a reference design for building and operating AI factories at maximum efficiency and profitability. It starts with DSX Sim. With the DSX Sim Omniverse Blueprint, partners design and validate an NVIDIA Vera Rubin AI factory before a single rack lands. They plan the layout. Simulate the power and cooling. Design the network. Validate every integration. Test every change in the digital twin. The factory powers on. DSX OS takes over and provisions, operates, monitors, and remediates the infrastructure, turning the installed systems into trusted, multi-tenant, resilient, AI-ready capacity. Today's AI factories over-provision power by up to 40%. DSX MaxLPS lets operators safely deploy more GPUs inside the same power budget, adding billions in annual revenue. Breakthrough hot liquid cooling at 45 degrees Celsius uses less water and energy. More power going to revenue-generating compute. Incredible. Dynamic power allocation steers power from rack to rack, recovering stranded watts, sending them where work is happening. In-rack power smoothing flattens peak current spikes and power surges. Throughout the factory, teams of AI agents work with DSX MaxLPS, continuously coordinating to balance cooling and power to meet workload demand. DSX AI factories are flexible energy assets that operate cooperatively with the grid. DSX Flex reads real-time grid signals and dynamically adjusts factory power when the grid needs relief. 100 GW of AI factories will come online before the end of the decade. NVIDIA DSX AI factories run at highest efficiency, produce the lowest-cost tokens, and make the grid stronger. I've shown you ecosystem slides of the past, where NVIDIA's computing layers and software and software and computing stacks are integrated into other people's platforms, third-party platforms and libraries that serves end markets. That was a computing ecosystem. This is an AI factory ecosystem. This is way downstream of all of you. Upstream of me is all of you, and downstream of us is this ecosystem. NVIDIA ultimately is not just building a GPU, not just building a system. We're helping customers build these AI factories, these AI infrastructure that is so immensely complex. Each one of these at 1 GW level started at $20 billion, $30 billion. It is at $50 billion, $60 billion, and soon it will be $80 billion, $100 billion per gigawatt. $100 billion into an AI factory. It must work the first time, and it must work right away. The cost of capital is incredible. The complexity is incredible. As you see, we used to design a chip inside a computer, and then we simulated a system inside a computer. Today, you saw just now everything was built in Omniverse. I've been working with Omniverse with all of you for a long time. This was the dream come true, so that we can build these gigantic systems as large as the world wants to build inside a digital framework, inside a digital simulator, in a digital world, long before we build the first break ground and put our money to work. This is our ecosystem. We call it DSX. RTX is for our GPU, DGX is for our systems, and now DSX, basically infrastructure. Because of the work that we do here across this entire stack, including our systems and software, it's the reason why we can work with small companies and enable them to be world-class AI clouds. Every one of these I'm about to show you are small companies just recently, now CoreWeave is worth $50 billion, $60 billion, $70 billion and growing incredibly fast. Recently, we worked with Nebius, again, they're growing incredibly fast. Each one of these clouds have incredible customers. Cursor, the software coding company, Black Forest Labs, Image Generation, World Labs, World Foundation Model, Revolut, the leading financial services AI company, Shopify. Here's another one. This is Nscale, their customers are British Telecom, Google. Google is using one of our AI clouds. Thinking Machines, a Frontier Labs company. We're super excited. Here's NAVER Cloud in Korea, Bank of Korea, Hyundai. Many incredible companies. Here's one in India, Yotta. Incredible companies. Here's one based in Singapore, building in Australia, Together AI Singapore. This is one in Indonesia. Each one of these companies are serving regional as well as global customers. AI is going to run everywhere. Every company will be powered by it. Every region will build it. Indosat here in Indonesia. Here in Taiwan, GMI. Here in Taiwan, GMI. It's okay to clap. Incredible companies, incredible opportunity, but all of them need several things. Of course, they need the computing stack, this entire stack underneath. This is what made NVIDIA famous. All of our hardware and software and libraries, our connection into the world's ecosystem of third-party developers makes it possible for anyone to stand up an AI cloud. However, the AI cloud is so complex now. This is the software version. This is the computer science version. The money version, the asset version is what I showed you earlier. It's a giant factory. Having this ability alone is not enough, which is the reason why NVIDIA has become an AI infrastructure company. Now, doing this well and becoming incredibly good at helping customers build AI factories and deploying AI factories is incredibly important, and the reason for that is this. Compute is revenue now. Compute is profit. The absence of revenues and profit is loss. It's really important to realize that this is an example of an AI infrastructure coming online. It could be coming online quickly, it could take a while, its throughput could be high, it could be low, its resilience or reliability could be good or bad, and its lifetime of usefulness could be long or short. This represents $50 billion, $60 billion, going to $100 billion, this curve matters greatly, which is the reason why NVIDIA is such a great partner. Working with us because of our fully integrated capability, we didn't just come up with a PowerPoint slide, we created the entire infrastructure. We connected everything together. We built out billions and billions of it ourselves to make sure that everything works well. As a result of that, our time to first token, our time to first inference, our time to training turned on is much faster. Second, because our throughput per watt, our tokens per watt is utterly world-class, and the reason for that is because we integrate everything, we design everything from the ground up, we simulate the entire system, and we use extreme co-design. Just like I showed you just now with the Vera Rubin rack, everything was designed in order to deliver on this incredible throughput. If your data center, if your factory has 1 GW, it will not have more. 1 GW means 1 GW. That's all the power generation you could do. If you have 1 GW of power, then throughput per watt is revenues because every token is profitable. Every token is revenues. This is the future. Compute is revenues. Performance per watt is your revenues. Choosing the wrong architecture just because the chips are cheaper doesn't translate, doesn't make sense. You need to make sure that your revenues per watt, the more you buy, the more you make. Tokens per watt. Third is reliability. If you ever get a chance to see these data centers, there are so many moving parts, millions of cables. The ability for all of those computers to work harmoniously, reliably is extremely low. It is just extremely difficult. We have now been operating very large scale for a very long time. That experience matters. That difference, mean time between interrupts, extremely important. Lastly, this is very hard. The lifetime of these systems, the software is changing all the time. four years ago, which is in the time of Hopper, AI has completely changed. six years ago, this is the timeframe of Ampere, AI has completely changed. We started out talking about CNNs. Here we are, then we talked about transformers, and then we talked about Mixture of Experts. Now we're talking about agentic systems. Every single generation, every single few months, the software industry is coming up with new technology. If your architecture is not flexible, if your ecosystem is not rich, then this curve cannot be long. You cannot predict how long your system can last. I can. NVIDIA systems is all over the world. Software developers start with NVIDIA CUDA, by definition, therefore, the life, the ecosystem, the useful asset is going to be much longer. The difference is essentially cost. You could think of it as revenues, the other side of revenues is cost. If the life of the asset is long, the TCO is low. This is the difference. This is what it looks like when compute. The more you buy, the more you make. All of you are experiencing this with me. Isn't that right? All of your demand, your factories are working so hard. Your people are working so hard all across Taiwan because everybody wants to make money. They realize that useful AI is here. Profitable AI is here. Compute demand is incredibly high, and compute demand is the constraint. Let's go work super hard and help the world stand up AI factories everywhere. This is why it's so important. I'm so happy. Here I am standing in front of you. Vera Rubin is in full production. Vera Rubin is in full production. The supply chain we created for Vera Rubin is twice as large as Grace Blackwell. Yeah, it's incredible. What used to take two hours to assemble one Grace Blackwell rack now only takes five minutes. Not only is the capacity higher, the throughput is a lot faster, and we need it all to support the demand. This ecosystem is extraordinary. Millions of square feet has been put online to support Grace Blackwell and preparing now, ramping up now, Vera Rubin. I want to thank all of you. Vera Rubin is now in full production. Thank you. Let's take a look. Large language models generate answers. AI agents can do work. Processing agentic AI is a whole different kind of problem. Agents observe, reason, plan, use tools. They manage massive context, juggling working memory and long-term memory. They spin up sub-agents, specialists on demand. NVIDIA Vera Rubin is a multi-rack pod scale system built to process agentic AI and is now in full production. The manufacturing, automation, and orchestration across the supply chain, a miracle to witness. Our journey started when we launched the first AI supercomputer, NVIDIA DGX-1. Over the next decade, we pushed every chip and system to the limit, from Pascal and the first NVLink to Grace Blackwell, the first rack-scale AI supercomputer. Now, Vera Rubin, the first multi-rack pod scale supercomputer built for the agentic age. It starts at TSMC. The seven new chips that make up Vera Rubin take shape through hundreds of processing steps. Three nanometer process, CoWoS-R and CoWoS-L packaging, HBM4 memory from Micron, SK hynix and Samsung. The Vera Rubin Compute Board, six trillion transistors with over 18,000 components on one board. Vera Rubin NVL72 does the thinking, prompt and context understanding, reasoning, and planning. Next, a new modular compute tray, streamlined with a new PCB mid-plane design. Superchips, ConnectX-9 SuperNICs, and BlueField-4 DPUs all mate in place with no cables for resiliency at AI factory scale. 18 compute trays, nine hot swappable NVLink switch trays. New high-efficiency manifolds, liquid-cooled busbars carrying over 5,000 amps, the equivalent of 20 electric cars at full acceleration. Together, 1.3 million components form this third-generation MGX rack design. Congratulations to Microsoft for their operational Vera Rubin NVL72 engineering rack. Congratulations to Dell and CoreWeave as well for standing up their Vera Rubin NVL72 engineering rack. The Vera CPU rack. 256 CPUs in a single liquid-cooled rack, orchestrating the models, shuffling memory, launching tools. At Foxconn and Quanta, Groq 3 LPX takes shape. 256 Groq 3 LPUs across 16 trays, 40 Pbps of SRAM bandwidth for ultra-low latency. While NVL72 generates tokens at the highest throughput, Groq LPX generates them at the lowest latency. Vera BlueField-4 STX, where AI keeps its memory. Storage processing accelerated by BlueField-4, connecting memory, storage, and in-silicon security. NVIDIA Spectrum-X Ethernet Photonics, the world's first Ethernet switch with 200 Gb co-packaged optics, TSMC's COUPE process, chip scale packaging, and ultra-high-powered laser dies on indium phosphide. Vera Rubin, five connected rack scale systems, a supercomputer for AI agents. 150 supply chain partners across Taiwan. Millions of square feet of factory floor. Hundreds of sites, chips, packages, systems, and data centers pushed to the limits of size, power, and scale. This is what we call extreme co-design. We did this with Taiwan. Together, we reinvented computing for the age of AI. Taiwan was with us at the beginning and here today as we bring Vera Rubin to the world. Thank you, Taiwan. Ladies and gentlemen, Vera Rubin. Vera Rubin was not just built for AI. Vera Rubin was not built just to run AI. Vera Rubin was built to run agents. This is an agentic system. Imagine the complexity, which is the reason why agents is the last computer science breakthrough. It has taken this many years for agents to realize its potential and become useful. It stands to reason that the computer that runs it is the most advanced in the world. This is Vera Rubin. Let's take a look. Can we bring out Vera Rubin, please? Janine, do we have the racks, the systems? It looks heavy. This is Vera Rubin. Vera Rubin NVL72. This is the Groq LPX. At the next GTC, I'm going to talk to you about a lot more of this. Today, we have so much to talk to you about. This is Vera CPU rack, 256 CPUs, all liquid-cooled. Let me tell you about Vera in just a moment. This is the Vera BlueField storage processing system and also security system. Of course, this is our Mellanox networking, the world's first CPO. This is Vera Rubin. Incredible technology all coming together. When we built Hopper, as you know, for pre-training. Pre-training was the most important application, the most important workload we were working on at the time. When we worked on Grace Blackwell, everybody said, "Jensen, NVIDIA is really good at pre-training. Inference is so easy." Do you remember that? People used to say, "Inference is so easy. We could do that, too." As you know, inference equals money, and the models, MoEs, are so complicated. To do it at incredibly high response time, fast interactivity, and high throughput at the same time is incredibly hard, which is the reason why we created NVLink 72. Today, NVIDIA's token cost is the lowest in the world, not by 10%, by X factors, orders of magnitude. All because we did extreme co-design, all because we understood the computing model, the computing pattern of inference, and we were able to create NVLink 72. With Vera Rubin, it is beyond inference. It is now inference in an agentic system. This is Vera Rubin. No cables, no hoses, no fans. What used to take the last time when I showed this to you, we had cables everywhere. The cables were amazing to look at. Now there's a PCB in the middle which connects both sides. What used to take two hours now takes five minutes. The reliability and the resilience of Vera Rubin is going to be off the charts. This is our Vera CPU tray, the most advanced CPUs that has ever been built. I'm going to show you that in just a second. This is our storage tray. Two Vera CPUs, four CX9, incredible amounts of software. This is our new LPX, LPU 30, the Groq system, designed for very low latency inference. The throughput is delivered by Vera Rubin and extended with NVLink 72. If you want to extend that even further, you can have Groq LPUs. Here, we have the Vera Rubin NVLink, the switch tray. This is the switches in the middle, and this is revolutionary. Because of Vera Rubin's, because of NVLink 72 and the NVLink switches that we created and invented. This is our Ethernet switches for scale-out. What's amazing is we introduced these two systems for Grace Blackwell. These two systems were created for Grace Blackwell. Today, NVIDIA is the largest networking company in the world. I'm so proud of the networking team. This is such an incredible enabler for everything that we do. I'm going to now talk to you about the next major industry we're going to be part of. Thank you. Janine. Thank you. I think there are 2,000 people back there pulling that. Let's talk about CPUs. Vera CPUs. CPUs built for the age of AI. All of the CPUs until now were created for people. We were the users, we were the renters. The way we use CPUs, we live in a world counted by seconds. The way we rent CPUs in the cloud, each one of them, more CPU cores you have, the more you can rent. The use case of the old CPU and the economics of the old CPU, fundamentally different than agents. Agents are impatient. They don't live in a world that is in seconds. They live in a world that's in nanoseconds. When it uses a tool, it wants the response time to be as fast as possible. When it access database, it has to come back as soon as possible. Every moment that the agent is waiting keeps it from going to the next step. It is vital that we make the CPUs as low latency as possible, as interactive as possible. We created Vera CPU for the age of AI. Now, inside our system, it's used for three different ways. The first way, of course, is Vera Rubin for thinking. Inside the Vera Rubin rack, there are already two CPUs. As you know, we are building and selling millions of Vera Rubins. We have sold millions of Grace Blackwells. NVIDIA already is one of the largest CPU makers in the world. In the Vera Rubin rack are two CPUs, one for orchestrating and managing the GPUs, managing the KV cache, dealing with all of the software that runs in the rack. We also have the Grace BlueField that is used for security and isolation. The Vera CPU is used for the harness, the orchestration of the AI models, tool use, accessing the database. The data servers are right here, Vera BlueField, the fastest storage servers, the fastest storage system the world has ever made. The reason why this is so vital is because agents are accessing memory so incredibly fast. These systems, the storage server, and the CPUs, are now the critical path of the most expensive part of the data center. This is the most expensive for a good reason. The economics of the AI factory is tokens. The tokens are created here. Of course you want to manufacture and generate as many tokens as possible. This is where you put all of your economics, and this has to not be in the way. Vera CPU has great pressure on the CPU architecture, which is the reason why we built a brand-new architecture from the ground up. A CPU the world has never seen before. We call it Vera. This is CPU for agents. All the CPUs of the past we built for humans. This CPU is built for agents. Well, there are four things to keep in mind. The four takeaways. The first takeaway is that the instructions per clock of Vera has to be incredibly good because we need the latency to be short. We need the processing time, single-threaded performance, not throughput, single-threaded performance has to be world-class. Absolutely the best single-threaded performance, which is the reason why the IPC, the instructions per clock of Vera, is so high. It's the highest in the world. Ten instructions fetched, decoded, and executed per clock. Number one. Number two. The bandwidth necessary to move data in and out for the CPU has to be utterly world-class. The second thing is bandwidth per core. The third is just bandwidth, period. Remember I said earlier, agentic systems is fundamentally disaggregated and distributed. When computing is disaggregated and distributed, networking becomes the problem. Therefore, we have to move the data around as fast as possible between the CPU cores and between the CPU and the storage, the CPU and the GPU. The bandwidth around the system and inside the CPU core has to be utterly world-class. This is the first CPU that has been built a long time that is literally at radical limits with a fabric that connects all of the CPU cores that is speed of light, 3.6 Tbps. No chiplet tax, no chip boundary crossings because the CPU cores are talking to each other with extremely high bandwidth. They are not rented core per core per core. They are all working together. The cross-sectional bandwidth of Vera is off the charts. It is the first one to be PCI Express Gen 6. It is also the first one to have LPDDR5 with 1.2 Tbps, 2x to 3x the bandwidth of the highest performance CPUs on the outside, 3x the bandwidth on the inside. The bandwidth per core and the bandwidth period is world-class. Remember, I showed you earlier, the number of CPU cores, the number of CPUs is going to be quite high. The reason for that is very simple. We created CPUs for humans in the past, and humans, there are only 1 billion of us. There will be billions of agents, and these agents are going to be using the CPUs with very little patience because the cost of the GPU they sit next to is too high, and therefore, too valuable, too precious. These CPUs are going to be both performant, but they also have to be extremely energy efficient so that we can cram as much CPU as we can into the factory without taking away power from the token generation, which we know is how we make money. These four properties, instructions per clock or single-threaded performance, bandwidth per core, the total bandwidth around the chip and inside the chip, and energy efficiency defines Vera. It is absolutely world-class. When you compare it to the highest performance x86, it is just off the charts. When you compare it in real single-threaded performance, real performance, it's off the charts. It is incredible to be able to deliver 5% improvement on CPUs. It is incredible to be able to deliver 10%. This kind of performance speed up is just unheard of. This is NVIDIA Vera. What do you think? Let's take a look. Agentic AI changes the role of the CPU. The CPU is now the conductor, and the GPU is the orchestra. Traditional CPUs were built for a different era, maximizing cores per socket. Slice them up, virtualize, rent by the hour. In the age of agents, the CPU is now a bottleneck to GPU utilization, directly affecting token throughput, latency, and user experience. NVIDIA Vera is the CPU built for the agentic loop, combining NVIDIA's custom data center CPU core with a scalable coherency fabric for the right balance of performance cores and bandwidth to maximize AI factory output. At the heart of Vera is the NVIDIA Olympus core, built for modern data center workloads, branch-heavy Python runtimes, tool calls, and sandboxed code execution. Each core is tuned for throughput. A neural branch predictor evaluating two taken branches per cycle. A 10-wide decode engine brings in more work each cycle. A large out-of-order engine keeps instructions moving. Advanced prefetchers with a novel graph engine anticipating the next data path. Fast cores only matter when data arrives correctly and on time. Vera is the first CPU to use LPDDR5X memory while correcting multiple errors simultaneously without compromising bandwidth. Vera achieves 40% lower peak memory latency versus x86, keeping cores fed on time through retrieval, analytics, and sandbox execution. NVIDIA's second-generation scalable coherency fabric unifies all 88 Olympus cores on a monolithic mesh with separate dies for memory and I/O. Cores are not split across chiplets, enabling 50% faster core-to-core communication than traditional CPUs. Memory-coherent NVLink chip-to-chip connects GPUs directly to the fabric. Beyond GPUs, NVLink chip-to-chip can scale Vera up to multiple sockets, enabling massive bandwidth between CPUs. Vera delivers 1.8x the agentic sandbox performance of x86 CPUs. Standalone Vera racks run agent sandboxes, tools, code, and data pipelines. Tightly coupled to Rubin GPUs, Vera keeps accelerated workflows moving. NVIDIA Vera BlueField-4 STX powers context memory and AI storage. Compute, networking, storage. Vera is the CPU for the age of agents. This is going to be our new major growth driver. The reviews are already coming out, it's pretty good. That's pretty good stuff. Remember, Grace and Vera are also the most highly qualified CPUs in the world of AI because every single data center, every single cloud, every single enterprise, every company that works with NVIDIA on AI has already qualified Grace. The entire software stack has already been optimized for Grace. Every company will be qualifying Vera. Vera will be the most optimized agentic CPU in the world simply because it's going to go with Vera Rubin, simply because we made the big hard switch. During Grace Blackwell transition, the biggest risk was going from external CPU x86 into Grace Blackwell. That transition was extremely dangerous, we did it with incredible execution. Grace is literally synonymous with Grace Blackwell. When people say Blackwell, they say Grace Blackwell, because it is utterly now everywhere. Every company's software stack has been optimized for it. Everybody's security stack has been optimized for it. Now here comes Vera. I'm super excited about that. Now look at some of the performance numbers. Speedups is one thing. It is extremely hard to speed up SQL. SQL, the most famous domain-specific language, DSL, that has ever been created. Before SQL, before CUDA, there was SQL. Before OpenGL, there was SQL, invented by IBM. Today, it is the structured database engine of the planet. Everybody uses SQL. This is SQL running 3x faster, not 10% faster, not 25% faster, 3x faster. Incredible. The next one is real-time stream processing. Remember, your AI is going to be not just reading documents. Your AI is going to be watching for telemetry, especially inside a factory, inside a stock exchange. You're going to be looking for telemetry continuously. The burst of data that's coming in goes into a CPU. This is Vera CPU running real-time stream processing for New York Stock Exchange. Lynn Martin, the president of New York Stock Exchange, has been so gracious to partner with us. This system is run all over the world in real-time stream processing. Vera CPU, 3x, all because of the bandwidth, the single-threaded instruction execution, the bandwidth inside between the cores, the bandwidth outside. Vera is completely revolutionary. That's Vera. X factors is something you talk about when you're talking about GPUs. It is quite rare that somebody talks about X factors on real workload that is associated with CPU. I'm so proud of the team. You guys did such a great job. We have an extraordinary roadmap coming. What's really exciting is almost everybody is supporting Vera. They're as excited as we are. This is Vera opening up. It's opened up a brand new market. Agents is a new workload. We built CPUs for humans in the past. We need CPUs for agents, agentic systems. Their properties are different. Why would the old CPUs be the same? We are building millions and millions of Veras, millions of Veras. To go to market with us, Taiwan's ODMs and computer makers, all the OEMs, and you could see the early adopters. The early adopters are the agentic companies. This is the beginning of a new market, a market that never existed before. It's not going to take away from the old markets, but this is a new market, CPU for agents. This market will surely be larger than the last, the reason for that is because there'll be a lot more agents than there are people, the agents are very impatient. NVIDIA Vera CPU. Thank you. This is the most important slide, really. This is the takeaway. The takeaway here is that this is the application pattern. This is the computing pattern of the next decade. Agents, harnesses, orchestrating large language models. Every company will run it. Every company will be an agent company. Every company will have agents running inside. Every company will see that agents will need its own operating system. Every company is asking us, "How do we run agents safely? How do we build agents for our own workloads?" We have the NVIDIA Agent Toolkit for Enterprise AI. You've seen me build this in plain sight. Almost everything that NVIDIA does, as you know, at every GTC, if you go back and look at my GTC five years ago or 10 years ago, you will see today. This, you've seen me talking about for several years now, because we've been building for this moment. There are four things that companies need in order to build agents as a service or build agents to operate. The first thing you need is you need models. Of course, large language models. The smarter, the better. The cheaper, the better. The faster, the better. The second is you need a harness to orchestrate the whole thing. The third, these models want to use tools, these tools come with its skills. I showed you CUDA-X libraries. Those are going to be amazing tools for the agents in the future. Lastly, you need a runtime. You need the operating system that holds it all together. This is the NVIDIA toolkit for agents. It includes models that you can modify, NVIDIA's world-class open models. I'm going to show you more. You could run agents from anybody. You could run Claude Code, incredible agent, Codex, incredible agent. You could run it inside this harness called OpenShell, which will be highly secure for your inside the enterprise. The shell protects the agent, keeps it grounded in security policies. Privacy is protected. Its rights and privileges are given. Its identity is protected. This OpenShell is being adopted all over the world. NVIDIA OpenShell is open source. You're going to see so many companies adopt it. Red Hat, Canonical, Microsoft. It's going to be adopted everywhere. This is important. This is the runtime. This runtime is fully optimized for the NVIDIA AI platform, which is everywhere. You can run OpenShell in any cloud, on-prem, and even on-device. You have now tools and libraries that they can use. You have models that you can modify or use as is, or you have agents. This would be OpenClaw, Hermes, another incredible harness. These agentic harnesses can now run on-prem or for you anywhere. Okay. Four things, and this represents the operating system of the modern enterprise. How do we use this? One of my favorite use cases of agents is chip designers. It is the single most important thing that NVIDIA does. Of course, we have to partner with Cadence to build Super Agent, a chip design super agent. It is orchestrated by Codex or Claude Code. It has RTL and architecture diagrams or schematics or specifications as input and whatever you need to fix. Together, we created some super agents that are optimized for the NVIDIA runtime with Nemotron, and let's take a look. It is really incredible. Cadence and NVIDIA are partnering to build chip design agents. Hundreds of thousands of NVIDIA chips come together to make the AI factories that power the world's frontier AI models. Designing these chips and the systems they run in is one of the hardest engineering challenges. Trillions of transistors, three-dimensional circuits, microscopic scale. Every gate, every wire, synchronized to picoseconds, must work in perfect harmony with no margin for error. Physical prototypes are too slow and too costly, so engineers work in the digital realm. Each chip begins as a set of architectural specifications, then translated into RTL, the language of chip design. RTL must be verified in simulation. A single bug can delay a chip by months. At NVIDIA, thousands of engineers, billions of compute hours per year, millions of tests written, run, and debugged. A cycle that takes teams weeks. To compress this cycle, Cadence and NVIDIA built a design verification agent. Codex orchestrates the process. Cadence ChipStack launches the RTL verification loop, powered by Nemotron and secured by NVIDIA OpenShell, calling on expert sub-agents in RTL generation, test bench creation, regression testing, and debug. The system drives itself. The ChipStack agents run hundreds of simulations with Cadence Xcelium, formal verification with Jasper. Design flaws revealed. Bugs in the code fixed. What once took weeks now takes hours. Verification cycles over 40x faster. Together, NVIDIA and Cadence are reinventing chip design with AI agents. From weeks to hours. NVIDIA has thousands of chip designers. We are going to hire hundreds of thousands of Cadence super agents that work with us so that we can accelerate our company, so that we can be even more ambitious, create even more amazing things, run even faster. You saw earlier that the toolkit with models, harness, tools. The tools in this case are Cadence simulators and verifiers, formal verification systems. It is the reason why we're working with Cadence so hard to accelerate all of their tools on CUDA. Because the agents are impatient. The agents want the answer immediately. Models, harnesses, accelerated CUDA, accelerated libraries and tools, and then the runtime. What you saw just now is all of that coming together. Now, one of the things that it starts with is a great model that Cadence could modify and tune to be expert at the Cadence workflow, at the Cadence expertise, so that they could create super agents that are proprietary to Cadence with their proprietary knowledge. They have to start with an excellent model. We call it Nemotron. NVIDIA is dedicated to build open models for the world so that all of you, all of us, could create our own agents. Today, we're announcing the Nemotron 3 Ultra. Yep. Our next open model, and it is smart. The Nemotron models not only give you the model, we give you all the data that we use to train the model, and because we have a coalition of incredible partners, you can see all of our partners down here, we work together, contribute data to each other. Nemotron is trained on one of the largest suites of long-running reasoning models, long-running tool task-solving, tool-using data sets in the world because of all of our great partnerships. All of this from the model, the training script, and the data made completely available to you. This is open models at its best, the best open model system policies in the world. Simple goal is that you can take all of it, add to it, make it even better, make it yours. Nemotron 3 Ultra is 5x faster. This is the world's first model based on a hybrid architecture of SSM, state-space models, with Mixture of Experts. The architecture is incredibly fast. We made it fast that you could think fast. When you think fast, you could think longer at the same cost. 5x faster. It is also 30% cheaper, 30% lower cost to run in total FLOPS and total inference time than even the most cost-effective in the world. We're comparing against the world's best open models. Frontier smart, 5x faster, 30% cheaper, completely open. We're completely dedicated to this. This is now Nemotron 3. We're currently working on Nemotron-4. This entire toolkit from models, harnesses, tools and skills, and runtimes is the reason why every enterprise company in the world has the ability now to create their own agents, just like Cadence did with their super agents. We're working with so many companies, Cadence and CrowdStrike and Dassault and Palantir, SAP and ServiceNow. People always said, "Jensen, the agents are going to disrupt these markets." I said completely opposite, and you can now see it. Agents is going to create the largest opportunity ever for my partners and friends. We have the NeMo, the NVIDIA agentic toolkit for enterprise AI to help them. There you go. First, Vera Rubin in full production. Two, Vera CPU built for a new generation for agents. Three, NVIDIA's enterprise AI toolkits, so that every enterprise and every enterprise software company can build agents. My relationship with you started here. Many of you, many of my friends and partners here in Taiwan, your companies started here. This is in a lot of ways, the beginning of the modern computer industry, 40 years now. NVIDIA's 33 years old. The PC industry was already starting to get to Windows 1 and Windows 2 and Apple 1 and Apple 2. By the time that we came along, Windows 3.1 was the PC. As you know, Windows 95 made PC personal. It took PC from enterprises, companies, and made it into a consumer electronics device. Everybody should have one, and everybody does. This is the beginning. This computing platform did several things incredibly smart. Windows was not just disaggregated, as you know. Windows was properly abstracted. It was architected just right. Systems BIOSes, open chipsets, the operating system with drivers that could be connected and installed at runtime, and an abstraction layer with a multimedia API that opened up the PC to what we all know today. Each one of these elements were essential in making the PC so popular. 40 years later, Microsoft and NVIDIA are going to reinvent the PC. This is going to be the new PC. Tomorrow night, I think it's tomorrow night our time, but I'm going to be with Satya, where we're going to talk a lot more about the work that we're doing together. Microsoft and NVIDIA, over the last three years, it took this long to completely reinvent how the PC's going to work so that we could be ready for this moment. As I mentioned earlier, that compute pattern called the agent is going to run in AI clouds. It's going to run inside enterprises. It is also going to run on your PC. What's going to happen to that PC when it has an autonomous agent? An agent that's helping you, that understands you. You could talk to it. It could look at you. You could ask it to read files, go help you do some research. It could do a lot more that I'll show you. The new operating system is, of course, the old operating system plus large language models. Large language models in a lot of ways is the modern version of DirectX. It has, of course, input and output, understands prompts, it understands computer vision, it can generate video, it can generate sounds. It is the modern extension, the intelligence extension of the PC, of a computer. On top of that, the application, as I mentioned before, is going to be replaced by now an agentic runtime, and that is the modern application, an agent. Let's now take a look at what it can do. It started with a spark. An idea to reimagine the PC for the first time in 40 years for the age of AI. What becomes of our personal computer in a world of agents? Agents running natively, connected to models, local or in the cloud. Our personal AI, sandboxed for security, running continuously, getting work done. The chips and the OS must evolve. Introducing RTX Spark. Everything we've learned over 33 years distilled into one chip. Blackwell RTX GPU with 6,144 CUDA cores, one petaflop of AI performance. A custom 20-core Grace CPU built in partnership with MediaTek, fused by NVLink. 128 GB of unified memory. TSMC 3nm process. 70 billion transistors. In close collaboration with Microsoft, a Windows platform for agents. We're reinventing the personal computer for creating, for gaming, for agents. This is the dawn of a new personal computing revolution, and it starts with NVIDIA RTX Spark. Here it is. Of course, I got to show you the most beautiful part, which is video games. It's also the closest to our heart. This is Forza. This is 007, by the way. The new 007 game. I'm looking forward to playing it. I look a little bit like him. Ladies and gentlemen, NVIDIA's RTX Spark laptops. I have too many things in my pocket. Okay. All right. This is the most amazing chip the world's ever built. This is the N1X that we built in partnership with MediaTek. I think I saw Rick earlier. This is N1X. This is a beautiful chip. This is a chip that, frankly, would take 33 years to build. The reason for that is because 100% of NVIDIA software stack runs here. If you want to run digital biology, no problem. If you want to do seismic processing, no problem. You want astrophysics, no problem. Everything associated with CUDA, all the physics, all the biology, all the genomics, all the AI, no problem. All the computer graphics, no problem. Every single application NVIDIA has ever created and every single application that Windows has ever run. Microsoft and NVIDIA meticulously optimized everything so that this computer literally runs everything the world has ever created. Plus, it now runs agents. An incredible computer. I'm so proud of it. Now, I want you to keep that in mind in the next video I'm going to show you. Just imagine everything here is going to run on your PC. That computer could have a local Nemotron 3 Ultra model or Nemotron 3 Super model, or it could have a Claude Code or Codex or some other model in the cloud, or something on the network, and it's going to work and do something amazing. Let's play it. Every house starts as an idea. Getting from idea to design takes a myriad of tools, expertise, and a lot of time. Now, an agent running locally on RTX Spark can help me design a house using the tools on my laptop with an OpenShell sandbox running the Hermes harness connected to Claude Sonnet in the cloud. I select the site, share my concept sketches and mood board of styles to inspire my design, and the prompt, a text description of the requirements and the design intent. My agent goes to work. Using the tools on my laptop, it opens Rhino and starts modeling the site, shaping terrain, setbacks, and the building envelope. It proposes building forms optimized for cost, comfort, and quality. With the form defined, my agent generates the interior layout. Walls, circulation, rooms begin to take shape. I jump in whenever I want to adjust, to change. Doors, windows, and structural elements are placed automatically. My agent detects its own mistakes and fixes them. When I approve, the agent exports the model from Rhino into Blender. Materials and object properties transfer with the design context intact. I fine-tune the materials, get the look just right. I pick the shots. Blender renders the house. My agent, using generative AI with the Flux.2 model, makes them photoreal. Multiple viewpoints, lighting conditions. What was once a complex workflow is now guided and simplified by my agent. Working with me on RTX Spark. Design at the speed of imagination. PC in the world of agents. The developers are so excited about it. This is an incredible computer. All of the acceleration, all the software capabilities associated with it, working with every developer to make it incredible for all of you. The next one, Adobe. Incredible tool suite, of course, used by tens of millions of people around the world. They have re-engineered the architecture, the core of Adobe Photoshop and Premiere. They'll release it for RTX Spark. It is twice as fast. It's already fast. Now it's going to be twice as fast. It's also designed to be agent-friendly. With its MCP server, it can now interact with agents on your laptop. The number of customers, the number of partners that are so excited to bring RTX Spark to the market is just incredible. This is the first across the lineup of PC reinvention for 40 years. I'm just so happy that all of you and the ecosystem around the world has joined us. This is basically everybody. Everybody will support RTX Spark and will be building incredibly smart and powerful and beautiful laptops with all of us. Thank you very much. That's not all. That's not all. RTX Spark is a reinvention of laptop. In fact, Microsoft NVIDIA is reinventing all of PC, and today we're announcing a whole new line. Three revolutionary Windows machines covering desktop, laptop, and workstations. All 100% Windows compatible, 100% CUDA, 100% NVIDIA AI Tensor Core. Everything that you see that runs on NVIDIA in all these different platforms around the world runs here. This is the first completely re-engineered, reinvented line of PCs that has happened in 40 years. What's really amazing is this. This is the RTX Spark laptop. This is the desktop. This one's from MSI. Joseph, this one's yours. Okay. Look how beautiful it is. This agent could run 24/7, meter-free. You could download your agent. You could raise your lobster in here. This is your claw. It is running all the time, no meter anxiety, and it is sitting here connected to your whole house, connected to your laptop, connected to your display, all the cameras, your dryer, your water cooler, your water heater, your everything, whatever you want. Your security system, all connected to this, and this becomes your personal AI, your personal AI agent. It gets smarter and smarter and smarter over time because today we have Nemotron 3 Ultra, tomorrow we have Nemotron 4, and then Nemotron 5, Nemotron 6, and we just keep getting it smarter and smarter and smarter. Meanwhile, this is sitting at home helping you do things. If you want to book a travel, no problem. If you want an incredible system, this is a DGX Station for Windows, compatible with Windows, runs everything in Windows, and it has 768 gigabytes of memory. You could run a trillion-parameter model. This is unbelievable. 20 petaflops, 8 Tbps of memory bandwidth, and this sits by your desk. If you're a developer of large language models, you're a developer of agents, having this sit by your desk gives you all the compute you need, and then when you deploy it, you put it into the cloud. There's something that if you look at this and think about this, something is happening here. Remember, 15, 20 years ago, we used to have an idea called a phone. Today, we have an idea called a PC. Today, when you think about your phone, the one thing you don't do with it is make phone calls. You do just about everything else. That phone means something very different to you than a phone of the past. I am certain what's going to happen here is that the PC 10 years from now and the PC that you think about today, a tool, whether you launch applications, click and type, and this PC is going to be completely different. Here's my theory. I can totally imagine, just as every house today has a home theater, or many houses have home theaters, big TVs, lawnmowers, dishwashers. I could totally imagine that someday there's actually an AI supercomputer in your house. It's running all of your agents, it's running all of your assistants, and they're doing all kinds of things for you all the time. You have to have it in your house, just like you have a home theater in your house, you have stereos in your house, you have game consoles in your house. You end up assist AI agent computers running in your house. These, in time, becomes a lot more like R2-D2 to you. It becomes more like C-3PO to you than it feels like a PC to you. There is no question this reinvention of the computer is as big of a deal as the reinvention of the phone into what we now know as the smartphone. This is the beginning of that journey. This is the beginning of a new line. We have a roadmap for this. This is a brand-new product family for us. Every single generation of architecture, we will have a desktop, a laptop, a workstation, and then a desktop, a laptop, and workstation. The thing that I am just incredibly pleased, incredibly honored, is that 100% of the world's PC industry has joined us to reinvent the PC. A new line, a new beginning. Thank you. As you know, agentic AI is just a digital robot. It understands, it reasons, it plans, and it acts and use tools. Agentic AI is going to run across all of these computers, and you've seen me talk about each and every one of these over time. We're working on human or robotics computers, robotics computers of all kinds. We're working on self-driving car computers. We're working on satellites. You have GeForce, which has tensor cores. I just talked about a whole new line of PCs. Agriculture equipment, manufacturing equipment, heavy industry equipment will all be agentic. You'll even have a little agentic helper for yourself. Even your base stations, the radio stations of the future are going to be agentic, understanding traffic and thinking about how to coordinate with the other base stations so that you could use as little energy as possible, increase the utilization, the efficiency of the spectral efficiency. Everything will run agents. Today, NVIDIA is largely in the center, but I am pretty certain that there will be tens of billions, hundreds of billions over time of agentic systems, agentic computers that are going to be running around the world. The biggest problem is data. In the case of language models, all the English and all the language that we have on the internet that we trained on was from the perspective of us. We wrote it and we're reading it. However, in order to create data for AI robotics, it has to be in the perception, the perspective of the robot. Most of the world's video data is from a third person, not first person. Agentic systems, robotic systems, physical AI, the data is the hardest problem. You've seen us move up this ladder. We started with teleoperations, which is basically human demonstration. This is no different than the big breakthrough of reinforcement learning human feedback. We use simulation. This is where Omniverse comes in. This is no different than reinforcement learning verifiable rewards. We use these systems to bootstrap the AI model, the physical AI model. Eventually, we're able to learn from third person, re-projecting it into first person, and now eventually, through bootstrapping, we have a world foundation model that can understand the physical world from any perspective you want. Third person, first person, outside in, inside out, doesn't matter. This is a big breakthrough indeed. Today, we're announcing Cosmos 3. Cosmos 3 is the frontier of physical AI. We are at the frontier with language models. There are so many people working on it. However, in physical AI, we are absolutely the world's best. I am so proud of the team for doing this. This is the foundation model for all of your work. Whenever you want to create a robot, whenever you want to create a factory robot or a robot that works in a factory, any kind of robot that involves physical world, you now have a companion, a Cosmos 3 that can understand and reason. It can generate. It can simulate in the loop. It can even be the policy itself. It is on the top of leaderboards all over the world. I am incredibly proud of Cosmos, and today we're announcing Cosmos 3. Let's take a look. The real world is infinite and unpredictable. Physical AI needs data, but real-world data is impossible to scale. For physical AI, compute is data. This is Cosmos, an open frontier omni-modal for physical AI, built on a new mixture of transformers architecture. Pixels, action, sound, and language flow into the autoregressive transformer, which reasons, plans, and instructs the diffusion transformer, which generates what comes next. Developers post-train Cosmos across embodiments and use cases. As a VLM, Cosmos watches the physical world, understands what's happening, describing scenes, and flagging what matters. As a world model, Cosmos generates physics-accurate synthetic video from an image, text, or video. As a simulator, Cosmos closes the loop for policy training and evaluation. As the foundation of NVIDIA Omnidreams, an action-conditioned world model, Cosmos predicts the future frame by frame. Post-train Cosmos, it becomes a world action model: perceiving, reasoning, planning, generating actions for robots of every kind, for everything that moves. A new kind of data, a new kind of teacher, generated by compute. Cosmos, the foundation for developers of the age of physical AI. It takes data plus compute, gives you AI. Now that we have AI, compute is data. Use Cosmos 3, train a whole bunch of AI models. Cosmos is such an incredible open model system. It's exactly the same as Nemotron. We open the model, we open the data, and we even open how we trained it so that you could enhance it for yourself and turn Cosmos into your proprietary model. We have such incredible partners working with us in so many different industries. Now, the model itself is, of course, the most understandable part of the AI stack. The AI stack is very complicated. It has generators, the model, simulators, and the runtime. Just as it is for agentic systems, these cars, or essentially a physical AI, a agentic robot that is a autonomous vehicle, has also this complicated stack. Today, we're announcing Alpamayo 2, an open model for self-driving cars. We're working with car companies across the world. If you look at these brands that have signed up for the NVIDIA Hyperion, that are building NVIDIA Hyperion cars, this represents about 80% of the world's cars. The manufacturers represent 80% of the world's cars. We are going to have a whole lot of NVIDIA Hyperion systems that are able to run Alpamayo or anybody else's AV stack. We are also connected into mobility services. Approximately 97% of the world's mobility services are connecting with us so that when we deploy Alpamayo on the Hyperion runtime with the Halos operating system, we will be able to connect to all of these services across the world. Let's take a look at this. Hey, Mercedes. Let's go to my favorite sandwich shop. Routing to your destination. Lane is clear. Pulling out to start drive. Nudge left due to the stationary lead vehicle ahead blocking our lane. Slow down to stop at the stop sign controlling the intersection. Stop to yield to the pedestrians since the person is in our lane. Yield to the cut-in vehicle from the left. Nudge left to clear the stopped vehicle blocking on the right. Keep distance to the cut-in vehicle since it is merging into our lane. Nudge left due to the stopped van blocking the right side of our lane. [crosstalk]. Stop to keep distance to the lead vehicle since it is stationary ahead. Keep distance to the vehicle directly ahead in our lane. Keep distance to the vehicle directly ahead in our lane. Stop for the stop sign since the intersection is controlled by a stop sign. Stop to yield to cross-traffic since a vehicle is crossing ahead. Keep distance to the lead vehicle. Nudge right due to the truck blocking the right side of our lane. Nudge right due to the truck blocking the left side of our lane. Nudge left due to the truck blocking the right side of our lane. Keep distance to the lead vehicle. Your destination is on the right. Alpamayo. The world's first reasoning autonomous vehicle. If you let it talk all the time, it will drive you crazy. We're very happy that it's talking to itself all the time. That's called thinking. Alpamayo is a reasoning car. The technology that we've created also applies to humanoids. Of course, there are many new breakthroughs that has to happen. The NVIDIA Isaac GR00T is our humanoid robotic stack. Model, data generation, simulation, the runtime, including the operating system. This represents GR00T platform, the Isaac GR00T platform. Every one of our systems, as you can see, the exact same pattern, whether it's agentic system for the cloud, Agentic system for the PC, a robotic system for a self-driving car, a robotic system for a human or robot, all the same. In every single case, we build everything completely. We build everything vertically, completely integrated with co-design, extreme co-design. Then we open it up for everybody to use whichever part you like. Whatever you want to use, we even help you modify. The one thing that is missing is we need a reference platform for robotic systems. These robotic systems are so complicated, so many motors, so many sensors, so fragile, and yet we need to have a way to deliver these reference platforms just like we do with PCs and DGXs and clouds and self-driving cars. We now are going to do it for robots. Today we're announcing the NVIDIA Isaac GR00T, a reference humanoid robot, all fully integrated, 25 degrees of freedom on each hand made by Sharpa. 31 degrees of freedom on the robot, 6 ft, 150 pounds, just like me. The first number is shorter, the second number is bigger. Otherwise, pretty close. This platform runs the new Thor and our entire software stack. Data generation stack, data simulation stack, the runtime, all integrated into a robot that is designed for everyone to use. We built this for higher education and university researchers because for them to build this is insanely hard to do. Let's take a look at that. The next leap in AI is general-purpose robots, humanoids. Building one is hard. Every team starts from scratch, stitching together simulators, teleop systems, data pipelines, and training infrastructure. Months of setup before research can start. NVIDIA Isaac GR00T, an open development platform for humanoid robots. Open models, simulation and training libraries, and data generators. Plus the robot computer. Fully pipe clean, ready to go in hours. First, set up the simulation environment in Isaac Lab. Capture demonstrations with Isaac Teleop on a real or simulated robot. Generate synthetic data with Omniverse and Cosmos, scaling one demonstration into thousands. Train policies. Evaluate them in Isaac Lab Arena. Deploy through Isaac ROS, running on Jetson Thor. Every element, modular, open. Use ours or swap in your own. GR00T is powering robotics research across every discipline for every domain, from research labs to factory floors. One open platform. Now a new addition, Isaac GR00T reference design robots. Built on NVIDIA's open platform, ready for frontier research for any lab, anywhere. The age of robotics starts here. NVIDIA Isaac GR00T. Many robots. We're working with just about everybody who's working on robots in the world or robotic systems in world. Let me tell you what I told you. The computer industry has been completely changed. In the last six months, everything changed. Everything changed because agents were realized, and it converged with the latest frontier models, and it made possible the AI to now do useful work. The computing pattern will repeat over and over again. This computing pattern of an agent that's a model, a harness that uses tools with skills and runs in a runtime, that runtime depends on whether it's in the cloud or on-prem, on a PC or in a robot. The computing pattern is exactly the same for all of them. You will use different harnesses because of your preference. You'll use different models because of your preference. You will improve them for your proprietary use. You would create super agents that you can rent to other people to help them do their work. This agentic platform, this agentic pattern, NVIDIA has an enterprise AI toolkit. This is a wonderful way for all of you to engage AIs, and for us, it's a wonderful growth opportunity. Vera Rubin is in full production. Whereas Grace Blackwell was created to process AI, particularly inference, Vera Rubin was created to run agents. It is in full production. It is much more than a GPU. It is an entire disaggregated, distributed agent processing system. NVIDIA has really become an infrastructure company, not just a GPU company, not just a systems company, but an infrastructure company to help you generate the maximum revenues, the maximum profit, and to get there as soon as possible. The agent world, this new way of doing computing, where you build CPUs now for agents, not for people. CPUs for agents has its own special requirement, and our NVIDIA Vera is revolutionary. I'm so happy about its ramp. The orders already is going to make it the fastest and the most successful product launch in our company's history. NVIDIA and Microsoft has created a whole new line of PCs. This is a new beginning. Of course, that exact same agentic processing pattern, computing pattern that I just described is also going to run on all kinds of devices. I mentioned PCs, in the future, it'll be robots and satellites and base stations and factories in the cloud, on-prem, at the edge. This pattern, agentic AI system, this agentic computing pattern, will be replicated in computers all over. How we think about the personal computer will very likely change. I want to thank all of you for your partnership, your friendship. We couldn't be here without everything that we do together. I am so proud of how you've been so successful this last year. The next year is going to be even more. I have one more thing for you. Let's take a look. You ready, Taiwan? Let's do this. The keynote's done at Computex. Jensen showed the world what's next. Useful AI has arrived. Agents working by your side. In case you missed things we said today. We're gonna break it all down for you, Taipei. Agents used to be misunderstood. Only movie stars had them in Hollywood. Now we all got teams making dreams come true. Building companies from living rooms. They need so much compute, we hear you. That's why we created Vera. Rubin stole the show, it's true. The cheapest tokens coming through. 10x faster, inference heaven. More special agents than 007. BlueField keeps agents' memory true. Now, let's talk about its CPU. 50% faster, that's outrageous. Not for Vera. It's built for agents. NVLink Fusion blends ASICs smartly. Everyone's welcome to the NVLink party. Well, if you like that introduction. Vera Rubin's in full production. Nemotron Ultra leads the run. 5x faster, work gets done. NeMo Cloud keeps the guardrails right. OpenShell keeps the sandbox tight. Your code migrated and reviewed. All before this song is through. AI is a five-layer cake. Compute's revenue, make no mistake. Global AI clouds build lots of gigawatts. DSX keeps power lean, connecting dots. Every watt optimized for you. You can have your cake. Eat it, too. RTX 40 is finally here. Biggest PC moment in 40 years. For you. Agents powering all workflows. Running anywhere Windows goes. Harnesses run on CPU. Models fly on GPU. Cosmos builds worlds that robots need. Turning compute into synthetic feed. Alpamayo sees and reasons through. Understands roads like people do. GR00T is how they learn to move. Learning skills and finding groove. JuliUs is powered by Thor. The future's humanoid. Count on more. [audio distortion] The future's bright. Come see what's next. Thank you, Taiwan. Welcome to Computex. Have a great Computex. Thanks for an amazing year. Thank you for all your friendship and support. Thank you. Take care.

Speaker 2: Welcome to the stage, NVIDIA Founder and CEO, Jensen Huang. Welcome to the stage, NVIDIA Founder and CEO, Jensen Huang. welcome to the stage nvidia founder and ceo jensen huang

Speaker 1: Welcome to GTC Taiwan. Great to see all of you. Very good to be home. I brought my parents home. Where are my parents? Everybody give a round of applause to my mum and dad. A round of applause for our pre-game show superstars, ladies and gentlemen. Look how adorable they are. The superstars of Taiwan. There are so many of you here today. We are broadcasting this right now to 70 other watch parties across Taiwan. 70 different conferences are going at the same time. Everybody is watching this keynote. We have so much to tell you, and I have so many partners to thank. It is incredible how large our ecosystem in Taiwan has become. Most of the time, when people think about ecosystem, they think about our software stack. Welcome to GTC Taiwan. welcome to gtc taiwan Great to see all of you. great to see all of you Very good to be home. very good to be home I brought my parents home. i brought my parents home Where are my parents? where are my parents Everybody give a round of applause to my mum and dad. everybody give a round of applause to my mum and dad A round of applause for our pre-game show superstars, ladies and gentlemen. a round of applause for our pre-game show superstars ladies and gentlemen Look how adorable they are. look how adorable they are The superstars of Taiwan. the superstars of taiwan There are so many of you here today. there are so many of you here today We are broadcasting this right now to 70 other watch parties across Taiwan. 70 different conferences are going at the same time. we are broadcasting this right now to 70 other watch parties across taiwan 70 different conferences are going at the same time Everybody is watching this keynote. everybody is watching this keynote We have so much to tell you, and I have so many partners to thank. we have so much to tell you and i have so many partners to thank It is incredible how large our ecosystem in Taiwan has become. it is incredible how large our ecosystem in taiwan has become Most of the time, when people think about ecosystem, they think about our software stack. most of the time when people think about ecosystem they think about our software stack They think about the developer ecosystem above the computing systems that NVIDIA builds. NVIDIA's ecosystem spans all the way upstream to all of our supply chain here in Taiwan, where it all begins, and downstream all the way to data centers, and eventually to end users. Today, we're going to talk about almost all of the ecosystem. There are so many people to thank. I love my ecosystem here. There are so many companies here, and some of my favorite ecosystem partners. So many. Taiwan's rich ecosystem, the richest ecosystem, the world's best supply chain ecosystem. Unbelievable. Well, thank you all for being here. This year, our businesses together are growing incredibly. In fact, somebody told me last night that the annual GDP of Taiwan is going to grow almost 10%. Unbelievable. Well, we have a lot to talk about. Let's get going. They think about the developer ecosystem above the computing systems that NVIDIA builds. they think about the developer ecosystem above the computing systems that nvidia builds NVIDIA's ecosystem spans all the way upstream to all of our supply chain here in Taiwan, where it all begins, and downstream all the way to data centers, and eventually to end users. nvidia's ecosystem spans all the way upstream to all of our supply chain here in taiwan where it all begins and downstream all the way to data centers and eventually to end users Today, we're going to talk about almost all of the ecosystem. today we're going to talk about almost all of the ecosystem There are so many people to thank. there are so many people to thank I love my ecosystem here. i love my ecosystem here There are so many companies here, and some of my favorite ecosystem partners. there are so many companies here and some of my favorite ecosystem partners So many. so many Taiwan's rich ecosystem, the richest ecosystem, the world's best supply chain ecosystem. taiwan's rich ecosystem the richest ecosystem the world's best supply chain ecosystem Unbelievable. unbelievable Well, thank you all for being here. well thank you all for being here This year, our businesses together are growing incredibly. this year our businesses together are growing incredibly In fact, somebody told me last night that the annual GDP of Taiwan is going to grow almost 10%. in fact somebody told me last night that the annual gdp of taiwan is going to grow almost 10% Unbelievable. unbelievable Well, we have a lot to talk about. well we have a lot to talk about Let's get going. let's get going Two years ago, when I was here, I started to talk to you about how AI has moved from generative AI and the other waves of AIs that are coming. The next wave of AI was agentic AI. Today, we can say that agentic AI has arrived, that useful AI has arrived. Now, what does this mean? This is GitHub. Of course, one of the first applications of agentic AI is software coding, one of the most valuable professions. Incredibly large ecosystem, 30 million, 40 million professional software developers, probably another couple of hundred who are students and enthusiasts and so on and so forth. Say 30 million, 40 million software developers in the world, code for a living. This represents most of them. This is GitHub. The pull request is when they download software, they modify it, and commit is when they push it back up. Okay? Two years ago, when I was here, I started to talk to you about how AI has moved from generative AI and the other waves of AIs that are coming. two years ago when i was here i started to talk to you about how ai has moved from generative ai and the other waves of ais that are coming The next wave of AI was agentic AI. the next wave of ai was agentic ai Today, we can say that agentic AI has arrived, that useful AI has arrived. today we can say that agentic ai has arrived that useful ai has arrived Now, what does this mean? now what does this mean This is GitHub. this is github Of course, one of the first applications of agentic AI is software coding, one of the most valuable professions. of course one of the first applications of agentic ai is software coding one of the most valuable professions Incredibly large ecosystem, 30 million, 40 million professional software developers, probably another couple of hundred who are students and enthusiasts and so on and so forth. incredibly large ecosystem 30 million 40 million professional software developers probably another couple of hundred who are students and enthusiasts and so on and so forth Say 30 million, 40 million software developers in the world, code for a living. say 30 million 40 million software developers in the world code for a living This represents most of them. this represents most of them This is GitHub. this is github The pull request is when they download software, they modify it, and commit is when they push it back up. the pull request is when they download software they modify it and commit is when they push it back up Okay? okay If you could look at this, in 2023, the number of commits was 300 million, 2024, 400 million, 2025, 500 million commits. In the first few months of 2026, it has nearly tripled. Now, what does that mean? 30 million software developers representing about $3 trillion worth of GDP. That's what they're paid. $3 trillion worth of salaries per year, which is generating economic growth for the rest of the industries. Say $100 trillion of the world's industries is generated by $3 trillion worth of salary. That $3 trillion worth of salary is now producing nearly 3x as much output. It's effectively a $9 trillion productivity from $3 trillion of salaries. Does that make any sense? The difference is absolutely extraordinary. This is the potential. This is the promise of AI. The number of software engineers is actually increasing. People talk about AI reducing jobs. Complete nonsense. If you could look at this, in 2023, the number of commits was 300 million, 2024, 400 million, 2025, 500 million commits. if you could look at this in 2023 the number of commits was 300 million 2024 400 million 2025 500 million commits In the first few months of 2026, it has nearly tripled. in the first few months of 2026 it has nearly tripled Now, what does that mean? 30 million software developers representing about $3 trillion worth of GDP. now what does that mean 30 million software developers representing about $3 trillion worth of gdp That's what they're paid. $3 trillion worth of salaries per year, which is generating economic growth for the rest of the industries. that's what they're paid $3 trillion worth of salaries per year which is generating economic growth for the rest of the industries Say $100 trillion of the world's industries is generated by $3 trillion worth of salary. say $100 trillion of the world's industries is generated by $3 trillion worth of salary That $3 trillion worth of salary is now producing nearly 3x as much output. that $3 trillion worth of salary is now producing nearly 3x as much output It's effectively a $9 trillion productivity from $3 trillion of salaries. it's effectively a $9 trillion productivity from $3 trillion of salaries Does that make any sense? does that make any sense The difference is absolutely extraordinary. the difference is absolutely extraordinary This is the potential. this is the potential This is the promise of AI. this is the promise of ai The number of software engineers is actually increasing. the number of software engineers is actually increasing People talk about AI reducing jobs. people talk about ai reducing jobs Complete nonsense. complete nonsense It's causing more software engineers to be hired. The reason for that is very simple. If you can hire a software engineer and you could generate $9 trillion worth of productive work, why wouldn't you want to hire more software engineers? If that line was flat, obviously people will hire fewer software engineers. Because the output is so incredible, people want to hire more software engineers. This is going to show up in our economy somehow, soon. The first thing is useful AI has arrived. Now, what does that mean from the industry's perspective? From the industry's perspective, that means that tokens are now in extraordinary demand, because if you could do this, you're going to want to produce more of it. Because tokens are now profitable units. Tokens are now profitable units of revenues. It's causing more software engineers to be hired. it's causing more software engineers to be hired The reason for that is very simple. the reason for that is very simple If you can hire a software engineer and you could generate $9 trillion worth of productive work, why wouldn't you want to hire more software engineers? If that line was flat, obviously people will hire fewer software engineers. if you can hire a software engineer and you could generate $9 trillion worth of productive work why wouldn't you want to hire more software engineers? if that line was flat obviously people will hire fewer software engineers Because the output is so incredible, people want to hire more software engineers. because the output is so incredible people want to hire more software engineers This is going to show up in our economy somehow, soon. this is going to show up in our economy somehow soon The first thing is useful AI has arrived. the first thing is useful ai has arrived Now, what does that mean from the industry's perspective? now what does that mean from the industry's perspective From the industry's perspective, that means that tokens are now in extraordinary demand, because if you could do this, you're going to want to produce more of it. from the industry's perspective that means that tokens are now in extraordinary demand because if you could do this you're going to want to produce more of it Because tokens are now profitable units. because tokens are now profitable units Tokens are now profitable units of revenues. tokens are now profitable units of revenues Because it is now profitable, the AI companies want to build a lot more tokens, generate a lot more tokens, build more AI factories, which is the reason why compute demand here in Taiwan has skyrocketed. It is precisely the reason why all of you are so busy and your businesses are doing so well. In fact, that looks like some of your stock price. The compute pattern has changed. Everything has changed. The first idea is that useful AI has arrived. AI is now a profit generator. AI is now a GDP generator. Behind it is a whole new kind of computing pattern, not just a large language model, but an agent. Today, almost everything we are going to talk about is going to be based on this. Let me take a quick moment and show you what I am talking about. Inside, this is an agent. Because it is now profitable, the AI companies want to build a lot more tokens, generate a lot more tokens, build more AI factories, which is the reason why compute demand here in Taiwan has skyrocketed. because it is now profitable the ai companies want to build a lot more tokens generate a lot more tokens build more ai factories which is the reason why compute demand here in taiwan has skyrocketed It is precisely the reason why all of you are so busy and your businesses are doing so well. it is precisely the reason why all of you are so busy and your businesses are doing so well In fact, that looks like some of your stock price. in fact that looks like some of your stock price The compute pattern has changed. the compute pattern has changed Everything has changed. everything has changed The first idea is that useful AI has arrived. the first idea is that useful ai has arrived AI is now a profit generator. ai is now a profit generator AI is now a GDP generator. ai is now a gdp generator Behind it is a whole new kind of computing pattern, not just a large language model, but an agent. behind it is a whole new kind of computing pattern not just a large language model but an agent Today, almost everything we are going to talk about is going to be based on this. today almost everything we are going to talk about is going to be based on this Let me take a quick moment and show you what I am talking about. let me take a quick moment and show you what i am talking about Inside, this is an agent. inside this is an agent It's an agent application. In the old days, this would be application, this would be code, and this would be operating system. Application code running inside an operating system. Today it is agent, which consists of a large language model or many, sitting inside a harness, and that harness helps it, orchestrates it to do productive work. This is the input. When that input comes, it has to understand, observe, reason, act, use tools. That tool could be a spreadsheet, web browser, a data processing engine, database engine, for example. This is orchestrated, this harness orchestrates this routing of information every single time it touches, either processing the context, understanding what is happening, reasoning about what to do, coming up with a plan that it acts on. That orchestration path is orchestrated by some software. This is fundamentally an agent. It's an agent application. it's an agent application In the old days, this would be application, this would be code, and this would be operating system. in the old days this would be application this would be code and this would be operating system Application code running inside an operating system. application code running inside an operating system Today it is agent, which consists of a large language model or many, sitting inside a harness, and that harness helps it, orchestrates it to do productive work. today it is agent which consists of a large language model or many sitting inside a harness and that harness helps it orchestrates it to do productive work This is the input. this is the input When that input comes, it has to understand, observe, reason, act, use tools. when that input comes it has to understand observe reason act use tools That tool could be a spreadsheet, web browser, a data processing engine, database engine, for example. that tool could be a spreadsheet web browser a data processing engine database engine for example This is orchestrated, this harness orchestrates this routing of information every single time it touches, either processing the context, understanding what is happening, reasoning about what to do, coming up with a plan that it acts on. this is orchestrated this harness orchestrates this routing of information every single time it touches either processing the context understanding what is happening reasoning about what to do coming up with a plan that it acts on That orchestration path is orchestrated by some software. that orchestration path is orchestrated by some software This is fundamentally an agent. this is fundamentally an agent It deals with short-term memory, called working memory, long-term memory, just like we do, we have long-term memory. The memory management system is incredibly important. This entire system is called an agent. The large language model is used to do the thinking, and the harness connects everything together, just like an operating system. Okay? This is the new computing model, and this is what an agent. It could do incredible things. This is the big breakthrough. The simultaneous convergence of large language models that are now able to do a really good job thinking, reasoning, planning, using tools, and the fact that we have now these harnesses that manages memory, the orchestration, uses tools, we can now do amazing things. Let me give you some example. This is the prompt. This is the code that is generated, and this comes out. This is the input. It deals with short-term memory, called working memory, long-term memory, just like we do, we have long-term memory. it deals with short-term memory called working memory long-term memory just like we do we have long-term memory The memory management system is incredibly important. the memory management system is incredibly important This entire system is called an agent. this entire system is called an agent The large language model is used to do the thinking, and the harness connects everything together, just like an operating system. the large language model is used to do the thinking and the harness connects everything together just like an operating system Okay? okay This is the new computing model, and this is what an agent. this is the new computing model and this is what an agent It could do incredible things. it could do incredible things This is the big breakthrough. this is the big breakthrough The simultaneous convergence of large language models that are now able to do a really good job thinking, reasoning, planning, using tools, and the fact that we have now these harnesses that manages memory, the orchestration, uses tools, we can now do amazing things. the simultaneous convergence of large language models that are now able to do a really good job thinking reasoning planning using tools and the fact that we have now these harnesses that manages memory the orchestration uses tools we can now do amazing things Let me give you some example. let me give you some example This is the prompt. this is the prompt This is the code that is generated, and this comes out. this is the code that is generated and this comes out This is the input. this is the input This is the input, and that's the output. What do you guys think? It's pretty amazing, right? We use Claude Code here, but Codex does an incredible job as well. Here's another example. This is the input, "Create a GIF, NVIDIA green dots on black scatter, form Taiwan 101 building, morph to GTC Taipei 2026, morph to NVIDIA AI logo, then scatter and repeat." Right? You saw that. That was the prompt. Here's the next one, "I lost my remote control battery clip. It looks like this. Create a CAD file. It uses a tool. Create a CAD file ready for 3D printing to create a new one." Make sense? This is now the new computing pattern. This is the input, and that's the output. this is the input and that's the output What do you guys think? what do you guys think It's pretty amazing, right? it's pretty amazing right We use Claude Code here, but Codex does an incredible job as well. we use claude code here but codex does an incredible job as well Here's another example. here's another example This is the input, "Create a GIF, NVIDIA green dots on black scatter, form Taiwan 101 building, morph to GTC Taipei 2026, morph to NVIDIA AI logo, then scatter and repeat." Right? this is the input "create a gif nvidia green dots on black scatter form taiwan 101 building morph to gtc taipei 2026 morph to nvidia ai logo then scatter and repeat." right You saw that. you saw that That was the prompt. that was the prompt Here's the next one, " I lost my remote control battery clip. here's the next one, " i lost my remote control battery clip It looks like this. it looks like this Create a CAD file. create a cad file It uses a tool. it uses a tool Create a CAD file ready for 3D printing to create a new one." Make sense? create a cad file ready for 3d printing to create a new one." make sense This is now the new computing pattern. this is now the new computing pattern Whereas we used to launch an application, click and type, we now replace that with explaining to the AI what we want, our intent, and the AI generates the code or uses tools and produce the necessary output. This is how computers are going to work in the future. This is agentic AI. For two years, we've been building towards this, and now it has arrived. One of the big breakthroughs, of course, is tool use. A lot of people have said, "Jensen, AI is coming, agentic AI is coming, therefore, all of the software companies are going to go out of business." I said, it's exactly the opposite, because there are going to be so many agents, the world is no longer limited by the number of people. Therefore, those agents are going to use more tools than ever. Whereas we used to launch an application, click and type, we now replace that with explaining to the AI what we want, our intent, and the AI generates the code or uses tools and produce the necessary output. whereas we used to launch an application click and type we now replace that with explaining to the ai what we want our intent and the ai generates the code or uses tools and produce the necessary output This is how computers are going to work in the future. this is how computers are going to work in the future This is agentic AI. this is agentic ai For two years, we've been building towards this, and now it has arrived. for two years we've been building towards this and now it has arrived One of the big breakthroughs, of course, is tool use. one of the big breakthroughs of course is tool use A lot of people have said, "Jensen, AI is coming, agentic AI is coming, therefore, all of the software companies are going to go out of business." I said, it's exactly the opposite, because t here are going to be so many agents, the world is no longer limited by the number of people. Therefore, th ose agents are going to use more tools than ever. a lot of people have said "jensen ai is coming agentic ai is coming therefore all of the software companies are going to go out of business." i said it's exactly the opposite, because t here are going to be so many agents the world is no longer limited by the number of people. therefore, th ose agents are going to use more tools than ever This is actually an incredible time to be a software company. The software has to be presented to the agent in a way that the agent can use it. This is a big breakthrough. In fact, what we have done, as you know, what NVIDIA's treasure is all of our CUDA libraries. I call them CUDA-X libraries. This is NVIDIA's treasure. Today, we're able to now present these CUDA-X libraries to agents who can use it much more effectively than even humans. So this is a wonderful time for CUDA-X libraries. Let's take a look. This is actually an incredible time to be a software company. this is actually an incredible time to be a software company The software has to be presented to the agent in a way that the agent can use it. the software has to be presented to the agent in a way that the agent can use it This is a big breakthrough. this is a big breakthrough In fact, what we have done, as you know, what NVIDIA's treasure is all of our CUDA libraries. in fact what we have done as you know what nvidia's treasure is all of our cuda libraries I call them CUDA- X libraries. i call them cuda- x libraries This is NVIDIA's treasure. this is nvidia's treasure Today, we're able to now present these CUDA- X libraries to agents who can use it much more effectively than even humans. today we're able to now present these cuda- x libraries to agents who can use it much more effectively than even humans So this is a wonderful time for CUDA- X libraries. so this is a wonderful time for cuda- x libraries Let's take a look. let's take a look

Speaker 3: 20 years ago, we built CUDA, a single architecture for accelerated computing. We reinvented computing. 1,000 CUDA-X libraries help developers make breakthroughs in every field of science and engineering. CUDA-X libraries are tools for agents; cuLitho for computational lithography, cuOpt for decision optimization, cuDSS for direct sparse solvers. AI-Q for deep research across structured and unstructured documents. Aerial for AI-RAN. Warp for differentiable physics. Parabricks for genomics. At their foundation are algorithms, and they are beautiful. 20 years ago, we built CUDA, a single architecture for accelerated computing. 20 years ago we built cuda a single architecture for accelerated computing We reinvented computing. 1,000 CUDA- X libraries help developers make breakthroughs in every field of science and engineering. we reinvented computing 1,000 cuda- x libraries help developers make breakthroughs in every field of science and engineering CUDA- X libraries are tools for agents; cuLitho for computational lithography, cuOpt for decision optimization, c uDSS for direct sparse solvers. cuda- x libraries are tools for agents culitho for computational lithography cuopt for decision optimization, c udss for direct sparse solvers AI-Q for deep research across structured and unstructured documents. ai-q for deep research across structured and unstructured documents Aerial for AI- RAN. aerial for ai- ran Warp for differentiable physics. warp for differentiable physics Parabricks for genomics. parabricks for genomics At their foundation are algorithms, and they are beautiful. at their foundation are algorithms and they are beautiful

Speaker 1: A round of applause for math. Math is beautiful. The computing pattern of software is going to change. Let's come back to this. This is the agent. It is the ultimate disaggregated and distributed computing model. Many different computers are going to be activated in order to process this agent. The agent consists of model, harness, tools and skills, and a runtime. All of that is running at different places in a data center. You can think of the model as the brain, the harness as the body. The tools that it uses, working in a runtime, think of it as a workshop. A round of applause for math. a round of applause for math Math is beautiful. math is beautiful The computing pattern of software is going to change. the computing pattern of software is going to change Let's come back to this. let's come back to this This is the agent. this is the agent It is the ultimate disaggregated and distributed computing model. it is the ultimate disaggregated and distributed computing model Many different computers are going to be activated in order to process this agent. many different computers are going to be activated in order to process this agent The agent consists of model, harness, tools and skills, and a runtime. the agent consists of model harness tools and skills and a runtime All of that is running at different places in a data center. all of that is running at different places in a data center You can think of the model as the brain, the harness as the body. you can think of the model as the brain the harness as the body The tools that it uses, working in a runtime, think of it as a workshop. the tools that it uses working in a runtime think of it as a workshop This is a person, a worker, working with tools in a workshop. Of course, this is being done at extraordinarily large scales, and each one of those steps are running in a different part of the computer. You could see the large language model is thinking, context processing, observing, understanding the environment, reasoning, coming up with a plan, and acting on the plan. This is a person, a worker, working with tools in a workshop. this is a person a worker working with tools in a workshop Of course, this is being done at extraordinarily large scales, and each one of those steps are running in a different part of the computer. of course this is being done at extraordinarily large scales and each one of those steps are running in a different part of the computer You could see the large language model is thinking, context processing, observing, understanding the environment, reasoning, coming up with a plan, and acting on the plan. you could see the large language model is thinking context processing observing understanding the environment reasoning coming up with a plan and acting on the plan Every single time that happens, an entire rack of Grace Blackwell NVL72 is activated. It's thinking with a large language model. Whenever it uses a tool, a CPU is used. That tool could be a C compiler, it could be Python, it could be JavaScript, or it could be accelerated computing. Today's agents are relatively simple users of tools. Tomorrow they're going to be very sophisticated users of tools, which is the reason why the CUDA-X libraries that I showed you are going to be incredibly popular with agents. They solve some of the most important problems the world knows, and all of our CUDA-X libraries are now going to come with skills that the AI could learn how to use. Every single time that happens, an entire rack of Grace Blackwell NVL72 is activated. every single time that happens an entire rack of grace blackwell nvl72 is activated It's thinking with a large language model. it's thinking with a large language model Whenever it uses a tool, a CPU is used. whenever it uses a tool a cpu is used That tool could be a C compiler, it could be Python, it could be JavaScript, or it could be accelerated computing. that tool could be a c compiler it could be python it could be javascript or it could be accelerated computing Today's agents are relatively simple users of tools. today's agents are relatively simple users of tools Tomorrow they're going to be very sophisticated users of tools, which is the reason why the CUDA-X libraries that I showed you are going to be incredibly popular with agents. tomorrow they're going to be very sophisticated users of tools which is the reason why the cuda-x libraries that i showed you are going to be incredibly popular with agents They solve some of the most important problems the world knows, and all of our CUDA-X libraries are now going to come with skills that the AI could learn how to use. they solve some of the most important problems the world knows and all of our cuda-x libraries are now going to come with skills that the ai could learn how to use The CUDA-X library, some skills, basically a manual, the AI reads it and go, "Uh-huh, that's how you use it." The ability to use these libraries by agents are going to be incredible. The tools run on CPUs and GPUs and large language models. The security harness runs on CPUs and a security processor called a DPU, NVIDIA's BlueField. The orchestration of all this runs on a CPU. This is the entire harness, and the CPU is orchestrating all of the work. One of the hardest parts is memory. You could just imagine the working memory is called KV Caching. What to remember, compaction, not just compression, but how to retrieve. Do you retrieve structured data? Do you retrieve unstructured data? What is the ontology, the relationship of all of these different data to itself? That entire processing is incredibly complicated. The CUDA-X library, some skills, basically a manual, the AI reads it and go, "Uh-huh, that's how you use it." The ability to use these libraries by agents are going to be incredible. the cuda-x library some skills basically a manual the ai reads it and go "uh-huh that's how you use it." the ability to use these libraries by agents are going to be incredible The tools run on CPUs and GPUs and large language models. the tools run on cpus and gpus and large language models The security harness runs on CPUs and a security processor called a DPU, NVIDIA's BlueField. the security harness runs on cpus and a security processor called a dpu nvidia's bluefield The orchestration of all this runs on a CPU. the orchestration of all this runs on a cpu This is the entire harness, and the CPU is orchestrating all of the work. this is the entire harness and the cpu is orchestrating all of the work One of the hardest parts is memory. one of the hardest parts is memory You could just imagine the working memory is called KV Caching. you could just imagine the working memory is called kv caching What to remember, compaction, not just compression, but how to retrieve. what to remember compaction not just compression but how to retrieve Do you retrieve structured data? do you retrieve structured data Do you retrieve unstructured data? do you retrieve unstructured data What is the ontology, the relationship of all of these different data to itself? what is the ontology the relationship of all of these different data to itself That entire processing is incredibly complicated. that entire processing is incredibly complicated The memory system of AIs is going to cause the storage system to be completely revolutionized. As you could see, every aspect of this computing model, this computing pattern, this new application called an agent, is fundamentally different than the way that applications used to run. A whole bunch of software sitting inside a binary, sitting inside an operating system. This is the reason, this disaggregated, this distributed, this heterogeneous computing problem is precisely the reason we built our next generation, Vera Rubin. Vera Rubin is not one chip. Vera Rubin is not a GPU only, it's starts with a GPU. Vera Rubin is incredible. This entire thing is Vera Rubin from end to end. It has GPUs, Vera Rubin NVLink 72. It is orchestrated by Vera CPUs that I'm going to tell you more about. The storage systems, revolutionary. The memory system of AIs is going to cause the storage system to be completely revolutionized. the memory system of ais is going to cause the storage system to be completely revolutionized As you could see, every aspect of this computing model, this computing pattern, this new application called an agent, is fundamentally different than the way that applications used to run. as you could see every aspect of this computing model this computing pattern this new application called an agent is fundamentally different than the way that applications used to run A whole bunch of software sitting inside a binary, sitting inside an operating system. a whole bunch of software sitting inside a binary sitting inside an operating system This is the reason, this disaggregated, this distributed, this heterogeneous computing problem is precisely the reason we built our next generation, Vera Rubin. this is the reason this disaggregated this distributed this heterogeneous computing problem is precisely the reason we built our next generation vera rubin Vera Rubin is not one chip. vera rubin is not one chip Vera Rubin is not a GPU only, it's starts with a GPU. vera rubin is not a gpu only, it's starts with a gpu Vera Rubin is incredible. vera rubin is incredible This entire thing is Vera Rubin from end to end. this entire thing is vera rubin from end to end It has GPUs, Vera Rubin NVLink 72. it has gpus vera rubin nvlink 72 It is orchestrated by Vera CPUs that I'm going to tell you more about. it is orchestrated by vera cpus that i'm going to tell you more about The storage systems, revolutionary. the storage systems revolutionary Vera, along with CX-9, our software stack called DOCA, the security processor that's inside so that everything is encrypted at rest, in motion, as well as in use. Everything across this is secure because the AI model is so precious. This is the reason why this entire system obeys confidential computing. Each one of these systems would be a complete revolution in itself. Vera Rubin is the most ambitious endeavor in the history of our company. The whole company worked on Vera Rubin across all 40,000 engineers. Not to mention all of you. All of you participated in the creation of this entire system. Vera Rubin is really a miracle, and it's not just one chip, it is so many. Well, it's even beyond that. A long time ago, NVIDIA used to be a GPU company. Over the years, we've evolved to become a systems company. Vera, along with CX-9, our software stack called DOCA, the security processor that's inside so that everything is encrypted at rest, in motion, as well as in use. vera along with cx-9 our software stack called doca the security processor that's inside so that everything is encrypted at rest in motion as well as in use Everything across this is secure because the AI model is so precious. everything across this is secure because the ai model is so precious This is the reason why this entire system obeys confidential computing. this is the reason why this entire system obeys confidential computing Each one of these systems would be a complete revolution in itself. each one of these systems would be a complete revolution in itself Vera Rubin is the most ambitious endeavor in the history of our company. vera rubin is the most ambitious endeavor in the history of our company The whole company worked on Vera Rubin across all 40,000 engineers. the whole company worked on vera rubin across all 40,000 engineers Not to mention all of you. not to mention all of you All of you participated in the creation of this entire system. all of you participated in the creation of this entire system Vera Rubin is really a miracle, and it's not just one chip, it is so many. vera rubin is really a miracle and it's not just one chip it is so many Well, it's even beyond that. well it's even beyond that A long time ago, NVIDIA used to be a GPU company. a long time ago nvidia used to be a gpu company Over the years, we've evolved to become a systems company. over the years we've evolved to become a systems company You're looking here now for the most complex system, most complex and ground-up system ever designed. Ultimately, our customers, our partners, don't want to buy computers, they want to build AI factories. The reason why NVIDIA has really started to transform ourself yet again. You could see so much of our technology is now at the entire infrastructure scale. Our partners are at infrastructure scale. Power generators, cooling systems, the grid providers. Many industrial companies are now part of our ecosystem because ultimately, we're trying to build an entire stack just like GPUs, just like when we were building Grace Blackwell NVL72, just like now, we are building a full stack system so that our customers could build amazing AI infrastructure. Let's take a look. You're looking here now for the most complex system, most complex and ground-up system ever designed. you're looking here now for the most complex system most complex and ground-up system ever designed Ultimately, our customers, our partners, don't want to buy computers, they want to build AI factories. ultimately our customers our partners don't want to buy computers they want to build ai factories The reason why NVIDIA has really started to transform ourself yet again. the reason why nvidia has really started to transform ourself yet again You could see so much of our technology is now at the entire infrastructure scale. you could see so much of our technology is now at the entire infrastructure scale Our partners are at infrastructure scale. our partners are at infrastructure scale Power generators, cooling systems, the grid providers. power generators cooling systems the grid providers Many industrial companies are now part of our ecosystem because ultimately, we're trying to build an entire stack just like GPUs, just like when we were building Grace Blackwell NVL72, just like now, we are building a full stack system so that our customers could build amazing AI infrastructure. many industrial companies are now part of our ecosystem because ultimately we're trying to build an entire stack just like gpus just like when we were building grace blackwell nvl72 just like now we are building a full stack system so that our customers could build amazing ai infrastructure Let's take a look. let's take a look

Speaker 3: The world is racing to build AI factories, the largest infrastructure build-out in human history. AI factories are incredibly complex. Every layer, chip, rack, network, power, cooling, grid, must be designed together from end to end, because compute is revenues. NVIDIA DSX is the blueprint, a reference design for building and operating AI factories at maximum efficiency and profitability. It starts with DSX Sim. With the DSX Sim Omniverse Blueprint, partners design and validate an NVIDIA Vera Rubin AI factory before a single rack lands. They plan the layout. Simulate the power and cooling. Design the network. Validate every integration. Test every change in the digital twin. The factory powers on. DSX OS takes over and provisions, operates, monitors, and remediates the infrastructure, turning the installed systems into trusted, multi-tenant, resilient, AI-ready capacity. Today's AI factories over-provision power by up to 40%. The world is racing to build AI factories, the largest infrastructure build-out in human history. the world is racing to build ai factories the largest infrastructure build-out in human history AI factories are incredibly complex. ai factories are incredibly complex Every layer, chip, rack, network, power, cooling, grid, must be designed together from end to end, because compute is revenues. every layer chip rack network power cooling grid must be designed together from end to end because compute is revenues NVIDIA DSX is the blueprint, a reference design for building and operating AI factories at maximum efficiency and profitability. nvidia dsx is the blueprint a reference design for building and operating ai factories at maximum efficiency and profitability It starts with DSX Sim. it starts with dsx sim With the DSX Sim Omniverse Blueprint, partners design and validate an NVIDIA Vera Rubin AI factory before a single rack lands. with the dsx sim omniverse blueprint partners design and validate an nvidia vera rubin ai factory before a single rack lands They plan the layout. Simulate the power and cooling. they plan the layout. simulate the power and cooling Design the network. design the network Validate every integration. validate every integration Test every change in the digital twin. test every change in the digital twin The factory powers on. the factory powers on DSX OS takes over and provisions, operates, monitors, and remediates the infrastructure, turning the installed systems into trusted, multi-tenant, resilient, AI-ready capacity. dsx os takes over and provisions operates monitors and remediates the infrastructure turning the installed systems into trusted multi-tenant resilient ai-ready capacity Today's AI factories over-provision power by up to 40%. today's ai factories over-provision power by up to 40% DSX MaxLPS lets operators safely deploy more GPUs inside the same power budget, adding billions in annual revenue. Breakthrough hot liquid cooling at 45 degrees Celsius uses less water and energy. More power going to revenue-generating compute. Incredible. Dynamic power allocation steers power from rack to rack, recovering stranded watts, sending them where work is happening. In-rack power smoothing flattens peak current spikes and power surges. Throughout the factory, teams of AI agents work with DSX MaxLPS, continuously coordinating to balance cooling and power to meet workload demand. DSX AI factories are flexible energy assets that operate cooperatively with the grid. DSX Flex reads real-time grid signals and dynamically adjusts factory power when the grid needs relief. 100 GW of AI factories will come online before the end of the decade. NVIDIA DSX AI factories run at highest efficiency, produce the lowest-cost tokens, and make the grid stronger. DSX MaxLPS lets operators safely deploy more GPUs inside the same power budget, adding billions in annual revenue. dsx maxlps lets operators safely deploy more gpus inside the same power budget adding billions in annual revenue Breakthrough hot liquid cooling at 45 degrees Celsius uses less water and energy. breakthrough hot liquid cooling at 45 degrees celsius uses less water and energy More power going to revenue-generating compute. more power going to revenue-generating compute Incredible. incredible Dynamic power allocation steers power from rack to rack, recovering stranded watts, sending them where work is happening. dynamic power allocation steers power from rack to rack recovering stranded watts sending them where work is happening In-rack power smoothing flattens peak current spikes and power surges. in-rack power smoothing flattens peak current spikes and power surges Throughout the factory, teams of AI agents work with DSX MaxLPS, continuously coordinating to balance cooling and power to meet workload demand. throughout the factory teams of ai agents work with dsx maxlps continuously coordinating to balance cooling and power to meet workload demand DSX AI factories are flexible energy assets that operate cooperatively with the grid. dsx ai factories are flexible energy assets that operate cooperatively with the grid DSX Flex reads real-time grid signals and dynamically adjusts factory power when the grid needs relief. 100 GW of AI factories will come online before the end of the decade. dsx flex reads real-time grid signals and dynamically adjusts factory power when the grid needs relief 100 gw of ai factories will come online before the end of the decade NVIDIA DSX AI factories run at highest efficiency, produce the lowest-cost tokens, and make the grid stronger. nvidia dsx ai factories run at highest efficiency produce the lowest-cost tokens and make the grid stronger

Speaker 1: I've shown you ecosystem slides of the past, where NVIDIA's computing layers and software and software and computing stacks are integrated into other people's platforms, third-party platforms and libraries that serves end markets. That was a computing ecosystem. This is an AI factory ecosystem. This is way downstream of all of you. Upstream of me is all of you, and downstream of us is this ecosystem. NVIDIA ultimately is not just building a GPU, not just building a system. We're helping customers build these AI factories, these AI infrastructure that is so immensely complex. Each one of these at 1 GW level started at $20 billion, $30 billion. It is at $50 billion, $60 billion, and soon it will be $80 billion, $100 billion per gigawatt. $100 billion into an AI factory. It must work the first time, and it must work right away. I've shown you ecosystem slides of the past, where NVIDIA's computing layers and software and software and computing stacks are integrated into other people's platforms, third-party platforms and libraries that serves end markets. i've shown you ecosystem slides of the past where nvidia's computing layers and software and software and computing stacks are integrated into other people's platforms third-party platforms and libraries that serves end markets That was a computing ecosystem. that was a computing ecosystem This is an AI factory ecosystem. this is an ai factory ecosystem This is way downstream of all of you. this is way downstream of all of you Upstream of me is all of you, and downstream of us is this ecosystem. upstream of me is all of you and downstream of us is this ecosystem NVIDIA ultimately is not just building a GPU, not just building a system. nvidia ultimately is not just building a gpu not just building a system We're helping customers build these AI factories, these AI infrastructure that is so immensely complex. we're helping customers build these ai factories these ai infrastructure that is so immensely complex Each one of these at 1 GW level started at $20 billion, $30 billion. each one of these at 1 gw level started at $20 billion $30 billion It is at $50 billion, $60 billion, and soon it will be $80 billion, $100 billion per gigawatt. $100 billion into an AI factory. it is at $50 billion $60 billion and soon it will be $80 billion $100 billion per gigawatt $100 billion into an ai factory It must work the first time, and it must work right away. it must work the first time and it must work right away The cost of capital is incredible. The complexity is incredible. As you see, we used to design a chip inside a computer, and then we simulated a system inside a computer. Today, you saw just now everything was built in Omniverse. I've been working with Omniverse with all of you for a long time. This was the dream come true, so that we can build these gigantic systems as large as the world wants to build inside a digital framework, inside a digital simulator, in a digital world, long before we build the first break ground and put our money to work. This is our ecosystem. We call it DSX. RTX is for our GPU, DGX is for our systems, and now DSX, basically infrastructure. The cost of capital is incredible. the cost of capital is incredible The complexity is incredible. the complexity is incredible As you see, we used to design a chip inside a computer, and then we simulated a system inside a computer. as you see we used to design a chip inside a computer and then we simulated a system inside a computer Today, you saw just now everything was built in Omniverse. today you saw just now everything was built in omniverse I've been working with Omniverse with all of you for a long time. i've been working with omniverse with all of you for a long time This was the dream come true, so that we can build these gigantic systems as large as the world wants to build inside a digital framework, inside a digital simulator, in a digital world, long before we build the first break ground and put our money to work. this was the dream come true so that we can build these gigantic systems as large as the world wants to build inside a digital framework inside a digital simulator in a digital world long before we build the first break ground and put our money to work This is our ecosystem. this is our ecosystem We call it DSX. we call it dsx RTX is for our GPU, DGX is for our systems, and now DSX, basically infrastructure. rtx is for our gpu dgx is for our systems and now dsx basically infrastructure Because of the work that we do here across this entire stack, including our systems and software, it's the reason why we can work with small companies and enable them to be world-class AI clouds. Every one of these I'm about to show you are small companies just recently, now CoreWeave is worth $50 billion, $60 billion, $70 billion and growing incredibly fast. Recently, we worked with Nebius, again, they're growing incredibly fast. Each one of these clouds have incredible customers. Cursor, the software coding company, Black Forest Labs, Image Generation, World Labs, World Foundation Model, Revolut, the leading financial services AI company, Shopify. Here's another one. This is Nscale, their customers are British Telecom, Google. Google is using one of our AI clouds. Thinking Machines, a Frontier Labs company. We're super excited. Here's NAVER Cloud in Korea, Bank of Korea, Hyundai. Many incredible companies. Because of the work that we do here across this entire stack, including our systems and software, it's the reason why we can work with small companies and enable them to be world-class AI clouds. because of the work that we do here across this entire stack including our systems and software it's the reason why we can work with small companies and enable them to be world-class ai clouds Every one of these I'm about to show you are small companies just recently, now CoreWeave is worth $50 billion, $60 billion, $70 billion and growing incredibly fast. every one of these i'm about to show you are small companies just recently now coreweave is worth $50 billion $60 billion $70 billion and growing incredibly fast Recently, we worked with Nebius, again, they're growing incredibly fast. recently we worked with nebius again they're growing incredibly fast Each one of these clouds have incredible customers. each one of these clouds have incredible customers Cursor, the software coding company, Black Forest Labs, Image Generation, World Labs, World Foundation Model, Revolut, the leading financial services AI company, Shopify. cursor the software coding company black forest labs image generation world labs world foundation model revolut the leading financial services ai company shopify Here's another one. here's another one This is Nscale, their customers are British Telecom, Google. this is nscale their customers are british telecom google Google is using one of our AI clouds. google is using one of our ai clouds Thinking Machines, a Frontier Labs company. thinking machines a frontier labs company We're super excited. we're super excited Here's NAVER Cloud in Korea, Bank of Korea, Hyundai. here's naver cloud in korea bank of korea hyundai Many incredible companies. many incredible companies Here's one in India, Yotta. Incredible companies. Here's one based in Singapore, building in Australia, Together AI Singapore. This is one in Indonesia. Each one of these companies are serving regional as well as global customers. AI is going to run everywhere. Every company will be powered by it. Every region will build it. Indosat here in Indonesia. Here in Taiwan, GMI. Here in Taiwan, GMI. It's okay to clap. Incredible companies, incredible opportunity, but all of them need several things. Of course, they need the computing stack, this entire stack underneath. This is what made NVIDIA famous. All of our hardware and software and libraries, our connection into the world's ecosystem of third-party developers makes it possible for anyone to stand up an AI cloud. However, the AI cloud is so complex now. This is the software version. This is the computer science version. Here's one in India, Yotta. here's one in india yotta Incredible companies. incredible companies Here's one based in Singapore, building in Australia, Together AI Singapore. here's one based in singapore building in australia together ai singapore This is one in Indonesia. this is one in indonesia Each one of these companies are serving regional as well as global customers. each one of these companies are serving regional as well as global customers AI is going to run everywhere. ai is going to run everywhere Every company will be powered by it. every company will be powered by it Every region will build it. every region will build it Indosat here in Indonesia. indosat here in indonesia Here in Taiwan, GMI. here in taiwan gmi Here in Taiwan, GMI. here in taiwan gmi It's okay to clap. it's okay to clap Incredible companies, incredible opportunity, but all of them need several things. incredible companies incredible opportunity but all of them need several things Of course, they need the computing stack, this entire stack underneath. of course they need the computing stack this entire stack underneath This is what made NVIDIA famous. this is what made nvidia famous All of our hardware and software and libraries, our connection into the world's ecosystem of third-party developers makes it possible for anyone to stand up an AI cloud. all of our hardware and software and libraries our connection into the world's ecosystem of third-party developers makes it possible for anyone to stand up an ai cloud However, the AI cloud is so complex now. however the ai cloud is so complex now This is the software version. this is the software version This is the computer science version. this is the computer science version The money version, the asset version is what I showed you earlier. It's a giant factory. Having this ability alone is not enough, which is the reason why NVIDIA has become an AI infrastructure company. Now, doing this well and becoming incredibly good at helping customers build AI factories and deploying AI factories is incredibly important, and the reason for that is this. Compute is revenue now. Compute is profit. The absence of revenues and profit is loss. It's really important to realize that this is an example of an AI infrastructure coming online. It could be coming online quickly, it could take a while, its throughput could be high, it could be low, its resilience or reliability could be good or bad, and its lifetime of usefulness could be long or short. The money version, the asset version is what I showed you earlier. the money version the asset version is what i showed you earlier It's a giant factory. it's a giant factory Having this ability alone is not enough, which is the reason why NVIDIA has become an AI infrastructure company. having this ability alone is not enough which is the reason why nvidia has become an ai infrastructure company Now, doing this well and becoming incredibly good at helping customers build AI factories and deploying AI factories is incredibly important, and the reason for that is this. now doing this well and becoming incredibly good at helping customers build ai factories and deploying ai factories is incredibly important and the reason for that is this Compute is revenue now. compute is revenue now Compute is profit. compute is profit The absence of revenues and profit is loss. the absence of revenues and profit is loss It's really important to realize that this is an example of an AI infrastructure coming online. it's really important to realize that this is an example of an ai infrastructure coming online It could be coming online quickly, it could take a while, its throughput could be high, it could be low, its resilience or reliability could be good or bad, and its lifetime of usefulness could be long or short. it could be coming online quickly it could take a while its throughput could be high it could be low its resilience or reliability could be good or bad and its lifetime of usefulness could be long or short This represents $50 billion, $60 billion, going to $100 billion, this curve matters greatly, which is the reason why NVIDIA is such a great partner. Working with us because of our fully integrated capability, we didn't just come up with a PowerPoint slide, we created the entire infrastructure. We connected everything together. We built out billions and billions of it ourselves to make sure that everything works well. As a result of that, our time to first token, our time to first inference, our time to training turned on is much faster. Second, because our throughput per watt, our tokens per watt is utterly world-class, and the reason for that is because we integrate everything, we design everything from the ground up, we simulate the entire system, and we use extreme co-design. This represents $50 billion, $60 billion, going to $100 billion, this curve matters greatly, which is the reason why NVIDIA is such a great partner. this represents $50 billion $60 billion going to $100 billion this curve matters greatly which is the reason why nvidia is such a great partner Working with us because of our fully integrated capability, we didn't just come up with a PowerPoint slide, we created the entire infrastructure. working with us because of our fully integrated capability we didn't just come up with a powerpoint slide we created the entire infrastructure We connected everything together. we connected everything together We built out billions and billions of it ourselves to make sure that everything works well. we built out billions and billions of it ourselves to make sure that everything works well As a result of that, our time to first token, our time to first inference, our time to training turned on is much faster. as a result of that our time to first token our time to first inference our time to training turned on is much faster Second, because our throughput per watt, our tokens per watt is utterly world-class, and the reason for that is because we integrate everything, we design everything from the ground up, we simulate the entire system, and we use extreme co-design. second because our throughput per watt our tokens per watt is utterly world-class and the reason for that is because we integrate everything we design everything from the ground up we simulate the entire system and we use extreme co-design Just like I showed you just now with the Vera Rubin rack, everything was designed in order to deliver on this incredible throughput. If your data center, if your factory has 1 GW, it will not have more. 1 GW means 1 GW. That's all the power generation you could do. If you have 1 GW of power, then throughput per watt is revenues because every token is profitable. Every token is revenues. This is the future. Compute is revenues. Performance per watt is your revenues. Choosing the wrong architecture just because the chips are cheaper doesn't translate, doesn't make sense. You need to make sure that your revenues per watt, the more you buy, the more you make. Tokens per watt. Third is reliability. If you ever get a chance to see these data centers, there are so many moving parts, millions of cables. Just like I showed you just now with the Vera Rubin rack, everything was designed in order to deliver on this incredible throughput. just like i showed you just now with the vera rubin rack everything was designed in order to deliver on this incredible throughput If your data center, if your factory has 1 GW , it will not have more. 1 GW means 1 GW . if your data center if your factory has 1 gw it will not have more 1 gw means 1 gw That's all the power generation you could do. that's all the power generation you could do If you have 1 GW of power, then throughput per watt is revenues because every token is profitable. if you have 1 gw of power then throughput per watt is revenues because every token is profitable Every token is revenues. every token is revenues This is the future. this is the future Compute is revenues. compute is revenues Performance per watt is your revenues. performance per watt is your revenues Choosing the wrong architecture just because the chips are cheaper doesn't translate, doesn't make sense. choosing the wrong architecture just because the chips are cheaper doesn't translate doesn't make sense You need to make sure that your revenues per watt, the more you buy, the more you make. you need to make sure that your revenues per watt the more you buy the more you make Tokens per watt. tokens per watt Third is reliability. third is reliability If you ever get a chance to see these data centers, there are so many moving parts, millions of cables. if you ever get a chance to see these data centers there are so many moving parts millions of cables The ability for all of those computers to work harmoniously, reliably is extremely low. It is just extremely difficult. We have now been operating very large scale for a very long time. That experience matters. That difference, mean time between interrupts, extremely important. Lastly, this is very hard. The lifetime of these systems, the software is changing all the time. four years ago, which is in the time of Hopper, AI has completely changed. six years ago, this is the timeframe of Ampere, AI has completely changed. We started out talking about CNNs. Here we are, then we talked about transformers, and then we talked about Mixture of Experts. Now we're talking about agentic systems. Every single generation, every single few months, the software industry is coming up with new technology. The ability for all of those computers to work harmoniously, reliably is extremely low. the ability for all of those computers to work harmoniously reliably is extremely low It is just extremely difficult. it is just extremely difficult We have now been operating very large scale for a very long time. we have now been operating very large scale for a very long time That experience matters. that experience matters That difference, mean time between interrupts, extremely important. that difference mean time between interrupts extremely important Lastly, this is very hard. lastly this is very hard The lifetime of these systems, the software is changing all the time. four years ago, which is in the time of Hopper, AI has completely changed. six years ago, this is the timeframe of Ampere, AI has completely changed. the lifetime of these systems the software is changing all the time four years ago which is in the time of hopper ai has completely changed six years ago this is the timeframe of ampere ai has completely changed We started out talking about CNNs. we started out talking about cnns Here we are, then we talked about transformers, and then we talked about Mixture of Experts. here we are then we talked about transformers and then we talked about mixture of experts Now we're talking about agentic systems. now we're talking about agentic systems Every single generation, every single few months, the software industry is coming up with new technology. every single generation every single few months the software industry is coming up with new technology If your architecture is not flexible, if your ecosystem is not rich, then this curve cannot be long. You cannot predict how long your system can last. I can. NVIDIA systems is all over the world. Software developers start with NVIDIA CUDA, by definition, therefore, the life, the ecosystem, the useful asset is going to be much longer. The difference is essentially cost. You could think of it as revenues, the other side of revenues is cost. If the life of the asset is long, the TCO is low. This is the difference. This is what it looks like when compute. The more you buy, the more you make. If your architecture is not flexible, if your ecosystem is not rich, then this curve cannot be long. if your architecture is not flexible if your ecosystem is not rich then this curve cannot be long You cannot predict how long your system can last. you cannot predict how long your system can last I can. i can NVIDIA systems is all over the world. nvidia systems is all over the world Software developers start with NVIDIA CUDA, by definition, therefore, the life, the ecosystem, the useful asset is going to be much longer. software developers start with nvidia cuda by definition therefore the life the ecosystem the useful asset is going to be much longer The difference is essentially cost. the difference is essentially cost You could think of it as revenues, the other side of revenues is cost. you could think of it as revenues the other side of revenues is cost If the life of the asset is long, the TCO is low. if the life of the asset is long the tco is low This is the difference. this is the difference This is what it looks like when compute. this is what it looks like when compute The more you buy, the more you make. the more you buy the more you make All of you are experiencing this with me. Isn't that right? All of your demand, your factories are working so hard. Your people are working so hard all across Taiwan because everybody wants to make money. They realize that useful AI is here. Profitable AI is here. Compute demand is incredibly high, and compute demand is the constraint. Let's go work super hard and help the world stand up AI factories everywhere. This is why it's so important. I'm so happy. Here I am standing in front of you. Vera Rubin is in full production. Vera Rubin is in full production. The supply chain we created for Vera Rubin is twice as large as Grace Blackwell. Yeah, it's incredible. What used to take two hours to assemble one Grace Blackwell rack now only takes five minutes. All of you are experiencing this with me. all of you are experiencing this with me Isn't that right? isn't that right All of your demand, your factories are working so hard. all of your demand your factories are working so hard Your people are working so hard all across Taiwan because everybody wants to make money. your people are working so hard all across taiwan because everybody wants to make money They realize that useful AI is here. they realize that useful ai is here Profitable AI is here. profitable ai is here Compute demand is incredibly high, and compute demand is the constraint. compute demand is incredibly high and compute demand is the constraint Let's go work super hard and help the world stand up AI factories everywhere. let's go work super hard and help the world stand up ai factories everywhere This is why it's so important. this is why it's so important I'm so happy. i'm so happy Here I am standing in front of you. here i am standing in front of you Vera Rubin is in full production. vera rubin is in full production Vera Rubin is in full production. vera rubin is in full production The supply chain we created for Vera Rubin is twice as large as Grace Blackwell. the supply chain we created for vera rubin is twice as large as grace blackwell Yeah, it's incredible. yeah it's incredible What used to take two hours to assemble one Grace Blackwell rack now only takes five minutes. what used to take two hours to assemble one grace blackwell rack now only takes five minutes Not only is the capacity higher, the throughput is a lot faster, and we need it all to support the demand. This ecosystem is extraordinary. Millions of square feet has been put online to support Grace Blackwell and preparing now, ramping up now, Vera Rubin. I want to thank all of you. Vera Rubin is now in full production. Thank you. Let's take a look. Not only is the capacity higher, the throughput is a lot faster, and we need it all to support the demand. not only is the capacity higher the throughput is a lot faster and we need it all to support the demand This ecosystem is extraordinary. this ecosystem is extraordinary Millions of square feet has been put online to support Grace Blackwell and preparing now, ramping up now, Vera Rubin. millions of square feet has been put online to support grace blackwell and preparing now ramping up now vera rubin I want to thank all of you. i want to thank all of you Vera Rubin is now in full production. vera rubin is now in full production Thank you. thank you Let's take a look. let's take a look

Speaker 3: Large language models generate answers. AI agents can do work. Processing agentic AI is a whole different kind of problem. Agents observe, reason, plan, use tools. They manage massive context, juggling working memory and long-term memory. They spin up sub-agents, specialists on demand. NVIDIA Vera Rubin is a multi-rack pod scale system built to process agentic AI and is now in full production. The manufacturing, automation, and orchestration across the supply chain, a miracle to witness. Large language models generate answers. large language models generate answers AI agents can do work. ai agents can do work Processing agentic AI is a whole different kind of problem. processing agentic ai is a whole different kind of problem Agents observe, reason, plan, use tools. agents observe reason plan use tools They manage massive context, juggling working memory and long-term memory. they manage massive context juggling working memory and long-term memory They spin up sub-agents, specialists on demand. they spin up sub-agents specialists on demand NVIDIA Vera Rubin is a multi-rack pod scale system built to process agentic AI and is now in full production. nvidia vera rubin is a multi-rack pod scale system built to process agentic ai and is now in full production The manufacturing, automation, and orchestration across the supply chain, a miracle to witness. the manufacturing automation and orchestration across the supply chain a miracle to witness Our journey started when we launched the first AI supercomputer, NVIDIA DGX-1. Over the next decade, we pushed every chip and system to the limit, from Pascal and the first NVLink to Grace Blackwell, the first rack-scale AI supercomputer. Now, Vera Rubin, the first multi-rack pod scale supercomputer built for the agentic age. It starts at TSMC. The seven new chips that make up Vera Rubin take shape through hundreds of processing steps. Three nanometer process, CoWoS-R and CoWoS-L packaging, HBM4 memory from Micron, SK hynix and Samsung. The Vera Rubin Compute Board, six trillion transistors with over 18,000 components on one board. Vera Rubin NVL72 does the thinking, prompt and context understanding, reasoning, and planning. Next, a new modular compute tray, streamlined with a new PCB mid-plane design. Superchips, ConnectX-9 SuperNICs, and BlueField-4 DPUs all mate in place with no cables for resiliency at AI factory scale. Our journey started when we launched the first AI supercomputer, NVIDIA DGX-1. our journey started when we launched the first ai supercomputer nvidia dgx-1 Over the next decade, we pushed every chip and system to the limit, from Pascal and the first NVLink to Grace Blackwell, the first rack-scale AI supercomputer. over the next decade we pushed every chip and system to the limit from pascal and the first nvlink to grace blackwell the first rack-scale ai supercomputer Now, Vera Rubin, the first multi-rack pod scale supercomputer built for the agentic age. now vera rubin the first multi-rack pod scale supercomputer built for the agentic age It starts at TSMC. it starts at tsmc The seven new chips that make up Vera Rubin take shape through hundreds of processing steps. the seven new chips that make up vera rubin take shape through hundreds of processing steps Three nanometer process, CoWoS-R and CoWoS-L packaging, HBM4 memory from Micron, SK hynix and Samsung. three nanometer process cowos-r and cowos-l packaging hbm4 memory from micron sk hynix and samsung The Vera Rubin Compute Board, six trillion transistors with over 18,000 components on one board. the vera rubin compute board six trillion transistors with over 18,000 components on one board Vera Rubin NVL72 does the thinking, prompt and context understanding, reasoning, and planning. vera rubin nvl72 does the thinking prompt and context understanding reasoning and planning Next, a new modular compute tray, streamlined with a new PCB mid-plane design. next a new modular compute tray streamlined with a new pcb mid-plane design Superchips, ConnectX-9 SuperNICs, and BlueField-4 DPUs all mate in place with no cables for resiliency at AI factory scale. superchips connectx-9 supernics and bluefield-4 dpus all mate in place with no cables for resiliency at ai factory scale 18 compute trays, nine hot swappable NVLink switch trays. New high-efficiency manifolds, liquid-cooled busbars carrying over 5,000 amps, the equivalent of 20 electric cars at full acceleration. Together, 1.3 million components form this third-generation MGX rack design. Congratulations to Microsoft for their operational Vera Rubin NVL72 engineering rack. Congratulations to Dell and CoreWeave as well for standing up their Vera Rubin NVL72 engineering rack. The Vera CPU rack. 256 CPUs in a single liquid-cooled rack, orchestrating the models, shuffling memory, launching tools. At Foxconn and Quanta, Groq 3 LPX takes shape. 256 Groq 3 LPUs across 16 trays, 40 Pbps of SRAM bandwidth for ultra-low latency. While NVL72 generates tokens at the highest throughput, Groq LPX generates them at the lowest latency. Vera BlueField-4 STX, where AI keeps its memory. Storage processing accelerated by BlueField-4, connecting memory, storage, and in-silicon security. 18 compute trays, nine hot swappable NVLink switch trays. 18 compute trays nine hot swappable nvlink switch trays New high-efficiency manifolds, liquid-cooled busbars carrying over 5,000 amps, the equivalent of 20 electric cars at full acceleration. new high-efficiency manifolds liquid-cooled busbars carrying over 5,000 amps the equivalent of 20 electric cars at full acceleration Together, 1.3 million components form this third-generation MGX rack design. together 1.3 million components form this third-generation mgx rack design Congratulations to Microsoft for their operational Vera Rubin NVL72 engineering rack. congratulations to microsoft for their operational vera rubin nvl72 engineering rack Congratulations to Dell and CoreWeave as well for standing up their Vera Rubin NVL72 engineering rack. congratulations to dell and coreweave as well for standing up their vera rubin nvl72 engineering rack The Vera CPU rack. 256 CPUs in a single liquid-cooled rack, orchestrating the models, shuffling memory, launching tools. the vera cpu rack 256 cpus in a single liquid-cooled rack orchestrating the models shuffling memory launching tools At Foxconn and Quanta, Groq 3 LPX takes shape. 256 Groq 3 LPUs across 16 trays, 40 Pbps of SRAM bandwidth for ultra-low latency. at foxconn and quanta groq 3 lpx takes shape 256 groq 3 lpus across 16 trays 40 pbps of sram bandwidth for ultra-low latency While NVL72 generates tokens at the highest throughput, Groq LPX generates them at the lowest latency. while nvl72 generates tokens at the highest throughput groq lpx generates them at the lowest latency Vera BlueField-4 STX, where AI keeps its memory. vera bluefield-4 stx where ai keeps its memory Storage processing accelerated by BlueField-4, connecting memory, storage, and in-silicon security. storage processing accelerated by bluefield-4 connecting memory storage and in-silicon security NVIDIA Spectrum-X Ethernet Photonics, the world's first Ethernet switch with 200 Gb co-packaged optics, TSMC's COUPE process, chip scale packaging, and ultra-high-powered laser dies on indium phosphide. Vera Rubin, five connected rack scale systems, a supercomputer for AI agents. 150 supply chain partners across Taiwan. Millions of square feet of factory floor. Hundreds of sites, chips, packages, systems, and data centers pushed to the limits of size, power, and scale. This is what we call extreme co-design. We did this with Taiwan. Together, we reinvented computing for the age of AI. Taiwan was with us at the beginning and here today as we bring Vera Rubin to the world. Thank you, Taiwan. NVIDIA Spectrum-X Ethernet Photonics, the world's first Ethernet switch with 200 Gb co-packaged optics, TSMC's COUPE process, chip scale packaging, and ultra-high-powered laser dies on indium phosphide. nvidia spectrum-x ethernet photonics the world's first ethernet switch with 200 gb co-packaged optics tsmc's coupe process chip scale packaging and ultra-high-powered laser dies on indium phosphide Vera Rubin, five connected rack scale systems, a supercomputer for AI agents. 150 supply chain partners across Taiwan. vera rubin five connected rack scale systems a supercomputer for ai agents 150 supply chain partners across taiwan Millions of square feet of factory floor. millions of square feet of factory floor Hundreds of sites, chips, packages, systems, and data centers pushed to the limits of size, power, and scale. hundreds of sites chips packages systems and data centers pushed to the limits of size power and scale This is what we call extreme co-design. this is what we call extreme co-design We did this with Taiwan. we did this with taiwan Together, we reinvented computing for the age of AI. together we reinvented computing for the age of ai Taiwan was with us at the beginning and here today as we bring Vera Rubin to the world. taiwan was with us at the beginning and here today as we bring vera rubin to the world Thank you, Taiwan. thank you taiwan

Speaker 1: Ladies and gentlemen, Vera Rubin. Vera Rubin was not just built for AI. Vera Rubin was not built just to run AI. Vera Rubin was built to run agents. This is an agentic system. Imagine the complexity, which is the reason why agents is the last computer science breakthrough. It has taken this many years for agents to realize its potential and become useful. It stands to reason that the computer that runs it is the most advanced in the world. This is Vera Rubin. Let's take a look. Can we bring out Vera Rubin, please? Janine, do we have the racks, the systems? It looks heavy. This is Vera Rubin. Vera Rubin NVL72. This is the Groq LPX. At the next GTC, I'm going to talk to you about a lot more of this. Today, we have so much to talk to you about. Ladies and gentlemen, Vera Rubin. ladies and gentlemen vera rubin Vera Rubin was not just built for AI. vera rubin was not just built for ai Vera Rubin was not built just to run AI. vera rubin was not built just to run ai Vera Rubin was built to run agents. vera rubin was built to run agents This is an agentic system. this is an agentic system Imagine the complexity, which is the reason why agents is the last computer science breakthrough. imagine the complexity which is the reason why agents is the last computer science breakthrough It has taken this many years for agents to realize its potential and become useful. it has taken this many years for agents to realize its potential and become useful It stands to reason that the computer that runs it is the most advanced in the world. it stands to reason that the computer that runs it is the most advanced in the world This is Vera Rubin. this is vera rubin Let's take a look. let's take a look Can we bring out Vera Rubin, please? can we bring out vera rubin please Janine, do we have the racks, the systems? janine do we have the racks the systems It looks heavy. it looks heavy This is Vera Rubin. this is vera rubin Vera Rubin NVL72. vera rubin nvl72 This is the Groq LPX. this is the groq lpx At the next GTC, I'm going to talk to you about a lot more of this. at the next gtc i'm going to talk to you about a lot more of this Today, we have so much to talk to you about. today we have so much to talk to you about This is Vera CPU rack, 256 CPUs, all liquid-cooled. Let me tell you about Vera in just a moment. This is the Vera BlueField storage processing system and also security system. Of course, this is our Mellanox networking, the world's first CPO. This is Vera Rubin. Incredible technology all coming together. When we built Hopper, as you know, for pre-training. Pre-training was the most important application, the most important workload we were working on at the time. When we worked on Grace Blackwell, everybody said, "Jensen, NVIDIA is really good at pre-training. Inference is so easy." Do you remember that? People used to say, "Inference is so easy. We could do that, too." As you know, inference equals money, and the models, MoEs, are so complicated. This is Vera CPU rack, 256 CPUs, all liquid-cooled. this is vera cpu rack 256 cpus all liquid-cooled Let me tell you about Vera in just a moment. let me tell you about vera in just a moment This is the Vera BlueField storage processing system and also security system. this is the vera bluefield storage processing system and also security system Of course, this is our Mellanox networking, the world's first CPO. of course this is our mellanox networking the world's first cpo This is Vera Rubin. this is vera rubin Incredible technology all coming together. incredible technology all coming together When we built Hopper, as you know, for pre-training. when we built hopper as you know for pre-training Pre-training was the most important application, the most important workload we were working on at the time. pre-training was the most important application the most important workload we were working on at the time When we worked on Grace Blackwell, everybody said, "Jensen, NVIDIA is really good at pre-training. when we worked on grace blackwell everybody said "jensen nvidia is really good at pre-training Inference is so easy." Do you remember that? inference is so easy." do you remember that People used to say, "Inference is so easy. people used to say "inference is so easy We could do that, too." As you know, inference equals money, and the models, MoEs, are so complicated. we could do that too." as you know inference equals money and the models moes are so complicated To do it at incredibly high response time, fast interactivity, and high throughput at the same time is incredibly hard, which is the reason why we created NVLink 72. Today, NVIDIA's token cost is the lowest in the world, not by 10%, by X factors, orders of magnitude. All because we did extreme co-design, all because we understood the computing model, the computing pattern of inference, and we were able to create NVLink 72. With Vera Rubin, it is beyond inference. It is now inference in an agentic system. This is Vera Rubin. No cables, no hoses, no fans. What used to take the last time when I showed this to you, we had cables everywhere. The cables were amazing to look at. Now there's a PCB in the middle which connects both sides. What used to take two hours now takes five minutes. To do it at incredibly high response time, fast interactivity, and high throughput at the same time is incredibly hard, which is the reason why we created NVLink 72. to do it at incredibly high response time fast interactivity and high throughput at the same time is incredibly hard which is the reason why we created nvlink 72 Today, NVIDIA's token cost is the lowest in the world, not by 10%, by X factors, orders of magnitude. today nvidia's token cost is the lowest in the world not by 10% by x factors orders of magnitude All because we did extreme co-design, all because we understood the computing model, the computing pattern of inference, and we were able to create NVLink 72. all because we did extreme co-design all because we understood the computing model the computing pattern of inference and we were able to create nvlink 72 With Vera Rubin, it is beyond inference. with vera rubin it is beyond inference It is now inference in an agentic system. it is now inference in an agentic system This is Vera Rubin. this is vera rubin No cables, no hoses, no fans. no cables no hoses no fans What used to take the last time when I showed this to you, we had cables everywhere. what used to take the last time when i showed this to you we had cables everywhere The cables were amazing to look at. the cables were amazing to look at Now there's a PCB in the middle which connects both sides. now there's a pcb in the middle which connects both sides What used to take two hours now takes five minutes. what used to take two hours now takes five minutes The reliability and the resilience of Vera Rubin is going to be off the charts. This is our Vera CPU tray, the most advanced CPUs that has ever been built. I'm going to show you that in just a second. This is our storage tray. Two Vera CPUs, four CX9, incredible amounts of software. This is our new LPX, LPU 30, the Groq system, designed for very low latency inference. The throughput is delivered by Vera Rubin and extended with NVLink 72. If you want to extend that even further, you can have Groq LPUs. Here, we have the Vera Rubin NVLink, the switch tray. This is the switches in the middle, and this is revolutionary. Because of Vera Rubin's, because of NVLink 72 and the NVLink switches that we created and invented. This is our Ethernet switches for scale-out. The reliability and the resilience of Vera Rubin is going to be off the charts. the reliability and the resilience of vera rubin is going to be off the charts This is our Vera CPU tray, the most advanced CPUs that has ever been built. this is our vera cpu tray the most advanced cpus that has ever been built I'm going to show you that in just a second. i'm going to show you that in just a second This is our storage tray. this is our storage tray Two Vera CPUs, four CX9, incredible amounts of software. two vera cpus four cx9 incredible amounts of software This is our new LPX, LPU 30, the Groq system, designed for very low latency inference. this is our new lpx lpu 30 the groq system designed for very low latency inference The throughput is delivered by Vera Rubin and extended with NVLink 72. the throughput is delivered by vera rubin and extended with nvlink 72 If you want to extend that even further, you can have Groq LPUs. if you want to extend that even further you can have groq lpus Here, we have the Vera Rubin NVLink, the switch tray. here we have the vera rubin nvlink the switch tray This is the switches in the middle, and this is revolutionary. this is the switches in the middle and this is revolutionary Because of Vera Rubin's, because of NVLink 72 and the NVLink switches that we created and invented. because of vera rubin's because of nvlink 72 and the nvlink switches that we created and invented This is our Ethernet switches for scale-out. this is our ethernet switches for scale-out What's amazing is we introduced these two systems for Grace Blackwell. These two systems were created for Grace Blackwell. Today, NVIDIA is the largest networking company in the world. I'm so proud of the networking team. This is such an incredible enabler for everything that we do. I'm going to now talk to you about the next major industry we're going to be part of. Thank you. Janine. Thank you. I think there are 2,000 people back there pulling that. Let's talk about CPUs. Vera CPUs. CPUs built for the age of AI. All of the CPUs until now were created for people. We were the users, we were the renters. The way we use CPUs, we live in a world counted by seconds. What's amazing is we introduced these two systems for Grace Blackwell. what's amazing is we introduced these two systems for grace blackwell These two systems were created for Grace Blackwell. these two systems were created for grace blackwell Today, NVIDIA is the largest networking company in the world. today nvidia is the largest networking company in the world I'm so proud of the networking team. i'm so proud of the networking team This is such an incredible enabler for everything that we do. this is such an incredible enabler for everything that we do I'm going to now talk to you about the next major industry we're going to be part of. i'm going to now talk to you about the next major industry we're going to be part of Thank you. thank you Janine. janine Thank you. thank you I think there are 2,000 people back there pulling that. i think there are 2,000 people back there pulling that Let's talk about CPUs. let's talk about cpus Vera CPUs. vera cpus CPUs built for the age of AI. cpus built for the age of ai All of the CPUs until now were created for people. all of the cpus until now were created for people We were the users, we were the renters. we were the users we were the renters The way we use CPUs, we live in a world counted by seconds. the way we use cpus we live in a world counted by seconds The way we rent CPUs in the cloud, each one of them, more CPU cores you have, the more you can rent. The use case of the old CPU and the economics of the old CPU, fundamentally different than agents. Agents are impatient. They don't live in a world that is in seconds. They live in a world that's in nanoseconds. When it uses a tool, it wants the response time to be as fast as possible. When it access database, it has to come back as soon as possible. Every moment that the agent is waiting keeps it from going to the next step. It is vital that we make the CPUs as low latency as possible, as interactive as possible. We created Vera CPU for the age of AI. Now, inside our system, it's used for three different ways. The way we rent CPUs in the cloud, each one of them, more CPU cores you have, the more you can rent. the way we rent cpus in the cloud each one of them more cpu cores you have the more you can rent The use case of the old CPU and the economics of the old CPU, fundamentally different than agents. the use case of the old cpu and the economics of the old cpu fundamentally different than agents Agents are impatient. agents are impatient They don't live in a world that is in seconds. they don't live in a world that is in seconds They live in a world that's in nanoseconds. they live in a world that's in nanoseconds When it uses a tool, it wants the response time to be as fast as possible. when it uses a tool it wants the response time to be as fast as possible When it access database, it has to come back as soon as possible. when it access database it has to come back as soon as possible Every moment that the agent is waiting keeps it from going to the next step. every moment that the agent is waiting keeps it from going to the next step It is vital that we make the CPUs as low latency as possible, as interactive as possible. it is vital that we make the cpus as low latency as possible as interactive as possible We created Vera CPU for the age of AI. we created vera cpu for the age of ai Now, inside our system, it's used for three different ways. now inside our system it's used for three different ways The first way, of course, is Vera Rubin for thinking. Inside the Vera Rubin rack, there are already two CPUs. As you know, we are building and selling millions of Vera Rubins. We have sold millions of Grace Blackwells. NVIDIA already is one of the largest CPU makers in the world. In the Vera Rubin rack are two CPUs, one for orchestrating and managing the GPUs, managing the KV cache, dealing with all of the software that runs in the rack. We also have the Grace BlueField that is used for security and isolation. The Vera CPU is used for the harness, the orchestration of the AI models, tool use, accessing the database. The data servers are right here, Vera BlueField, the fastest storage servers, the fastest storage system the world has ever made. The first way, of course, is Vera Rubin for thinking. the first way of course is vera rubin for thinking Inside the Vera Rubin rack, there are already two CPUs. inside the vera rubin rack there are already two cpus As you know, we are building and selling millions of Vera Rubins. as you know we are building and selling millions of vera rubins We have sold millions of Grace Blackwells. we have sold millions of grace blackwells NVIDIA already is one of the largest CPU makers in the world. nvidia already is one of the largest cpu makers in the world In the Vera Rubin rack are two CPUs, one for orchestrating and managing the GPUs, managing the KV cache, dealing with all of the software that runs in the rack. in the vera rubin rack are two cpus one for orchestrating and managing the gpus managing the kv cache dealing with all of the software that runs in the rack We also have the Grace BlueField that is used for security and isolation. we also have the grace bluefield that is used for security and isolation The Vera CPU is used for the harness, the orchestration of the AI models, tool use, accessing the database. the vera cpu is used for the harness the orchestration of the ai models tool use accessing the database The data servers are right here, Vera BlueField, the fastest storage servers, the fastest storage system the world has ever made. the data servers are right here vera bluefield the fastest storage servers the fastest storage system the world has ever made The reason why this is so vital is because agents are accessing memory so incredibly fast. These systems, the storage server, and the CPUs, are now the critical path of the most expensive part of the data center. This is the most expensive for a good reason. The economics of the AI factory is tokens. The tokens are created here. Of course you want to manufacture and generate as many tokens as possible. This is where you put all of your economics, and this has to not be in the way. Vera CPU has great pressure on the CPU architecture, which is the reason why we built a brand-new architecture from the ground up. A CPU the world has never seen before. We call it Vera. This is CPU for agents. All the CPUs of the past we built for humans. The reason why this is so vital is because agents are accessing memory so incredibly fast. the reason why this is so vital is because agents are accessing memory so incredibly fast These systems, the storage server, and the CPUs, are now the critical path of the most expensive part of the data center. these systems the storage server and the cpus are now the critical path of the most expensive part of the data center This is the most expensive for a good reason. this is the most expensive for a good reason The economics of the AI factory is tokens. the economics of the ai factory is tokens The tokens are created here. the tokens are created here Of course you want to manufacture and generate as many tokens as possible. of course you want to manufacture and generate as many tokens as possible This is where you put all of your economics, and this has to not be in the way. this is where you put all of your economics and this has to not be in the way Vera CPU has great pressure on the CPU architecture, which is the reason why we built a brand-new architecture from the ground up. vera cpu has great pressure on the cpu architecture which is the reason why we built a brand-new architecture from the ground up A CPU the world has never seen before. a cpu the world has never seen before We call it Vera. we call it vera This is CPU for agents. this is cpu for agents All the CPUs of the past we built for humans. all the cpus of the past we built for humans This CPU is built for agents. Well, there are four things to keep in mind. The four takeaways. The first takeaway is that the instructions per clock of Vera has to be incredibly good because we need the latency to be short. We need the processing time, single-threaded performance, not throughput, single-threaded performance has to be world-class. Absolutely the best single-threaded performance, which is the reason why the IPC, the instructions per clock of Vera, is so high. It's the highest in the world. Ten instructions fetched, decoded, and executed per clock. Number one. Number two. The bandwidth necessary to move data in and out for the CPU has to be utterly world-class. The second thing is bandwidth per core. The third is just bandwidth, period. Remember I said earlier, agentic systems is fundamentally disaggregated and distributed. When computing is disaggregated and distributed, networking becomes the problem. This CPU is built for agents. this cpu is built for agents Well, there are four things to keep in mind. well there are four things to keep in mind The four takeaways. the four takeaways The first takeaway is that the instructions per clock of Vera has to be incredibly good because we need the latency to be short. the first takeaway is that the instructions per clock of vera has to be incredibly good because we need the latency to be short We need the processing time, single-threaded performance, not throughput, single-threaded performance has to be world-class. we need the processing time single-threaded performance not throughput single-threaded performance has to be world-class Absolutely the best single-threaded performance, which is the reason why the IPC, the instructions per clock of Vera, is so high. absolutely the best single-threaded performance which is the reason why the ipc the instructions per clock of vera is so high It's the highest in the world. it's the highest in the world Ten instructions fetched, decoded, and executed per clock. ten instructions fetched decoded and executed per clock Number one. number one Number two. number two The bandwidth necessary to move data in and out for the CPU has to be utterly world-class. the bandwidth necessary to move data in and out for the cpu has to be utterly world-class The second thing is bandwidth per core. the second thing is bandwidth per core The third is just bandwidth, period. the third is just bandwidth period Remember I said earlier, agentic systems is fundamentally disaggregated and distributed. remember i said earlier agentic systems is fundamentally disaggregated and distributed When computing is disaggregated and distributed, networking becomes the problem. when computing is disaggregated and distributed networking becomes the problem Therefore, we have to move the data around as fast as possible between the CPU cores and between the CPU and the storage, the CPU and the GPU. The bandwidth around the system and inside the CPU core has to be utterly world-class. This is the first CPU that has been built a long time that is literally at radical limits with a fabric that connects all of the CPU cores that is speed of light, 3.6 Tbps. No chiplet tax, no chip boundary crossings because the CPU cores are talking to each other with extremely high bandwidth. They are not rented core per core per core. They are all working together. The cross-sectional bandwidth of Vera is off the charts. It is the first one to be PCI Express Gen 6. Therefore, we have to move the data around as fast as possible between the CPU cores and between the CPU and the storage, the CPU and the GPU. therefore we have to move the data around as fast as possible between the cpu cores and between the cpu and the storage the cpu and the gpu The bandwidth around the system and inside the CPU core has to be utterly world-class. the bandwidth around the system and inside the cpu core has to be utterly world-class This is the first CPU that has been built a long time that is literally at radical limits with a fabric that connects all of the CPU cores that is speed of light, 3.6 Tbps . this is the first cpu that has been built a long time that is literally at radical limits with a fabric that connects all of the cpu cores that is speed of light 3.6 tbps No chiplet tax, no chip boundary crossings because the CPU cores are talking to each other with extremely high bandwidth. no chiplet tax no chip boundary crossings because the cpu cores are talking to each other with extremely high bandwidth They are not rented core per core per core. they are not rented core per core per core They are all working together. they are all working together The cross-sectional bandwidth of Vera is off the charts. the cross-sectional bandwidth of vera is off the charts It is the first one to be PCI Express Gen 6. it is the first one to be pci express gen 6 It is also the first one to have LPDDR5 with 1.2 Tbps, 2x to 3x the bandwidth of the highest performance CPUs on the outside, 3x the bandwidth on the inside. The bandwidth per core and the bandwidth period is world-class. Remember, I showed you earlier, the number of CPU cores, the number of CPUs is going to be quite high. The reason for that is very simple. We created CPUs for humans in the past, and humans, there are only 1 billion of us. There will be billions of agents, and these agents are going to be using the CPUs with very little patience because the cost of the GPU they sit next to is too high, and therefore, too valuable, too precious. It is also the first one to have LPDDR5 with 1.2 Tbps , 2x to 3x the bandwidth of the highest performance CPUs on the outside, 3x the bandwidth on the inside. The bandwidth per core and the bandwidth period is world-class. it is also the first one to have lpddr5 with 1.2 tbps 2x to 3x the bandwidth of the highest performance cpus on the outside 3x the bandwidth on the inside. the bandwidth per core and the bandwidth period is world-class Remember, I showed you earlier, the number of CPU cores, the number of CPUs is going to be quite high. remember i showed you earlier the number of cpu cores the number of cpus is going to be quite high The reason for that is very simple. the reason for that is very simple We created CPUs for humans in the past, and humans, there are only 1 billion of us. we created cpus for humans in the past and humans there are only 1 billion of us There will be billions of agents, and these agents are going to be using the CPUs with very little patience because the cost of the GPU they sit next to is too high, and therefore, too valuable, too precious. there will be billions of agents and these agents are going to be using the cpus with very little patience because the cost of the gpu they sit next to is too high and therefore too valuable too precious These CPUs are going to be both performant, but they also have to be extremely energy efficient so that we can cram as much CPU as we can into the factory without taking away power from the token generation, which we know is how we make money. These four properties, instructions per clock or single-threaded performance, bandwidth per core, the total bandwidth around the chip and inside the chip, and energy efficiency defines Vera. It is absolutely world-class. When you compare it to the highest performance x86, it is just off the charts. When you compare it in real single-threaded performance, real performance, it's off the charts. It is incredible to be able to deliver 5% improvement on CPUs. It is incredible to be able to deliver 10%. This kind of performance speed up is just unheard of. This is NVIDIA Vera. What do you think? Let's take a look. These CPUs are going to be both performant, but they also have to be extremely energy efficient so that we can cram as much CPU as we can into the factory without taking away power from the token generation, which we know is how we make money. these cpus are going to be both performant but they also have to be extremely energy efficient so that we can cram as much cpu as we can into the factory without taking away power from the token generation which we know is how we make money These four properties, instructions per clock or single-threaded performance, bandwidth per core, the total bandwidth around the chip and inside the chip, and energy efficiency defines Vera. these four properties instructions per clock or single-threaded performance bandwidth per core the total bandwidth around the chip and inside the chip and energy efficiency defines vera It is absolutely world-class. it is absolutely world-class When you compare it to the highest performance x86, it is just off the charts. when you compare it to the highest performance x86 it is just off the charts When you compare it in real single-threaded performance, real performance, it's off the charts. when you compare it in real single-threaded performance real performance it's off the charts It is incredible to be able to deliver 5% improvement on CPUs. it is incredible to be able to deliver 5% improvement on cpus It is incredible to be able to deliver 10%. it is incredible to be able to deliver 10% This kind of performance speed up is just unheard of. this kind of performance speed up is just unheard of This is NVIDIA Vera. this is nvidia vera What do you think? what do you think Let's take a look. let's take a look

Speaker 3: Agentic AI changes the role of the CPU. The CPU is now the conductor, and the GPU is the orchestra. Traditional CPUs were built for a different era, maximizing cores per socket. Slice them up, virtualize, rent by the hour. In the age of agents, the CPU is now a bottleneck to GPU utilization, directly affecting token throughput, latency, and user experience. NVIDIA Vera is the CPU built for the agentic loop, combining NVIDIA's custom data center CPU core with a scalable coherency fabric for the right balance of performance cores and bandwidth to maximize AI factory output. At the heart of Vera is the NVIDIA Olympus core, built for modern data center workloads, branch-heavy Python runtimes, tool calls, and sandboxed code execution. Each core is tuned for throughput. A neural branch predictor evaluating two taken branches per cycle. A 10-wide decode engine brings in more work each cycle. Agentic AI changes the role of the CPU. agentic ai changes the role of the cpu The CPU is now the conductor, and the GPU is the orchestra. the cpu is now the conductor and the gpu is the orchestra Traditional CPUs were built for a different era, maximizing cores per socket. traditional cpus were built for a different era maximizing cores per socket Slice them up, virtualize, rent by the hour. slice them up virtualize rent by the hour In the age of agents, the CPU is now a bottleneck to GPU utilization, directly affecting token throughput, latency, and user experience. in the age of agents the cpu is now a bottleneck to gpu utilization directly affecting token throughput latency and user experience NVIDIA Vera is the CPU built for the agentic loop, combining NVIDIA's custom data center CPU core with a scalable coherency fabric for the right balance of performance cores and bandwidth to maximize AI factory output. nvidia vera is the cpu built for the agentic loop combining nvidia's custom data center cpu core with a scalable coherency fabric for the right balance of performance cores and bandwidth to maximize ai factory output At the heart of Vera is the NVIDIA Olympus core, built for modern data center workloads, branch-heavy Python runtimes, tool calls, and sandboxed code execution. at the heart of vera is the nvidia olympus core built for modern data center workloads branch-heavy python runtimes tool calls and sandboxed code execution Each core is tuned for throughput. each core is tuned for throughput A neural branch predictor evaluating two taken branches per cycle. a neural branch predictor evaluating two taken branches per cycle A 10-wide decode engine brings in more work each cycle. a 10-wide decode engine brings in more work each cycle A large out-of-order engine keeps instructions moving. Advanced prefetchers with a novel graph engine anticipating the next data path. Fast cores only matter when data arrives correctly and on time. Vera is the first CPU to use LPDDR5X memory while correcting multiple errors simultaneously without compromising bandwidth. Vera achieves 40% lower peak memory latency versus x86, keeping cores fed on time through retrieval, analytics, and sandbox execution. NVIDIA's second-generation scalable coherency fabric unifies all 88 Olympus cores on a monolithic mesh with separate dies for memory and I/O. Cores are not split across chiplets, enabling 50% faster core-to-core communication than traditional CPUs. Memory-coherent NVLink chip-to-chip connects GPUs directly to the fabric. Beyond GPUs, NVLink chip-to-chip can scale Vera up to multiple sockets, enabling massive bandwidth between CPUs. Vera delivers 1.8x the agentic sandbox performance of x86 CPUs. A large out-of-order engine keeps instructions moving. a large out-of-order engine keeps instructions moving Advanced prefetchers with a novel graph engine anticipating the next data path. advanced prefetchers with a novel graph engine anticipating the next data path Fast cores only matter when data arrives correctly and on time. fast cores only matter when data arrives correctly and on time Vera is the first CPU to use LPDDR5X memory while correcting multiple errors simultaneously without compromising bandwidth. vera is the first cpu to use lpddr5x memory while correcting multiple errors simultaneously without compromising bandwidth Vera achieves 40% lower peak memory latency versus x86, keeping cores fed on time through retrieval, analytics, and sandbox execution. vera achieves 40% lower peak memory latency versus x86 keeping cores fed on time through retrieval analytics and sandbox execution NVIDIA's second-generation scalable coherency fabric unifies all 88 Olympus cores on a monolithic mesh with separate dies for memory and I/O. nvidia's second-generation scalable coherency fabric unifies all 88 olympus cores on a monolithic mesh with separate dies for memory and i/o Cores are not split across chiplets, enabling 50% faster core-to-core communication than traditional CPUs. cores are not split across chiplets enabling 50% faster core-to-core communication than traditional cpus Memory-coherent NVLink chip-to-chip connects GPUs directly to the fabric. memory-coherent nvlink chip-to-chip connects gpus directly to the fabric Beyond GPUs, NVLink chip-to-chip can scale Vera up to multiple sockets, enabling massive bandwidth between CPUs. beyond gpus nvlink chip-to-chip can scale vera up to multiple sockets enabling massive bandwidth between cpus Vera delivers 1.8x the agentic sandbox performance of x86 CPUs. vera delivers 1.8x the agentic sandbox performance of x86 cpus Standalone Vera racks run agent sandboxes, tools, code, and data pipelines. Tightly coupled to Rubin GPUs, Vera keeps accelerated workflows moving. NVIDIA Vera BlueField-4 STX powers context memory and AI storage. Compute, networking, storage. Vera is the CPU for the age of agents. Standalone Vera racks run agent sandboxes, tools, code, and data pipelines. standalone vera racks run agent sandboxes tools code and data pipelines Tightly coupled to Rubin GPUs, Vera keeps accelerated workflows moving. tightly coupled to rubin gpus vera keeps accelerated workflows moving NVIDIA Vera BlueField-4 STX powers context memory and AI storage. nvidia vera bluefield-4 stx powers context memory and ai storage Compute, networking, storage. compute networking storage Vera is the CPU for the age of agents. vera is the cpu for the age of agents

Speaker 1: This is going to be our new major growth driver. The reviews are already coming out, it's pretty good. That's pretty good stuff. Remember, Grace and Vera are also the most highly qualified CPUs in the world of AI because every single data center, every single cloud, every single enterprise, every company that works with NVIDIA on AI has already qualified Grace. The entire software stack has already been optimized for Grace. Every company will be qualifying Vera. Vera will be the most optimized agentic CPU in the world simply because it's going to go with Vera Rubin, simply because we made the big hard switch. During Grace Blackwell transition, the biggest risk was going from external CPU x86 into Grace Blackwell. That transition was extremely dangerous, we did it with incredible execution. Grace is literally synonymous with Grace Blackwell. This is going to be our new major growth driver. this is going to be our new major growth driver The reviews are already coming out, it's pretty good. the reviews are already coming out it's pretty good That's pretty good stuff. that's pretty good stuff Remember, Grace and Vera are also the most highly qualified CPUs in the world of AI because every single data center, every single cloud, every single enterprise, every company that works with NVIDIA on AI has already qualified Grace. remember grace and vera are also the most highly qualified cpus in the world of ai because every single data center every single cloud every single enterprise every company that works with nvidia on ai has already qualified grace The entire software stack has already been optimized for Grace. the entire software stack has already been optimized for grace Every company will be qualifying Vera. every company will be qualifying vera Vera will be the most optimized agentic CPU in the world simply because it's going to go with Vera Rubin, simply because we made the big hard switch. vera will be the most optimized agentic cpu in the world simply because it's going to go with vera rubin simply because we made the big hard switch During Grace Blackwell transition, the biggest risk was going from external CPU x86 into Grace Blackwell. That transition was extremely dangerous, we did it with incredible execution. during grace blackwell transition the biggest risk was going from external cpu x86 into grace blackwell. that transition was extremely dangerous we did it with incredible execution Grace is literally synonymous with Grace Blackwell. grace is literally synonymous with grace blackwell When people say Blackwell, they say Grace Blackwell, because it is utterly now everywhere. Every company's software stack has been optimized for it. Everybody's security stack has been optimized for it. Now here comes Vera. I'm super excited about that. Now look at some of the performance numbers. Speedups is one thing. It is extremely hard to speed up SQL. SQL, the most famous domain-specific language, DSL, that has ever been created. Before SQL, before CUDA, there was SQL. Before OpenGL, there was SQL, invented by IBM. Today, it is the structured database engine of the planet. Everybody uses SQL. This is SQL running 3x faster, not 10% faster, not 25% faster, 3x faster. Incredible. The next one is real-time stream processing. Remember, your AI is going to be not just reading documents. When people say Blackwell, they say Grace Blackwell, because it is utterly now everywhere. when people say blackwell they say grace blackwell because it is utterly now everywhere Every company's software stack has been optimized for it. every company's software stack has been optimized for it Everybody's security stack has been optimized for it. everybody's security stack has been optimized for it Now here comes Vera. now here comes vera I'm super excited about that. i'm super excited about that Now look at some of the performance numbers. now look at some of the performance numbers Speedups is one thing. speedups is one thing It is extremely hard to speed up SQL. it is extremely hard to speed up sql SQL, the most famous domain-specific language, DSL, that has ever been created. sql the most famous domain-specific language dsl that has ever been created Before SQL, before CUDA, there was SQL. before sql before cuda there was sql Before OpenGL, there was SQL, invented by IBM. before opengl there was sql invented by ibm Today, it is the structured database engine of the planet. today it is the structured database engine of the planet Everybody uses SQL. everybody uses sql This is SQL running 3x faster, not 10% faster, not 25% faster, 3x faster. this is sql running 3x faster not 10% faster not 25% faster 3x faster Incredible. incredible The next one is real-time stream processing. the next one is real-time stream processing Remember, your AI is going to be not just reading documents. remember your ai is going to be not just reading documents Your AI is going to be watching for telemetry, especially inside a factory, inside a stock exchange. You're going to be looking for telemetry continuously. The burst of data that's coming in goes into a CPU. This is Vera CPU running real-time stream processing for New York Stock Exchange. Lynn Martin, the president of New York Stock Exchange, has been so gracious to partner with us. This system is run all over the world in real-time stream processing. Vera CPU, 3x, all because of the bandwidth, the single-threaded instruction execution, the bandwidth inside between the cores, the bandwidth outside. Vera is completely revolutionary. That's Vera. X factors is something you talk about when you're talking about GPUs. It is quite rare that somebody talks about X factors on real workload that is associated with CPU. I'm so proud of the team. Your AI is going to be watching for telemetry, especially inside a factory, inside a stock exchange. your ai is going to be watching for telemetry especially inside a factory inside a stock exchange You're going to be looking for telemetry continuously. you're going to be looking for telemetry continuously The burst of data that's coming in goes into a CPU. the burst of data that's coming in goes into a cpu This is Vera CPU running real-time stream processing for New York Stock Exchange. this is vera cpu running real-time stream processing for new york stock exchange Lynn Martin, the president of New York Stock Exchange, has been so gracious to partner with us. lynn martin the president of new york stock exchange has been so gracious to partner with us This system is run all over the world in real-time stream processing. this system is run all over the world in real-time stream processing Vera CPU, 3x , all because of the bandwidth, the single-threaded instruction execution, the bandwidth inside between the cores, the bandwidth outside. vera cpu 3x all because of the bandwidth the single-threaded instruction execution the bandwidth inside between the cores the bandwidth outside Vera is completely revolutionary. vera is completely revolutionary That's Vera. that's vera X factors is something you talk about when you're talking about GPUs. x factors is something you talk about when you're talking about gpus It is quite rare that somebody talks about X factors on real workload that is associated with CPU. it is quite rare that somebody talks about x factors on real workload that is associated with cpu I'm so proud of the team. i'm so proud of the team You guys did such a great job. We have an extraordinary roadmap coming. What's really exciting is almost everybody is supporting Vera. They're as excited as we are. This is Vera opening up. It's opened up a brand new market. Agents is a new workload. We built CPUs for humans in the past. We need CPUs for agents, agentic systems. Their properties are different. Why would the old CPUs be the same? We are building millions and millions of Veras, millions of Veras. To go to market with us, Taiwan's ODMs and computer makers, all the OEMs, and you could see the early adopters. The early adopters are the agentic companies. This is the beginning of a new market, a market that never existed before. It's not going to take away from the old markets, but this is a new market, CPU for agents. You guys did such a great job. you guys did such a great job We have an extraordinary roadmap coming. we have an extraordinary roadmap coming What's really exciting is almost everybody is supporting Vera. what's really exciting is almost everybody is supporting vera They're as excited as we are. they're as excited as we are This is Vera opening up. this is vera opening up It's opened up a brand new market. it's opened up a brand new market Agents is a new workload. agents is a new workload We built CPUs for humans in the past. we built cpus for humans in the past We need CPUs for agents, agentic systems. we need cpus for agents agentic systems Their properties are different. their properties are different Why would the old CPUs be the same? why would the old cpus be the same We are building millions and millions of Veras, millions of Veras. we are building millions and millions of veras millions of veras To go to market with us, Taiwan's ODMs and computer makers, all the OEMs, and you could see the early adopters. to go to market with us taiwan's odms and computer makers all the oems and you could see the early adopters The early adopters are the agentic companies. the early adopters are the agentic companies This is the beginning of a new market, a market that never existed before. this is the beginning of a new market a market that never existed before It's not going to take away from the old markets, but this is a new market, CPU for agents. it's not going to take away from the old markets but this is a new market cpu for agents This market will surely be larger than the last, the reason for that is because there'll be a lot more agents than there are people, the agents are very impatient. NVIDIA Vera CPU. Thank you. This is the most important slide, really. This is the takeaway. The takeaway here is that this is the application pattern. This is the computing pattern of the next decade. Agents, harnesses, orchestrating large language models. Every company will run it. Every company will be an agent company. Every company will have agents running inside. Every company will see that agents will need its own operating system. Every company is asking us, "How do we run agents safely? How do we build agents for our own workloads?" We have the NVIDIA Agent Toolkit for Enterprise AI. You've seen me build this in plain sight. This market will surely be larger than the last, the reason for that is because there'll be a lot more agents than there are people, the agents are very impatient. this market will surely be larger than the last the reason for that is because there'll be a lot more agents than there are people the agents are very impatient NVIDIA Vera CPU. nvidia vera cpu Thank you. thank you This is the most important slide, really. this is the most important slide really This is the takeaway. this is the takeaway The takeaway here is that this is the application pattern. the takeaway here is that this is the application pattern This is the computing pattern of the next decade. this is the computing pattern of the next decade Agents, harnesses, orchestrating large language models. agents harnesses orchestrating large language models Every company will run it. every company will run it Every company will be an agent company. every company will be an agent company Every company will have agents running inside. every company will have agents running inside Every company will see that agents will need its own operating system. every company will see that agents will need its own operating system Every company is asking us, "How do we run agents safely? every company is asking us "how do we run agents safely How do we build agents for our own workloads?" We have the NVIDIA Agent Toolkit for Enterprise AI. how do we build agents for our own workloads?" we have the nvidia agent toolkit for enterprise ai You've seen me build this in plain sight. you've seen me build this in plain sight Almost everything that NVIDIA does, as you know, at every GTC, if you go back and look at my GTC five years ago or 10 years ago, you will see today. This, you've seen me talking about for several years now, because we've been building for this moment. There are four things that companies need in order to build agents as a service or build agents to operate. The first thing you need is you need models. Of course, large language models. The smarter, the better. The cheaper, the better. The faster, the better. The second is you need a harness to orchestrate the whole thing. The third, these models want to use tools, these tools come with its skills. I showed you CUDA-X libraries. Those are going to be amazing tools for the agents in the future. Lastly, you need a runtime. Almost everything that NVIDIA does, as you know, at every GTC, if you go back and look at my GTC five years ago or 10 years ago, you will see today. almost everything that nvidia does as you know at every gtc if you go back and look at my gtc five years ago or 10 years ago you will see today This, you've seen me talking about for several years now, because we've been building for this moment. this you've seen me talking about for several years now because we've been building for this moment There are four things that companies need in order to build agents as a service or build agents to operate. there are four things that companies need in order to build agents as a service or build agents to operate The first thing you need is you need models. the first thing you need is you need models Of course, large language models. of course large language models The smarter, the better. the smarter the better The cheaper, the better. the cheaper the better The faster, the better. the faster the better The second is you need a harness to orchestrate the whole thing. the second is you need a harness to orchestrate the whole thing The third, these models want to use tools, these tools come with its skills. the third these models want to use tools these tools come with its skills I showed you CUDA-X libraries. i showed you cuda-x libraries Those are going to be amazing tools for the agents in the future. those are going to be amazing tools for the agents in the future Lastly, you need a runtime. lastly you need a runtime You need the operating system that holds it all together. This is the NVIDIA toolkit for agents. It includes models that you can modify, NVIDIA's world-class open models. I'm going to show you more. You could run agents from anybody. You could run Claude Code, incredible agent, Codex, incredible agent. You could run it inside this harness called OpenShell, which will be highly secure for your inside the enterprise. The shell protects the agent, keeps it grounded in security policies. Privacy is protected. Its rights and privileges are given. Its identity is protected. This OpenShell is being adopted all over the world. NVIDIA OpenShell is open source. You're going to see so many companies adopt it. Red Hat, Canonical, Microsoft. It's going to be adopted everywhere. This is important. This is the runtime. You need the operating system that holds it all together. you need the operating system that holds it all together This is the NVIDIA toolkit for agents. this is the nvidia toolkit for agents It includes models that you can modify, NVIDIA's world-class open models. it includes models that you can modify nvidia's world-class open models I'm going to show you more. i'm going to show you more You could run agents from anybody. you could run agents from anybody You could run Claude Code, incredible agent, Codex, incredible agent. you could run claude code incredible agent codex incredible agent You could run it inside this harness called OpenShell, which will be highly secure for your inside the enterprise. you could run it inside this harness called openshell which will be highly secure for your inside the enterprise The shell protects the agent, keeps it grounded in security policies. the shell protects the agent keeps it grounded in security policies Privacy is protected. privacy is protected Its rights and privileges are given. its rights and privileges are given Its identity is protected. its identity is protected This OpenShell is being adopted all over the world. this openshell is being adopted all over the world NVIDIA OpenShell is open source. nvidia openshell is open source You're going to see so many companies adopt it. you're going to see so many companies adopt it Red Hat, Canonical, Microsoft. red hat canonical microsoft It's going to be adopted everywhere. it's going to be adopted everywhere This is important. this is important This is the runtime. this is the runtime This runtime is fully optimized for the NVIDIA AI platform, which is everywhere. You can run OpenShell in any cloud, on-prem, and even on-device. You have now tools and libraries that they can use. You have models that you can modify or use as is, or you have agents. This would be OpenClaw, Hermes, another incredible harness. These agentic harnesses can now run on-prem or for you anywhere. Okay. Four things, and this represents the operating system of the modern enterprise. How do we use this? One of my favorite use cases of agents is chip designers. It is the single most important thing that NVIDIA does. Of course, we have to partner with Cadence to build Super Agent, a chip design super agent. It is orchestrated by Codex or Claude Code. This runtime is fully optimized for the NVIDIA AI platform, which is everywhere. this runtime is fully optimized for the nvidia ai platform which is everywhere You can run OpenShell in any cloud, on-prem, and even on-device. you can run openshell in any cloud on-prem and even on-device You have now tools and libraries that they can use. you have now tools and libraries that they can use You have models that you can modify or use as is, or you have agents. you have models that you can modify or use as is or you have agents This would be OpenClaw, Hermes, another incredible harness. this would be openclaw hermes another incredible harness These agentic harnesses can now run on-prem or for you anywhere. these agentic harnesses can now run on-prem or for you anywhere Okay. okay Four things, and this represents the operating system of the modern enterprise. four things and this represents the operating system of the modern enterprise How do we use this? how do we use this One of my favorite use cases of agents is chip designers. one of my favorite use cases of agents is chip designers It is the single most important thing that NVIDIA does. it is the single most important thing that nvidia does Of course, we have to partner with Cadence to build Super Agent, a chip design super agent. of course we have to partner with cadence to build super agent a chip design super agent It is orchestrated by Codex or Claude Code. it is orchestrated by codex or claude code It has RTL and architecture diagrams or schematics or specifications as input and whatever you need to fix. Together, we created some super agents that are optimized for the NVIDIA runtime with Nemotron, and let's take a look. It is really incredible. It has RTL and architecture diagrams or schematics or specifications as input and whatever you need to fix. it has rtl and architecture diagrams or schematics or specifications as input and whatever you need to fix Together, we created some super agents that are optimized for the NVIDIA runtime with Nemotron, and let's take a look. together we created some super agents that are optimized for the nvidia runtime with nemotron and let's take a look It is really incredible. it is really incredible

Speaker 3: Cadence and NVIDIA are partnering to build chip design agents. Hundreds of thousands of NVIDIA chips come together to make the AI factories that power the world's frontier AI models. Designing these chips and the systems they run in is one of the hardest engineering challenges. Trillions of transistors, three-dimensional circuits, microscopic scale. Every gate, every wire, synchronized to picoseconds, must work in perfect harmony with no margin for error. Physical prototypes are too slow and too costly, so engineers work in the digital realm. Each chip begins as a set of architectural specifications, then translated into RTL, the language of chip design. RTL must be verified in simulation. A single bug can delay a chip by months. At NVIDIA, thousands of engineers, billions of compute hours per year, millions of tests written, run, and debugged. A cycle that takes teams weeks. Cadence and NVIDIA are partnering to build chip design agents. cadence and nvidia are partnering to build chip design agents Hundreds of thousands of NVIDIA chips come together to make the AI factories that power the world's frontier AI models. hundreds of thousands of nvidia chips come together to make the ai factories that power the world's frontier ai models Designing these chips and the systems they run in is one of the hardest engineering challenges. designing these chips and the systems they run in is one of the hardest engineering challenges Trillions of transistors, three-dimensional circuits, microscopic scale. trillions of transistors three-dimensional circuits microscopic scale Every gate, every wire, synchronized to picoseconds, must work in perfect harmony with no margin for error. every gate every wire synchronized to picoseconds must work in perfect harmony with no margin for error Physical prototypes are too slow and too costly, so engineers work in the digital realm. physical prototypes are too slow and too costly so engineers work in the digital realm Each chip begins as a set of architectural specifications, then translated into RTL, the language of chip design. each chip begins as a set of architectural specifications then translated into rtl the language of chip design RTL must be verified in simulation. rtl must be verified in simulation A single bug can delay a chip by months. a single bug can delay a chip by months At NVIDIA, thousands of engineers, billions of compute hours per year, millions of tests written, run, and debugged. at nvidia thousands of engineers billions of compute hours per year millions of tests written run and debugged A cycle that takes teams weeks. a cycle that takes teams weeks To compress this cycle, Cadence and NVIDIA built a design verification agent. Codex orchestrates the process. Cadence ChipStack launches the RTL verification loop, powered by Nemotron and secured by NVIDIA OpenShell, calling on expert sub-agents in RTL generation, test bench creation, regression testing, and debug. The system drives itself. The ChipStack agents run hundreds of simulations with Cadence Xcelium, formal verification with Jasper. Design flaws revealed. Bugs in the code fixed. What once took weeks now takes hours. Verification cycles over 40x faster. Together, NVIDIA and Cadence are reinventing chip design with AI agents. To compress this cycle, Cadence and NVIDIA built a design verification agent. to compress this cycle cadence and nvidia built a design verification agent Codex orchestrates the process. codex orchestrates the process Cadence ChipStack launches the RTL verification loop, powered by Nemotron and secured by NVIDIA OpenShell, calling on expert sub-agents in RTL generation, test bench creation, regression testing, and debug. cadence chipstack launches the rtl verification loop powered by nemotron and secured by nvidia openshell calling on expert sub-agents in rtl generation test bench creation regression testing and debug The system drives itself. the system drives itself The ChipStack agents run hundreds of simulations with Cadence Xcelium, formal verification with Jasper. the chipstack agents run hundreds of simulations with cadence xcelium formal verification with jasper Design flaws revealed. design flaws revealed Bugs in the code fixed. bugs in the code fixed What once took weeks now takes hours. what once took weeks now takes hours Verification cycles over 40x faster. verification cycles over 40x faster Together, NVIDIA and Cadence are reinventing chip design with AI agents. together nvidia and cadence are reinventing chip design with ai agents

Speaker 1: From weeks to hours. NVIDIA has thousands of chip designers. We are going to hire hundreds of thousands of Cadence super agents that work with us so that we can accelerate our company, so that we can be even more ambitious, create even more amazing things, run even faster. You saw earlier that the toolkit with models, harness, tools. The tools in this case are Cadence simulators and verifiers, formal verification systems. It is the reason why we're working with Cadence so hard to accelerate all of their tools on CUDA. Because the agents are impatient. The agents want the answer immediately. Models, harnesses, accelerated CUDA, accelerated libraries and tools, and then the runtime. What you saw just now is all of that coming together. From weeks to hours. from weeks to hours NVIDIA has thousands of chip designers. nvidia has thousands of chip designers We are going to hire hundreds of thousands of Cadence super agents that work with us so that we can accelerate our company, so that we can be even more ambitious, create even more amazing things, run even faster. we are going to hire hundreds of thousands of cadence super agents that work with us so that we can accelerate our company so that we can be even more ambitious create even more amazing things run even faster You saw earlier that the toolkit with models, harness, tools. you saw earlier that the toolkit with models harness tools The tools in this case are Cadence simulators and verifiers, formal verification systems. the tools in this case are cadence simulators and verifiers formal verification systems It is the reason why we're working with Cadence so hard to accelerate all of their tools on CUDA. it is the reason why we're working with cadence so hard to accelerate all of their tools on cuda Because the agents are impatient. because the agents are impatient The agents want the answer immediately. the agents want the answer immediately Models, harnesses, accelerated CUDA, accelerated libraries and tools, and then the runtime. models harnesses accelerated cuda accelerated libraries and tools and then the runtime What you saw just now is all of that coming together. what you saw just now is all of that coming together Now, one of the things that it starts with is a great model that Cadence could modify and tune to be expert at the Cadence workflow, at the Cadence expertise, so that they could create super agents that are proprietary to Cadence with their proprietary knowledge. They have to start with an excellent model. We call it Nemotron. NVIDIA is dedicated to build open models for the world so that all of you, all of us, could create our own agents. Today, we're announcing the Nemotron 3 Ultra. Yep. Our next open model, and it is smart. The Nemotron models not only give you the model, we give you all the data that we use to train the model, and because we have a coalition of incredible partners, you can see all of our partners down here, we work together, contribute data to each other. Now, one of the things that it starts with is a great model that Cadence could modify and tune to be expert at the Cadence workflow, at the Cadence expertise, so that they could create super agents that are proprietary to Cadence with their proprietary knowledge. now one of the things that it starts with is a great model that cadence could modify and tune to be expert at the cadence workflow at the cadence expertise so that they could create super agents that are proprietary to cadence with their proprietary knowledge They have to start with an excellent model. they have to start with an excellent model We call it Nemotron. we call it nemotron NVIDIA is dedicated to build open models for the world so that all of you, all of us, could create our own agents. nvidia is dedicated to build open models for the world so that all of you all of us could create our own agents Today, we're announcing the Nemotron 3 Ultra. today we're announcing the nemotron 3 ultra Yep. yep Our next open model, and it is smart. our next open model and it is smart The Nemotron models not only give you the model, we give you all the data that we use to train the model, and because we have a coalition of incredible partners, you can see all of our partners down here, we work together, contribute data to each other. the nemotron models not only give you the model we give you all the data that we use to train the model and because we have a coalition of incredible partners you can see all of our partners down here we work together contribute data to each other Nemotron is trained on one of the largest suites of long-running reasoning models, long-running tool task-solving, tool-using data sets in the world because of all of our great partnerships. All of this from the model, the training script, and the data made completely available to you. This is open models at its best, the best open model system policies in the world. Simple goal is that you can take all of it, add to it, make it even better, make it yours. Nemotron 3 Ultra is 5x faster. This is the world's first model based on a hybrid architecture of SSM, state-space models, with Mixture of Experts. The architecture is incredibly fast. We made it fast that you could think fast. When you think fast, you could think longer at the same cost. 5x faster. Nemotron is trained on one of the largest suites of long-running reasoning models, long-running tool task-solving, tool-using data sets in the world because of all of our great partnerships. nemotron is trained on one of the largest suites of long-running reasoning models long-running tool task-solving tool-using data sets in the world because of all of our great partnerships All of this from the model, the training script, and the data made completely available to you. all of this from the model the training script and the data made completely available to you This is open models at its best, the best open model system policies in the world. this is open models at its best the best open model system policies in the world Simple goal is that you can take all of it, add to it, make it even better, make it yours. simple goal is that you can take all of it add to it make it even better make it yours Nemotron 3 Ultra is 5x faster. nemotron 3 ultra is 5x faster This is the world's first model based on a hybrid architecture of SSM, state-space models, with Mixture of Experts. this is the world's first model based on a hybrid architecture of ssm state-space models with mixture of experts The architecture is incredibly fast. the architecture is incredibly fast We made it fast that you could think fast. we made it fast that you could think fast When you think fast, you could think longer at the same cost. 5x faster. when you think fast you could think longer at the same cost 5x faster It is also 30% cheaper, 30% lower cost to run in total FLOPS and total inference time than even the most cost-effective in the world. We're comparing against the world's best open models. Frontier smart, 5x faster, 30% cheaper, completely open. We're completely dedicated to this. This is now Nemotron 3. We're currently working on Nemotron-4. This entire toolkit from models, harnesses, tools and skills, and runtimes is the reason why every enterprise company in the world has the ability now to create their own agents, just like Cadence did with their super agents. We're working with so many companies, Cadence and CrowdStrike and Dassault and Palantir, SAP and ServiceNow. People always said, "Jensen, the agents are going to disrupt these markets." I said completely opposite, and you can now see it. It is also 30% cheaper, 30% lower cost to run in total FLOPS and total inference time than even the most cost-effective in the world. it is also 30% cheaper 30% lower cost to run in total flops and total inference time than even the most cost-effective in the world We're comparing against the world's best open models. we're comparing against the world's best open models Frontier smart, 5x faster, 30% cheaper, completely open. frontier smart 5x faster 30% cheaper completely open We're completely dedicated to this. we're completely dedicated to this This is now Nemotron 3. this is now nemotron 3 We're currently working on Nemotron-4. we're currently working on nemotron-4 This entire toolkit from models, harnesses, tools and skills, and runtimes is the reason why every enterprise company in the world has the ability now to create their own agents, just like Cadence did with their super agents. this entire toolkit from models harnesses tools and skills and runtimes is the reason why every enterprise company in the world has the ability now to create their own agents just like cadence did with their super agents We're working with so many companies, Cadence and CrowdStrike and Dassault and Palantir, SAP and ServiceNow. we're working with so many companies cadence and crowdstrike and dassault and palantir sap and servicenow People always said, "Jensen, the agents are going to disrupt these markets." I said completely opposite, and you can now see it. people always said "jensen the agents are going to disrupt these markets." i said completely opposite and you can now see it Agents is going to create the largest opportunity ever for my partners and friends. We have the NeMo, the NVIDIA agentic toolkit for enterprise AI to help them. There you go. First, Vera Rubin in full production. Two, Vera CPU built for a new generation for agents. Three, NVIDIA's enterprise AI toolkits, so that every enterprise and every enterprise software company can build agents. My relationship with you started here. Many of you, many of my friends and partners here in Taiwan, your companies started here. This is in a lot of ways, the beginning of the modern computer industry, 40 years now. NVIDIA's 33 years old. The PC industry was already starting to get to Windows 1 and Windows 2 and Apple 1 and Apple 2. By the time that we came along, Windows 3.1 was the PC. Agents is going to create the largest opportunity ever for my partners and friends. agents is going to create the largest opportunity ever for my partners and friends We have the NeMo, the NVIDIA agentic toolkit for enterprise AI to help them. we have the nemo the nvidia agentic toolkit for enterprise ai to help them There you go. there you go First, Vera Rubin in full production. first vera rubin in full production Two, Vera CPU built for a new generation for agents. two vera cpu built for a new generation for agents Three, NVIDIA's enterprise AI toolkits, so that every enterprise and every enterprise software company can build agents. three nvidia's enterprise ai toolkits so that every enterprise and every enterprise software company can build agents My relationship with you started here. my relationship with you started here Many of you, many of my friends and partners here in Taiwan, your companies started here. many of you many of my friends and partners here in taiwan your companies started here This is in a lot of ways, the beginning of the modern computer industry, 40 years now. this is in a lot of ways the beginning of the modern computer industry 40 years now NVIDIA's 33 years old. nvidia's 33 years old The PC industry was already starting to get to Windows 1 and Windows 2 and Apple 1 and Apple 2. the pc industry was already starting to get to windows 1 and windows 2 and apple 1 and apple 2 By the time that we came along, Windows 3.1 was the PC. by the time that we came along windows 3.1 was the pc As you know, Windows 95 made PC personal. It took PC from enterprises, companies, and made it into a consumer electronics device. Everybody should have one, and everybody does. This is the beginning. This computing platform did several things incredibly smart. Windows was not just disaggregated, as you know. Windows was properly abstracted. It was architected just right. Systems BIOSes, open chipsets, the operating system with drivers that could be connected and installed at runtime, and an abstraction layer with a multimedia API that opened up the PC to what we all know today. Each one of these elements were essential in making the PC so popular. 40 years later, Microsoft and NVIDIA are going to reinvent the PC. This is going to be the new PC. As you know, Windows 95 made PC personal. as you know windows 95 made pc personal It took PC from enterprises, companies, and made it into a consumer electronics device. it took pc from enterprises companies and made it into a consumer electronics device Everybody should have one, and everybody does. everybody should have one and everybody does This is the beginning. this is the beginning This computing platform did several things incredibly smart. this computing platform did several things incredibly smart Windows was not just disaggregated, as you know. windows was not just disaggregated as you know Windows was properly abstracted. windows was properly abstracted It was architected just right. it was architected just right Systems BIOSes, open chipsets, the operating system with drivers that could be connected and installed at runtime, and an abstraction layer with a multimedia API that opened up the PC to what we all know today. systems bioses open chipsets the operating system with drivers that could be connected and installed at runtime and an abstraction layer with a multimedia api that opened up the pc to what we all know today Each one of these elements were essential in making the PC so popular. 40 years later, Microsoft and NVIDIA are going to reinvent the PC. each one of these elements were essential in making the pc so popular 40 years later microsoft and nvidia are going to reinvent the pc This is going to be the new PC. this is going to be the new pc Tomorrow night, I think it's tomorrow night our time, but I'm going to be with Satya, where we're going to talk a lot more about the work that we're doing together. Microsoft and NVIDIA, over the last three years, it took this long to completely reinvent how the PC's going to work so that we could be ready for this moment. As I mentioned earlier, that compute pattern called the agent is going to run in AI clouds. It's going to run inside enterprises. It is also going to run on your PC. What's going to happen to that PC when it has an autonomous agent? An agent that's helping you, that understands you. You could talk to it. It could look at you. You could ask it to read files, go help you do some research. It could do a lot more that I'll show you. Tomorrow night, I think it's tomorrow night our time, but I'm going to be with Satya, where we're going to talk a lot more about the work that we're doing together. tomorrow night i think it's tomorrow night our time but i'm going to be with satya where we're going to talk a lot more about the work that we're doing together Microsoft and NVIDIA, over the last three years, it took this long to completely reinvent how the PC's going to work so that we could be ready for this moment. microsoft and nvidia over the last three years it took this long to completely reinvent how the pc's going to work so that we could be ready for this moment As I mentioned earlier, that compute pattern called the agent is going to run in AI clouds. as i mentioned earlier that compute pattern called the agent is going to run in ai clouds It's going to run inside enterprises. it's going to run inside enterprises It is also going to run on your PC. it is also going to run on your pc What's going to happen to that PC when it has an autonomous agent? An agent that's helping you, that understands you. what's going to happen to that pc when it has an autonomous agent? an agent that's helping you that understands you You could talk to it. you could talk to it It could look at you. it could look at you You could ask it to read files, go help you do some research. you could ask it to read files go help you do some research It could do a lot more that I'll show you. it could do a lot more that i'll show you The new operating system is, of course, the old operating system plus large language models. Large language models in a lot of ways is the modern version of DirectX. It has, of course, input and output, understands prompts, it understands computer vision, it can generate video, it can generate sounds. It is the modern extension, the intelligence extension of the PC, of a computer. On top of that, the application, as I mentioned before, is going to be replaced by now an agentic runtime, and that is the modern application, an agent. Let's now take a look at what it can do. The new operating system is, of course, the old operating system plus large language models. the new operating system is of course the old operating system plus large language models Large language models in a lot of ways is the modern version of DirectX. large language models in a lot of ways is the modern version of directx It has, of course, input and output, understands prompts, it understands computer vision, it can generate video, it can generate sounds. it has of course input and output understands prompts it understands computer vision it can generate video it can generate sounds It is the modern extension, the intelligence extension of the PC, of a computer. it is the modern extension the intelligence extension of the pc of a computer On top of that, the application, as I mentioned before, is going to be replaced by now an agentic runtime, and that is the modern application, an agent. on top of that the application as i mentioned before is going to be replaced by now an agentic runtime and that is the modern application an agent Let's now take a look at what it can do. let's now take a look at what it can do

Speaker 3: It started with a spark. An idea to reimagine the PC for the first time in 40 years for the age of AI. What becomes of our personal computer in a world of agents? Agents running natively, connected to models, local or in the cloud. Our personal AI, sandboxed for security, running continuously, getting work done. The chips and the OS must evolve. Introducing RTX Spark. Everything we've learned over 33 years distilled into one chip. Blackwell RTX GPU with 6,144 CUDA cores, one petaflop of AI performance. A custom 20-core Grace CPU built in partnership with MediaTek, fused by NVLink. 128 GB of unified memory. It started with a spark. it started with a spark An idea to reimagine the PC for the first time in 40 years for the age of AI. an idea to reimagine the pc for the first time in 40 years for the age of ai What becomes of our personal computer in a world of agents? what becomes of our personal computer in a world of agents Agents running natively, connected to models, local or in the cloud. agents running natively connected to models local or in the cloud Our personal AI, sandboxed for security, running continuously, getting work done. our personal ai sandboxed for security running continuously getting work done The chips and the OS must evolve. the chips and the os must evolve Introducing RTX Spark. introducing rtx spark Everything we've learned over 33 years distilled into one chip. everything we've learned over 33 years distilled into one chip Blackwell RTX GPU with 6,144 CUDA cores, one petaflop of AI performance. blackwell rtx gpu with 6,144 cuda cores one petaflop of ai performance A custom 20-core Grace CPU built in partnership with MediaTek, fused by NVLink. 128 GB of unified memory. a custom 20-core grace cpu built in partnership with mediatek fused by nvlink 128 gb of unified memory TSMC 3nm process. 70 billion transistors. In close collaboration with Microsoft, a Windows platform for agents. We're reinventing the personal computer for creating, for gaming, for agents. This is the dawn of a new personal computing revolution, and it starts with NVIDIA RTX Spark. TSMC 3 nm process. 70 billion transistors. tsmc 3 nm process 70 billion transistors In close collaboration with Microsoft, a Windows platform for agents. in close collaboration with microsoft a windows platform for agents We're reinventing the personal computer for creating, for gaming, for agents. we're reinventing the personal computer for creating for gaming for agents This is the dawn of a new personal computing revolution, and it starts with NVIDIA RTX Spark. this is the dawn of a new personal computing revolution and it starts with nvidia rtx spark

Speaker 1: Here it is. Of course, I got to show you the most beautiful part, which is video games. It's also the closest to our heart. This is Forza. This is 007, by the way. The new 007 game. I'm looking forward to playing it. I look a little bit like him. Ladies and gentlemen, NVIDIA's RTX Spark laptops. I have too many things in my pocket. Okay. All right. This is the most amazing chip the world's ever built. This is the N1X that we built in partnership with MediaTek. I think I saw Rick earlier. This is N1X. This is a beautiful chip. This is a chip that, frankly, would take 33 years to build. The reason for that is because 100% of NVIDIA software stack runs here. If you want to run digital biology, no problem. Here it is. here it is Of course, I got to show you the most beautiful part, which is video games. of course i got to show you the most beautiful part which is video games It's also the closest to our heart. it's also the closest to our heart This is Forza. this is forza This is 007, by the way. this is 007 by the way The new 007 game. the new 007 game I'm looking forward to playing it. i'm looking forward to playing it I look a little bit like him. i look a little bit like him Ladies and gentlemen, NVIDIA's RTX Spark laptops. I have too many things in my pocket. ladies and gentlemen nvidia's rtx spark laptops i have too many things in my pocket Okay. okay All right. all right This is the most amazing chip the world's ever built. this is the most amazing chip the world's ever built This is the N1X that we built in partnership with MediaTek. this is the n1x that we built in partnership with mediatek I think I saw Rick earlier. i think i saw rick earlier This is N1X. this is n1x This is a beautiful chip. this is a beautiful chip This is a chip that, frankly, would take 33 years to build. this is a chip that frankly would take 33 years to build The reason for that is because 100% of NVIDIA software stack runs here. the reason for that is because 100% of nvidia software stack runs here If you want to run digital biology, no problem. if you want to run digital biology no problem If you want to do seismic processing, no problem. You want astrophysics, no problem. Everything associated with CUDA, all the physics, all the biology, all the genomics, all the AI, no problem. All the computer graphics, no problem. Every single application NVIDIA has ever created and every single application that Windows has ever run. Microsoft and NVIDIA meticulously optimized everything so that this computer literally runs everything the world has ever created. Plus, it now runs agents. An incredible computer. I'm so proud of it. Now, I want you to keep that in mind in the next video I'm going to show you. Just imagine everything here is going to run on your PC. If you want to do seismic processing, no problem. if you want to do seismic processing no problem You want astrophysics, no problem. you want astrophysics no problem Everything associated with CUDA, all the physics, all the biology, all the genomics, all the AI, no problem. everything associated with cuda all the physics all the biology all the genomics all the ai no problem All the computer graphics, no problem. all the computer graphics no problem Every single application NVIDIA has ever created and every single application that Windows has ever run. every single application nvidia has ever created and every single application that windows has ever run Microsoft and NVIDIA meticulously optimized everything so that this computer literally runs everything the world has ever created. microsoft and nvidia meticulously optimized everything so that this computer literally runs everything the world has ever created Plus, it now runs agents. plus it now runs agents An incredible computer. an incredible computer I'm so proud of it. i'm so proud of it Now, I want you to keep that in mind in the next video I'm going to show you. now i want you to keep that in mind in the next video i'm going to show you Just imagine everything here is going to run on your PC. just imagine everything here is going to run on your pc That computer could have a local Nemotron 3 Ultra model or Nemotron 3 Super model, or it could have a Claude Code or Codex or some other model in the cloud, or something on the network, and it's going to work and do something amazing. Let's play it. That computer could have a local Nemotron 3 Ultra model or Nemotron 3 Super model, or it could have a Claude Code or Codex or some other model in the cloud, or something on the network, and it's going to work and do something amazing. that computer could have a local nemotron 3 ultra model or nemotron 3 super model or it could have a claude code or codex or some other model in the cloud or something on the network and it's going to work and do something amazing Let's play it. let's play it

Speaker 3: Every house starts as an idea. Getting from idea to design takes a myriad of tools, expertise, and a lot of time. Now, an agent running locally on RTX Spark can help me design a house using the tools on my laptop with an OpenShell sandbox running the Hermes harness connected to Claude Sonnet in the cloud. I select the site, share my concept sketches and mood board of styles to inspire my design, and the prompt, a text description of the requirements and the design intent. My agent goes to work. Using the tools on my laptop, it opens Rhino and starts modeling the site, shaping terrain, setbacks, and the building envelope. It proposes building forms optimized for cost, comfort, and quality. With the form defined, my agent generates the interior layout. Walls, circulation, rooms begin to take shape. Every house starts as an idea. every house starts as an idea Getting from idea to design takes a myriad of tools, expertise, and a lot of time. getting from idea to design takes a myriad of tools expertise and a lot of time Now, an agent running locally on RTX Spark can help me design a house using the tools on my laptop with an OpenShell sandbox running the Hermes harness connected to Claude Sonnet in the cloud. now an agent running locally on rtx spark can help me design a house using the tools on my laptop with an openshell sandbox running the hermes harness connected to claude sonnet in the cloud I select the site, share my concept sketches and mood board of styles to inspire my design, and the prompt, a text description of the requirements and the design intent. My agent goes to work. i select the site share my concept sketches and mood board of styles to inspire my design and the prompt a text description of the requirements and the design intent. my agent goes to work Using the tools on my laptop, it opens Rhino and starts modeling the site, shaping terrain, setbacks, and the building envelope. using the tools on my laptop it opens rhino and starts modeling the site shaping terrain setbacks and the building envelope It proposes building forms optimized for cost, comfort, and quality. it proposes building forms optimized for cost comfort and quality With the form defined, my agent generates the interior layout. with the form defined my agent generates the interior layout Walls, circulation, rooms begin to take shape. walls circulation rooms begin to take shape I jump in whenever I want to adjust, to change. Doors, windows, and structural elements are placed automatically. My agent detects its own mistakes and fixes them. When I approve, the agent exports the model from Rhino into Blender. Materials and object properties transfer with the design context intact. I fine-tune the materials, get the look just right. I pick the shots. Blender renders the house. My agent, using generative AI with the Flux.2 model, makes them photoreal. Multiple viewpoints, lighting conditions. What was once a complex workflow is now guided and simplified by my agent. Working with me on RTX Spark. Design at the speed of imagination. I jump in whenever I want to adjust, to change. i jump in whenever i want to adjust to change Doors, windows, and structural elements are placed automatically. doors windows and structural elements are placed automatically My agent detects its own mistakes and fixes them. my agent detects its own mistakes and fixes them When I approve, the agent exports the model from Rhino into Blender. when i approve the agent exports the model from rhino into blender Materials and object properties transfer with the design context intact. materials and object properties transfer with the design context intact I fine-tune the materials, get the look just right. i fine-tune the materials get the look just right I pick the shots. i pick the shots Blender renders the house. blender renders the house My agent, using generative AI with the Flux.2 model, makes them photoreal. my agent using generative ai with the flux.2 model makes them photoreal Multiple viewpoints, lighting conditions. multiple viewpoints lighting conditions What was once a complex workflow is now guided and simplified by my agent. what was once a complex workflow is now guided and simplified by my agent Working with me on RTX Spark. working with me on rtx spark Design at the speed of imagination. design at the speed of imagination

Speaker 1: PC in the world of agents. The developers are so excited about it. This is an incredible computer. All of the acceleration, all the software capabilities associated with it, working with every developer to make it incredible for all of you. The next one, Adobe. Incredible tool suite, of course, used by tens of millions of people around the world. They have re-engineered the architecture, the core of Adobe Photoshop and Premiere. They'll release it for RTX Spark. It is twice as fast. It's already fast. Now it's going to be twice as fast. It's also designed to be agent-friendly. With its MCP server, it can now interact with agents on your laptop. The number of customers, the number of partners that are so excited to bring RTX Spark to the market is just incredible. PC in the world of agents. pc in the world of agents The developers are so excited about it. the developers are so excited about it This is an incredible computer. this is an incredible computer All of the acceleration, all the software capabilities associated with it, working with every developer to make it incredible for all of you. all of the acceleration all the software capabilities associated with it working with every developer to make it incredible for all of you The next one, Adobe. the next one adobe Incredible tool suite, of course, used by tens of millions of people around the world. incredible tool suite of course used by tens of millions of people around the world They have re-engineered the architecture, the core of Adobe Photoshop and Premiere. they have re-engineered the architecture the core of adobe photoshop and premiere They'll release it for RTX Spark. they'll release it for rtx spark It is twice as fast. it is twice as fast It's already fast. it's already fast Now it's going to be twice as fast. now it's going to be twice as fast It's also designed to be agent-friendly. it's also designed to be agent-friendly With its MCP server, it can now interact with agents on your laptop. with its mcp server it can now interact with agents on your laptop The number of customers, the number of partners that are so excited to bring RTX Spark to the market is just incredible. the number of customers the number of partners that are so excited to bring rtx spark to the market is just incredible This is the first across the lineup of PC reinvention for 40 years. I'm just so happy that all of you and the ecosystem around the world has joined us. This is basically everybody. Everybody will support RTX Spark and will be building incredibly smart and powerful and beautiful laptops with all of us. Thank you very much. That's not all. That's not all. RTX Spark is a reinvention of laptop. In fact, Microsoft NVIDIA is reinventing all of PC, and today we're announcing a whole new line. Three revolutionary Windows machines covering desktop, laptop, and workstations. All 100% Windows compatible, 100% CUDA, 100% NVIDIA AI Tensor Core. This is the first across the lineup of PC reinvention for 40 years. this is the first across the lineup of pc reinvention for 40 years I'm just so happy that all of you and the ecosystem around the world has joined us. i'm just so happy that all of you and the ecosystem around the world has joined us This is basically everybody. this is basically everybody Everybody will support RTX Spark and will be building incredibly smart and powerful and beautiful laptops with all of us. everybody will support rtx spark and will be building incredibly smart and powerful and beautiful laptops with all of us Thank you very much. thank you very much That's not all. that's not all That's not all. that's not all RTX Spark is a reinvention of laptop. rtx spark is a reinvention of laptop In fact, Microsoft NVIDIA is reinventing all of PC, and today we're announcing a whole new line. in fact microsoft nvidia is reinventing all of pc and today we're announcing a whole new line Three revolutionary Windows machines covering desktop, laptop, and workstations. three revolutionary windows machines covering desktop laptop and workstations All 100% Windows compatible, 100% CUDA, 100% NVIDIA AI Tensor Core. all 100% windows compatible 100% cuda 100% nvidia ai tensor core Everything that you see that runs on NVIDIA in all these different platforms around the world runs here. This is the first completely re-engineered, reinvented line of PCs that has happened in 40 years. What's really amazing is this. This is the RTX Spark laptop. This is the desktop. This one's from MSI. Joseph, this one's yours. Okay. Look how beautiful it is. This agent could run 24/7, meter-free. You could download your agent. You could raise your lobster in here. This is your claw. It is running all the time, no meter anxiety, and it is sitting here connected to your whole house, connected to your laptop, connected to your display, all the cameras, your dryer, your water cooler, your water heater, your everything, whatever you want. Everything that you see that runs on NVIDIA in all these different platforms around the world runs here. everything that you see that runs on nvidia in all these different platforms around the world runs here This is the first completely re-engineered, reinvented line of PCs that has happened in 40 years. this is the first completely re-engineered reinvented line of pcs that has happened in 40 years What's really amazing is this. what's really amazing is this This is the RTX Spark laptop. this is the rtx spark laptop This is the desktop. this is the desktop This one's from MSI. this one's from msi Joseph, this one's yours. joseph this one's yours Okay. okay Look how beautiful it is. look how beautiful it is This agent could run 24/7, meter-free. this agent could run 24/7 meter-free You could download your agent. you could download your agent You could raise your lobster in here. you could raise your lobster in here This is your claw. this is your claw It is running all the time, no meter anxiety, and it is sitting here connected to your whole house, connected to your laptop, connected to your display, all the cameras, your dryer, your water cooler, your water heater, your everything, whatever you want. it is running all the time no meter anxiety and it is sitting here connected to your whole house connected to your laptop connected to your display all the cameras your dryer your water cooler your water heater your everything whatever you want Your security system, all connected to this, and this becomes your personal AI, your personal AI agent. It gets smarter and smarter and smarter over time because today we have Nemotron 3 Ultra, tomorrow we have Nemotron 4, and then Nemotron 5, Nemotron 6, and we just keep getting it smarter and smarter and smarter. Meanwhile, this is sitting at home helping you do things. If you want to book a travel, no problem. If you want an incredible system, this is a DGX Station for Windows, compatible with Windows, runs everything in Windows, and it has 768 gigabytes of memory. Your security system, all connected to this, and this becomes your personal AI, your personal AI agent. your security system all connected to this and this becomes your personal ai your personal ai agent It gets smarter and smarter and smarter over time because today we have Nemotron 3 Ultra, tomorrow we have Nemotron 4, and then Nemotron 5, Nemotron 6, and we just keep getting it smarter and smarter and smarter. it gets smarter and smarter and smarter over time because today we have nemotron 3 ultra tomorrow we have nemotron 4 and then nemotron 5 nemotron 6 and we just keep getting it smarter and smarter and smarter Meanwhile, this is sitting at home helping you do things. meanwhile this is sitting at home helping you do things If you want to book a travel, no problem. if you want to book a travel no problem If you want an incredible system, this is a DGX Station for Windows, compatible with Windows, runs everything in Windows, and it has 768 gigabytes of memory. if you want an incredible system this is a dgx station for windows compatible with windows runs everything in windows and it has 768 gigabytes of memory You could run a trillion-parameter model. This is unbelievable. 20 petaflops, 8 Tbps of memory bandwidth, and this sits by your desk. If you're a developer of large language models, you're a developer of agents, having this sit by your desk gives you all the compute you need, and then when you deploy it, you put it into the cloud. There's something that if you look at this and think about this, something is happening here. Remember, 15, 20 years ago, we used to have an idea called a phone. Today, we have an idea called a PC. Today, when you think about your phone, the one thing you don't do with it is make phone calls. You could run a trillion-parameter model. you could run a trillion-parameter model This is unbelievable. 20 petaflops, 8 Tbps of memory bandwidth, and this sits by your desk. this is unbelievable 20 petaflops 8 tbps of memory bandwidth and this sits by your desk If you're a developer of large language models, you're a developer of agents, having this sit by your desk gives you all the compute you need, and then when you deploy it, you put it into the cloud. if you're a developer of large language models you're a developer of agents having this sit by your desk gives you all the compute you need and then when you deploy it you put it into the cloud There's something that if you look at this and think about this, something is happening here. there's something that if you look at this and think about this something is happening here Remember, 15, 20 years ago, we used to have an idea called a phone. remember 15 20 years ago we used to have an idea called a phone Today, we have an idea called a PC. today we have an idea called a pc Today, when you think about your phone, the one thing you don't do with it is make phone calls. today when you think about your phone the one thing you don't do with it is make phone calls You do just about everything else. That phone means something very different to you than a phone of the past. I am certain what's going to happen here is that the PC 10 years from now and the PC that you think about today, a tool, whether you launch applications, click and type, and this PC is going to be completely different. Here's my theory. I can totally imagine, just as every house today has a home theater, or many houses have home theaters, big TVs, lawnmowers, dishwashers. I could totally imagine that someday there's actually an AI supercomputer in your house. It's running all of your agents, it's running all of your assistants, and they're doing all kinds of things for you all the time. You do just about everything else. you do just about everything else That phone means something very different to you than a phone of the past. that phone means something very different to you than a phone of the past I am certain what's going to happen here is that the PC 10 years from now and the PC that you think about today, a tool, whether you launch applications, click and type, and this PC is going to be completely different. i am certain what's going to happen here is that the pc 10 years from now and the pc that you think about today a tool whether you launch applications click and type and this pc is going to be completely different Here's my theory. here's my theory I can totally imagine, just as every house today has a home theater, or many houses have home theaters, big TVs, lawnmowers, dishwashers. i can totally imagine just as every house today has a home theater or many houses have home theaters big tvs lawnmowers dishwashers I could totally imagine that someday there's actually an AI supercomputer in your house. i could totally imagine that someday there's actually an ai supercomputer in your house It's running all of your agents, it's running all of your assistants, and they're doing all kinds of things for you all the time. it's running all of your agents it's running all of your assistants and they're doing all kinds of things for you all the time You have to have it in your house, just like you have a home theater in your house, you have stereos in your house, you have game consoles in your house. You end up assist AI agent computers running in your house. These, in time, becomes a lot more like R2-D2 to you. It becomes more like C-3PO to you than it feels like a PC to you. There is no question this reinvention of the computer is as big of a deal as the reinvention of the phone into what we now know as the smartphone. This is the beginning of that journey. This is the beginning of a new line. We have a roadmap for this. This is a brand-new product family for us. You have to have it in your house, just like you have a home theater in your house, you have stereos in your house, you have game consoles in your house. you have to have it in your house just like you have a home theater in your house you have stereos in your house you have game consoles in your house You end up assist AI agent computers running in your house. you end up assist ai agent computers running in your house These, in time, becomes a lot more like R2-D2 to you. these in time becomes a lot more like r2-d2 to you It becomes more like C-3PO to you than it feels like a PC to you. it becomes more like c-3po to you than it feels like a pc to you There is no question this reinvention of the computer is as big of a deal as the reinvention of the phone into what we now know as the smartphone. there is no question this reinvention of the computer is as big of a deal as the reinvention of the phone into what we now know as the smartphone This is the beginning of that journey. this is the beginning of that journey This is the beginning of a new line. this is the beginning of a new line We have a roadmap for this. we have a roadmap for this This is a brand-new product family for us. this is a brand-new product family for us Every single generation of architecture, we will have a desktop, a laptop, a workstation, and then a desktop, a laptop, and workstation. The thing that I am just incredibly pleased, incredibly honored, is that 100% of the world's PC industry has joined us to reinvent the PC. A new line, a new beginning. Thank you. As you know, agentic AI is just a digital robot. It understands, it reasons, it plans, and it acts and use tools. Agentic AI is going to run across all of these computers, and you've seen me talk about each and every one of these over time. We're working on human or robotics computers, robotics computers of all kinds. We're working on self-driving car computers. We're working on satellites. You have GeForce, which has tensor cores. I just talked about a whole new line of PCs. Every single generation of architecture, we will have a desktop, a laptop, a workstation, and then a desktop, a laptop, and workstation. every single generation of architecture we will have a desktop a laptop a workstation and then a desktop a laptop and workstation The thing that I am just incredibly pleased, incredibly honored, is that 100% of the world's PC industry has joined us to reinvent the PC. the thing that i am just incredibly pleased incredibly honored is that 100% of the world's pc industry has joined us to reinvent the pc A new line, a new beginning. a new line a new beginning Thank you. thank you As you know, agentic AI is just a digital robot. as you know agentic ai is just a digital robot It understands, it reasons, it plans, and it acts and use tools. it understands it reasons it plans and it acts and use tools Agentic AI is going to run across all of these computers, and you've seen me talk about each and every one of these over time. agentic ai is going to run across all of these computers and you've seen me talk about each and every one of these over time We're working on human or robotics computers, robotics computers of all kinds. we're working on human or robotics computers robotics computers of all kinds We're working on self-driving car computers. we're working on self-driving car computers We're working on satellites. we're working on satellites You have GeForce, which has tensor cores. you have geforce which has tensor cores I just talked about a whole new line of PCs. i just talked about a whole new line of pcs Agriculture equipment, manufacturing equipment, heavy industry equipment will all be agentic. You'll even have a little agentic helper for yourself. Even your base stations, the radio stations of the future are going to be agentic, understanding traffic and thinking about how to coordinate with the other base stations so that you could use as little energy as possible, increase the utilization, the efficiency of the spectral efficiency. Everything will run agents. Today, NVIDIA is largely in the center, but I am pretty certain that there will be tens of billions, hundreds of billions over time of agentic systems, agentic computers that are going to be running around the world. The biggest problem is data. In the case of language models, all the English and all the language that we have on the internet that we trained on was from the perspective of us. Agriculture equipment, manufacturing equipment, heavy industry equipment will all be agentic. agriculture equipment manufacturing equipment heavy industry equipment will all be agentic You'll even have a little agentic helper for yourself. you'll even have a little agentic helper for yourself Even your base stations, the radio stations of the future are going to be agentic, understanding traffic and thinking about how to coordinate with the other base stations so that you could use as little energy as possible, increase the utilization, the efficiency of the spectral efficiency. even your base stations the radio stations of the future are going to be agentic understanding traffic and thinking about how to coordinate with the other base stations so that you could use as little energy as possible increase the utilization the efficiency of the spectral efficiency Everything will run agents. everything will run agents Today, NVIDIA is largely in the center, but I am pretty certain that there will be tens of billions, hundreds of billions over time of agentic systems, agentic computers that are going to be running around the world. today nvidia is largely in the center but i am pretty certain that there will be tens of billions hundreds of billions over time of agentic systems agentic computers that are going to be running around the world The biggest problem is data. the biggest problem is data In the case of language models, all the English and all the language that we have on the internet that we trained on was from the perspective of us. in the case of language models all the english and all the language that we have on the internet that we trained on was from the perspective of us We wrote it and we're reading it. However, in order to create data for AI robotics, it has to be in the perception, the perspective of the robot. Most of the world's video data is from a third person, not first person. Agentic systems, robotic systems, physical AI, the data is the hardest problem. You've seen us move up this ladder. We started with teleoperations, which is basically human demonstration. This is no different than the big breakthrough of reinforcement learning human feedback. We use simulation. This is where Omniverse comes in. This is no different than reinforcement learning verifiable rewards. We use these systems to bootstrap the AI model, the physical AI model. We wrote it and we're reading it. we wrote it and we're reading it However, in order to create data for AI robotics, it has to be in the perception, the perspective of the robot. however in order to create data for ai robotics it has to be in the perception the perspective of the robot Most of the world's video data is from a third person, not first person. most of the world's video data is from a third person not first person Agentic systems, robotic systems, physical AI, the data is the hardest problem. agentic systems robotic systems physical ai the data is the hardest problem You've seen us move up this ladder. you've seen us move up this ladder We started with teleoperations, which is basically human demonstration. we started with teleoperations which is basically human demonstration This is no different than the big breakthrough of reinforcement learning human feedback. this is no different than the big breakthrough of reinforcement learning human feedback We use simulation. we use simulation This is where Omniverse comes in. this is where omniverse comes in This is no different than reinforcement learning verifiable rewards . this is no different than reinforcement learning verifiable rewards We use these systems to bootstrap the AI model, the physical AI model. we use these systems to bootstrap the ai model the physical ai model Eventually, we're able to learn from third person, re-projecting it into first person, and now eventually, through bootstrapping, we have a world foundation model that can understand the physical world from any perspective you want. Third person, first person, outside in, inside out, doesn't matter. This is a big breakthrough indeed. Today, we're announcing Cosmos 3. Cosmos 3 is the frontier of physical AI. We are at the frontier with language models. There are so many people working on it. However, in physical AI, we are absolutely the world's best. I am so proud of the team for doing this. This is the foundation model for all of your work. Eventually, we're able to learn from third person, re-projecting it into first person, and now eventually, through bootstrapping, we have a world foundation model that can understand the physical world from any perspective you want. eventually we're able to learn from third person re-projecting it into first person and now eventually through bootstrapping we have a world foundation model that can understand the physical world from any perspective you want Third person, first person, outside in, inside out, doesn't matter. third person first person outside in inside out doesn't matter This is a big breakthrough indeed. this is a big breakthrough indeed Today, we're announcing Cosmos 3. today we're announcing cosmos 3 Cosmos 3 is the frontier of physical AI. cosmos 3 is the frontier of physical ai We are at the frontier with language models. we are at the frontier with language models There are so many people working on it. there are so many people working on it However, in physical AI, we are absolutely the world's best. however in physical ai we are absolutely the world's best I am so proud of the team for doing this. i am so proud of the team for doing this This is the foundation model for all of your work. this is the foundation model for all of your work Whenever you want to create a robot, whenever you want to create a factory robot or a robot that works in a factory, any kind of robot that involves physical world, you now have a companion, a Cosmos 3 that can understand and reason. It can generate. It can simulate in the loop. It can even be the policy itself. It is on the top of leaderboards all over the world. I am incredibly proud of Cosmos, and today we're announcing Cosmos 3. Let's take a look. Whenever you want to create a robot, whenever you want to create a factory robot or a robot that works in a factory, any kind of robot that involves physical world, you now have a companion, a Cosmos 3 that can understand and reason. whenever you want to create a robot whenever you want to create a factory robot or a robot that works in a factory any kind of robot that involves physical world you now have a companion a cosmos 3 that can understand and reason It can generate. It can simulate in the loop. it can generate. it can simulate in the loop It can even be the policy itself. it can even be the policy itself It is on the top of leaderboards all over the world. it is on the top of leaderboards all over the world I am incredibly proud of Cosmos, and today we're announcing Cosmos 3. i am incredibly proud of cosmos and today we're announcing cosmos 3 Let's take a look. let's take a look

Speaker 3: The real world is infinite and unpredictable. Physical AI needs data, but real-world data is impossible to scale. For physical AI, compute is data. This is Cosmos, an open frontier omni-modal for physical AI, built on a new mixture of transformers architecture. Pixels, action, sound, and language flow into the autoregressive transformer, which reasons, plans, and instructs the diffusion transformer, which generates what comes next. Developers post-train Cosmos across embodiments and use cases. As a VLM, Cosmos watches the physical world, understands what's happening, describing scenes, and flagging what matters. As a world model, Cosmos generates physics-accurate synthetic video from an image, text, or video. As a simulator, Cosmos closes the loop for policy training and evaluation. As the foundation of NVIDIA Omnidreams, an action-conditioned world model, Cosmos predicts the future frame by frame. The real world is infinite and unpredictable. the real world is infinite and unpredictable Physical AI needs data, but real-world data is impossible to scale. physical ai needs data but real-world data is impossible to scale For physical AI, compute is data. for physical ai compute is data This is Cosmos, an open frontier omni-modal for physical AI, built on a new mixture of transformers architecture. this is cosmos an open frontier omni-modal for physical ai built on a new mixture of transformers architecture Pixels, action, sound, and language flow into the autoregressive transformer, which reasons, plans, and instructs the diffusion transformer, which generates what comes next. pixels action sound and language flow into the autoregressive transformer which reasons plans and instructs the diffusion transformer which generates what comes next Developers post-train Cosmos across embodiments and use cases. developers post-train cosmos across embodiments and use cases As a VLM, Cosmos watches the physical world, understands what's happening, describing scenes, and flagging what matters. as a vlm cosmos watches the physical world understands what's happening describing scenes and flagging what matters As a world model, Cosmos generates physics-accurate synthetic video from an image, text, or video. as a world model cosmos generates physics-accurate synthetic video from an image text or video As a simulator, Cosmos closes the loop for policy training and evaluation. as a simulator cosmos closes the loop for policy training and evaluation As the foundation of NVIDIA Omnidreams, an action-conditioned world model, Cosmos predicts the future frame by frame. as the foundation of nvidia omnidreams an action-conditioned world model cosmos predicts the future frame by frame Post-train Cosmos, it becomes a world action model: perceiving, reasoning, planning, generating actions for robots of every kind, for everything that moves. A new kind of data, a new kind of teacher, generated by compute. Cosmos, the foundation for developers of the age of physical AI. Post-train Cosmos, it becomes a world action model: perceiving, reasoning, planning, generating actions for robots of every kind, for everything that moves. post-train cosmos it becomes a world action model perceiving reasoning planning generating actions for robots of every kind for everything that moves A new kind of data, a new kind of teacher, generated by compute. a new kind of data a new kind of teacher generated by compute Cosmos, the foundation for developers of the age of physical AI. cosmos the foundation for developers of the age of physical ai

Speaker 1: It takes data plus compute, gives you AI. Now that we have AI, compute is data. Use Cosmos 3, train a whole bunch of AI models. Cosmos is such an incredible open model system. It's exactly the same as Nemotron. We open the model, we open the data, and we even open how we trained it so that you could enhance it for yourself and turn Cosmos into your proprietary model. We have such incredible partners working with us in so many different industries. Now, the model itself is, of course, the most understandable part of the AI stack. The AI stack is very complicated. It has generators, the model, simulators, and the runtime. Just as it is for agentic systems, these cars, or essentially a physical AI, a agentic robot that is a autonomous vehicle, has also this complicated stack. It takes data plus compute, gives you AI. it takes data plus compute gives you ai Now that we have AI, compute is data. now that we have ai compute is data Use Cosmos 3, train a whole bunch of AI models. use cosmos 3 train a whole bunch of ai models Cosmos is such an incredible open model system. cosmos is such an incredible open model system It's exactly the same as Nemotron. it's exactly the same as nemotron We open the model, we open the data, and we even open how we trained it so that you could enhance it for yourself and turn Cosmos into your proprietary model. we open the model we open the data and we even open how we trained it so that you could enhance it for yourself and turn cosmos into your proprietary model We have such incredible partners working with us in so many different industries. we have such incredible partners working with us in so many different industries Now, the model itself is, of course, the most understandable part of the AI stack. now the model itself is of course the most understandable part of the ai stack The AI stack is very complicated. the ai stack is very complicated It has generators, the model, simulators, and the runtime. it has generators the model simulators and the runtime Just as it is for agentic systems, these cars, or essentially a physical AI, a agentic robot that is a autonomous vehicle, has also this complicated stack. just as it is for agentic systems these cars or essentially a physical ai a agentic robot that is a autonomous vehicle has also this complicated stack Today, we're announcing Alpamayo 2, an open model for self-driving cars. We're working with car companies across the world. If you look at these brands that have signed up for the NVIDIA Hyperion, that are building NVIDIA Hyperion cars, this represents about 80% of the world's cars. The manufacturers represent 80% of the world's cars. We are going to have a whole lot of NVIDIA Hyperion systems that are able to run Alpamayo or anybody else's AV stack. We are also connected into mobility services. Approximately 97% of the world's mobility services are connecting with us so that when we deploy Alpamayo on the Hyperion runtime with the Halos operating system, we will be able to connect to all of these services across the world. Let's take a look at this. Today, we're announcing Alpamayo 2, an open model for self-driving cars. today we're announcing alpamayo 2 an open model for self-driving cars We're working with car companies across the world. we're working with car companies across the world If you look at these brands that have signed up for the NVIDIA Hyperion, that are building NVIDIA Hyperion cars, this represents about 80% of the world's cars. if you look at these brands that have signed up for the nvidia hyperion that are building nvidia hyperion cars this represents about 80% of the world's cars The manufacturers represent 80% of the world's cars. the manufacturers represent 80% of the world's cars We are going to have a whole lot of NVIDIA Hyperion systems that are able to run Alpamayo or anybody else's AV stack. we are going to have a whole lot of nvidia hyperion systems that are able to run alpamayo or anybody else's av stack We are also connected into mobility services. we are also connected into mobility services Approximately 97% of the world's mobility services are connecting with us so that when we deploy Alpamayo on the Hyperion runtime with the Halos operating system, we will be able to connect to all of these services across the world. approximately 97% of the world's mobility services are connecting with us so that when we deploy alpamayo on the hyperion runtime with the halos operating system we will be able to connect to all of these services across the world Let's take a look at this. let's take a look at this

Speaker 3: Hey, Mercedes. Let's go to my favorite sandwich shop. Hey, Mercedes. hey mercedes Let's go to my favorite sandwich shop. let's go to my favorite sandwich shop

Speaker 4: Routing to your destination. Lane is clear. Pulling out to start drive. Nudge left due to the stationary lead vehicle ahead blocking our lane. Slow down to stop at the stop sign controlling the intersection. Stop to yield to the pedestrians since the person is in our lane. Yield to the cut-in vehicle from the left. Nudge left to clear the stopped vehicle blocking on the right. Keep distance to the cut-in vehicle since it is merging into our lane. Nudge left due to the stopped van blocking the right side of our lane. [crosstalk]. Stop to keep distance to the lead vehicle since it is stationary ahead. Keep distance to the vehicle directly ahead in our lane. Keep distance to the vehicle directly ahead in our lane. Stop for the stop sign since the intersection is controlled by a stop sign. Routing to your destination. routing to your destination Lane is clear. lane is clear Pulling out to start drive. pulling out to start drive Nudge left due to the stationary lead vehicle ahead blocking our lane. nudge left due to the stationary lead vehicle ahead blocking our lane Slow down to stop at the stop sign controlling the intersection. slow down to stop at the stop sign controlling the intersection Stop to yield to the pedestrians since the person is in our lane. stop to yield to the pedestrians since the person is in our lane Yield to the cut-in vehicle from the left. yield to the cut-in vehicle from the left Nudge left to clear the stopped vehicle blocking on the right. nudge left to clear the stopped vehicle blocking on the right Keep distance to the cut-in vehicle since it is merging into our lane. keep distance to the cut-in vehicle since it is merging into our lane Nudge left due to the stopped van blocking the right side of our lane. nudge left due to the stopped van blocking the right side of our lane [crosstalk] . [crosstalk] Stop to keep distance to the lead vehicle since it is stationary ahead. stop to keep distance to the lead vehicle since it is stationary ahead keep Keep distance to the vehicle directly ahead in our lane. keep distance to the vehicle directly ahead in our lane Keep distance to the vehicle directly ahead in our lane. keep distance to the vehicle directly ahead in our lane Stop for the stop sign since the intersection is controlled by a stop sign. stop for the stop sign since the intersection is controlled by a stop sign Stop to yield to cross-traffic since a vehicle is crossing ahead. Keep distance to the lead vehicle. Nudge right due to the truck blocking the right side of our lane. Nudge right due to the truck blocking the left side of our lane. Nudge left due to the truck blocking the right side of our lane. Keep distance to the lead vehicle. Your destination is on the right. Stop to yield to cross-traffic since a vehicle is crossing ahead. stop to yield to cross-traffic since a vehicle is crossing ahead Keep distance to the lead vehicle. keep distance to the lead vehicle Nudge right due to the truck blocking the right side of our lane. nudge right due to the truck blocking the right side of our lane Nudge right due to the truck blocking the left side of our lane. nudge right due to the truck blocking the left side of our lane Nudge left due to the truck blocking the right side of our lane. nudge left due to the truck blocking the right side of our lane keep Keep distance to the lead vehicle. of our lane keep distance to the lead vehicle Your destination is on the right. your destination is on the right

Speaker 1: Alpamayo. The world's first reasoning autonomous vehicle. If you let it talk all the time, it will drive you crazy. We're very happy that it's talking to itself all the time. That's called thinking. Alpamayo is a reasoning car. The technology that we've created also applies to humanoids. Of course, there are many new breakthroughs that has to happen. The NVIDIA Isaac GR00T is our humanoid robotic stack. Model, data generation, simulation, the runtime, including the operating system. This represents GR00T platform, the Isaac GR00T platform. Every one of our systems, as you can see, the exact same pattern, whether it's agentic system for the cloud, Agentic system for the PC, a robotic system for a self-driving car, a robotic system for a human or robot, all the same. In every single case, we build everything completely. Alpamayo . alpamayo The world's first reasoning autonomous vehicle. the world's first reasoning autonomous vehicle If you let it talk all the time, it will drive you crazy. if you let it talk all the time it will drive you crazy We're very happy that it's talking to itself all the time. we're very happy that it's talking to itself all the time That's called thinking. that's called thinking Alpamayo is a reasoning car. alpamayo is a reasoning car The technology that we've created also applies to humanoids. the technology that we've created also applies to humanoids Of course, there are many new breakthroughs that has to happen. of course there are many new breakthroughs that has to happen The NVIDIA Isaac GR00T is our humanoid robotic stack. the nvidia isaac gr00t is our humanoid robotic stack Model, data generation, simulation, the runtime, including the operating system. model data generation simulation the runtime including the operating system This represents GR00T platform, the Isaac GR00T platform. this represents gr00t platform the isaac gr00t platform Every one of our systems, as you can see, the exact same pattern, whether it's agentic system for the cloud, Agentic system for the PC, a robotic system for a self-driving car, a robotic system for a human or robot, all the same. every one of our systems as you can see the exact same pattern whether it's agentic system for the cloud, agentic system for the pc a robotic system for a self-driving car a robotic system for a human or robot all the same In every single case, we build everything completely. in every single case we build everything completely We build everything vertically, completely integrated with co-design, extreme co-design. Then we open it up for everybody to use whichever part you like. Whatever you want to use, we even help you modify. The one thing that is missing is we need a reference platform for robotic systems. These robotic systems are so complicated, so many motors, so many sensors, so fragile, and yet we need to have a way to deliver these reference platforms just like we do with PCs and DGXs and clouds and self-driving cars. We now are going to do it for robots. Today we're announcing the NVIDIA Isaac GR00T, a reference humanoid robot, all fully integrated, 25 degrees of freedom on each hand made by Sharpa. 31 degrees of freedom on the robot, 6 ft, 150 pounds, just like me. The first number is shorter, the second number is bigger. We build everything vertically, completely integrated with co-design, extreme co-design. we build everything vertically completely integrated with co-design extreme co-design Then we open it up for everybody to use whichever part you like. then we open it up for everybody to use whichever part you like Whatever you want to use, we even help you modify. whatever you want to use we even help you modify The one thing that is missing is we need a reference platform for robotic systems. the one thing that is missing is we need a reference platform for robotic systems These robotic systems are so complicated, so many motors, so many sensors, so fragile, and yet we need to have a way to deliver these reference platforms just like we do with PCs and DGXs and clouds and self-driving cars. these robotic systems are so complicated so many motors so many sensors so fragile and yet we need to have a way to deliver these reference platforms just like we do with pcs and dgxs and clouds and self-driving cars We now are going to do it for robots. we now are going to do it for robots Today we're announcing the NVIDIA Isaac GR00T, a reference humanoid robot, all fully integrated, 25 degrees of freedom on each hand made by Sharpa. 31 degrees of freedom on the robot, 6 ft , 150 pounds, just like me. today we're announcing the nvidia isaac gr00t a reference humanoid robot all fully integrated 25 degrees of freedom on each hand made by sharpa 31 degrees of freedom on the robot 6 ft 150 pounds just like me The first number is shorter, the second number is bigger. the first number is shorter the second number is bigger Otherwise, pretty close. This platform runs the new Thor and our entire software stack. Data generation stack, data simulation stack, the runtime, all integrated into a robot that is designed for everyone to use. We built this for higher education and university researchers because for them to build this is insanely hard to do. Let's take a look at that. Otherwise, pretty close. otherwise pretty close This platform runs the new Thor and our entire software stack. this platform runs the new thor and our entire software stack Data generation stack, data simulation stack, the runtime, all integrated into a robot that is designed for everyone to use. data generation stack data simulation stack the runtime all integrated into a robot that is designed for everyone to use We built this for higher education and university researchers because for them to build this is insanely hard to do. we built this for higher education and university researchers because for them to build this is insanely hard to do Let's take a look at that. let's take a look at that

Speaker 3: The next leap in AI is general-purpose robots, humanoids. Building one is hard. Every team starts from scratch, stitching together simulators, teleop systems, data pipelines, and training infrastructure. Months of setup before research can start. NVIDIA Isaac GR00T, an open development platform for humanoid robots. Open models, simulation and training libraries, and data generators. Plus the robot computer. Fully pipe clean, ready to go in hours. First, set up the simulation environment in Isaac Lab. The next leap in AI is general-purpose robots, humanoids. the next leap in ai is general-purpose robots humanoids Building one is hard. building one is hard Every team starts from scratch, stitching together simulators, teleop systems, data pipelines, and training infrastructure. every team starts from scratch stitching together simulators teleop systems data pipelines and training infrastructure Months of setup before research can start. months of setup before research can start NVIDIA Isaac GR00T, an open development platform for humanoid robots. nvidia isaac gr00t an open development platform for humanoid robots Open models, simulation and training libraries, and data generators. open models simulation and training libraries and data generators Plus the robot computer. plus the robot computer Fully pipe clean, ready to go in hours. fully pipe clean ready to go in hours First, set up the simulation environment in Isaac Lab. first set up the simulation environment in isaac lab Capture demonstrations with Isaac Teleop on a real or simulated robot. Generate synthetic data with Omniverse and Cosmos, scaling one demonstration into thousands. Train policies. Evaluate them in Isaac Lab Arena. Deploy through Isaac ROS, running on Jetson Thor. Every element, modular, open. Use ours or swap in your own. GR00T is powering robotics research across every discipline for every domain, from research labs to factory floors. One open platform. Now a new addition, Isaac GR00T reference design robots. Built on NVIDIA's open platform, ready for frontier research for any lab, anywhere. The age of robotics starts here. NVIDIA Isaac GR00T. Capture demonstrations with Isaac Teleop on a real or simulated robot. capture demonstrations with isaac teleop on a real or simulated robot Generate synthetic data with Omniverse and Cosmos, scaling one demonstration into thousands. generate synthetic data with omniverse and cosmos scaling one demonstration into thousands Train policies. train policies Evaluate them in Isaac Lab Arena. evaluate them in isaac lab arena Deploy through Isaac ROS, running on Jetson Thor. deploy through isaac ros running on jetson thor Every element, modular, open. every element modular open Use ours or swap in your own. use ours or swap in your own GR00T is powering robotics research across every discipline for every domain, from research labs to factory floors. gr00t is powering robotics research across every discipline for every domain from research labs to factory floors One open platform. one open platform Now a new addition, Isaac GR00T reference design robots. now a new addition isaac gr00t reference design robots Built on NVIDIA's open platform, ready for frontier research for any lab, anywhere. built on nvidia's open platform ready for frontier research for any lab anywhere The age of robotics starts here. the age of robotics starts here NVIDIA Isaac GR00T. nvidia isaac gr00t

Speaker 1: Many robots. We're working with just about everybody who's working on robots in the world or robotic systems in world. Let me tell you what I told you. The computer industry has been completely changed. In the last six months, everything changed. Everything changed because agents were realized, and it converged with the latest frontier models, and it made possible the AI to now do useful work. The computing pattern will repeat over and over again. This computing pattern of an agent that's a model, a harness that uses tools with skills and runs in a runtime, that runtime depends on whether it's in the cloud or on-prem, on a PC or in a robot. The computing pattern is exactly the same for all of them. You will use different harnesses because of your preference. Many robots. many robots We're working with just about everybody who's working on robots in the world or robotic systems in world. we're working with just about everybody who's working on robots in the world or robotic systems in world Let me tell you what I told you. let me tell you what i told you The computer industry has been completely changed. the computer industry has been completely changed In the last six months, everything changed. in the last six months everything changed Everything changed because agents were realized, and it converged with the latest frontier models, and it made possible the AI to now do useful work. everything changed because agents were realized and it converged with the latest frontier models and it made possible the ai to now do useful work The computing pattern will repeat over and over again. the computing pattern will repeat over and over again This computing pattern of an agent that's a model, a harness that uses tools with skills and runs in a runtime, that runtime depends on whether it's in the cloud or on-prem, on a PC or in a robot. this computing pattern of an agent that's a model a harness that uses tools with skills and runs in a runtime that runtime depends on whether it's in the cloud or on-prem on a pc or in a robot The computing pattern is exactly the same for all of them. the computing pattern is exactly the same for all of them You will use different harnesses because of your preference. you will use different harnesses because of your preference You'll use different models because of your preference. You will improve them for your proprietary use. You would create super agents that you can rent to other people to help them do their work. This agentic platform, this agentic pattern, NVIDIA has an enterprise AI toolkit. This is a wonderful way for all of you to engage AIs, and for us, it's a wonderful growth opportunity. Vera Rubin is in full production. Whereas Grace Blackwell was created to process AI, particularly inference, Vera Rubin was created to run agents. It is in full production. It is much more than a GPU. It is an entire disaggregated, distributed agent processing system. You'll use different models because of your preference. you'll use different models because of your preference You will improve them for your proprietary use. you will improve them for your proprietary use You would create super agents that you can rent to other people to help them do their work. you would create super agents that you can rent to other people to help them do their work This agentic platform, this agentic pattern, NVIDIA has an enterprise AI toolkit. this agentic platform this agentic pattern nvidia has an enterprise ai toolkit This is a wonderful way for all of you to engage AIs, and for us, it's a wonderful growth opportunity. this is a wonderful way for all of you to engage ais and for us it's a wonderful growth opportunity Vera Rubin is in full production. vera rubin is in full production Whereas Grace Blackwell was created to process AI, particularly inference, Vera Rubin was created to run agents. whereas grace blackwell was created to process ai particularly inference vera rubin was created to run agents It is in full production. it is in full production It is much more than a GPU. it is much more than a gpu It is an entire disaggregated, distributed agent processing system. it is an entire disaggregated distributed agent processing system NVIDIA has really become an infrastructure company, not just a GPU company, not just a systems company, but an infrastructure company to help you generate the maximum revenues, the maximum profit, and to get there as soon as possible. The agent world, this new way of doing computing, where you build CPUs now for agents, not for people. CPUs for agents has its own special requirement, and our NVIDIA Vera is revolutionary. I'm so happy about its ramp. NVIDIA has really become an infrastructure company, not just a GPU company, not just a systems company, but an infrastructure company to help you generate the maximum revenues, the maximum profit, and to get there as soon as possible. nvidia has really become an infrastructure company not just a gpu company not just a systems company but an infrastructure company to help you generate the maximum revenues the maximum profit and to get there as soon as possible The agent world, this new way of doing computing, where you build CPUs now for agents, not for people. CPUs for agents has its own special requirement, and our NVIDIA Vera is revolutionary. the agent world this new way of doing computing where you build cpus now for agents not for people. cpus for agents has its own special requirement and our nvidia vera is revolutionary I'm so happy about its ramp. i'm so happy about its ramp The orders already is going to make it the fastest and the most successful product launch in our company's history. NVIDIA and Microsoft has created a whole new line of PCs. This is a new beginning. Of course, that exact same agentic processing pattern, computing pattern that I just described is also going to run on all kinds of devices. I mentioned PCs, in the future, it'll be robots and satellites and base stations and factories in the cloud, on-prem, at the edge. This pattern, agentic AI system, this agentic computing pattern, will be replicated in computers all over. The orders already is going to make it the fastest and the most successful product launch in our company's history. the orders already is going to make it the fastest and the most successful product launch in our company's history NVIDIA and Microsoft has created a whole new line of PCs. nvidia and microsoft has created a whole new line of pcs This is a new beginning. this is a new beginning Of course, that exact same agentic processing pattern, computing pattern that I just described is also going to run on all kinds of devices. of course that exact same agentic processing pattern computing pattern that i just described is also going to run on all kinds of devices I mentioned PCs, in the future, it'll be robots and satellites and base stations and factories in the cloud, on-prem, at the edge. i mentioned pcs in the future it'll be robots and satellites and base stations and factories in the cloud on-prem at the edge This pattern, agentic AI system, this agentic computing pattern, will be replicated in computers all over. this pattern agentic ai system this agentic computing pattern will be replicated in computers all over How we think about the personal computer will very likely change. I want to thank all of you for your partnership, your friendship. We couldn't be here without everything that we do together. I am so proud of how you've been so successful this last year. The next year is going to be even more. I have one more thing for you. Let's take a look. How we think about the personal computer will very likely change. how we think about the personal computer will very likely change I want to thank all of you for your partnership, your friendship. i want to thank all of you for your partnership your friendship We couldn't be here without everything that we do together. we couldn't be here without everything that we do together I am so proud of how you've been so successful this last year. i am so proud of how you've been so successful this last year The next year is going to be even more. the next year is going to be even more I have one more thing for you. i have one more thing for you Let's take a look. let's take a look

Speaker 5: You ready, Taiwan? Let's do this. The keynote's done at Computex. Jensen showed the world what's next. Useful AI has arrived. Agents working by your side. In case you missed things we said today. We're gonna break it all down for you, Taipei. Agents used to be misunderstood. Only movie stars had them in Hollywood. Now we all got teams making dreams come true. Building companies from living rooms. They need so much compute, we hear you. That's why we created Vera. Rubin stole the show, it's true. The cheapest tokens coming through. 10x faster, inference heaven. More special agents than 007. BlueField keeps agents' memory true. Now, let's talk about its CPU. 50% faster, that's outrageous. Not for Vera. It's built for agents. NVLink Fusion blends ASICs smartly. You ready, Taiwan? you ready taiwan Let's do this. let's do this The keynote's done at Computex. the keynote's done at computex Jensen showed the world what's next. jensen showed the world what's next Useful AI has arrived. useful ai has arrived Agents working by your side. agents working by your side In case you missed things we said today. in case you missed things we said today We're gonna break it all down for you, Taipei. we're gonna break it all down for you taipei Agents used to be misunderstood. agents used to be misunderstood Only movie stars had them in Hollywood. only movie stars had them in hollywood Now we all got teams making dreams come true. now we all got teams making dreams come true Building companies from living rooms. building companies from living rooms They need so much compute, we hear you. they need so much compute we hear you That's why we created Vera. that's why we created vera Rubin stole the show, it's true. rubin stole the show it's true The cheapest tokens coming through. 10x faster, inference heaven. the cheapest tokens coming through 10x faster inference heaven More special agents than 007. more special agents than 007 BlueField keeps agents' memory true. bluefield keeps agents' memory true Now, let's talk about its CPU. 50% faster, that's outrageous. now let's talk about its cpu 50% faster that's outrageous Not for Vera. not for vera It's built for agents. it's built for agents NVLink Fusion blends ASICs smartly. nvlink fusion blends asics smartly Everyone's welcome to the NVLink party. Well, if you like that introduction. Vera Rubin's in full production. Nemotron Ultra leads the run. 5x faster, work gets done. NeMo Cloud keeps the guardrails right. OpenShell keeps the sandbox tight. Your code migrated and reviewed. All before this song is through. AI is a five-layer cake. Compute's revenue, make no mistake. Global AI clouds build lots of gigawatts. DSX keeps power lean, connecting dots. Every watt optimized for you. You can have your cake. Eat it, too. RTX 40 is finally here. Biggest PC moment in 40 years. For you. Agents powering all workflows. Everyone's welcome to the NVLink party. everyone's welcome to the nvlink party Well, if you like that introduction. well if you like that introduction Vera Rubin's in full production. vera rubin's in full production Nemotron Ultra leads the run. 5x faster, work gets done. nemotron ultra leads the run 5x faster work gets done NeMo Cloud keeps the guardrails right. nemo cloud keeps the guardrails right OpenShell keeps the sandbox tight. openshell keeps the sandbox tight Your code migrated and reviewed. your code migrated and reviewed All before this song is through. all before this song is through AI is a five-layer cake. ai is a five-layer cake Compute's revenue, make no mistake. compute's revenue make no mistake Global AI clouds build lots of gigawatts. global ai clouds build lots of gigawatts DSX keeps power lean, connecting dots. dsx keeps power lean connecting dots Every watt optimized for you. every watt optimized for you You can have your cake. you can have your cake Eat it, too. eat it too RTX 40 is finally here. rtx 40 is finally here Biggest PC moment in 40 years. For you. biggest pc moment in 40 years. for you Agents powering all workflows. agents powering all workflows Running anywhere Windows goes. Harnesses run on CPU. Models fly on GPU. Cosmos builds worlds that robots need. Turning compute into synthetic feed. Alpamayo sees and reasons through. Understands roads like people do. GR00T is how they learn to move. Learning skills and finding groove. JuliUs is powered by Thor. The future's humanoid. Count on more. [audio distortion] The future's bright. Come see what's next. Thank you, Taiwan. Welcome to Computex. Running anywhere Windows goes. running anywhere windows goes Harnesses run on CPU. harnesses run on cpu Models fly on GPU. models fly on gpu Cosmos builds worlds that robots need. cosmos builds worlds that robots need Turning compute into synthetic feed. turning compute into synthetic feed Alpamayo sees and reasons through. alpamayo sees and reasons through Understands roads like people do. understands roads like people do GR00T is how they learn to move. gr00t is how they learn to move Learning skills and finding groove. learning skills and finding groove JuliUs is powered by Thor. julius is powered by thor The future's humanoid. the future's humanoid Count on more. count on more [audio distortion] The future's bright. [audio distortion] the future's bright Come see what's next. come see what's next Thank you, Taiwan. thank you taiwan Welcome to Computex. welcome to computex

Speaker 1: Have a great Computex. Thanks for an amazing year. Thank you for all your friendship and support. Thank you. Take care. Have a great Computex. have a great computex Thanks for an amazing year. thanks for an amazing year Thank you for all your friendship and support. thank you for all your friendship and support Thank you. thank you Take care. take care