SPEAKER_00: you look at the existing cloud infrastructure that was built over the last decade, it was built for serializable workloads. It wasn't built for parallelizable workloads. And it's like you're having to rebuild the cloud, so to say, and you're having to rebuild physical infrastructure at the pace of AI software adoption. It's a mind-blowing concept, right? Because AI software is being adopted at the most rapid scale of any technology that we've ever observed. I mean, we're building at, I think it's 28 data centers this year across North America. We're one of the largest operators SPEAKER_01: of this infrastructure in the world. And we are unable to keep up with demand. And we really don't SPEAKER_03: see that subsiding for years to come. This Week in Startups is brought to you by OpenPhone. Create business phone numbers for you and your team that work through an app on your smartphone or desktop. Twist listeners can get an extra 20% off any plan for your first six months at openphone.com slash twist. Gusto is easy online payroll, benefits, and HR built for modern small businesses. Get three months free when you run your first payroll at gusto.com slash twist. And Northwest Registered Agent. When starting your business, it's important to use a service that will actually help you. Northwest Registered Agent is that service. They'll form your company fast, give you the documents you need to open a business bank account, and even provide you with mail scanning and a business address to keep your personal privacy intact. Visit NorthwestRegisteredAgent.com slash twist to get a 60% discount on your next LLC. All right, everybody. Welcome back to this SPEAKER_08: Week in Startups. We've got a great guest for you today. You may have been wondering who's buying all of SPEAKER_09: these Nvidia H100s and how are people getting access to all of this hardware? Well, there's a couple of companies that got to the hosting of AI and GPUs early. One of those companies is CoreWeave. They started in the space, I believe, doing a lot of crypto where miners were renting GPUs from their cluster. And fortune favors the bold. They were in the catbird seat when the AI revolution happened. And everybody decided, well, they got to train their own models, going to need a bunch of Nvidia's hardware and other people's hardware. And we'll talk about that today. And they've since grown to a massive scale. For those of you who don't know, Nvidia's market cap has increased more than 7x from 300 billion to 2.3 trillion. If you've been living under a buck and you haven't been watching this, it's because people want access to these chips. SPEAKER_12: Welcome to the program. Brandon McBee, who is the CDO and co-founder. What does CDO stand for? SPEAKER_00: Chief development officer. So my role is raising capital for the business. I interface with equity SPEAKER_16: and debt participants for the company and help fuel the growth of the business. SPEAKER_17: And this is a very capital intensive business. You spent a lot of money on GPUs and setting up SPEAKER_09: infrastructure. The company's been around for just under a decade. Am I correct? SPEAKER_22: Yes, we founded the company in 2018. Got it. And am I also correct that you were supplying GPUs largely to crypto and Bitcoin miners and this cohort of individuals? Or that was the beachhead market? SPEAKER_00: It was absolutely the beachhead market. So I can take a couple steps back on our founding story. So it's myself, my two co-founders. We're from the institutional commodity trading sector. So we're SPEAKER_14: risk managers by background. We're from hedge funds, finance guys, is probably the best way to look at it. But we were finance guys who were heavily data oriented. We worked in this commodity sector that you can actually solve for price. You can figure out supply demand. And there's a dollar per barrel of oil, so to say, that solves market dynamics. So we've always worked with a lot of compute. We've worked with a lot of software. And the crypto space was interesting because it was an arbitrage opportunity. There's a very discrete input price, cost of power. And you could model the revenue very efficiently because there was no customers. You were just participating in this network. And thus, if you were just to sell the revenue from the crypto mining proceeds every day, you could effectively qualify it as an arbitrage opportunity. And that was interesting to us. SPEAKER_00: But it wasn't as compelling as a large business because at the end of the day, all you're going to do is chase the price of power lower. That's the only advantage you can really extract unless you expand into other markets, right? So when we were looking at the cryptocurrency space, it wasn't Bitcoin mining that we were interested in. It was Ethereum mining and GPU oriented mining because a Bitcoin miner, ASIC, things that are produced by entities like Bitmain, they can only do that one thing. They can only participate in Bitcoin and they're very good at it. But a GPU, well, it can do lots of things, including running AI workloads. So we started in the crypto space, but it was always with this idea. And we had no idea how complicated an idea was at the time, but started with the idea that, well, you could do crypto and other things. Right. It started there. SPEAKER_09: The other use of these is, I guess, running video games in the cloud. Is that correct? Yes, cloud video games become a real market or do people who are into video games just buy themselves an Alienware, Dell, whatever, and be done with it? SPEAKER_00: I think it's more the latter that they use it more for that. We certainly don't see that demand SPEAKER_14: for video game streaming. I think a few of the hyperscalers tried to launch into that market and I don't believe that there's been a substantial demand. SPEAKER_36: Got it. And so when we look at crypto, that market kind of fizzled right as AI was starting SPEAKER_09: to boom. So you were able to sort of just navigate that or is the crypto still going on and people are SPEAKER_38: still using your services for Ethereum and being part of that network? Or is that just too hard of an arbitrage now because people in China have stolen electricity, we hear, or their friend runs the hydro dam. So they run an extension cord, so to speak, over to their warehouse with a bunch of servers in it. And you're up against people getting zero cost of input electricity. SPEAKER_40: It's a great question. And yes, I'm extremely excited that we're not involved in that cryptocurrency market anymore. We haven't been involved for a number of years at this point. We actually started making this transition into the cloud infrastructure market in 2019. We hired Peter Selenke, who was recently elevated to our chief technology officer position in 2019, to build a cloud for us. And it's not just plugging in GPUs and having users come SPEAKER_00: access it. It's this really complicated software stack that runs the cloud. Or in other words, it's an orchestration environment that enables users to access and use our infrastructure. And it's one that it's very different than the way that the hyperscalers built it because they built for hosting websites and storing data lakes. And we built our cloud from a no compromises engineering solution for running AI workloads and highly parallelizable workloads. And there's engineering SPEAKER_14: decisions you make in doing that that you wouldn't make for hosting websites. And that's allowed for us to, I would say, overperform in market with a product that really doesn't have competition. SPEAKER_40: We started in 2019. SPEAKER_09: Yeah, the software layer to provision these H100s, A100s, whatever people are using, SPEAKER_38: that's a key part of the puzzle that you have to build, you have to master. And AWS is Google's cloud, everybody's Azure, they're all slightly different in using their own provisioning software. Or is there some open source standard there for doing all that? SPEAKER_14: That's exactly correct. We use and contribute to open source as much as we can, but we have a SPEAKER_40: proprietary orchestration solution that looks different than the hyperscalers do. My favorite SPEAKER_14: analogy for this actually comes out of the automobile sector, where at the end of the day, everyone produces vehicles the same way. From research, design, scaling, servicing, it's the same sort of product with different badges and different colors on it, right? And it's been that way for 60 plus years. And then in the 2000s, a company came along and said, well, what if we started with a blank slate and design this process today? And ultimately, Ford might have to produce vehicles like Tesla does, but I think we can all appreciate the SPEAKER_00: foundational difference in the way that those vehicles have brought to market and the challenges that Ford will have to go through to get there. And I think it's a lot of the same for the hyperscalers, right? I'm not going to tell you a trillion dollar company with tens of thousands of engineers can't do what we do, but I will highlight the innovators dilemma that sits there because there's an existing product, you have to, you know, change everything underneath to run infrastructure like SPEAKER_14: we do and that it's going to be hard to get there. And these, uh, we go through the economics of let's SPEAKER_38: just say NH 100. This is Nvidia's state of the art. It's not just a GPU. It's a rack. Essentially, it's a platform to put many GPUs on. I'm not sure exactly how many of the H 100 holds, SPEAKER_08: um, but it holds a number of GPUs. They go for like 30, 40 grand each is my understanding. SPEAKER_14: So it's a, um, it's a server or node as we call it within a network fabric and a server has SPEAKER_40: typically eight GPUs within it. Um, and then you put those into a cabinet and you put those into a data center and bring power into the data center and you connect it to the internet. Right. And then, yes, those things, um, that, that price range was, was accurate for each one of those GPUs in a SPEAKER_14: server. So a server can cost upwards of a quarter million dollars. And so you rent out one of those SPEAKER_38: H 100 GPUs, an individual GPU in a server like that, a cluster, a node for four bucks an hour, SPEAKER_57: something to that effect. Yeah. Yeah. Yeah. That that's right. Um, where we specialize though, SPEAKER_14: is doing it at scale, right? Like we don't have many clients who just use one at a time. Our clients will use 10,000 at a time and a single contiguous fabric, which makes it a supercomputer. And it's, it's interesting. Like this actually has become where we operate some of the world's largest supercomputers at this point. I think, you know, several of the top 10 now sit on our platform because SPEAKER_30: of how large these fabrics are and how performant these GPUs are at these specific tasks. SPEAKER_09: So somebody fires up 10,000 of those, I'm assuming they get some kind of volume discount. So if it was two or $3 an hour, they're spending 20, $30,000 an hour on a job at one of those, correct? Something SPEAKER_00: in that range. Yes. But I, I will correct the, the discounts. It actually works inverse, right? Cause it's, uh, it's extremely surprising, but it's because building of a single fabric of this size is so engineering intensive that not many people are able to do it. Uh, right. Like there's not a template. Not a lot of companies have gone out and done it. It's actually, you know, maybe three or four on the planet. We're actually building fabrics of this scale. So you actually make a more scarce resource through scaling or you're, you're kind of like decommoditizing the market through scale. SPEAKER_14: Got it. Like it's not one GPU or 10,000 GPUs. It's, oh, 10,000 GPUs. That's a totally different SPEAKER_65: engineering solution. SPEAKER_67: Juggling multiple devices and apps to run your business is a mess. Open phone is here to make it simple by simplifying your business communications with one easy to use app. Open phone has rethought every detail of what a modern business phone should be. And here's the magic. It works through a beautiful, elegant app on your phone, or you can just use it on your desktop, making it super easy to get a business phone number for your entire team. And you know how brilliant open phone is? My teams use it every single day. My sales team loves it. My ops team, they use it all day long. And here's the features that we love. You can create a shared phone number like customer support with multiple employees fielding all the calls and all the texts to that one number. At my investment firm launch, we pride ourselves on replying to every single call or email instantly. And open phone is the number one rated business phone on G2 for customer satisfaction. So here's your call to action. Super easy. Open phone is already affordable. Starts at just 13 bucks a month, but twist listeners get an extra 20% off any plan for the first six months at open phone.com slash twist. And if you have existing numbers with other services, no problem. Open phone is going to port them over easy peasy lemon squeezy, no extra cost. Head over to open phone.com slash twist to start your free trial and SPEAKER_09: get 20% off. And how much of that costs when we look at a $4 an hour cost, would you say energy is because these things are extremely energy reliant, right? They consume a lot of energy. So I'm curious how much of all this is energy? And then where do you put your this will take us down the energy rabbit hole, but where do you put your data centers? And then where are data centers going to be in the future? Because these GPUs are taking a multiple of what CPUs SPEAKER_73: would use. Yeah. So maybe you could explain that to us. Yeah. Yeah. It actually causes, SPEAKER_00: you know, a pretty substantial bottleneck that exists in the market now is data center capacity, right? And it's not square footage of data centers. It's data centers that have enough SPEAKER_14: power brought into them and adding more power to those data centers. And arguably that's where like the next bottleneck in this cloud infrastructure or GPU cloud infrastructure market sits is how do we access enough data center space to accommodate the volume of demand that's coming in? So power, you know, it's roughly about 10% of our cost to deliver this infrastructure. The infrastructure itself is actually where most of the cost sits from a depreciation perspective, depreciate over a six year life on the infrastructure. The power side is, that's my background, right? It was in power markets and trading different electricity markets. So it consumes a lot of power. It has an immense amount of efficiency over CPU infrastructure as well for what it's doing there, you know, to run the same workloads on GPU versus CPU. It's actually more power efficient to run it on a GPU, right? If you're trying to achieve the same outcome, right? Cause you would have to use so many CPU cores to get to the same solution. So SPEAKER_01: yes, they consume more power on a density basis, but on a workload basis, they're more efficient. SPEAKER_38: Yeah. And this is, I mean, it's staggering. I was, I was looking at one study that said like SPEAKER_09: one of these GPUs, um, at 60, 70% capacity years, like the average American households, energy consumption. And that's just one of them. So this would be the equivalent of like, if somebody is using 10,000 of these, or I think Zuckerberg's going to be using low millions of these, SPEAKER_38: it's like putting a million households online or something to that effect. Yeah. SPEAKER_40: Yeah. Um, so yes, it's an immense amount, but it's also a, you know, it's a transformational SPEAKER_14: technology. I'd say that we're looking at and it's, and it's ability to unlock value from data is something that we've never observed before. Um, it going to, yeah. So that's justifiable. Like, SPEAKER_38: yeah, I mean, I'm not even, I'm not even looking at this judgmentally. Like, is it worth SPEAKER_09: the amount of energy it's consuming? I was looking pragmatically where let's assume it is worth it. It comes from. Yeah. Let's say it's going to cure cancer. It's, it's going to find, uh, solutions for renewables or fusion that like we didn't even conceive of, or the gains from it will be so SPEAKER_38: extraordinary. It will obviously pay for itself and create an energy independent future. But what's happening in the industry today, as people are buying these and looking for places to store them, you're looking to build up your infrastructure. Are we just out of energy and where are people, where are the nooks and crannies where people are looking to locate these facilities? I heard that SPEAKER_09: nuclear power plants are going to become like a place where people put these, uh, uh, the plan is SPEAKER_86: to put a nuclear power plant and these GPU data centers next to each other. Yeah. Is there any truth to SPEAKER_00: that? Yeah. Yeah. Look, I believe that was Microsoft or Amazon is, is effectively taking SPEAKER_40: that nuclear plant and citing a data center next to it power. I think, you know, beyond data center space, right. It is a national concern. I'd, I'd, so to say there's been an immense amount building out, uh, renewable capacity over the past decade, which is fantastic, but it's, it's also a, SPEAKER_00: not necessarily the, the right kind of capacity, what you need for consistent demand growth, right. As you know, solar works when the sun's out, wind works, when the wind blows, neither of those things work for a data center, um, or even necessarily for electric vehicles, um, for all these like kind of demand areas, right. We need more baseload power. That's SPEAKER_14: traditionally come from coal and natural gas. Um, fortunately it's been more so from natural gas SPEAKER_00: over the last decade. Uh, cause coal is quite dirty from an emissions perspective. And you know, my personal hope is that it's more nuclear going forward, uh, but it takes time to build nuclear SPEAKER_16: sites, right. So I think it's a decade for, you know, citing and build and in the United States. SPEAKER_38: And we haven't built one in a long time. I mean, I think the last one broke ground in the late sixties or early seventies, and we haven't had one since. So this would lead to lead one to believe that somebody with a lot of nuclear power and a lot of GPUs would have a massive advantage. SPEAKER_14: It, it certainly helps. Um, you know, we've, we've contracted a substantial amount of capacity, uh, for looking to ensure the growth profile of, of our business, but it, it is going to be a bottleneck, uh, for, for all other participants in the market. What about heat? Uh, these things throw SPEAKER_38: off a lot of heat. Um, and you know, some areas in the country are warmer than others. Is it, are people moving these data centers north in order to get the cold air to just, you know, you know, we've, we've seen pictures of, you know, data centers that have open sides where cold air just blows right in or, you know, open doors essentially. Um, because in other places, if you were to put these GPUs in Texas, I think you're going to be air conditioning them. Um, SPEAKER_09: which seems doubly inefficient. So maybe talk a little bit about the heat these things generate SPEAKER_11: today and there's any hope of cooling them down, um, without air conditioning. SPEAKER_14: Yeah. So it's, it's a great question. Um, it's funny, like it takes me vividly and visually back to my crypto mining days where we did run those warehouses with the open sides and the giant fans. And we were up north could never run this infrastructure in those environments, right. From, from a security perspective, from a reliability perspective, like it is mandatory to run this infrastructure and what's called a tier four data center environment, where sometimes even a tier five data center, and that's the highest classification in terms of SPEAKER_40: reliability, redundancy, security, and environmental handling, right? So these are sites that you would SPEAKER_14: see like Amazon or Google or Microsoft running within, right? Like true data centers that are meant for cloud infrastructure and the way to think about the heat output, the other variable in there, which is wild is actually the sound is there extremely loud in these environments, you know, about a hundred decibels, right? Which is, you know, a direct derivative of the heat they're consuming, right? And then to move the air. But, um, the way you look at it is the critical load around the infrastructure, right? So if it takes one unit of energy to run the infrastructure, it takes another 0.2 to 0.3 units of energy to cool the infrastructure and run the networking and everything else around it. The way that the world's leading sites handle this is just through forced air, right? Just move tens of thousands of cubic feet per minute of air through these highly contained pods. So you have all the hot air in a really small area and you're just jamming air through it, right? And sometimes it's conditioned, sometimes it's just air. Um, but eventually it's going to be liquid, right? We're going to move to an environment where you have direct chip liquid cooling instead. And that efficiency ratio call it 1.3 will drop to about 1.1 instead. So your GPU infrastructure will inherently become more energy efficient as we move to a liquid cooled environment. And we're working with, you know, leading data center operators such as switch to, to facilitate and implement that movement for these SPEAKER_01: upcoming generations of GPUs. When people say liquid cooled, um, most people have not actually SPEAKER_09: physically seen that. Yeah. Unless maybe you're a gamer and you've seen your chip in a, a tube run to the chip and there's literally liquid on the top of the chip that's cooling it. Um, how do these, uh, when you say liquid cooled, what could people envision, uh, of how these solutions are going to work? Are they going to just be like a bunch of racks in a, in a, in an Olympic sized swimming pool, or is it just like little contained amounts of water on top of the GPUs? Yeah. So you're qualifying SPEAKER_14: that. Correct. There's two broad categories of liquid cooling. There's immersion cooling, which is the Olympic swimming pool method downsized obviously. And then there's direct to chip liquid cooling, which is running the pipes to the chips. Um, we will sit on the direct chip liquid cooling side cause it's, uh, operationally more efficient for us. Um, that that's where we think that the sector is broadly going to go. If you think of immersion cooling, right? Like you're, you're literally dunking a server into a vat of liquid that liquid has its own problems with it as well. Uh, but let's say you had to go service that, uh, that server, there's a node or component was wrong with it. SPEAKER_00: Well, you got to lift it out of the liquid, right? And what happens then, well, you got to wait probably an hour for all the liquid to drain out now before the tech can even get into it. So you're, you're extending these service times materially in response times versus direct to chip liquid SPEAKER_14: cooling, uh, you know, pop it out and you don't have water containment issues. You know, things splashing across data center. Like it's, it's a mess. These are highly sterile and contained environments that even let us bring cardboard inside of the data center area. Cause it's combustible. Right. And you can have little particulates that float around the site and can accrete into the nodes. SPEAKER_09: Um, bad. You don't want fires and data centers. I mean, when you are talking about that much air being pushed around, that means any particulates in the air are going to get pushed around. And so if you just had some very small amount of particulates floating around the room now, imagine that room is changing the air every X amount of time, the number of particulates is going SPEAKER_38: to grow. Uh, and then you're going to have a small fire on a chip, which is just absolutely crazy. Bad, bad. Do you think the demand is going to keep up? Are you seeing any signs of demand? SPEAKER_09: People saying, okay, we, we built our 10,000 GPUs. We're making more efficient software, making more efficient use of the chips. Okay. Yeah. We're, we're getting to a steady state. We bought enough. So are you starting to see that with your customers saying, you know what, we've got enough GPUs. We've got enough infrastructure right now, or are they still, you know, in the begging, pleading and, you know, uh, doubling, you know, what are you seeing from top customers? Are they doubling their capacity every year? Are they tripling it? What's the, SPEAKER_00: what's the field report? Yeah, I'll, I'll, I'll qualify it a couple of ways. So, you know, one way we will increase our revenue by about tenfold this year, and we're already sold out of all of our capacity through the end of the year. Right. So I have a build schedule. We have about 500 employees SPEAKER_14: today. I'll be closer to 800 by the end of this year. That build schedule is fully booked this SPEAKER_00: year already. We see that broadly across the sector. There's just, uh, an immovable wall of demand SPEAKER_14: for this compute. A lot of it is being driven from this move from training the models to, to inference, SPEAKER_40: right? And inference is, you know, actually bringing the commercial value out of training, right? So you want to go train a foundation model that takes compute to be built in the configuration that, that we build it right in these 10,000, 30,000 GPU clusters. Uh, and then you got to go, you know, SPEAKER_14: make it action, right? Like drive revenue off it and bring a product. And you know, what we're observing is it might take 10,000 GPUs to train a model, but inference, inference is linked to the number of users, right? If you go into chat GPT, for example, and query that's spinning up a GPU. And now there's a million of you doing it 5 million, 10 million that informs the size of inference. So inference will really, truly be linked to the growth of this market. And we're seeing users who were using, you know, 10,000 GPUs for training need hundreds of thousands for their, their early stage inference products. Right. So we, we don't see demand going anywhere, but up to the right. Got it. SPEAKER_09: For this infrastructure, while they may not need exponential use of, uh, you know, GPUs for training, they'll get more and more efficient at that. Um, and what, what they will need is those inference. When people ask the query that's inference, not training the model, but asking a question of the model that is massively compute intensive. And in each 100, if we were to look at that unit in an hour at full capacity, how many queries, you know, I know it depends is always the answer, but an average query, like these things were costing a couple of pennies per query. Is that SPEAKER_78: correct? Ballpark? Yes. Yeah, that that's correct. And I think that's the right, right, right way to SPEAKER_14: qualify it is, is, uh, cents per query or dollars per query. Right. So you're getting in, you know, hundreds of queries within that period. And as you said, like that will become more efficient over time as well. Uh, but it's, it's just an unbelievable volume of demand. And like, when you step back and think about it, right, like you look at the existing cloud infrastructure that was built over SPEAKER_00: the last decade, it wasn't built for this use case, right? It was built for serializable workloads. It wasn't built for parallelizable workloads. And it's like, you're having to rebuild the cloud, so to say, and you're having to rebuild it. You have to rebuild physical infrastructure at the pace of AI software adoption. It's a mind blowing concept, right? Because AI software is being adopted at the most rapid scale of any technology that we've ever observed. And you're asking people to build, I mean, we're building at, I think it's 28 data centers this year across North America. We're SPEAKER_01: one of the largest operators of this infrastructure in the world. And we are unable to keep up with demand. And we really don't see that subsiding for years to come. SPEAKER_115: Listen, as a founder, there are things I love doing, like building products or meeting with partners, hanging out with my team and dreaming up new ideas. And then there are chores that I don't want to do. I don't want to do HR. I don't want to do payroll. I don't want to deal with all that. So I use Gusto. Gusto is the best for payroll, for HR services, and for running a small business. It makes everything so much easier. Even a midsize business, man. I get a lot of portfolio companies that are pretty sizable using Gusto because it is designed for you, the small business owner. And payroll is something you definitely do not want to mess up. You got to get it right. And Gusto is going to make it perfect for you by calculating paychecks perfectly. Also payroll taxes. You got to get your taxes right. You can't make mistakes there. And you want to set up open enrollment. You want to be good to your people. Gusto handles onboarding, health insurance, 401k, time tracking, commuter benefits, off the letters, and they even give you access to HR experts. So Gusto takes all of this off your hand and lets you focus on important stuff, your product and your customers. It's super easy to set up and get started. And if you're moving from another provider, Gusto will transfer all your data for you. Here's your call to action. Because you're a Twist listener and you're part of the family, you're going to get three months free. Incredibly generous. Totally unnecessary. Thank you so much to our friends at Gusto.com slash twist. You must go to Gusto again, Gusto.com slash TWIST to get three SPEAKER_38: months free. Thank you, Gusto team. So then this would lead us to, um, LPUs, uh, obviously using a GPU, very expensive, right? Um, but Grok, uh, my friend Chamath's company, um, has this, uh, uh, inference engine and these LPUs. Are you starting to see those and that hardware stack emerge SPEAKER_09: these language processing, um, units? And do you think that'll have, um, a good effect on the SPEAKER_33: industry in terms of lowering costs and having purpose-built hardware for the inference, uh, SPEAKER_00: moment? Yeah. So, so as opposed to the Lord of the rings, right, where there's one ring to rule them all, I don't think that there's going to be one GPU, LPU, one accelerator to rule them all, nor do I think there's going to be one model to rule them all either. I think there's going to be SPEAKER_14: lots of different models with different, uh, objectives, right? Like models that do different things, right? Whether it's helping drive a car or cure cancer or, uh, be an AI character, right? Models will do different things. And then there will be infrastructure that is most efficient for each SPEAKER_00: different type of model, right? And I think that's why you're seeing entities like Microsoft, Meta, et cetera, who are focused on building their own Silicon, right? They're not trying to replace the GPU. They're just trying to solve for different models that they're running internally. So I think, SPEAKER_14: you know, the Groks of the world will absolutely have a place somewhere, right? But I also think SPEAKER_00: that you'll see GPUs have this place and what we're observing their places is at foundation models at latest generation models, the most demanding and complex workloads will continue to sit on, on GPUs. And NVIDIA just has this unbelievable solution for iterating continually better generations SPEAKER_14: of GPUs. And we think that those models will continue to accrete to NVIDIA's platform. SPEAKER_91: So they're going to win the day, no doubt. Uh, NVIDIA, when it comes to training the models SPEAKER_33: inference, you might see other folks carve a niche for themselves is how you would bet this emerges. SPEAKER_00: Yeah. Yeah. Yeah. I think inference will, will have various levels of infrastructure that provide SPEAKER_14: solutions for it. Um, I will say, you know, if a model is trained on a one hundreds, it'll probably run inference on a one hundreds as well, like, like kind of tough to make that architecture shift. And it's tough because of the software that NVIDIA has, right? Like their driver solution, CUDA, um, NVIDIA very thoughtfully open source that driver solution in the early 2010s to, to support SPEAKER_00: this sector and the engineers who wanted to work on these products. And it has become SPEAKER_14: effectively a default solution, um, across the market, right? It's, it's similar to, uh, drivers for CPU, right? Everything was x86 for decades, right? And it didn't really matter if something was, was better or not than it. It's just what people use, right? Cause there's, uh, uh, an efficiency loss. If you say, well, I'm going to go learn this other thing and, you know, you just hope other people will, will use it, or you could just use the thing that everyone else uses. And that's what SPEAKER_00: dominates the market. And NVIDIA has a amazing moat that they've developed out of the superiority of their software solution for their infrastructure. Um, and I think that's going to keep people using SPEAKER_36: their platform for a long time to come. Now are people using CUDA yet to address other GPUs and, SPEAKER_09: and cause it's open source and it's obviously being used for parallel computing here. When you've got a supercomputer, you need to, you know, send a job across many different GPUs or are people for CUDA or have they adapted CUDA in order to, you know, have it send a job to some Intel servers, some NVIDIA ones in, is that opening up possibilities? I think for more open source future. And then I'm SPEAKER_38: curious what you think of open source chips and chip architecture. And if you think that is ever going SPEAKER_00: to have some sort of an impact here on the space. Sure. So this will get a little bit, you know, beyond my domain expertise, but yes, there has been forks, so to say, and software that enables CUDA to run on different infrastructure, but it comes at a hefty cost, right? It comes at performance loss. It comes with configurability loss, like so much so that none of our clients are requesting that, right? SPEAKER_14: And we're talking, you know, 30, 60, 80% performance loss, right? So the most natural thing for the largest consumers of this compute is to stick on NVIDIA infrastructure with NVIDIA software. And that goes to your second question, which is around open source. You know, it's tough for me to say, but I would highlight the behemoth that's driving the research and the path forward on NVIDIA GPUs, you know, they just have so much capital they're putting to work to ensure that they have the most performant piece of infrastructure in the market that, you know, sure, there might be some use SPEAKER_00: cases for that open source infrastructure to be applied, similar to how rock can be there or other SPEAKER_14: custom silicon chips. I think that the vast majority of workloads are going to accrete and stay with, SPEAKER_08: uh, but GPU infrastructure of which who's, who's number two or three in the space. Does anybody have a chance of closing the gap? And because obviously people are watching NVIDIA print money, you know, SPEAKER_38: and obviously that's, I don't know what percentage of your infrastructure that you provide is NVIDIA, SPEAKER_09: but I'm guessing it's 90% plus. So, but is there a number two or three in this space? And, you know, do they have a chance of gaining market share? Or do you think this is SPEAKER_38: fait accompli? We're just, we're going to live in an NVIDIA world for the next decade. SPEAKER_14: I think we're in NVIDIA world for a while, right? You know, you have AMD out there, but AMD doesn't have a performant training fabric, right? Like that's something that's proprietary SPEAKER_00: with an NVIDIA is in fit and band, right? So you can build this, a comparatively performant SPEAKER_14: training fabric with AMD infrastructure, right? So you can kind of only use it then for inference and it's sort of, well, if you've, if you've already trained your model on NVIDIA, it's, it's a tough leap to want to move your software, move your infrastructure over to AMD compliant. So it's, it's, it's certainly a market I would expect that, that AMD is allocating their time to, but we're not seeing the customer demand for it at scale, right? And we, we really serve as scale consumers of compute. Certainly there's, you know, your, your guys who want ones or tens of GPUs out there who will say, Oh, I'd love to work with, you know, in my 300s. But it's, it's those entities who want tens of thousands of GPUs that are sticking with NVIDIA. SPEAKER_40: And we haven't really seen any deviation from that. SPEAKER_09: And for folks who don't know, InfiniBand is kind of a contemporary or a competitor to ethernet or fiber in a data center. If you had a bunch of storage in one location, or even in a GPU between GPUs, passing data between them, there has to be some way to move data from one cluster to the other, if they were, you know, passing, you know, training data or something, the speed at which the training data can get on the GPU to be processed, that is a bottleneck. And InfiniBand SPEAKER_38: is the solution to moving large amounts of data. Am I correct in my description? SPEAKER_78: That's exactly right. That's exactly right. It's, it's, it's infrastructure that NVIDIA acquired SPEAKER_14: and have integrated into their solution called the DGX, you know, solution, and it is the most SPEAKER_152: performant fabric solution, other words, network solution for this infrastructure for data throughput. SPEAKER_67: Hey, startups, you're a new company, and you're looking to form your business, but navigating through a maze of hidden fees and legal jargon, it's complicated. It's going to eat up all your time. Well, Northwest Registered Agent will form your business quickly and easily, and it only takes 10 clicks and 10 minutes. They provide you with a full business identity setup. That means they'll give you everything you need to start and to maintain your business. When you hire a registered agent to form your company, they take care of everything. You get a registered agent service, a business address, their corporate guide service, a phone line, mail scanning, a free domain, a website, and hosting. Northwest Registered Agent makes the whole process transparent, quick, and enjoyable. Whether you're setting up an LLC, a corporation, or a nonprofit, they've got you covered. Here's your call to action. For just $39 plus state fees, Northwest Registered Agent will form your company and launch your business in minutes. Visit NorthwestRegisteredAgent.com slash twist today. That's NorthwestRegisteredAgent.com slash twist today. SPEAKER_09: Is this still one of the key challenges in terms of training large language models is the throughput of the InfiniBand or Ethernet solutions to just move the data around? This is the bottleneck over GPUs in many SPEAKER_14: of these jobs. Yes. It's critical to build with a non-blocking InfiniBand fabric. So non-blocking means that every component can operate at the same performance and efficiency as everything else. There's just nothing blocking that performance. No bottlenecks. SPEAKER_00: Yes. No bottlenecks. And it's really interesting because it's a physical engineering problem. SPEAKER_01: So a 16,000 GPU fabric, which is about 2,000 nodes or individual servers with eight GPUs per server, SPEAKER_14: it has 48,000 discrete connections that have to be made across the fabric. So you plug in InfiniBand into each GPU in the server, then that goes out to a switch and you're part of this fabric. Right. Every connection has to be made correctly. And you're doing this with 500 miles of fiber optic cabling within that 16,000 GPU fabric. Wow. And we run a number of those of larger and smaller size, right? So we built a lot of these things, SPEAKER_00: run a lot of fiber in our days, but it's a complex physical problem that no one's really been presented with before, right? This wasn't a problem when you're running with Ethernet or hosting websites and storing data lakes. You didn't have to build fabric this way. It'd be SPEAKER_14: one connection per server, not eight connections per server. And to be doing it in this contiguous, non-blocking fabric in a single footprint, right? And so it's just lots of new things that are happening at the same time with an immense amount of capital at risk and SPEAKER_00: immense amount of capital that's being consumed in the fastest-paced technology environment that we've ever been in. And it's creating problems all over the market. Yeah. SPEAKER_14: And where we've found ourselves is having a software solution and a company that's only focused on these types of workloads. And accordingly, we accrete clients into our platform for having that best SPEAKER_01: engineering solution and actually being able to deliver it to end consumers. SPEAKER_09: Yeah. And we've never seen at-scale companies like a Microsoft, like a Meta, like a Google. SPEAKER_38: These companies are at scale. They have massive amounts of capital, which they can't deploy in M&A anymore, right? We have a framework in the West where you're not allowed to buy companies. And I made this point on All In a couple of months ago. SPEAKER_09: Instead of like, if you were Apple or you're Googling, you're sitting on tens of billions, hundreds of billions of dollars in cash. You can't buy Uber, Airbnb. You can't buy Coinbase. You're not allowed to buy even Figma for $20 billion. You can't even make a small purchase like that without getting blocked. Well, what's the next best thing you can do with that capital? SPEAKER_38: You could build infrastructure, infrastructure as a weapon, you know, and now you've got this massive SPEAKER_17: infrastructure. Will you have jobs for it? I'm sure there'll be some. Will those jobs turn into SPEAKER_127: commercial products? Some will, some won't, but it's a better use than sitting on the cash SPEAKER_38: or it's a better bet. It's a better, you know, use of capital rather than trying to make a couple of points on it and, you know, or buying back your shares. It feels like, gosh, if you have this infrastructure, you could have induced jobs, which is to say some crazy person on the meta team is going to be like, what if we did X and having that infrastructure allows somebody with a crazy idea to then go give it a shot and spend a million dollars running a job across this infrastructure, whatever the pro rata version of it is. And maybe they find something really interesting. Who knows? Yeah. What, what people are going to do with this infrastructure? You do, you know, you're watching them. What's, what is the interesting jobs you're starting to see and use cases? I mean, some of it's public and some is private. So I'm obviously don't want you to betray anybody's trust here, but just what are people doing with this infrastructure that you find interesting when they come to you and they say, Hey, we need a solution for this, or here's what we're building. What, what are some of the things that you think are most promising on certain verticals, SPEAKER_00: sectors that are most promising? Sure. So I, I think the areas where AI will be adopted first and fastest and do it at scale, right? You can always find like five users to do something, SPEAKER_14: right? But how do you get 5 million users? Yeah. It's going to be within products that the user doesn't have to learn something new. SPEAKER_00: It might not even be a new, but right. It just comes naturally to them. It feels organic. It doesn't require, you know, a new app to be somewhere it's integrated into existing products. And I think SPEAKER_14: that's largely going to be co-pilot, right? Like various co-pilot solutions, not the name to one product, but just the idea that you're integrating AI into apps to assist a user with a preexisting process, right? That's something that we're seeing scale right now. And those, the ability for that, those products to scale are limited by the amount of cloud infrastructure that's able to handle those users. Again, remember it takes each time you come in and query that SPEAKER_00: co-pilot product, it's using a GPU. So cloud infrastructure inherently limits the pace at which those products can grow. And I think you've seen some products delayed even because there wasn't enough cloud infrastructure available to power their, their launch even. SPEAKER_09: Yeah. If you look at search engines like Bing, Bing kind of was doing the custom answers. You'd have to like click a second button to get it, right? Go to another experience. Whereas some, you know, search engines powered by AI were doing it automatically because they didn't have a large flow of it. If every single Google search resulted in a query to a GPU, they would actually bankrupt Google right now because they have so many queries and a three or four cents extra per query. SPEAKER_38: There's not enough infrastructure in the world to convert all of those queries today. SPEAKER_14: Yeah. Look, it brings up a question of will AI be a tax or a margin expander for software products? I think some of them, it will be a tax, right? It'll become mandatory and they might not be able to drive incremental direct revenue off those products. But you know, the other, uh, outcome, SPEAKER_01: if you didn't integrate that AI at that tax could be, you lose users and you lose market share to SPEAKER_38: someone else, right? If you look at search, that would be the perfect example. If Bing offers this to, you know, their four or 5% market share, they're kind of, they can lose money on it because they're, they're building that business. Whereas Google, it's their core business. They would, if they put it SPEAKER_09: on all 90% and they start losing money, they could just flip their business upside down, right? SPEAKER_14: Yes, that's right. And you know, the other interesting point to that is, you know, SPEAKER_00: Google might not have the option to integrate AI if it doesn't have the infrastructure available at the volume that's required. And I think that's why you're seeing some companies that up, uh, uh, Microsoft throw so much capex into ensuring they have the volume of infrastructure necessary because SPEAKER_30: having the compute at scale, going back to my point earlier, it decommoditizes compute, right? Like that, that in and of itself is a strategic advantage. So I'd say the, the other area that SPEAKER_62: I personally think will accrete AI rapidly is, is in the advertising sector. SPEAKER_177: Oh, really? I thought you were going to say healthcare or biology or something. SPEAKER_14: I agree. I think that's the second half of this decade thing that we're extremely excited about. I mean, I can't wait to have infrastructure that supports directly supports the advancement of healthcare solutions, but advertising, I mean, think of the way that ads work, right? SPEAKER_00: Like you throw an ad, you hope it reaches an audience and then a subset of that audience will actually identify with the rights, probably a pretty small sliver of it. Yeah. SPEAKER_14: Instead, if you could use generative AI to create on demands, always on ads for people that are SPEAKER_00: a hundred percent specified to the metadata associated with that user, those are going to be much more highly effective. Yeah. SPEAKER_14: Here's the example, right? Like you were, you live in Utah, you have a green kayak and you're searching for a new Kia, right? It's, it's a blue Kia and you've been looking at it for a few days. And instead of just receiving the general Kia ad, uh, of, you know, a gray Kia somewhere, you now get an ad that is a blue Kia with a green kayak on top, SPEAKER_00: Yep. Driving through the desert in Utah to a river, right? And you can get several different iterations of that until you go buy that Kia. Yeah. That that's a area where the user doesn't know that it's generative AI, but it'll be so accretive and disruptive to the advertising sector. That'll just be mandatory for them to use it, right? Because SPEAKER_01: the ads effectiveness will increase that much. It's fascinating. You know, you, you look at what SPEAKER_38: happened with Meta. There was this idea that when they lost access to smartphone data, when Apple anonymized it, you know, they would have a really hard time doing targeted advertising. It actually kicked them in the ass and made them implement AI and they have now recovered and gone further. I think in terms of personalization, you're exactly right. If it knows you have two kids, SPEAKER_09: it's going to put, you know, in that Jeep Wrangler, you know, with your kayak on the top or whichever car it is, two car seats and it's going to show kids in it. And the message will have something about how great it is for toddlers or young kids. And here are some, you know, here's SPEAKER_38: the media center that puts the TVs on the back so they can watch Netflix. It's just going to have so SPEAKER_09: much information to customize the ad that you get the gap between the aspiration of the ad and the SPEAKER_186: reality of your life is going to close, right? Because ads are aspirational. So it's my, it's like SPEAKER_188: minority report, if you remember, and everything goes back to minority report, you know, the SPEAKER_38: customization of the ads will be absolutely phenomenal to a level that, yeah, it's like SPEAKER_08: beyond creepy. It's just like mind reading ads. SPEAKER_40: Yeah. And you've, you've never had that before, right? And what's important there is it, SPEAKER_00: it took X amount of resources to generate that, that one kind of mass media ad previously. Well, now each time you have that iterative always on ad that that's querying infrastructure, right? So the infrastructure demand for this new type of advertising will be voluminous. Yeah. Right. It'll be more effective. And I think it will actually like be better dollars spent in advertising, but it'll be an immense amount of infrastructure demand behind it. So I, I think co-pilot's there today and scaling, but the next big thing in there to really scale will be within the advertising space. SPEAKER_41: Yeah. It makes a lot of sense. It's a huge business and yeah, you know, in anywhere, SPEAKER_38: there's a lot of data and a frequent transaction. I mean, that's just a great place for GPUs and this AI revolution to take part of because it's frequent and there's a transaction. This is where like Amazon and how Amazon sells you stuff and Walmart and target my Lord e-commerce in the last mile. SPEAKER_09: It's already been impacted in a way. I mean, if you look at the amount of advertising revenue for Uber and Instacart, which are, you know, for Uber eats essentially, you know, that's like being in line at the checkout counter and then they're doing a billion each, I think a year roughly. And then Amazon might be doing 30 or 40 billion in advertising. Now, like they're those three businesses, which are seemingly transaction-based businesses, right? Shopping for groceries, food SPEAKER_08: delivery, mobile transportation, then Amazon buy anything. They're all becoming advertising SPEAKER_38: businesses. It's like pure profit for them. It's going to be wild, the AI impact on those businesses. SPEAKER_00: And all roads lead back to generative AI, right? It's just all converging here, right? And it all needs this type of infrastructure. And that goes back to my point of like, how much demand is there? We just don't see the path to resolve the amount of infrastructure that needs to be built for the demand that there is within, you know, at minimum the next few years, right? There's just so much needs to be built because last generations clouds aren't designed for this. And it's not like you're swapping out a UI and saying like, oh, you like tweak some software here and there. And all of a sudden it SPEAKER_14: works, right? It's no, it's the foundational difference. It's the Tesla versus Ford manufacturing SPEAKER_200: process. Yeah. And it's just never going to take these workloads. CPUs will be for serving up images, SPEAKER_202: light work. It's not going to ever compete with this level of audacity. It's completely different. SPEAKER_09: Um, any worries about overbuilding this infrastructure at this point, if we were going SPEAKER_38: to start talking about a slowdown or a certain amount of infrastructure, is there somewhere on the chart that you started thinking, yeah, this is gonna, we'll fill the demand, um, five years out, 10 years out. Where do you think we have enough capacity enough, you know, and, and the SPEAKER_206: supply, supply demand becomes normalized right now. It's abnormal, obviously. When do you think this SPEAKER_152: normalizes? When do we catch up? So between the infrastructure demand and the data center demand, right? So it's multiple components in here. And then, you know, it's all the, the infrastructure SPEAKER_00: pieces that go into a data center. Yep. And then it's the power that goes into the, like, it's this really complicated physical stack. Yeah. Uh, to serve it, it honestly could be the end of this decade until you see this, this rebalancing of supply and demand and not to say that that's overbuilding, right? Like that's just still on a, on a heavy growth trajectory. That's just when infrastructure may have had an ability to catch up to where demand is. And, you know, I, I geek out over this stuff because that that's my background and my co-founders were all from this commodity trading sector where all we did was assess supply and demand and understand physical disruption of commoditized markets. And that's exactly what we're looking at here. SPEAKER_43: Yeah. But these aren't commodities yet, you know, like H 100s trade at a price and I guess they're commodities, but they feel like a very resource constrained commodity right now. So I guess they are SPEAKER_212: commodities. Yeah. And they're, they're, they're sort of, you know, if you think about like cloud SPEAKER_00: infrastructure for hosting websites, right, it was, it was fungible, right? It didn't really matter if you were on AWS, GCP, Azure to host your website. It all felt like the same thing. It was the same product to host your website, right? Yeah. What's, what's changing is that lack of fungibility, right? An H 100 hosted at AWS is very different than an H 100 hosted at Coreweave because of the way that we run that infrastructure differently from a software perspective and the way we build it from a, from a physical perspective. Yeah. Right. So like that, that's the, the SPEAKER_14: commoditization that did exist. That's now being decommoditized through software and infrastructure SPEAKER_217: disruption. Right. Yeah. It's amazing. What a moment, what a time to be alive. Well, SPEAKER_38: it's so much fun. It was absolutely fascinating to talk to you for an hour and I'll let you get back to SPEAKER_70: uh, racking and stacking. I'm sure you've got tons of H 100s and a 100s to unbox. I mean, just unboxing and racking stuff. I mean, you have hundreds of people doing that at this very moment, SPEAKER_152: hundreds and, and semi semi trucks arriving to our 28 data centers across the U S. Um, it's, SPEAKER_14: it's a, it's an operational feat. I think we're hiring 20 people a week right now. Yeah. Um, it's, SPEAKER_38: and these are like system operations people, high level people to come in and configure this infrastructure. Uh, well, massive success and thanks for building out the infrastructure. Let's solve some huge problems and we'll see you all next time. And this week in startups, bye-bye.