SPEAKER_00: All right, everybody, welcome to Twist. It's April 8th, 2026. My co-host Alex is with me. We've got a bunch of guests today, and we've got a major breaking news story, which is that Anthropic released a promo video and a thread that their new model, Alex, is so powerful that they cannot release it. We knew this day would come. The day is here. Why can't they release it? They believe when they tested it, or they found out when they tested it, that it would try to escape. That was one issue. But a bigger issue was it could find exploits in 10, 20, 30-year-old software projects, and it could thread together multiple security vulnerabilities. Catch everybody up in the audience on this incredibly important story. Big model race going on between the major SPEAKER_02: AI labs. Anthropic is now very far out in front with its new model called Mythos. It is a general purpose LLM, so it's not tuned for one specific task. It is currently in preview. You cannot use it. A consortium of companies, Jason, are working with Anthropic to basically use it in a defensive capacity because, as you said, Mythos is incredible at both finding, exploiting, and patching security vulnerabilities in software that humans have often missed. This goes back decades, as you said, to things like OpenBSD, a famously secure piece of software, it found something there. It found something in FFMPEG, which is an important part of the open source infrastructure of online video, SPEAKER_04: for example. Basically, the gist is with this model, anyone can go to any piece of software and find zero-day exploits quickly, and then basically go to war with them. So, Anthropic cannot let this out of the bag because if they did, then North Korea and China and everyone else could use it to essentially break the modern digital infrastructure that we depend on. So, today, Jason, Project Glasswing is the goal. A bunch of companies, your NVIDIAs, your AWSs, your Azures, are all going to work with Anthropic to basically take the model, Mythos Preview, and harden everything. Anthropic has also put together a $100 million credit fund, essentially saying, here, use the model up to $100 million of compute, to essentially harden these systems. So, I think that Anthropic is doing the right thing here by saying, hey, we're not going to release something this dangerous, you know, from just off the cuff, but it does create a situation in which we now have a very much two-tier economy. There are the companies that are sufficiently important that Anthropic is letting them have access to Mythos, and that means they can be ahead on both defense and offense. And also, we're now seeing a world in which smaller companies are just stuck outside the glass looking in. And I think it's a little bit of a change because previously, every AI lab was so focused SPEAKER_05: on having their newest and best model out in the world that it was very democratic. SPEAKER_06: I'm going to play the video that they shot, and we can get into, with our guests today, SPEAKER_07: whether this is cynical and this is showmanship, and they want everybody to understand, hey, SPEAKER_00: it's this week in Anthropic basically here. We're going to rename the show because they're dropping so much incredible content. But before we do, I have to take a moment to applaud, applaud. You see my applaud pen here. They are our sponsor of the show today. You have the applaud wristband. I hold and press. I get a nice little haptic. Boom, the red light goes on. And now, when I put it into my charging tray, here's my little charging tray, and here's my backup, applaud. You should buy two. That's my best advice. Have two. One in your bag, and then one at your desk. Any meeting you do, SPEAKER_08: you're just one click away from having the note taker on. And this note taker is independent of all other platforms. It does beautiful summaries. I am addicted to this. You know the big test for me, SPEAKER_00: Alex? It's like this is called the key and the wallet test. The key. Oh, I know where you're going with this. If I leave my house without my keys or my wallet, now in today's day, it's your smartphone and your pouches, your nicotine pouches. Those are the two things I got to turn. No, I don't turn SPEAKER_12: back for nicotine pouches. But I do turn back for my applaud. I literally put the car in park, SPEAKER_14: walked back to the house, put the car in reverse, went back into the house, and I got my blood pen. SPEAKER_04: I think it's fantastic. If you want to have a good recording of all your conversations and automatically taken notes, then you should go to plaud.ai slash twist. P-L-A-U-D dot A-I slash twist. Use the code twist, Jason. Save 10% on your purchase. We're big Plaud fans. Shout out to SPEAKER_08: them. Let's get started with the show. I want to go deeper on this. And you know, our pillars here on the show, show don't tell. We like to have great demos. And we like experts, experts only. No offense to our journalist friends. The top 20% do a great job. The other 80%, we all know, you know, maybe they're, they're trying their best, but we like to have experts on experts directly from the horse's mouth. Here we go. We got Rob May from Neurometric. Neurometric SPEAKER_02: builds small language models in contrast, Jason, to large language models or LLM things we talk about SPEAKER_04: so very much. We're going to talk about SLMs in a little bit of time, but Rob, I presume you've had SPEAKER_21: a chance to read through the anthropic mythos card, read the red team report and, uh, and chew on it. SPEAKER_23: So Rob, Rob has opinions. So people know Rob was one of my first angel investments. One of my first SPEAKER_00: five, uh, for a company called Backupify. This was a genius idea he had. I cold emailed him. SPEAKER_14: Is that correct, Rob? I cold emailed you and said, Hey, I love your product. SPEAKER_27: Yeah, you did. This was back before you were famous, Jason. So you had to explain to me who you were. SPEAKER_08: Yeah. Yeah. I was like, Hey, my name is Jason Calacanis. I use Backupify. And this was the best idea ever. If you lost your Gmail account, Alex, you know, in the days before, SPEAKER_00: this is before Google suite, I think existed or was just coming out. If you lost your Gmail account, Google would say, okay, you got hacked and you'd lose your entire archive. Rob figured out a way to use the API to back up your Gmail box and your G drive and, and, and every other service you can think of. It was an amazing vision and it was a great company. And we had a nice little exit. SPEAKER_31: I backed every company Rob's done. Uh, and he's got a new company. So I'm excited to hear about it. Catch up with Rob. Rob also ran the open angel forum for me in Boston. Yeah. For a little bit. SPEAKER_34: Yeah. And I actually, I think I was one of the first people to talk to you about AI, Jason. Back like 10 years ago when I was, uh, I was talking to the incubator and telling the story and you and I had dinner with one of your friends, Jeff with a G and we, uh, stayed up late talking about AI. SPEAKER_38: Yeah. Good memories. And, uh, yeah, this is, I'm going to be in a old age home, Alex, SPEAKER_00: and I'm going to be like, welcome to episode 20,122 of this week in startups. We're here on twist and here's Rob May. He's going to be like in a wheelchair, like professor X, but it's going to be a levitating floating one. And it'd be my brain in a vat of a back to tank. And I'm just going SPEAKER_07: to be 120 years old. Uh, still doing this and loving it. Let's get into it. Um, mythos, this is the preview of the new anthropic model. Here is Dario and this video. I don't know if SPEAKER_03: you see this video yet. I don't think so. No. Okay. We'll get you to react to it for the first SPEAKER_31: time here. Here is the anthropic team talking about why they're withholding the model. SPEAKER_42: There's a kind of accelerating exponential, but along that exponential, there are, there are points of significance. Claude Mythos preview is a particularly big jump along that point. We SPEAKER_45: haven't trained it specifically to be good at cyber. We trained it to be good at code, but as a side effect of being good at code, it's also good at cyber. SPEAKER_49: The model that we're experimenting with is by and large as good as a professional human at identifying bugs. It's good for us because we can find more vulnerabilities sooner and we can fix SPEAKER_50: them. It has the ability to chain together vulnerabilities. So what this means is you SPEAKER_52: find two vulnerabilities, either of which doesn't really get you very much independently. But this model is able to create exploits out of three, four, sometimes five vulnerabilities that in sequence give you some kind of very sophisticated end outcome. And we think that this model can do this really SPEAKER_53: well because we notice that this model is very autonomous. It's just generally better at pursuing really long range tasks that are kind of like the tasks that a human security researcher would do throughout the course of an entire day. Obviously capabilities in a model like this could do harm if SPEAKER_55: in the wrong hands. And so we won't be releasing this model widely. More powerful models are going to come SPEAKER_58: from us and from others. And so we do need a plan to respond to this. That's why we're launching what SPEAKER_53: we're calling Project Glasswing, where we partner with a number of the organizations that power some of the world's most critical code to put the model into their hands, to allow them to look at how they can use models like this to bring down risk and protect everyone. And by giving these software developers SPEAKER_60: advanced tools before anyone else, it gives all of us a collective head start. It allows us to find SPEAKER_61: things that we couldn't find before. And it helps us fix these things much more quickly. Working with SPEAKER_53: our partners, we've been finding vulnerabilities across essentially every major platform. I found SPEAKER_63: more bugs in the last couple of weeks than I found in the rest of my life combined. Dario said SPEAKER_67: he's got very big concerns. And the reason he left OpenAI was he felt Sam Waltman was not trustworthy. SPEAKER_00: I'm not, you know, piling on Sam here. There was a big New Yorker story. But the truth is, Sam drove a lot of people out of OpenAI. And there were trust issues, according to those people. Again, SPEAKER_69: I'm not editorializing here. Everybody kind of thinks I'm Team Elon, which is fair enough. But I don't have investments in any of these companies. So I'm not talking my book or anything. But the truth is, he drove Dario out. Now Dario is passing OpenAI in models, profile, PR, influence, and dare I say, SPEAKER_70: the revenue here. So your take, Rob. If you've got an engineering team at your company, I'm betting there's a solid chance they're spending far too much time on infrastructure. You need your team building your product to delight your customers, not configuring your virtual network. Render is the all-in-one cloud platform for developers that allows you to deploy, scale, and secure your apps and agents with zero ops. Most cloud platforms ask you to split your focus between product and infrastructure, or they force you into platform constraints that you know you'll outgrow in six months. But just connect your GitHub repo to render and you are live, L-I-V-E, web services, cron jobs, manage Postgres, the whole stack in one platform. It's time to find out why five million developers are already using Render. Go to render.com slash twist and apply for the Render startup program. You'll get anywhere from $500 to $100,000 in free credits, depending on your stage and who your backers are. That's render.com slash twist. SPEAKER_72: A lot of stuff in there. Take it wherever you want to go. SPEAKER_34: Yeah. I think Anthropic has significantly passed OpenAI on a lot of things. And I think the reason is that they've been more focused. Like what's happened to OpenAI, to focus on them for a second, is that because they were first and they've had to raise so much money to build these models, Sam had to go out and sell that they were going to enter. They needed a $10 trillion TAM, which means you've got to be in every market, right? And so I think that's their problem. They've been trying to do a little bit too much. With respect to this model, I do think it's interesting and I like their approach. I think we were talking a little bit about at the beginning of the show about, you know, they're going to IPO this year. So this was a very well put together video that I'm sure they had the IPO in mind as they started to do this. But that said, I actually, I like the approach and what I like about it is I think we've been a little bit, the Silicon Valley vibe on AI for the last couple of years has been like, whoever hits AGI first wins, right? Because that machine extrapolates and gets better. And I think what we've seen over the last 18 months that surprised everybody is that the parody amongst the top labs, you know, and including Google and everybody is really what it means is when we hit AGI, like we're all going to have access to super intelligence for free and open source three to five months later. And so I like what they're doing because they're preparing us for that day. And I think that's, I think that's super important. Rob, do you think that SPEAKER_23: the open source? Alex, I want your opinion, actually. Having watched this and handicapping it, you've SPEAKER_06: been deep in this space and Cautious Optimism is a place to go look at that. CautiousOptimism.news? Yes, sir. Yeah. Okay. So I read the newsletter, everybody should subscribe. You've been obsessing over this. Your take on Dario being right, Dario being driven out of open AI, Dario focusing on code first, as opposed to side quest and consumer and them surging ahead of open AI. SPEAKER_84: I'm surprised at the speed we went from, oh, Anthropic is catching open AI to they seem to be SPEAKER_04: tied to Anthropic is winning October, January, and now April. It's pretty crazy how fast things have changed. Anthropic has gone from probably around 10 billion ARR in, you know, last October to like 30 now, which is just incredible. Unprecedented. Unprecedented to the point at which I don't even really understand the numbers. What I will say though, about their decisions and product terms is that I think product market fit is the ultimate arbiter of entrepreneur success. And clearly no one has more PMF than Anthropic today. Regarding the future, and Rob mentioned that the video is aimed kind of at investors in the IPO talking about future capabilities, future capacities. There was an interesting interview between Greg Brockman and I think it was Alex Kanterwitz over at Big Technology talking about models. This is right before Mythos came out. And they were discussing takeoff and how these models are now getting a little bit better at self-improvement, working on themselves, writing their own code. And so I think we're seeing the tailwinds of that. Though the thing that I'm unsure of back to Rob's point is how quickly the open source world will catch up to Mythos. Because today, Meta dropped their latest model and their new family. Shout out to them for pulling that off. And it looks pretty good compared to everything that came before Mythos, but it's not now state of the art. So if the proprietary labs are here and then the second tier players are here and open source is here, I'm curious if it's more than three to five months until they catch up. And if so, we have more time to fix the world, Jason, to find all these vulnerabilities and patch them. Because to me, this is at once the best tool in the world for cyber defense, find the bugs, patch them, and also the best tool for cyber offense, find the flaws and then abuse SPEAKER_84: them. And so we're going to be in an arms race, I think, until we're all dead. SPEAKER_06: Let's take a look at the anthropic polymarkets. This is a way for us to really understand how polymarkets actually doing. You know, there's so many to choose from. Here's the URL of all anthropic polymarkets. So let's pull this up. The first polymarket that comes up under the anthropic tag over at polymarket. And, you know, you can basically find a keyword there and it's polymarket.com slash prediction slash anthropic, right? The first one, SPEAKER_07: anthropic $500 billion valuation in 2026, 95% chance. Everybody realizes that's going to happen. SPEAKER_00: Anthropic quad score on frontier math benchmark by June 30th, 72%. And then there's the anthropic IPO closing market. Just tons to go here. I think the one that we want to know is when are they going to release this model? And this model is called Mythos. And so to walk us through this one here, because I don't see a chart on this one. Yeah, it's not charted. Some of them that have SPEAKER_88: fewer kind of like endpoints aren't charted per se, but in this case, we can see, Jason, that there was people betting that it might come out before the end of March. That has now, of course, lost, SPEAKER_04: because there was reporting about this model via fortune and a leaked blog post a couple weeks ago, if people remember. Now, April 30th, I would say zero percent chance. Polymarket Sharps say seven percent, Jason. But the thing that really caught my attention is this. June 30th, they're only handicapping a 28% chance, which means three to four times we won't have Mythos out in the market David Friedberg: by the start of July. That's a couple of months from now. Or two thirds. Yeah, two thirds, because you got to put the first two numbers together, I think. Oh, right. Of course. SPEAKER_67: So two thirds chance, you know, it's not out until, yeah, sometime in the summer or the fall. SPEAKER_04: Which in AI terms is years. That is so much time in this present moment. And the thing that I'm not sure about, and I'm really curious to know is all these companies that have Mythos now and can play with it and can use it, can put it to work hardening their software. How much progress can we make in that interval? And will it be half the work we need to get done? Will it be all of it? I don't know, Rob, do you have an idea of how fast we could put a model like this to work in fixing SPEAKER_27: all the code we depend on day to day? The problem is we're generating code SPEAKER_34: faster now, right? So with all the vibe coding, I think I read that 20 or more than 25% of GitHub commits now are vibe coded. And so as that's going to go up, it's like this is all going to boil down to, so it's, I know you love poker, Jason. And so like the world is going to boil down to poker. Running a business in this AI future is going to be about estimating probabilities of things happening, knowing the cost to go run those probabilities to ground and figure out what they really are. And like, so this is going to boil down to using your compute to figure out like, you know, code generation versus code checking. I mean, Anthropic has the main code generation platform, sort of the number one right now. So maybe they combine these together in ways that write better code. But, you know, if you're talking about it not coming out until June 30th, I would not be surprised to see an open source model like Quinn or Kimmy come out with something that's similar, GLM maybe, like before that SPEAKER_33: date from somebody who doesn't care as much about- Okay, so this is a key point that you're making, Rob. SPEAKER_06: What if a bad actor has already achieved this? Like China may have already accomplished this, and they have no incentive since every company in China is owned by the CCP, which would be the SPEAKER_07: equivalent of like the CIA having a board seat in every single company. And there's like four CIA people and FBI agents and Department of Justice, whatever, inside of Anthropic and OpenAI saying, okay, not only are you not making that promotional video, we're taking this and we're going to hack North Korea, Iran, you know, whatever bad actor we want. China might already have this. Everything could be completely compromised at this point. And I think that's when Americans who were debating, SPEAKER_38: you know, David Sachs and, you know, the administration, and are we hand-wringing too much SPEAKER_00: here that America has to win the AI race. Folks, it is existential who wins this. It's 100% existential if these tools give you the ability to hack everything. Here's a hot take. If your bank moves slower than a startup, that's a problem. I see this process behind the scenes every day, SPEAKER_70: and I know that bad banking can kill your company's momentum. That's why I'm so glad to introduce our newest partner, Grasshopper. Grasshopper's a real federally chartered digital bank, not some fintech rapper sitting atop some mystery institution. Nope. It was built just for founders like you. You want fast? You want easy? Open an account in just minutes and start earning yields that can top 5%. Wow. That's a big number. Plus you'll get unlimited 1% cash back on purchases, free ACH, free domestic wires, and no monthly fees. Plus, if you're sitting on some real runway, grasshopper's treasury product hits 5% plus with same day liquidity. As a twist listener, grasshopper wants to give you a $500 cash bonus just for opening account. And you can open an account really quick. So go right now to grasshopper.bank slash twist and use the promo code SPEAKER_04: twist to get started. I think we should consider this mythos model, this mythos preview to be essentially a cyber weapon and perhaps a cyber weapon of mass destruction. I mean, maybe we need a new term for this because I can't recall apart from certain hacks in history, a point at which people were worried about all of software having potential vulnerabilities. And I think Jason, it's a little bit weird that we're talking about a three to five month gap here of times we might have ahead of China, but I would so much rather have us have that time to work to get things secure than to be catching up for three to five months until our companies were of a sufficient quality to match SPEAKER_07: them. This sheds a new light, Alex, on the Emil Michael's appearance on All In three weeks ago, SPEAKER_00: maybe, where he was talking about the anthropic conflict and like, are they going to ban it or whatever? These guys must have made peace right now. And I wonder if Dario and Emil, when they broke bread and we're trying to work this out, I wonder if Dario said, by the way, our new model is would give the CIA the ability to hack North Korea, to hack China. And we are patriots at Anthropic. And instead of giving the tool to a bunch of security researchers, we're going to give it to the US government. Because if you made this the equivalent, Alex, of the race for the atomic bomb, SPEAKER_08: no American citizen would be like, you know what we need to do? We need to keep the atomic bomb to SPEAKER_07: ourselves. And we're just going to delay making it. Oppenheimer would say, we need to have this for ourselves before the Nazis get it. Period. Stop. I don't think I'm out of line here to say that this is SPEAKER_08: becoming the equivalent. This is becoming the equivalent. This might not seem as much because a nuclear bomb can cause, you know, such a mass destruction of life. But this could cause a massive SPEAKER_04: financial, you know, devastation across the economy. We talk about the GPT-3 moment a lot. I think that's like nuclear fission in your analogy. And then mythos would be the introduction of the hydrogen bomb. But Rob, I think I may have cut you off there. Yeah. Like, like, can this thing SPEAKER_34: figure out how to hack bank software, right? Like Swift and international transfers? Like, who knows? Bitcoin? Like, you know, it's going to be crazy to see how, because we do know, like, if this capability exists, I think you're right, Jason. I think probably some other places have it. Like, my guess is Google has something like this that they haven't announced or released. Like, I still think they're in front of Anthropic, from what I can tell the people I know there and the tools and technologies that they have. And so, and we also know that last year, I think for the first time, China passed the US in research papers on AI accepted into top tier conferences and journals. And so, like, there's a really, really good chance, Jason, that I think your point of view might, like, this might be going on behind the scenes that we don't even know. And it's interesting to think about the run-in that they had with the DOD a couple of weeks ago, five, six weeks ago, whatever, with Anthropic, right? Was sort of like, why don't they just use GPT-5 or whatever? Like, you don't need Claude's model, but it might've had to do with this and they knew it was coming and were discussing it. SPEAKER_07: And this was one month ago. So now we start doing game theory. I think the CIA and the government are inside of Gemini, Anthropic, OpenAI, XAI, talking to each of these model folks. I think they've probably been in there for a year or two and they are saying, hey, are you a patriot or not? Are these tools, are you going to hold them close to the vest or are you going to tell us exactly what's going down here? And then you turn over another card, Alex. SPEAKER_08: If this is in fact true, what they're saying, and if they present it as such, if Dario is presenting this as it's cataclysmic, you know, the entire economy could go down and we believe him, we take him at his word, there's an argument you have to nationalize this technology. There's an argument it's too powerful for a private company to own this. It would be like if a private company were to stumble on a bio weapon of, you know, such or, you know, this weapon that we were using, the disorientating weapon. If a private company gets to that level, SPEAKER_07: hey, we've got a weapon that makes you come in like a superhero and you could just Professor X the entire other army and just make them all grab their ears and blood starts coming out of their SPEAKER_00: nose. I hate to be graphic. You have an obligation to go to the president and say, Mr. President, this private company has this. He's okay, great. Let's go to Venezuela. Let's let's handle that problem. There's some equivalent here. I don't mean to be hyperbolic. I'm just telling you game theory. I don't think Dario is lying. I think Dario is being sincere. Now, I know everybody hates Dario. The right hates Dario. Dario wouldn't bend the knee to Trump. Dario did not donate to Trump like 25 million dollars that one of the OpenAI CEOs or co-founders did. He's not loved by this administration, but they love his tech. And then is he a patriot? Is he like Alex Karp and Palantir? Or is he a hippy dippy who says, I don't want my tool to be used for this? This sheds it in a totally different light. And there's some conversation going on here that we're not privy to between the president of the United States, Emil Dario, the CIA, the Department of War. This is cataclysmic in its severity. Yes. And this is a super weapon. I want to point out just one thing to SPEAKER_04: back up what you're saying, Jason. Thomas Friedman, not someone we usually quote here on the show, but he wrote a post for The Times and he says, this is not a publicity stunt referring to Anthropik's position here. Quote, in the run up to this announcement, reps of leading tech companies have been in private conversation with the Trump administration about the implications for the security of the U.S. and other countries. So not only is Anthropik saying this, people they've shared the model with are going to the White House and saying, holy crap, ring the alarm because suddenly everything is potentially insecure now. Thomas Friedman then is confirming, SPEAKER_38: and I didn't see that story, so great Paul. He's confirming what I believe to be occurring right now. And by the way, this is not, I'm not in the Trump administration. I probably could have been, SPEAKER_06: and I could have been on one of these projects. I didn't, I demurred. So yeah, well, it's just a SPEAKER_118: choice. I'm an independent. Well, how do you think, Jason, how do you, but how do you think about SPEAKER_119: the idea that the government needs to take this over at a time when the public's trust in government SPEAKER_34: is probably the lowest it's ever been on both sides of the aisle, right? Republican and Democrat. I mean, doesn't that present a really interesting sort of wrench in the game theory? How do you think SPEAKER_123: about it? Nobody trusts anybody because we're in literally, what was the show at David Duchovny? SPEAKER_08: The X-Files. Trust no one. We've caught up to the X-Files. We've caught up to the X-Files. Those people believe that the truth is out there and you should trust no one. Those were like the two signature concepts of that. That show was 30 years ago. You know what? Here we are, folks. This could be the thing that galvanizes three different groups of people who have the least trust in the world, journalists, AI, and the government. These three groups are not trusted anymore. And for good reason. Uh, you know, you can debate it, but the X-Files in 1993, they nailed this. The truth is out there. And the truth is this maybe could galvanize these three groups of people. Thomas Freeman, the New York Times should be saying, Hey, if we have any information about this, SPEAKER_07: do we put this out as a public story or do we go to the White House? Do we go to President Trump and Emil Michael and say, Hey, by the way, uh, we think this is real. Dario, Elon, SPEAKER_00: you know, uh, Sam Altman, uh, Sergey Brin, they all must be having some sort of zoom or private conference here to unpack this. And how could this be used to stop North Korea and their inter-ballistic missiles, which they have and their nukes, which they already have, as opposed to Iran, which has SPEAKER_38: none of that. Uh, you know, they're attempting, but North Korea is far ahead. There is a nuclear SPEAKER_07: threat out there from a bad actor. Who's insane. It's Kim Jong-un. He's insane. Um, and he has a SPEAKER_69: ballistic missile that could reach California and he has 30 or 40 nuclear warheads. They estimate. SPEAKER_04: I mean, yeah, but what's, what's more effective now, the threat of nuclear deterrence or the ability to take a model and break everyone's entire nation. So like, to me, like it's funny that I think I'm actually more worried about the second category than the first, because the first seems to be more about saber rattling and keeping your country from being, uh, you know, attacked, but this SPEAKER_06: gives them more useful offensive capacity. Jason, that's, we need to, we need to take 20, 30% of the cycles of these countries, of these companies, the government needs to say, Hey, we'll pay you for your time and not releasing these models. And we're going to create an Oppenheimer, uh, SPEAKER_23: Manhattan project. Uh, we're going to create a Manhattan project to, uh, sturdy the infrastructure SPEAKER_04: of the United States. Okay. Now let's take the other side of this coin. We're all kind of in broad agreement here, but here's the, here's the downside to all this caution. SPEAKER_132: And I'm going to pull it up right now. Hiring can be its own full-time job. SPEAKER_08: And Hey, guess what? I already have a full-time job. I make podcasts and I invest, but when you're running a small company, we both know every hire matters. You don't want to waste any of the seats you have at your company. And the best part that you can have is LinkedIn hiring pro. Why? There's a billion people using LinkedIn. All the great talent are there. If you're proud of your work, you build a LinkedIn page and you update it. LinkedIn hiring pro is going to streamline and simplify the entire process for you. Nearly 60% of companies using LinkedIn hiring pro. You're going to get an incredible candidate to interview in the first week. And you know, we're looking for a new producer for the pod. We did shout outs here on the show. We posted it on my social media. We asked friends, you know, where we found our next great hire LinkedIn. And it was competitive. We had like three or four really good choices. So hire, write the first time, post your first job and get a hundred dollars off towards your post at linkedin.com slash hiring pro offer. That's SPEAKER_135: linkedin.com slash hiring pro offer terms and conditions apply. This is, uh, an excerpt from SPEAKER_84: the mythos previews, uh, systems card, I think. And essentially what it shows if you're on the audio SPEAKER_04: version is a dramatic increase, uh, in score for mythos across a number of very important benchmarks. These in particular, Jason are software coding AI benchmarks. And as you can tell, you know, on SWE bench multimodal, we went from 27% from cloud. It was 4.6 to 59% with mythos preview. So what we're doing is we're also saying we're going to slow the, uh, pace at which AI models that we use day to day for day to day tasks down. We're going to retard that function dramatically. And I don't think anyone else in the world is going to do that. SPEAKER_07: Well, I don't know that you have to slow down the model creation to just make a SWAT team from each company and say to each company, we want your five best brains, uh, on cybersecurity available. We're taking five from each company. It's going to be 25 of these. Maybe the smaller models can SPEAKER_06: contribute one or two people. And they're going to be the brain trust that then builds a system. This is my proposal. If this is all true, which I'm 98% sure that Dario is not, you know, SPEAKER_07: being hyperbolic to raise the stock price. I'm going to take him on, um, you know, uh, his, uh, word. SPEAKER_08: If it's true, um, uh, we could create a piece of software. We could create a new unit of government that then goes and says, we're going to do our own red teams to try to turn off or, uh, overpower SPEAKER_07: this nuclear power plant and have it melt down. And we're going to use these tools to see if it's SPEAKER_142: possible and then we're going to fix it. And we're not going to talk about it. We're not going to talk SPEAKER_144: about it. So we're going to, you know, the test group exists. Yeah. I mean, just a, a group of people SPEAKER_07: who goes and says, what are the biggest vulnerabilities? And this is where immigration SPEAKER_00: of the most talented hackers in the world matters. We should also be trying to get every single AI researcher in China to defect, and we should pay them a million dollars to defect with their families. We should figure out when they're on vacation in Singapore, or if they can get out of the country to Australia to go on a holiday, we should figure out how to pick them up and let them defect, which we did during the cold war with Russia. We need to get them to defect and bring them to our team and then put them in a air gapped kind of space. So we know they're not double agents, whatever, monitor them like crazy. We should be in a talent war to try to recruit these people out of these countries and then put them on our, uh, cyber security team. Yeah. I bet you this is happening covertly. I bet you this is happening right now. If it's only a million dollars a pop, SPEAKER_105: that's the cheapest thing we can ever spend as a nation. But I, I, I, you know, look, I don't like SPEAKER_04: to talk about it much, but I mean, Jason, there was a lot of rich people in tech, right? So why doesn't one of them just say the first hundred millions on me? Well, sure. I mean, this could SPEAKER_07: come any number of ways. Uh, if the government's adding a 1.5 trillion in, uh, spending for the SPEAKER_38: military, you know, you gotta think like, uh, we could put a hundred billion from that spend into this SPEAKER_69: very, uh, delicate area. It might mean that we need to have the government building this tool. We need to have a fork of it, uh, you know, given to, and this is why that discussion where Emil said, anything legal, we should be able to do with the tool. We'll buy the tool from you. And, you know, kind of Alex Karp's position is, Hey, we make a tool. You decide how to do it. I think, you know, um, Andrew has the same position. Hey, we make these tools. It's up to the government to deploy them. It's not for us to tell them they can or cannot use it. Uh, and you know, then we have to trust the government to not use this in some kind of crazy, abusive way. This is going to be the story of the next year. I'm going to predict it right now. Um, and it's going to get political, SPEAKER_96: but this should be a galvanizing moment for America. Uh, I'm getting a call though. Uh, SPEAKER_04: we're getting a call to the show from a dear friend of ours. Um, I think we can pull, ah, SPEAKER_08: there he is. Nick. Nick, are you okay? Nick, are you okay? It's my guy, our correspondent, Nick in, he's in the Brickle. I think he's in South Beach. Edgewater. What is going on? I see SPEAKER_160: you have something on your forehead here. What's going on yesterday. Uh, I heard a mention that I SPEAKER_163: was, uh, sponsored by somebody and that I wasn't disclosing it. And I just wanted to say, I don't know what you're talking about. It was brought up about Higgs field paying for these things. I, I generally don't know what that is. I am in edgewater though, by the way, not Brickle. SPEAKER_08: Okay. So you're in edgewater. Um, and you can confirm for us that as much as Higgs Feld would love to partner with an influence of your, you know, international fame, dare I say, crypto circles, technology circles, finance. I mean, even in the artistic community with your, uh, you know, adjacency to NFTs and the Miami art scene, SPEAKER_00: you are confirming Higgs field is not sponsoring you, Nick. You can confirm right now. SPEAKER_163: Definitely not sponsoring me, despite the release of their new model, which integrates with seed dance 2.0. They are definitely not partnering with me, uh, by the way, that doesn't work in the U S uh, but they have not partnered with me, not paid me anything, nor has actually. Um, and I appreciate you asking me about, you know, what neighborhood I was in because that brings up to me to relate, uh, you know, LinkedIn jobs, which has nothing to do with me. And they're at linkedin.com slash twist. For example, I've never worked with them and I've never been paid for any of these things. SPEAKER_08: Okay. Great. And just to be clear, the promo code, Nick. Oh, getting 25% off. Not a true story. SPEAKER_163: Not sure. Never worked with them ever. Never worked with them. Never worked with them. Yeah. I, I just happen to be a big fan of these sort of products and services, but I just happen to be SPEAKER_175: independently wealthy and don't need any cash flow from any of these brands. Right, right, right. Nick, SPEAKER_04: I gotta ask though. Um, I'm thinking about getting a new tattoo. I only have one and I see you have a tattoo on your forehead that looks quite sharp. Um, but where'd you get it done? I think it's a little SPEAKER_178: dirt. You might have a little smudge over here. Oh, I didn't even, I didn't even notice that. SPEAKER_170: That's crazy. I had no idea. What? Oh, I had no idea. Nick, we appreciate you and thank you for SPEAKER_38: clearing this up. Everybody wanted to know, uh, everybody use the promo code NICO for 25% off your Roman sparks. If you need Roman sparks, Nick doesn't need it. He's all man. Oh, I need Roman sparking. Fully functioning. You can actually get that Roman sparks 25 free Roman sparks. If you use the NICO code. Well, well done, Nick. Appreciate you. SPEAKER_04: Our roving correspondent. Yeah. Coming hot from Miami. Also, I just Googled Roman sparks on my work computer. Am I going to get fired? I didn't know what that was. SPEAKER_38: No, it's totally fine. Roman sparks is a delightful little lozenge that, uh, Rob will fill us in. SPEAKER_187: Rob, you're, you're a man of a certain age. I don't know what that is either, but I'm making SPEAKER_04: some guesses. It turns out it's a, it's a high. Yeah. I Googled it. Anyways, uh, I'm going to grab this by the tail and drag it back on topic here. Uh, Rob, very glad to have you here. Neurometric is the company and you made me go out and learn stuff. I'd heard of SLMs, small language models, but I wasn't sure what the parameter cap was, what they're good for. So first of all, what is the parameter cap? The differentiation point between a small language model and a large language model. SPEAKER_84: And, um, to put it politely, why do we care? Yeah. So the, the cap keeps sliding, SPEAKER_33: right? Just like the LLM cap keeps sliding. So mythos, which we were just talking about, SPEAKER_34: it's a 10 trillion parameter model. So when that goes up, what's small in comparison is, you know, uh, changes, I would say in general, people think about small language models as something you could probably run on a high-end laptop. Uh, and that's sort of a rough. SPEAKER_198: Yeah. Is that 10 billion parameters? Is that 20 billion? SPEAKER_34: 20 billion is normally sort of the, the cutoff these days. Um, my guess is pretty soon people say anything under a hundred billion parameters is, is small, but, but here's why you should care because just as the intelligence density. So, so the intelligence density is the amount of things the model can know for the parameter size it is, right? That keeps going up for LLMs. Part of the way that these models learn more is they, they develop, you know, synthetic data training techniques and reinforcement learning techniques and new, you know, architecture tweaks. Um, those filter down to the small models. So like an 8 billion parameter model can do more this year than it could last year. And so if you're not trying to hack the world's entire code base, right? Right. If, if you're trying, if you're trying to, you know, build a model that reconciles a bank statement with your QuickBooks or, um, you know, predicts customer churn and your customer success funnel or things like that, these small models, as they climb up in their capabilities, we predict that by 2030, 90% of common work tasks will be able to be done by like a 10 billion SPEAKER_08: parameter model for smaller. So what impact does everybody having SPEAKER_07: the equivalent of today's Mac studio, or I just upgraded to a Mac pro 14 inch with 48 gigs of RAM. I could clearly run an SLM on a Mac studio. When you get to 256 gig, 512 gig, you can start running Kimi and another things, I guess, um, open seek those open seeks and the Kimi's, those are not SLMs. Those are straight up open source LLMs. They require server level memory, server level CPUs, but people are starting to do SLMs on their latest iPhone. They're doing it on their latest laptop. SPEAKER_00: So this will be embedded into every device eventually. Yeah. SPEAKER_33: Yeah. Yeah. It definitely will. And, and what's interesting about them is, is so if you look at what's happening with large enterprises that are AI forward, which is not very many people right now, SPEAKER_34: but if you look at companies that, uh, have deployed AI for a couple of years now, what they typically do is they start with a frontier model, right? You know, anthropic or open AI, or maybe one of the open source ones. And they run all their tasks through it. And as their inference charges climb up, uh, they start to go, huh, well, what are these tasks can we put to smaller, cheaper models? Because the bigger models are more expensive to run because when you run a layer of a neural network, you have to shuffle it in from memory, into the compute, calculate it, write it back out to memory and save it. So bigger models take more memory, they're longer to run and everything else. These smaller models, you can run on lower hardware, you know, uh, hardware that's six, seven, eight years old. Uh, that's refurbished. I mean, we even joked at Neurometric about like, can we just buy a bunch of phones and like line them up, connect them to the internet and run small models on them and, and serve those for people. Right. But, but these small models, there are things you can do to tweak them per task. So if you need to do a finance task or a sales task or whatever, you can, you can make them really, SPEAKER_206: really good at that one task. How do you tweak it? So if I wanted to say, take my photo archive and have all of that managed locally, I didn't trust Apple, Google photos, SPEAKER_08: whatever, but I just wanted all these photos on my desktop, on my laptop, and I wanted to tag them all. And I wanted to say, Hey, this is this person's face without having that go up to a frontier model. So now they know, Oh, that's lawn in your photos. Um, that seems to me to be like a great way to tweak this. Or if you just wanted it for writing, uh, or you were an Excel jockey, you would want to just have one tweak for Excel. Are those models already pre-tweak? Can you get a flavor of that model on GitHub or on, um, hugging face? Yeah. For some of them, you can get a flavor SPEAKER_33: of it, but the, the flavors tend not to be work task related. They tend to be things like Quinn has a, SPEAKER_34: uh, a series of SLMs that are called instruct. Um, and so they do instruction related following tasks better than say generative writing tasks or something like that. But there's two main ways to, to sort of like make these specialized. One is to augment your prompts with a whole bunch of stuff that can get hard, uh, with the context window limits and everything else. But you know, you add the stuff like imagine you're a CPA and you have all this knowledge of it, you know, and you can throw the whole gap, you know, accounting standards in there into your prompt or whatever, and pass it to the model. The other way is to sort of fine tune it. The best way to do this is what they call distillation and distillation is when a bigger model teaches a smaller model something. So you can think of us as like setting up a system where let's say you have a task that you're having, you know, Anthropic or open AI do. And maybe that task is like, Hey, I have a thousand pages here from industry reports on the energy industry and I'm an investor. So GPT-5, take this report and pull out all the stock symbols in here. Cause I want to go research them, right? And see what they should do. So if I, if you see that information, go in with a prompt and you see GPT-5's response, you get that back enough times, right? That prompt and response is a dataset you can use to distill a small language model, just to do that given task. Now, depending on how much you do, if you only do that task once in a while, it doesn't make any sense. But if you're like, we do this every day, a thousand times a day for all these pages of reports, you could probably save 90% by doing that. And in fact, there's a story in venture beat in February, uh, about AT&T got to the point where they were spending 8 billion tokens a day on their AI infrastructure. So probably a couple hundred thousand dollars a day. They re-architected everything using frontier models for 10% of the tasks, SLMs for 90% of the tasks and like dramatic, dramatic improvement in both speed and cost. SPEAKER_04: Yes. The headline from the venture beat story is 8 billion tokens a day forced AT&T to rethink AI orchestration and cut costs by 90%. Uh, Rob, can I just narrow down and focus on one of your SLMs really quick? Cause I think it's a good example. So here is one you guys built called deal Civ and uh, it's a task specific model that compares target company profiles against a firm's investment thesis, pretty pertinent to what Jason does here at launch. And it's based off of Quinn, uh, three to 4 billion instructor. Okay. So is this one that you made via distillation? Uh, SPEAKER_27: how did this one in particular come together? So it's a good question. We have, we have an automated process behind the scenes that does a lot of this. So I don't know on this SPEAKER_34: specific one. Um, but look, our, our bigger goal as a company is just make intelligence free, right? Why would we want to do that? Because it's, it's the Jevons paradox thing, right? The more the cost comes down, the more people are going to use it for more tasks. I think the better the world will be if, you know, good people control AI. Um, and then, and so people are like, well, then what do people pay for? Well, you're going to pay for an SLA. You're going to pay for testing. You're going to pay for analytics. You're going to like, you're going to need stuff to manage all this intelligence. So, um, yeah, so, so like you can go and you can download these models. You can, you know, we give them to you for free. If you, if you want to, what's the website again? SPEAKER_215: So people who are listening can go there now and try it. It's just a marketplace.neurometric.ai. And then today, do you want me to talk, Alex, about the claw pack? SPEAKER_04: Oh, we're gonna talk about the claw pack. Well, one second before we get to that, Jason, because you said something very important, which is you want to make intelligence free. And what's incredible Jason about that comment from Rob is that he's not that far off from it already. Um, they offer a hundred million free tokens a month. If you want to have neurometric handle your inference and it's two bucks a month per model per month, no token limits. So Ron, SPEAKER_84: how the hell can you afford to do that? Are these so cheap that effectively they're already free? SPEAKER_33: Well, you can run, you can run these on old hardware. Uh, and then the second thing is, you know, we're, we're a seed stage company, so we're learning the real usage distribution, but you know, at the beginning of the show, Jason talked about Backupify. Backupify offered SPEAKER_34: unlimited storage for, you know, Google apps, Salesforce, Office 365, and people be like, how can you offer unlimited storage for $3 a month? Because we knew the distribution of what people actually used. And I think what companies want here and particularly prosumers and open cloud users and cloud code users, but, but also enterprises is like, they want to not have to think about all these innovations that keep coming to drive down the costs and how they apply them and how they roll them out. So I think if, if we can manage all that and we can bear the token risk, SPEAKER_26: I think it's a great business model, but you're not doing the hosting or you are doing the hosting. SPEAKER_34: We will do the hosting. We don't have to, you can download it. We can deploy it wherever you want, manage it for you. But if you want us to host it, we'll do it. SPEAKER_12: I think this is going to become the key. Uh, anybody who goes deep down the open claw SPEAKER_07: or even perplexity computer, which is awesome. Or, uh, Claude co-work at some point you're like, am I getting enough value from spending a thousand dollars a day, $365,000 a year? Now, if your business is printing money and you don't have to hire the fourth developer, SPEAKER_08: you know, okay, fine. But in other circumstances, you're going to be like, you know what, uh, this task, you know, I have tasks I want to do, uh, which I'll talk to you offline, Rob, that are not necessary for my business, but you know, you know, we have to judge the startups that are coming in, but I would love to be running back testing on every startup I've ever met with in my life, the founders, where they wound up and be saying, okay, let's examine every 2009 startup, every 2010 SPEAKER_07: startup, every 2011 startup, and tell me, look for some patterns, look for the talent there, which talent created diasporas of, you know, Googlers or people who worked at Uber who went on to do other companies. There's all kinds of intelligence that I could see myself using, but it's not a priority. And I would look at the a hundred thousand dollars in token costs and say, not worth it, but at a thousand dollars or $10,000. Yeah. It might be worth it. I think a lot of people who are the tip of the spears here playing with this technology, SPEAKER_08: they're starting to come to that realization. There's things I want to do that don't make economic sense today, but if the tokens were cheaper, I would do it. Why not? SPEAKER_222: You know, Rob, if only there was a brand new product out there called claw pack that for only $8 a month, got you a loaded inference. You want to tell us about it? SPEAKER_34: So as soon as we put this marketplace out, one of the things people started using it for was obviously open claw because that is, man, you go on Reddit and people are complaining like crazy about their open claw prices. So we said, okay, let's, let's do some research on what are the top open claw use cases. And let's take 39 small language models that we package together under one API. So we host 39 models for you. They do common tasks like social media posting, right? Like you don't need Claude Mythos to write an email headline, right? Subject line. So it does all this stuff. You get a hundred million tokens for free on it. And then after that, you pay $8 a month, unlimited tokens. We manage the token risk. And there's a lot of ways to do this. The other thing, Alex, that we haven't talked about and you didn't bring up, but there's this new emerging concept in tech called harness engineering and harness engineering is the thing you wrap around the model to make it do things. And one of the things we noticed in our research on small language models is one of the reasons you can't use them for complex tasks is they kind of tend to go off and lose track of what they're doing where the big models don't. But if you put a nice harness around it, that helps it stay on, but it says stuff like, Hey, every time you do a step of a task, check back in with me and make sure you're on the right task. Like those kinds of harnesses are easy to write. Claude code can write you one for a specific task and you plug in a small model and run it. And that harness keeps the model on task. So we, we have a bunch of innovations like that, that, um, enable us to, and you know, we'll just keep driving down the cost and other people SPEAKER_224: are doing stuff to drive down the cost. So Rob on, on the harness point, uh, regarding SLMs, SPEAKER_84: is this like a dog sled? Do I have like one chain of harness that has multiple dogs, AKA SLMs that it can work with, or is it like one harness per SLM? SPEAKER_27: It's both today. It's mostly one harness per SLM because they're task specific, but you can SPEAKER_34: already see with the claw pack, right? Yeah. They're starting to work together. And what you're going to start to see is, uh, swarms, I think of SLMs that'll do common, common work tasks. Um, and you'll, you'll always need 20% of your workloads to fall back over to the frontier models because they're one-offs or you, you, you get lost. You don't know what it's doing. So you're always going to need a combo ensemble system, but I think we can take people's core workloads SPEAKER_04: and drive that cost way, way down. Yeah. Uh, Jason, the power of branding. Remember that moment in time in which we all said, oh, if you're an AI rapper, you're doomed. But now if you're an AI SPEAKER_186: harness, you're the tip of the spear. Yeah. Well, you know, it's, there's like two interesting SPEAKER_08: observations here for startups. One is, um, every Reddit bitch thread where people are bitching about something they hate is a potential startup. This is like a super important thing. Somebody needs to create for me, a skill, uh, for my open SPEAKER_00: claw to just look for these startup opportunities from people saying this product sucks. SPEAKER_38: And I would like, uh, you know, why can't they solve this very simple product problem? I would pay a lot of money for a tool that did just that. Um, because it would be like the request for startups SPEAKER_07: that we do, YC does other people do. It'd be like, these aren't my requests for a startup and my intuition as a VC, 55 year old guy in Austin, Texas. Like what I want is irrelevant SPEAKER_00: compared to what the world wants as demonstrated in a subreddit about music, where somebody wants a specific tool to do a specific task. And a thousand people participated in that thread. Like that's really an interesting, um, and it's an interesting go to market movement, SPEAKER_69: because you can go to that thread and say, Hey, I built something. Would you test it for me? Would you try? Uh, uh, and then people are like, wait, you made me a custom piece of software. Okay, great. And then startups now are not about who can build the product. It's about who doesn't stop building SPEAKER_07: the product. Startups aren't about who has the resources to build the product. It's about who will not stop building that product and refining it. In other words, it's now a test of your resiliency, your passion for the vertical. Will you just keep working because anybody can build anything then SPEAKER_00: well, who's going to build the next three or four features of this, you know, uh, meditation app or this, you know, um, fitness app, whatever it happens to be, you know, this enterprise piece of SPEAKER_69: software. It's really just who's willing to go on that product march for 10 years and not give up. SPEAKER_222: I think Aaron Levy is the example of that from the SAS era, but I think we're going to need to SPEAKER_04: find out who that is for the AI era. We're talking a lot though, Jason, kind of around this, uh, this Mark Andreessen tweet, uh, he agrees that you can't spend a thousand dollars a day on open claw. And he says, it's actually heading to $10,000 a day. If you really want to have a magical experience. And then he says the future shape of the entire tech industry will be how to drive SPEAKER_84: that to 20 bucks a month. So Rob. Yeah. How far can SLMs get us? Can they, can they reduce our open cost spend by 80% in time, 90% of time? And how quickly can you cut my bill down? Because my SPEAKER_34: wife's not happy. Depends on what you use it for, obviously. Right. Because there's some tasks SLMs can do and there's some that they can't, but the important thing is people are going to use more tokens every year and SLMs capabilities are going to increase. And there's a bunch of costs that are also making it more efficient to run those every year. So I would say today for an average open claw user, probably 70% reduction, you'll still have to use quad for some things or open AI or whatever. Um, but you're going to, you're going to have this weird paradox, right? Which is you're going to do so much more with AI that like your per unit costs are going to come down, but SPEAKER_77: you're going to spend more because you're going to turn over more parts of your life to it. And you're going to do more things. And, you know, I have a prediction. Let's hear it. I have a prediction SPEAKER_06: there. And I'll unpack it. There is a possibility that if we are reaching AGI right now, which I've SPEAKER_08: said, Hey, we've reached AGI. It's just not distributed yet. It's just not implemented yet there. And you believe in super intelligence and recursive learning, which everybody here believes, then there is no doubt in my mind that LLMs will get so small and there'll be so many verticalized ones because building a legal accounting design topography. I mean, I don't know how many steps down you can double click. There is a possibility that these SLMs being fractured, being essentially like skills, you know, instead of building a skill in open claw, there's an SLM that does that skill it's trained specifically. And somebody makes it better and makes it so good SPEAKER_00: that this could collapse the value of the frontier models, because how does a frontier model sell into a law firm or an accounting firm? Like, Hey, we've got this new amazing thing to help you solve these legal or accounting or design issues. When there's an SLM that the person goes good enough, that's a good enough logo. That's a good enough font. That's a good enough non-disclosure agreement. Like good enough happens. And then the ability to sell a note taking app, which we saw there were dozens of note taking apps. This was like a great business to be in dozens of mail. It was a unicorn. SPEAKER_08: Evernote was a unicorn. It's a perfect example. Like, and now it's like, well, notes, uh, you know, I can make one notion. It's just these things eventually become deflationary. This could be the deflationary moment for the frontier models. They may not realize it, but they might SPEAKER_00: have created their own demise. Yeah. They may have just created their own demise. Chamath Palihapitiya: Dario is going to have to buy Neurometric. I mean, I don't know, you know, we'll see. SPEAKER_00: Um, but I, but who's going to, who's going to pay for these things? If an SLM exists, that does it, or a Tau subnet solves that problem for you in a distributed computing way. This is so deflationary that we need a new word for deflationary. There's deflationary, but what is hyper deflation? Like it's, it's not just deflation where it's like, but it's going to get 10% cheaper year. What if it's, it's going to get 90% cheaper every month? Like then the, the compounding deflation, the hyper deflation could be so acute that just things like you're saying, get to free Rob. And that's just kind of a mind blowing exercise to do in your mind. Well, you're going to, I mean, SPEAKER_34: the price of compute is collapsing pretty fast as well as people on a per, you know, petaflop basis or however you want to measure it. Um, but I will say you will always come up at least against the cost of buying the, I wouldn't even say GPU, cause there's other chips coming, but that plus the energy. So, so there will be some floor to a unit of intelligence that'll start to advance, decline slower because of the laws of physics, right? Yes. But I think as the floor SPEAKER_02: rises in the quality of open source models and cheap SLMs for specific tasks, as Jason points out, SPEAKER_04: yes, there is less room for frontier models to charge, but I think also that the companies that are desperate to have an edge will, because everyone else is going to be defaulting to the cheap stuff or cheaper stuff. And so I think there's still going to be some market there, Jason. I just don't think that everyone's going to be paying Opus 4.6 level pricing for tokens in five years. But the question that just becomes will job on paradox drive token usage up enough in that same time period to have these still be growth businesses. But I'm still pretty bullish on improving the, the frontier of intelligence, because as mythos proves, there's still so much left to come. Like, I don't want to start thinking about ending that run, Jason. I want to go up the damn mountain SPEAKER_23: to the top. Compute performance per dollar has improved roughly 40% per year across 20 plus AI SPEAKER_08: accelerators released between 2012 and 2025. The GB300 costs nearly 9x the P100's release price, but delivers 24 times the performance per dollar. So that's all that matters is performance per dollar, SPEAKER_06: 24 times. That goes to this hyper deflation concept. SPEAKER_119: If you think about it as an investor, Jason, it's going to change the way that you think about SPEAKER_34: defensibility. And I think one of the possible outcomes of this is it, you know, it starts to enable a type of business that I don't know if it's like a venture funded, but it enables the type of business run by one, two, three people. Sure. It can do, they can, they can work in small TAMs, make $30 million a year, drop 9 million to the bottom line. And we're 29 million to the bottom line. Yeah, exactly. I, it's, it's going to be super, super interesting to see like how you adapt your angel. And I mean, the way I, I, I, I've been SPEAKER_08: through this before when storage became free or was trending towards free, um, YouTube and Netflix became viable. Mark Cuban was like, Netflix will never work. The infrastructure is not there. Um, it would, if everybody had Netflix and was streaming HD, it just wouldn't work today. SPEAKER_07: And he was right. But the compounding effects of deflationary, you know, technology in that case, it was the rollout of fiber in the case of hard drives. It was, you know, what those discs could store and at what price and how it would scale, you know, in hardware scales differently than software, SPEAKER_08: but you have two compounding effects here. The AI is so good that it's making the models more dense to your point earlier, the density of the model and the specificity of the model that could have more dramatic effect than the hardware curve, but then you add the hardware curve to it. This is where I think, you know, this 40% cheaper or tokens are 90% cheaper. We might be greatly underestimating what super intelligence does. It might be that super intelligence just rams this down 99% a year. SPEAKER_02: Maybe, but with the humans are doing quite well as well. So, uh, Meta, their new model, SPEAKER_04: they dropped today, Muse, Spark, uh, in their post talking about this, Jason, they talked about how they're getting more efficient with their compute. Um, they said they rebuilt their pre-training stack, and they said these advancements increase the capability we can extract from every unit of SPEAKER_85: compute. So it feels like every possible vector. But why does Meta still suck at doing anything with AI? SPEAKER_83: Like they built their AI search sucks on like Instagram. Nothing works. It's a disaster. SPEAKER_84: Their new model is not hot garbage. I have the, I have the chart here. This came out right before SPEAKER_04: the show. So no one's seen it really, uh, artificial intelligence, sorry, artificial analysis on their intelligence index points out that the last Meta model we got, which was long before Maverick, which is now complete garbage compared to the state of the market has now been replaced by Muse, Spark, which is now the fourth best model, um, out there. Now we don't, according to who, what is that? Artificial analysis on AI, they run, um, a series of very good, uh, benchmarks. And this is kind of a Meta benchmark that I pay a lot of attention to. Now we don't know how it's going to perform in the market. I'm not trying to say this is the best model since last spread, but I'm saying that it's a really pleasant surprise. SPEAKER_08: But what is their strategy? Like, okay, they leapfrogged their last model. They're trying to SPEAKER_00: catch up, but what is the business case here? Like, what is the user case here? I don't see a user case for what they're doing right now. Like what, what is it actually, I mean, I understand SPEAKER_261: you could serve better ads, but go ahead, Rob, what, I think it's mostly, yeah, I think for them, SPEAKER_34: I think initially Zuckerberg thought he would hurt his competitors. He would hurt, you know, OpenAI, he would hurt Google by open sourcing stuff. Now I think they're looking at Cogs improvement because I also know they're building a special chip, right? They announced this in a conference late last year that, so it's just an AI recommender chip. Probably, probably Meta, Netflix and Amazon might be the only three companies in the world that could spend the money to do a recommendation specific chip and have it make economic sense. So maybe that's SPEAKER_07: how they're thinking about it. Just make people that much more addicted to a product that's already so addictive that it's being regulated and banned for teens under 16. Okay, great. What a stupid idea. SPEAKER_00: Like the worst possible thing they could do is use this technology to make an already smoking level addiction, a heroin level addiction worse for children and adults. Like, that's what I'm saying is where's the vision here? What is he trying to accomplish? I haven't heard from Zuck. SPEAKER_267: There's no big change the world thing that I see. There's nothing. And you know what, SPEAKER_00: this proves my point about him. He's great at like copying other people's ideas, Snapchat, LinkedIn, whatever, Facebook marketplace, eBay and Craigslist. But where is the original idea from that cooperation of how to deploy something unique in the world? I just, SPEAKER_84: it's pathetic. On that vector, nothing to add, Jason. Nothing. But I did ask, I asked Alexander SPEAKER_04: Wang, formerly of Scale AI at the Meta Super Intelligence Lab, is it going to come out that I can use it in my open clause setup? And he told me that they will release the API soon and it will power some clause. So at a minimum, we can put it to test soon enough and see if it's any good, SPEAKER_272: see if it's, you know. If Zuckerberg were to put his entire energy on copying open clause, I would be very nervous because he's so good at stealing other people's ideas and doing them SPEAKER_38: better. Yeah. Like he would create open clause times 10. Uh, so yeah. Oh, that's a challenge. Let no, don't give him any ideas. All right, Mark, you heard it here first. We need, SPEAKER_04: we need Meta Claw. All right, let's bring up Giani. Giani, uh, is the man behind probably my favorite and least favorite tool online today, Jason. It's called death by Claude. And well, you know what, I'll let the man himself explain it. Giani, what is death by Claude? And why, why did you build a tool that insults me? Well, hey, thank you for having me on the podcast. So, SPEAKER_280: uh, what is it by Claude? You put a URL in there. It can be a person, it can be a company and it'll critique whether it's an AI rap or something that can be replaced. Why did I build it? Uh, I'm running a company myself and we have been more philanthropic a few times. You're building something like co-work and then co-work came out and their margins are bad. If you try selling co-work at and proper space, you can not do it. So we've been killed by Claude a few times. We read the Cetrini piece, uh, that was making the rounds. I did the tweet by Iran Peterson, where he SPEAKER_281: said, uh. Okay. Show us this, show it to us here. You are going to savage somebody personally, SPEAKER_00: like a roast comic, or you're going to savage a startup or a business idea. What's the best example SPEAKER_04: of this at work? Gianni, should I bring up the, the one I did about my own newsletter? Yes. Okay. I roasted myself for everyone's enjoyment. So here is the, here's what death by Claude kicked out for cautious optimism. Uh, I have an 89 of a hundred already dead score. Uh, Gianni, how do you calculate the scores that you're giving to each company or product? SPEAKER_280: So if you were like a hardware company or a science company, you're basically a model. And if you're a blog, then you're probably dead. So if you can get a place and you score high, SPEAKER_179: and if you cannot, then you score low. Yeah. So I'm pretty much doomed. And Jason, SPEAKER_04: what this service does is not only does it find different ways to tell you that you're replaceable, it creates an entire skill MD file to replace the thing that you're doing. And then also it gives you a death certificate. And it says that, uh, my newsletter was cautiously optimistic about its own survival, but it taught us that the real disruption was the sub stack fees we paid along the way before it died. Brutal. But the question is really like, like how much of the stock market do you think, Gianni is really at risk from being replaced because you can do Uber, do Uber, SPEAKER_00: do door. What would be one that would be hit me? No, let's do, um, I'm trying to think of, SPEAKER_07: well, I gotta be careful here. Cause I don't want to sling mud at anybody, but, oh, okay. Let's do SPEAKER_83: an end roll. No, no, no, no, don't start that up. I think those sound like serious companies to me. SPEAKER_04: Yeah. Cause the hardware, right? I mean, basically companies that rank, uh, SPEAKER_38: well, have a low score. Yeah, but let's not do Andrew. Let's do, um, let's do, um, do Slack. Slack. Okay. Or Chegg is a good one. The textbook company. Well, Chegg's already dead, SPEAKER_00: but okay. We'll do Chegg. Well, Peloton or Chegg, like these are things that there's a group of SPEAKER_272: people who are trying to pump Peloton right now. It's an interesting one. Chegg, 92% are already dead. SPEAKER_302: Yes. Let's see what it has to say. Um, it's just crud. It's an AI rapper. It has no moat depth. SPEAKER_04: It's marked down, replaceable, and it's expensive. And the, the skill file to replace it, 31 lines. And yeah, brutal, just absolutely mugged. I feel like every founder should be trying this to a lot. David Friedberg: Do Peloton. Peloton's a good one. I'm on it. What else you got, Rob? Who else SPEAKER_07: is do we want to figure out how dead they are? Cause Peloton to me is not dead. They, SPEAKER_08: they basically had all the suckage taken out and Peloton is actually defensible now because people SPEAKER_00: love their brand. Uh, they love the hardware and they love the people who are the teachers. They're addicted to those teachers and that brand and those teachers is defensible. Yes. Uh, and in this case, SPEAKER_88: 32 out of a hundred, much better than my blog and much better than other services that we've looked SPEAKER_84: at. Uh, it's not just crud. It's got reasonable moat depth. Yes, it is expensive, but you can't replace a bicycle, uh, with code, I suppose. Let's do one more for fun. Um, lovable, lovable. SPEAKER_00: Everybody's saying that vibe coding. Yeah. People say vibe coding is dead and like, yeah, it's a, but lovable keeps having revenue grow, but then there was some folks saying maybe revenue was stalled or a SPEAKER_83: lot of people canceling their subs. All right. 78 out of a hundred score on death by Claude for SPEAKER_317: lovable. Why? Why? I thought it would be 60. Okay. Go ahead. Lovable built an AI that writes code for SPEAKER_04: you, which is adorable because Claude already writes code for you without the $20 per month middleman tax. Brutal. And then a 31 line prompt for a full stack app generator. Um, yeah, brutal cause of death, terminal AI rapper syndrome. The underlying model got too good and ate the rapper alive. SPEAKER_156: All right. G-man, G-man, why did you build this? Who are you? Where, where are you located right now? SPEAKER_280: And why did you build this? I live in London. I am building an AI startup myself. I was building something like Claude co-work and then Enthropic released before us. It's very expensive to serve in France. So Rob is a friend, I guess in the future. So I was reading the Sertini piece. I read this piece by Ryan Peterson, where he said, Harvey is like replaceable by Claude for legal. And that's like a real feeling. Like we have been worked by Enthropic a few times. So we decided to SPEAKER_286: run our startup and YC's entire portfolio and be, the results are very funny. So put it on the internet. SPEAKER_00: This is becoming a reoccurring thing, Rob, is using AI to give a defensibility ranking. That's basically what you're giving. And I think it's genius. A defensibility ranking is a great idea. We do this with founders organically, Rob, in the founding university, you know, or the launch accelerator. And we tell them, hey, here are your competitors. How are you different than these competitors? Or, hey, if Claude releases this, what would you do? G-man, if I may call you G-man, G-man, you got your ass kicked and you said, I want to create a tool that prevents me from getting my SPEAKER_69: ass kicked again. What are the two or three things that prevent an ass kicking of a startup based on what SPEAKER_280: you've learned? G-man, a few things come to mind. So if you're doing hardware, we don't have physical models yet. So when Travis succeeds with our teams, I think that's a space that goes away. SPEAKER_12: But hardware is number one. Got it. And that, by the way, is becoming now, we went from hardware is hard to hardware is a great way to be defensible. And hard means defensible now. SPEAKER_144: G-man, hard means moat. G-man, hard means moat today. Whereas hardware meant SPEAKER_00: death and not fundable, it now means moat, highly fundable. Very interesting. SPEAKER_329: G-man, what's your number two? G-man, network effects are very, SPEAKER_280: very helpful. So if you're WhatsApp, you will not be replaced by Claude. I don't think any AI company at the moment has network effects, so I would be on the look for it. SPEAKER_08: G-man, so network effect number two. Network effect being, hey, how many people are participating in this, SPEAKER_00: like a marketplace like Uber? Hey, they've got 20 different AV companies putting cars into the system. They're in, you know, 10,000 cities. They've got 7 million drivers. They've got, SPEAKER_08: you know, 10 million restaurants. Okay, that makes it more defensible network effect. Great. What's SPEAKER_280: number three? G-man, if you're doing something deeply scientific or in the regulated industry, which requires talking to people, I think you're pretty safe as well. SPEAKER_185: So yeah, that's the third thing. G-man, please, G-man. Sorry, G-man. He's the G-man. SPEAKER_321: You're behind TabTabTab.ai, right? That's your main company? Oh, that's the thing that we were working SPEAKER_282: before we started pivoting because our score is 89 on the, it's like Claude said. Oh no, I was, SPEAKER_04: it's actually 92 now. Your score has gone up. I just thought it was very funny that your own product told you that you need to pivot and off you go. So there you have it. Rob, you got a question for the SPEAKER_272: G-man or do you have an observation here about this as it relates to being a serial entrepreneur SPEAKER_338: yourself? Well, defensibility has always been hard and I do think it's particularly hard now. SPEAKER_34: But I mean, I think he's on the, I think he's on the right path, right? I think this is making everybody a little more like a CEO because as a CEO, you realize like you have to review stuff, you have to coordinate stuff and there's some limit to what you can do. Like at some point you have to decide like, what are you doing? Right? And these tools help you do stuff. But yeah, I think it's, I don't know if we talked about this Jason, I stopped, you were one of my mentors for angel investing, you know, 11 years ago when I sold my first company, I've quit because of this. I think SPEAKER_12: it's too hard now. There is something to it. It is very hard to be an early stage investor. Even if you SPEAKER_08: look at people doing angel investing and getting into a big round, you got to get in at a higher evaluation. So even, you know, the idea of hitting a seven, six or 7,000 X, which I think is where SPEAKER_07: Uber might've peaked for me. Um, you know, or hitting a 500 X, which is, I think what Robinhood did, you know, a hundred X, whatever it is, calm a hundred X, like it's going to get harder and harder to do that. Cause the entry level valuation is a magnitude higher dilution might be higher. It's just hard. It's never been easy. I'll say that. Yeah. Yeah. SPEAKER_345: Tricky. Oh, no, you're not doing you. Oh, I don't know. Should we do neurometric? That's SPEAKER_349: weird. Oh, you want to, Oh, I don't, I don't care. All right. There's something you learn here. We'll figure out if we have to, we'll figure out if we have to pivot. SPEAKER_302: All right. Well, this is the meanest thing we've done to a guest on the show in at least a week. SPEAKER_07: Here we go. I, I, I know the score already and I know why it's giving that score, but I can, I can explain why this score will get better over time. Still calculating the exact level of doom. SPEAKER_356: Here it goes. I, I love that Alex. What does it say? SPEAKER_04: It's in counting the lives of marked I needed to replace cross-referencing with the obituary database. Uh, Guiani built one of the funniest tools I've ever seen. It's just consistently, SPEAKER_32: uh, not a great score, but not terrible. 72 out of a hundred. I think you can get that number SPEAKER_00: down because the, uh, the, a network effect would be an interesting one. If you could make your platform, Rob into a place where anybody can put up a, this is my idea. If you became a marketplace SPEAKER_08: of SLMs where anybody could post an SLM and put a price on it, and you would be the reseller hoster of it, then you become more like eBay or, you know, a marketplace of these idea ideas like SPEAKER_00: hugging faces. Right? So if you were the hugging face of SLMs, like put a hugging face in there. I wonder how it thinks about hugging face. It might also think that hugging faces, you know, SPEAKER_34: uh, not. Yeah. The, the thing we've been working on. So there's a, there's an open source project called Harbor and Harbor is a collection of thousands of, um, tasks with their evals. And currently people in the world think about benchmarks, not tasks, right? But a benchmark is a series of tasks. And what we've shown, and a lot of the thesis behind neurometric is that you want to choose the right model for the task. And even amongst frontier models, they don't all perform the same on a given task. So even the model that wins the benchmark doesn't win all the tasks within the benchmark. So ensembles of models will 100% of the time beat single models. Like you just, you can't make one model the best at every single thing. Um, and so, um, so to your point, we're kind of going in that direction that you're talking about, Jason, but it hasn't all been released yet, but it's a really great insight. It's why you're such a good early SPEAKER_00: stage investor, man. Yeah. I provide a little value on the margins. All right. This has been another amazing episode of twist G man. Thanks for coming on the show and sharing this with us. SPEAKER_08: Thanks for having me, Alex. Great job. Uh, Lon Harris. They say he's 72%. He's a 72 square unreplaceable. I disagree. I disagree. Personality SPEAKER_07: counts for a lot. Yeah, exactly. That's why, that's why he's a 99 out of a hundred 69. I put SPEAKER_353: him at 69.420. Uh, do me, do me, do my LinkedIn. If my newsletter doesn't make enough money by the end SPEAKER_222: of the year, I'll go work for a VC. I'll make, I'll make you a deal. I mean, it's hard. Oh boy. Uh, SPEAKER_00: this is going to be scary. I think it would probably say venture. If it says venture capital and podcaster, man, what am I? Maybe I'm like, uh, 65, 64. Uh, okay. So I have, I have the image SPEAKER_88: here from law. Okay. Here we go. Replaced by Claude. Oh, I should not have the full URL. SPEAKER_83: All right. 58. Yeah. You are at 58 and, uh, we can cut this nearby. No, leave it in. I'm a 58. SPEAKER_00: I like it. Um, yeah, look at this Brooklyn startup rooster in a Fordham blazer pacing the cap table at SPEAKER_38: 6am and yelling, ship it into three microphones while stuffing another safe into his jacket. I mean, it's, it's got a sense of humor. That's the thing. These are hilarious. I mean, I like it. I like it. Do. Okay. I could do this all day. It's fun. Yeah. We'll see you all on fr fr Friday. Bye.