Chamath Palihapitiya: It's really important to benchmark LLMs to figure out where they are good at, where are gaps in the models so that you can generate data that could help the LLMs get better at those specific tasks. One big challenge, Alex, that I see today in the overall evaluation and benchmarking market is that a lot of the evaluations are somewhat academic and somewhat synthetic and don't connect to real world applications or real world use. Ideally, you want AGI to progress in a way where the models get better at useful tasks. SPEAKER_02: This Week in Startups is brought to you by Northwest Registered Agent. Starting your business should be simple. With Northwest Registered Agent, you can form your entire business identity in just 10 clicks and 10 minutes. From LLCs to trademarks, domains to custom websites, they've got you covered. Get more privacy, more options, and more done. Visit NorthwestRegisteredAgent.com slash twist today. Dot tech. Say it without saying it. Head to get.tech slash twist or your favorite registrar to get a clean, sharp dot tech domain today. And AWS Activate. AWS Activate helps startups bring their ideas to life. As you build and scale your business, Activate Credits grow with you to support your changing needs. Apply to AWS Activate today and receive up to $100,000 in credits. Visit aws.amazon.com slash startups slash credits. SPEAKER_04: Hey, welcome to This Week in Startups. You may notice I am not Jason Calacanis. He is on the road this week. We're going to do a quick show without him. I am your host, Alon Harris. Joining me as always, Alex Wilhelm. Hello, hello. We have three incredible, amazing Twist 500 interviews coming up right now that Alex did all week. Don't move from that seat. But first, I do think we've got to talk about GPT-5 a little bit, Alex. SPEAKER_08: Yes. Yes, GPT-5. SPEAKER_10: So this has gotten a lot of play, but Sam Altman, one of the founders of OpenAI and kind of what you might call the face of AI from the software side, Jensen being the hardware face, if you will, went on a podcast with a man named Theo Vaughn. And they were having a pretty candid conversation about what they're building. And Sam says, you know, I was playing with GPT-5 and I gave it a challenge and it did it very quickly. And he talks about how it sat him back and like made him think about what they're building. And the quote that really went viral was he was like, you know, there are moments in time in science when people sit back and go, what have we done? And, you know, on one hand, Lon, I have three things here. First of all, Sam Altman loves to hype things up and then get mad at everyone when they get a little overhyped. He's a master at PR. So yeah, part one, two, if you go back in time to GPT-3, even maybe even GPT-2, there was a lot of worries that as this technology advanced and progressed and got better, it will be misused. And we're seeing that now when people are using voice cloning to trick elders out of their money and AI is being used in more and more nefarious circumstances. But we also at the time had a lot of people thinking that there was going to be a lot of improvement and things were going to get much better more quickly. So when I when I think about his GPT-5 comments, they don't actually feel that different to me than what we've seen with GPT-4, GPT-3, GPT-2. So that's to me, this probably does feel that impressive to him as he felt blown away before. So to me, I don't think he's being cynical. SPEAKER_12: I don't think he's overhyping here. I think we're just seeing improvement and I don't think we know exactly where it's all going. SPEAKER_09: I mean, I think that sounds reasonable enough. SPEAKER_04: And look, I'm sure he's playing around with GPT-5 and is genuinely impressed at what it can do. Like, I think there's two levels here. Like, even like I was very cynical about AI early on. Even I have had experiences like that where I'm impressed by what it can do and it sort of raises the bar on what I thought was possible. I'm not discounting that experience at all. I'm sure he does have that experience. But I also think this is a weird and unfortunate narrative that we've sort of fallen in. Like, and Jason has talked about this many times that like real science and real technology kind of chases sci-fi. And I feel like sci-fi introduced this idea that AI is scary and that AI is going to be thinking for itself. And all of those things like from Skynet and Terminator to Scarlett Johansson and her evolving SPEAKER_19: beyond Joaquin Phoenix and wanting to go off on her own. SPEAKER_09: Like, I think that's always spoilers for her. SPEAKER_04: Sorry, everybody. I feel like that's always the fictional narrative around AI. And I do feel like that has an impact on real world AI CEOs. And they're trying to be part of that narrative and play into that narrative because it's a helpful narrative for them. We're creating iconic, legendary history shaping technology that's akin to the atomic bomb or the discovery of electricity. SPEAKER_19: And I mean, like it plays into a friendly narrative if you're the CEO of an AI company. SPEAKER_10: That's absolutely true. No, you're not wrong. I just think that AI has put up enough in terms of real results, real progress, real products, real services, that to me, we're just talking about how correct they are SPEAKER_26: versus how incorrect they are when they make those large pronouncements. SPEAKER_10: But I do think the context here is that OpenAI is expected to release their open weights model, their kind of open source-ish projects sometime soon. And then GPT-5 is anticipated, we think sometime in August, which means that we're going to get past the hype and more into the testing. So, you know, will GPT-5 be as good as it's expected to be? Well, we're going to have to find out pretty soon. And that's encouraging. SPEAKER_04: Pretty soon. In fact, the media is sort of saying, you know, early August or sometime in August, the Sharps at Polymarket, of course. Yeah. Yeah. We were looking at this. Yeah. Not content to wait and see what happens. There's already, of course, wagering going on. There's already a market going on for this. 25% are betting that by August 5th, which is early next week, that would be, SPEAKER_19: we are right on the verge of GPT-5 being here. Yeah. One in four. SPEAKER_10: Yeah. One in four. And then by August 15th, people think it's a two thirds chance that it's out. So my read of this lawn is that basically middle August is kind of the over under, and that after that will be a surprise earlier than that will be a surprise. But we're only a couple of weeks away. That's what's exciting to me. That's the cool thing. Remember how exciting it was when XAI dropped Grok 4, and they talked about Grok 4 Heavy, and then the Grok 4 Coding. Then we got the Kimmy K2 family of models from Moonshot. Then we got the Z.AI, I forget, the GLM 4.5 series. The point is, it's been a flurry of models, and Alibaba's Quinn models are improving. Yeah. SPEAKER_19: We're in the middle of that meme that you are here, the world's greatest AI model. Like we're right in the middle of another cycle. Yeah. SPEAKER_10: But the question is, will this be just another iterative improvement? And if so, OpenAI is going to take a lot of egg on its face, or is it as big of a jump as maybe GPT three to four was or larger? And if that's the case, it does change the entire fabric and landscape that the technology world plays on. But Lon, let's talk to some founders. I love it. Let's do it. All right. So we have three interviews for you today. One, Cortical, then we're going to talk to Turing, and then we're going to talk to Mercore. Each one of these was an absolute treat. Cortical is probably the most out there in that we're talking about kind of like hybrid biologic and digital computers. It's an interview I chased down for a long time. It's super fun. And then Turing and Mercore, a bit more traditional in the software space, but incredibly interesting founders, incredibly interesting companies. And especially, I think, in the case of Mercore, a vision for a company that could become incredibly instrumental to how labor functions in the future. So these are a treat. I had a blast. Please enjoy them and say hello to three of our founders from the Twist 500. SPEAKER_12: Hey, welcome back to Twist. This is Alex. Today, we have another Twist 500 interview with a company that has been my white whale for several months, but I'm so glad to have them on the show. SPEAKER_10: We're going to talk about Cortical Labs. Now, here at Twist, we love chips. We love talking about what Grok is building to handle AI inference. We love that Etched is building ASICs just for LLMs. We've talked to Xtropic and their thermodynamic computing paradigm, but all that's entirely silicon-based, more or less. Cortical wants to bring the world of biology to the world of computing in a way that when I first heard about it, I thought was science fiction, but they've taken it out of the lab and they're commercializing it. So please join me in welcoming to the show. It's Cortical Labs CEO and co-founder Han Wing Chong. Han, how are you doing? SPEAKER_40: Hi, good, thanks. Thanks for having me on the show, Alex. I'm so excited. SPEAKER_10: Okay, so you guys are fusing biology and I would say digital computing in a way that's really interesting, but I think to help everyone understand how you got to where you are today, we should go back to the earlier days of the company when you guys were experimenting with your technology and you used it not to play Grand Theft Auto as everyone wants to do with generative AI, but instead to play Pong. So Han, can you tell me about why you chose that and then how the experiment went to prove out what Cortical Labs wants to build? SPEAKER_44: Yeah. So, you know, if you take a few steps back to, I guess, 2017, 2018. There was, I guess, the first AI boom that happened. That was, I would call it the Convolutional Neural Network Reinforcement Learning boom. I had just exited from my last business and as part of that, it was a medtech business that gathered a lot of data. And I was looking at using machine learning to do automated diagnostics and really got deep into the rabbit hole of machine learning and artificial intelligence. And I came across a paper written by Sir Demas Hedasabes from DeepMind who wrote that the machine learning and AI community had to go back to its roots in neuroscience, because that's where most of the initial discoveries were made. David Friedberg: Founders, if you've got a product, maybe you've got some customers or even just a little bit of SPEAKER_50: traction, guess what? You've got yourself a startup and it's time to make things legit. You want to be official, tighten it up. Investors don't want to wire money to a Gmail account with a PO box. No, they want to know you're a serious person and that your company is incorporated. In 10 clicks and 10 minutes, you can file your LLC, get an actual domain name, launch your official website, claim your business email, and even start fast tracking your trademark application. That's right. You need Northwest Registered Agent, the service that's going to do all of that for you. With NWRA's privacy by default option, they're going to use their address for all your public documents, not yours. So you're not going to get a ton of spam or junk mail. And you can also get a legit address and working phone number with their virtual office setup. So get more than just an LLC, get your entire business identity. Go to NorthwestRegisteredAgent.com slash twist and show SPEAKER_44: the world you're in business. So I did exactly that. I went to the neuroscience department at the University of Melbourne, which is my alumni, and spoke to the researchers there and asked them, what's exciting in your world? What are you working on these days? You know, I'm thinking of applying machine learning and AI to the neuroscience world. And one thing that really kind of stood out to me was they said, oh, we have this device called the multi electrode array. It's not a new device. It's been around, I guess, since the early 2000s, where you can actually itch in electrodes into petri dishes where you can grow neurons on top of them. And because neurons communicate electrically, and because chips run on electricity, you now have this common language SPEAKER_54: between the biological system and a compute system. Electricity is the shared language between both SPEAKER_57: neurons and chips. Absolutely. I mean, if you think about it, that's how, that's how Neuralink works, right? I mean, the joke we have at Cortical Labs is that we are the inverted Neuralink. SPEAKER_60: Yes, you take the brain out and put it into the computer versus putting the computer into the brain. SPEAKER_57: Correct. Exactly. I had read the paper from Demis. I was inspired by it. And the first thing that Demis tried to do at DeepMind was he tried to get the machine to play punk. SPEAKER_62: Ah, okay. So we're just repeating. SPEAKER_44: Yeah, yeah, yeah. Exactly. So quite a lot of the first few attempts weren't really successful. But, you know, fortunately, these neurons are very malleable and they're very dynamic and they self-organize. And when we realized that, you know, if we were trying to tweak the system while they were learning, you would end up with these very weird so-called attractive dynamics where, because, you know, it's like, how to describe it, when you're trying to find somebody in a shopping mall, but the other person's also looking for you and you're looking for them. And so you almost never meet because they're always moving around in that same pattern. That's what we were doing. So what we ended up doing was we said, we're not going to change some parameters. We're just going to leave it fixed. But we are going to, you know, ration with different, like, signals and so forth. Maybe hopefully that way, SPEAKER_64: they wouldn't actually be continuously, like, we wouldn't be wandering around trying to find it, SPEAKER_12: find each other. So Han, this is actually when I fell in love with what Corticol was doing, because according to at least NPR's reporting of this, you guys provided kind of a positive SPEAKER_10: and negative response electrically to the neurons that were interacting with the chip. And you gave a burst of, I guess, unfortunate white noise to the device if it was the wrong action. And then a positive organized burst of electrical activity if it got it right. So essentially, you very politely shocked it if it did bad Pong and you rewarded it if it did good Pong? SPEAKER_44: Kind of. I mean, I wouldn't say it's not really a, you know, a reward or a shock kind of thing, because at the end of the day, we're still giving them electrical signal. It's just a structure of it. And this was actually driven by us coming across a theory developed by a neuroscientist based in UCL in London by the name of Carl Friston, which is actually really interesting, really fascinating, called the free energy principle. The free energy principle kind of posits is that the brain or all biologically intelligent systems are actually generative models. What we do in what was now believed is happening in our brains is that we're not reactive machines, we're actually predictive machines. And so we're actually generating hypotheses about the world that is external to the brain in the brain as simulations. And then we're using the census that we receive and the ability to SPEAKER_57: affect the world as mini experiments to prove and disprove the hypotheses that are generated by the SPEAKER_10: brain. Okay, so you guys did the Pong experiment and showed that you could, in fact, create a hybrid biologic and digital computer that could play Pong reasonably well. It wasn't going to set world records for Pong playing. Now, take me from there to the CL1, your first commercial device that kind of brings this technology out into the world. How hard was it to go from successful experiments to SPEAKER_12: commercialization and creating an actual shippable hardware product? SPEAKER_54: Oh, really hard. I would say this is probably one of the hardest things I've ever had to do because SPEAKER_44: there's some element of doing something really cool, something like really novel and so forth. But more importantly is that if you're doing it as a commercial offering, it has to work a lot, like at least most of the time rather than some random time. So it has to work very reliably. You do a lot of things that are very boring that no one really cares about, but it's important and vital for the product and the machine to work. So for instance, life support plumbing, you can't publish that. That's kind of boring. Everyone knows what you need to do. But to actually go through the weeds of developing, prototyping, and testing is quite difficult. So it did take a lot of work, but what was actually the motivation for it was that we had a lot of our colleagues reach out to us after we had published the work and they were asking us, where can they buy this machine or not this machine, but the machine that we had used, you know, what software did we had to write, whether we could help them if there were any questions. And we, you know, we saw this recurring pattern. I said, look, you know, people shouldn't be continuously reinventing the wheel, right? Like a round wheel works really well and we know how to make it. Maybe you should focus on building the rest of the car and we can build the wheel. And then you just use that wheel. If you think about the explosion in AI, right, it had a very fertile ground because the groundwork had been laid SPEAKER_57: by decades of gamers buying GPUs, keeping the likes of AMD and Nvidia alive until the AI systems came online, SPEAKER_69: right? The venture capitalists, you're welcome for all my purchases to play games. SPEAKER_51: Exactly. So I think this is a, you know, something that we, you know, we saw that and we're like, well, you know, somebody is going to take that leap and try to support the research community. And we decided SPEAKER_44: that that was going to be our thing going forward. Mind you, we still do a lot of research work, you know, and we have quite a few publications that came out. We have some very interesting stuff that's happening in the lab at the moment. But we think that this is more important that no one company can do this by ourselves. And we need to actually expand this out to everyone who's interested and, and try to, to reduce the barrier of entry for people to get into the space. SPEAKER_10: Okay. I want to talk about commercial applications in a second, but first I want to give people a quick look at this. So if you're on the video version of the podcast, you'll see what I'm about to show you. Uh, first of all, Han, what, what is, what is this here? I believe this is from the Pong era. SPEAKER_12: It's a Petri dish with what appears to be a small dish in the middle. So what is this? SPEAKER_44: Yeah. So this is what I was talking about. The multi-electric array it's, um, got a well. So that round thing that, that cylinder there is used to contain the, um, what we call the cell culture media. So the cell culture media is that liquid that bathes the neurons to keep them alive. It's very similar to, I guess, uh, cerebral spinal fluid, you know, that bathed, uh, our own brain, um, as these neurons are processing, they're consuming glucose. Um, but at the same time, they're also producing lactic acid, just like running, right? You end up with the, with the cramps. Um, so, you know, one thing to be, we try to do is prevent it from getting too acidic. SPEAKER_78: And we have to use, uh, the CO2 as a, as a buffer for that. SPEAKER_56: Han, I'm just imagining now, like, you know, the dog didn't eat my homework. Instead, my computer got the cramps, so I couldn't, you know, it couldn't do anymore. SPEAKER_10: Yes. I want to connect this to this image right here. This, I believe, shows, uh, a closer up of the neurons actually spread across, uh, a similar, uh, electric interface. SPEAKER_79: Yeah. Yeah. So, um, if we, if we, um, remember that image of the multi-electric array with that SPEAKER_44: cylinder thing in the center of a cylinder, there is actually, um, a CCD or CMOS sensor. This is actually the same sensor that you have in your cameras. Um, and each of those square grids are actually a sensor that captures the electrical productivity. It's also somewhat SPEAKER_78: photoreceptive. So that's why we have to keep them in a dark space. SPEAKER_50: We all understand the importance of a crisp, memorable, easy to spell domain name. One of those names you can say over the phone and people know how to type it in without asking you the spelling, but let's get real. The good ones are either taken or there's some poacher who's holding it and waiting for some huge payday and they don't reply to you. If you, even if you want to pay for a premium domain, you don't want to use up all your runway on a domain name. That's just the truth for a startup. You want to put that valuable cash back into your startup's operations. So you should consider this dot tech domain. You can get a clean, crisp, super memorable name for your website and company and signal out loud to your customers and investors, we're a tech company that's instant branding for you. That's why over 500,000 founders have collectively raised over 5 billion in investment, building their companies on dot tech. So skip the hassle, head to www.get.tech.twist or go to your favorite registrar and grab your dot tech domain today. And so you'll see things like, you know, SPEAKER_57: the large long piece that's going across, that's an axon and all the stringy bits are the dendrites. SPEAKER_44: So if you think about that as just one segment and 50, that's microns. That's actually the width of a human hair, 50 microns. And so, yeah, that's how small it is. And the level of connectivity, SPEAKER_87: the level of self-organization as well is actually pretty phenomenal. SPEAKER_89: It's so, this is, this image just, I've actually stared at this for probably 20 minutes, SPEAKER_12: just straight because it's such a, it's literally to me like 1980s science fiction, but just actually real and now turning into a commercial product. Like this is so cool. And I just want to show everyone, the culmination of all this work that you guys have done at Cortical Labs is this bad boy. This is the CL1. If you're watching, sorry, if you're listening to the audio version, imagine a very long toaster with a clear plastic top and lots of tubes. And I believe this bad boy can keep neurons alive for SPEAKER_91: six full months on going back to your point about the, the plumbing for life support. Yeah. SPEAKER_87: We say six months because it's, it's just easier to explain it that way, but actually you can keep, SPEAKER_44: you know, and we, as a lab and other people have kept them alive for, for, you know, months, even years kind of thing. When you, when you're talking about trying to do large scale experiments, you need a lot of samples. You need a lot of units to get, you know, statistical significance in power through a, through a large end. We wanted that scale. We wanted that longitudinal ability to, to study the neurons, but we also didn't want to have to spend a lot of effort doing it. So, um, a lot of, a bit of engineering went into a bit of plumbing, uh, to, to, to build this system that encapsulates the neurons in a, in a sealed environment, which is really important because these neurons don't have an immune system. And so, you know, if you had COVID and you coughed into it, SPEAKER_89: they would actually just also get COVID. So your computer could literally catch the cold. SPEAKER_78: That's really funny. Yes. Yeah. Pretty much. Um, you know, don't eat bread next to it because SPEAKER_44: you'll get the yeast and mold that goes into it and they'll kind of kill the, the, the cells, but we also built in, you know, our own neural interfacing system. So there's actually, you don't really see it at the base of the CL1 has a, a compute unit. So we have, uh, you know, FPGA, uh, and, um, what do you call it? Uh, general purpose, uh, compute units in there as well. That's where we, we digitize the signal. Um, we, we then, uh, put it through, uh, our very tight processing loops and, you know, we're developing APIs and SDKs for this so that, you know, as a, as a programmer, you don't have to go deep into getting these things to stimulate the neurons, um, with high precision, but, you know, a lot of code you can just use, um, our API and we've developed a DSL domain specific language to express, how do you want to stimulate these neurons at individual SPEAKER_56: channels at low latency loops? That, that brings me to kind of, I think the question on everyone's mind who's listening to us right now, which is, okay, this is awesome. You've made a, a digital SPEAKER_10: biologic brain. What's it good for? And I, I think that this is actually where my, my knowledge really kind of runs short because I'm not sure what, uh, compute loads, what experiment types, what commercial applications, the CL1 and its successors, uh, are going to be best at. And so SPEAKER_98: when we think about the market that Cortical Labs is going after, um, how big is it? SPEAKER_57: Uh, I think the two most interesting things that we've learned along the way about the system here SPEAKER_44: is that they do, uh, two things really well. Well, actually three things, but that third thing is attached to the second one. Firstly, they use significantly less energy than your traditional compute. Um, they're not a von Neumann based architecture. Uh, the, the energy is derived from the glucose in the actual cell culture media. So doing a back on the envelope calculation, we came across some really phenomenal numbers. So for our original upon experiment, we grew about 800,000 to a million neurons. Um, and, uh, we did a very rough estimate where we just said, what's the glucose content in the media? How often are we changing the media? So therefore, every change in the year, you know, X number of times a week. And there's this much of glucose and they're combusting all of it, assuming because they never combust 100%, you know, how much energy were they using? It turned out it was actually 10 to the minus four watts of energy. So that's 0.0001 watt of energy. Uh, it's probably even less than that because that's a massive SPEAKER_56: overestimation. But, but still, even with massive overestimation, it's effectively zero. And glucose, last time I checked is, uh, is sugar water. So it's nice and cheap. Exactly. I mean, SPEAKER_57: if you think about it, you and I, uh, as walking reference, you use only 20 watts in our brain. SPEAKER_10: The guys at Xtropic were telling me this, they're like, you know, we have this amazing probabilistic computer that runs on 20 watts and it's literally inside your head. And they're working on thermodynamic computing, which is different, but I've been thinking about that ever since they brought it up. Okay. So clearly this is not going to be used right now to train the next major LLM, but, um, the people who are purchasing the CL ones that will come out sometime next year. What, what do they want to do with it? I think the thing you understand here SPEAKER_44: as well is that LLMs, um, could not exist 20 years ago, right? Uh, even if we had the compute available, we could not make it work because LLMs required large data sets. I mean, they're called large for a reason, right? We, we, we've gotten to this point of what is it? 40 years. What are we now? 20, 25, let's say 40 years of, of the public internet being this repository of, you know, free language training data. There are lots of other domains that are not language based, that do not SPEAKER_57: have large data sets that are publicly available. And this is the second point that we, we've, uh, discovered along the way is that if you do a hit to hit comparison, and there's a paper coming out, SPEAKER_44: actually, this was in Europe. It's about two years ago where we put it in the RL workshop. It's one of the posters. Uh, if you took reinforcement learning agents, so we, we did an experiment where we took three of them, uh, DQN, which is your classic alpha go. It's kind of, you know, the, I guess the gold standard, but it's kind of old now. There are better algorithms out there that are more what's so-called sample efficient. And we put that against the biological SPEAKER_57: system and we gather the same game, but we said, here's the one thing we're going to accept. We're SPEAKER_44: going to constrain the reinforcement learning agents so that they get the same amount of data. That the biological system gets right. Because if you think about reinforcement learning systems, nobody really talks about this, but the way they actually learn in a simulation is that they spawn millions of parallel, uh, processes and they speed up the time of the game by, you know, 200, 300 fold. And so if you read the, the alpha star paper or something like that, uh, deed mind kind of said that, you know, if you were to train a thing or human to play at the amount of data that alpha star was playing, it would take 400 years of continuous gameplay to get to that level of performance. And so what we did was we said, we're going to sample constraint it because if you think about this, um, the real world, you can't speed up time. Time ticks at the same rate that you and I have. Um, and if data is a factor of time, right. It's a bit rate and so forth. Then if you wanted to get, you know, say 10 minutes of data, you would have to spend 10 minutes collecting that data. Knowing that, you know, these systems use far less data, uh, when, when we did the comparison with the reinforcement agents, um, and you know, data being a factor of time, we realized these things are actually potentially very good for problem sets that don't problems that don't have a data set, problems that are real time specific. And that is actually the vast majority of the domains SPEAKER_105: that have not actually been touched by language. This is when I get really excited and I don't, SPEAKER_70: I don't want to make you project the future too much because that's not really fair, SPEAKER_10: but to me with, with neurons that we can interact with the software, it seems like we've managed to take care of the best of both worlds, the programmability of, of computers as we understand them. And also the probabilistic low power consumption and kind of generalized intelligence of brains and brought them together on. So to me in time, it feels like we should be able to use both, you know, sure our mass GPU clusters that everyone's building, but also we should have probably large tanks of biologic, you know, neurons that are also helping us do a lot of, a lot of work. Is that too kind of science fictiony or am I on the right path? SPEAKER_108: We're all familiar with AWS, Amazon Web Services. That's the cloud platform that powers so many of your favorite brands. But do you know about AWS Activate? That's their program for startups where they provide up to a hundred thousand dollars in AWS credits for all startups. Whether you're backed SPEAKER_50: by an investor or you're bootstrapping, money is time to keep innovating, time to delight customers and time for you to keep gaining traction. You need runway and AWS's Activate is going to help you with that runway. We hear this story from so many of our founding university companies. They're finding product market fit, the words getting out about their product, just a little boost, finding some savings here or there to get a little extra time can be the difference between getting traction and bringing in revenue or, hey, let's call it what it is, running out of money, okay? And shutting down. AWS knows this. That's why they've created the ultimate toolkit for early stage startups looking to boost growth. With AWS Activate, you're going to get up to a hundred thousand dollars in AWS credits, hands-on support and training, plus exclusive discounts with some of our favorite companies and tools. So start getting the support you need at every stage of your startup journey. To learn more, visit aws.amazon.com slash startups slash credits. That's right. aws.amazon.com slash startups slash SPEAKER_57: credits. I think you're on the right path. I mean, the thing about it is that it's we're good at some things that machines aren't and machines are good at things that we aren't, right? Right. So this is SPEAKER_44: called more of X paradox, particularly in robotics. And I think this is the motivating factor for why, you know, Elon started up things like Neuralink, right? And the BCI people, which is that, you know, as AI gets smarter and better, we also need to skill up, right? They're getting, so if you think about it, the machines are getting better at what we reserve as traditionally human or biologically centered tasks. It's not beyond the realms of possibility that we could also go the other way, assuming that BCI has really become a thing, where not only do we utilize our brains for the load-powered stochastic computation, the probabilistic computation. But we can also just, SPEAKER_78: you know, have a, I don't know, an iPhone chip in there that can just give us the square root SPEAKER_10: number at the speed of thought. Right. So I can essentially allow compute to handle the math that I don't want to do myself, but all the learning I can use my brain for. So, correct. Once again, you guys are kind of Neuralink in reverse. So look ahead of just a couple of years, not super long term, but the CL1 gets into the market next year, customers, you guys are also going to have an API to let people access the technology remotely if they don't have the lab set up necessary to run this thing. What's the next generation? Like, what are you guys going to SPEAKER_43: build after the CL1? What is the CL2? Well, I think the CL2 is, we've been thinking about it, SPEAKER_44: and there are a lot of features that didn't make it into the CL1 that we'll probably, you know, start thinking about putting into the CL2. I mean, the things, your standard things, right, make it like smaller, make it, you know, probably cheaper to obtain, easier to program, you know, maybe more modular units. There are quite a few things that, you know, we've been thinking about. Now, the question is, does it make it to CL2 or does it get pushed out to the CL3? Is, you know, a discussion for the team. But again, you know, I think it's, it's still a little bit premature. What we really want to do is get the CL1 out to our partners, our research collaborators, and really get their feedback books, what doesn't work. So that way, we can sort of triage and prioritize what we want to do for the CL2. No point of putting in features that no one else wants. I mean, we might want it, but you know, we have the ability to just spin up whatever CL3 for whatever internally. But you know, if there are features that are, that are missing, that we have on our list, that's kind of lower down our sort of priority. But it would, you know, tremendously help the community get going faster, or do more work for less, you know, I think that's something that we want to be SPEAKER_27: prioritizing. You talked about using mice neurons at the top of the show. And anyone who knows anything SPEAKER_12: about biology and experiments knows that a lot of a lot of mice and rats die. So that is kind of SPEAKER_10: par for the course for how we currently handle ethics. Are mice neurons and human neurons SPEAKER_67: radically different? Are they relatively fungible? Silly question, but I'm not sure about the answer. Before I get to an ethics question, I thought we'd start there. SPEAKER_57: So, and this is the really interesting thing. We, we used to think that there was no difference SPEAKER_44: between a human and a mouse neuron. And if you read any of the papers before, I think 2022, 2020, everyone said, yeah, there's no difference. All mammalian neurons are the same. And it turns out that I think on average, human dendrites are actually longer than mouse dendrites, but they have the same number of ion channels. So these are the gates that allow the potassium and sodium to go back and forth between the cell and the external environment. And that's what generates the actual potential. But having them spaced out more, you actually have the ability to hold more so-called electrical states. Well, that's the theory that's been pushed by some of the researchers at Harvard, MIT. And so, um, yeah, it turns out there are different and that probably contributes to the differential in the overall performance. And we've seen that as well. The human neurons are better at SPEAKER_56: playing and processing information. But, but for now, but for now though, Han, the, the mice neurons are sufficient for the state of technology and the progress you guys want to make. SPEAKER_10: So we're not, there's going to eventually become a new story entitled startups takes human neurons, puts them into computer. Are we creating a new form of consciousness? Oh, oh, panic. But it sounds like. SPEAKER_44: So, well, we actually do use human neurons now alongside mouse neurons. Human neurons is the other really big area of research, because that's where a lot of the biomedical stuff happens, where we're looking for new drugs, um, you know, understanding disease models and so forth. We stand at the shoulder of giants. And so we use a lot of the techniques that have been developed, um, by the, um, uh, what you call the synthetic biology, the biological, uh, engineering, uh, field to grow stem cells obtained from, uh, adult cells. So, um, if we went back about 10 years ago, um, in Japan, there was a researcher by the name of Yamakana. Professor Yamakana won the Nobel Prize, I think a few years back, uh, for this discovery of, uh, inducible pluripotent stem cells. And what he discovered was that if you took anybody's cells, your cells, my cells, you know, skin or blood, anything with a nuclei, um, and you expose them to four, uh, compounds, um, you can actually reverse the clock back of these cells where they go from a blood cell or a skin cell back into a naive state stem cell that can then become turned into anything. We, um, utilize that functionality. Yeah. To, to, to essentially grow human neurons, uh, from stem cells. So this way we don't kill any animals. It's continuously renewable, uh, as long as you provide the right conditions. And yeah, SPEAKER_43: that is, uh, that's how we do it now at Cortical House, because, you know, we, we don't really particularly like killing the mice as well. And so we think that this is actually SPEAKER_12: an easy approach. It's absolutely brutal. No, I'm totally in favor of that. I think there's going to come in time, some, some questions that appear ethically interesting with what you guys are building, but I'm also of the opinion that they're not actually going to be ethically serious. I think SPEAKER_10: they're going to be mostly good for not making fun of the media, my industry, but like great for a headline, but less dicey in practice. So I'm, I'm very bullish on the company. Um, just before I let you go, uh, how has a demand been for the CL one? I know you guys have a form up. I know you guys are working on deliveries. Um, has demand exceeded expectations? Is it about what you thought? Um, SPEAKER_129: just where's the business side of things going? Yeah, demand has actually far exceeded what we, SPEAKER_44: we, we had expected. Um, uh, suffice to say we are now struggling to try to figure out our logistics, um, and supply chain for, to balance supply and demand. Um, you know, we, we've had a lot of signups, particularly for the cortical cloud. Um, I think over 3000 signups, we have no ability to service more than 20 or 30 people on the, on the cloud system. Um, we have a hundred X more demand. We have a hundred X. Yeah. And so they're on the waiting list and so forth. And then, you know, for, for the hardware, uh, we have had a lot of labs, uh, and research groups reach out to us. Um, we are, you know, very, uh, excited about this, but now we're looking at the challenger of, oh God, now we actually have to ship this thing and we're gonna have to ship lots of them. And, you know, if they break in the field, we're kind of screwed. So, you know, we better make sure they don't break out there. Um, so there's a lot of like, we're, we're doing our best at the moment to a, uh, finalize, uh, all of the software finalize the hardware, you know, getting our partners ready. It's also very, very expensive CapEx wise, right? Because we have to, we, we don't charge our customers until we ship them, but we still have to build the units first. And so, you know, trying to figure out where the next source of funding is going to come from is going to be a bit of a challenge. SPEAKER_91: Han, why don't just charge them half up front and resolve some of your cashflow issues? SPEAKER_51: Well, we, we could, but you know, it's, um, for us, yeah, we could do that. But, you know, SPEAKER_44: uh, I, I think for us, we, we want to make sure that, you know, we, we, um, I don't want to do a Kickstarter kind of thing, right? And at the start of the, SPEAKER_89: I don't think that's, I don't think that's Kickstarter, but I, I respect where you're at. SPEAKER_12: I guess then the correct question to close with is this, uh, sometimes people say that venture capital is a little bit conservative. I think the VC is that bet on cortical were clearly being relatively adventurous, which is good. Um, so as you guys do need more, uh, capital is the venture capital ecosystem in Australia, uh, Europe and the West at large, is it interested in, in putting more money into the business? Um, or are you guys still a little bit outside of what you might call adventure SPEAKER_71: norms? Yeah, I think we're outside of venture norms. Uh, I don't know what it is, but the venture SPEAKER_44: industry, um, likes to gravitate to thematics. So in this case, it's been, what is it? We went from LL, transformer LLMs to agents and something, something wherever next year brings, um, it's very hard. Uh, I think unless you fall into specific thematics to attract that. So, you know, I think we have some of the best investors in the world backing us because they're just the ones who, um, are willing SPEAKER_78: to take bets, which is essentially what the industry initially started out with. Right. SPEAKER_44: Blackbird ventures and horizons ventures. Correct. Yeah. Um, and you know, we've had, you know, uh, a new round that was just put together to get us a little bit more fuel to, to deliver our products by kind of people like three CAGI. They've just come into the game. And then, you know, we've had players like in Qtel as well, who, who have backed us. It's, it's really about, I think, trying to figure out who, who are the Mavericks in the space, who are the ones who, who have the imagination, right. And the sci-fi background, because at the end of the day, SPEAKER_57: if you have 200, like maybe not 200, maybe like say even 10 fund foundational model companies, SPEAKER_78: but they're all pretty much, I couldn't really tell you the difference between Claude versus Gemini SPEAKER_44: versus ChatGPT versus, I don't know, whatever else, like Quinn or Geepseek, they are mostly the same. And so if they're all mostly the same and I can jump between one another, I don't really see any, you know, uh, significant, uh, stickiness or moat to it. So anyway, that's a problem the VCs do to resolve, but, um, you know, rather than all jumping in into one space, I think we should try to spread it out. So for us, you know, we're very focused on, on delivering the product. And, you know, uh, we welcome anyone who's interested, who's listening to your podcast, who understands the fundamental nature of building hardware first, before you can get the software, uh, you know, to, to, to look into the space, because the, the thing, as we really talked about the AI space already had that groundwork put in by the gamers, right? You didn't, we didn't need to invest in the hardware because the gamers just bought it already. So I think, uh, that's something that we have to bear in mind, um, with any new compute space, uh, there is always going to be a hardware component SPEAKER_104: and we, we, we shouldn't shy away from having to, uh, invest in and build into that space. SPEAKER_56: Well, I do, I do look forward to the eventual future when we have Canva in Sydney and we have Cortical in Melbourne, and we'll have a good old fashioned competition about who's going to be the next future leader of Australian technology. Han, thank you so much for coming on. Thank you for SPEAKER_10: answering all of my questions. And, uh, when you do have so much extra production capacity, I'll give you my address to ship my CL1 and you just tell me where to send the check. Okay. SPEAKER_91: All right. Thanks, Alex. Thanks, Han. SPEAKER_12: When Meta invested in Scale AI, absorbing its CEO and co-founder in the process, startups that competed with Scale saw an absolutely huge opportunity. Now, mostly people focused on the SPEAKER_10: data labeling side of what Scale had done historically as the place that startups might actually capture the most market share, but Scale also offered AI evaluation tools to folks who build LLMs. So, startups that offer AI evaluation tools may also be in line to benefit from Scale aligning with Meta, a company that competes with other firms to build the next great AI model. Why work with Scale if Meta owns about half of it? One startup CEO, Turing's Jonathan Siddharth, told Reuters after Scale's partial exit to the social giant that leading AI labs now realize that neutrality is no longer optional amongst service providers. It's, quote, essential. So, to help us understand the LLM evaluation market, just how big it is and how data comes into play in 2025 to build those next models, please welcome to the show, it's Turing CEO, Jonathan Siddharth. Jonathan, Chamath Palihapitiya: hey, welcome to the show. Thank you, Alex, for having me. It's great to be here. SPEAKER_10: I'm very impressed. Also, we're both in our home offices today, and I think this just goes to show that remote work, not entirely dead, even though everyone seems to claim that it is. I'm glad that SPEAKER_146: you're here. So, starting for folks who are less aware of what LLM evaluation is, Jonathan, can you Chamath Palihapitiya: just give us the working definition from your side of the fence? Yeah, so it's really important to benchmark LLMs to figure out where they are good at, where are gaps in the models, so that you can generate data that could help the LLMs get better at those specific tasks. One big challenge, Alex, that I see today in the overall evaluation and benchmarking market, is that a lot of the evaluations are somewhat academic and somewhat synthetic and don't connect to real-world applications or real-world use. Ideally, you want AGI to progress in a way where the models get better at useful tasks. So, when we evaluate models, the three dimensions to look at are complexity, like are you evaluating them on really hard tasks? Real-world use, are you evaluating them on something that a human would actually care about? And third is diversity, you want like a wide breadth of test cases that you're evaluating the models on. It's such an exciting space, and I think of evaluation as step zero of data generation. That's why we do both evaluation and data generation. SPEAKER_12: Okay, so let's talk about the benchmarks, because there's been some commentary, I think Apple had a set of paper that came out and they said that, you know, one of the problems we have with a lot of SPEAKER_10: the benchmarks that everyone likes to trot out their new LLM and put them up against and have the charts is that there's data contamination issues, there's overfitting. So, based on what you just said about trying to solve real-world problems, and also the fact that we know that these benchmarks are getting a little bit, I don't know, dicey to use, is the standard way that companies announce how their LLMs perform an effective form of evaluation? Or is that mostly window dressing to get more social media hits, SPEAKER_67: because you're one point higher on one test or another? Chamath Palihapitiya: So, I'd say the labs care about the benchmarks, because it's good for bragging rights, it's good for recruiting, like who has the best model for coding, who has the best model for STEM, etc. So benchmarks serve some useful purpose, and you can argue that in today's $100 million sort of talent wars, or maybe it's closer to $300, $400 million talent wars, that those bragging rights help, because the best researchers ideally want to join, like the winning team, like who's already kind of close to being number one. But the labs care about two things, public benchmarks and private evals. And I would argue that private evals are actually more important, because then you're evaluating how well is your model doing relative to the competition on the prompt distribution that you care about? Meaning if you're Google, or if you're, if you're, or if you're meta, or if you're OpenAI or Anthropic, the type of queries you might get inside in a coding context might look different from what you might get on the phone in like a general chat assistant context, or what a human might put when you're searching through like a desktop, right? Sure. So you want to optimize for your own query distribution that your users are testing your model on. So that's why these private evals are helpful. So you need public benchmarks and private evals. The challenge with public benchmarks, Alex, is if you look at coding, for example, there's this, there is this good benchmark called Sweebench, which is created by this lab at Princeton. Last year, we went from 2% to 50% in Sweebench. This year, the best models are already north of 60%. We'll probably saturate Sweebench this year. And then where do you go? Right? Like we haven't, clearly we haven't automated all of software engineering yet. The benchmarks are getting saturated. Yeah. Can you double click on what you mean by saturated in that context? I think that's an important point. Yes. So if a model does 90% plus in on a benchmark, you've kind of aced the test to some degree. And researchers call that saturating a test because now the test no longer tells you whether your models are improving or not. Right? So the obvious answer is you have to create a harder benchmark where the models would have even more headroom to climb, right? Even more room to improve, which is why at Turing, we're actually creating a benchmark for coding that's even harder than Sweebench. And it gives us deep satisfaction SPEAKER_148: because we managed to have all the models start at zero. SPEAKER_12: Oh, okay. So everyone's failing 100% of this new test. Ah, okay. Let's talk about private evals, SPEAKER_10: because this is a core thing of what Turing offers to its customers. So I'm curious. Let's say that I'm open AI. I have a new model. I want you guys to take a look at it. You find a couple of places where it's not quite up to snuff, not quite where we expected it to be. So then do you turn around and say, Hey guys, here are the places where there are gaps and then help them solve those issues? Or do you guys just point out, here's the spot where, you know, you might want to do some more work? Chamath Palihapitiya: Uh, so, uh, we do both. We evaluate the model and we generate data to help the models improve. But let me take a step back, Alex, and, uh, let's look at this landscape, which is super interesting right now. So as these models, what's happening is, uh, as these models have gotten smarter and smarter, the data that's needed to advance the models has become increasingly harder to generate, right? The models are advancing in depth. They're improving in coding, STEM, reasoning, et cetera. They're advancing in breadth, in multimodality, multilinguality, multi-industry, SPEAKER_148: et cetera. And the models are becoming agentic, meaning the models can now execute complex Chamath Palihapitiya: multi-step workflows in a real world business context. Now with the models advancing like this, what's needed in a platform is, uh, you don't just need Iron Man, you need the Avengers. Frontier models need frontier data. Frontier data needs a team of frontier humans, right? So we have like PhDs from physics, chemistry, math, biology, expert Olympiad level coders, uh, people who are literally at the pinnacle of their fields. And sometimes you have to have them working together to break the model. For example, in physics, if you, if an expert is SPEAKER_148: evaluating the model in physics to test some theory in physics, you might want to, uh, a PhD in physics might ask the question to test the theory. A software engineer might need to build a simulation to test that theory. And then a data scientist might have to, uh, analyze the results of that simulation. So you need to daisy chain really smart humans together to generate data. SPEAKER_10: Actually, Jonathan, can I talk to you about that? Because what you're describing to me sounds like, um, very interesting tests, having people with different domain expertises work together to find places where the model might have a weakness or might not be able to answer something. Uh, and then you say that they're generating data. Is the data they're generating simply where the LLM in question, uh, misses or doesn't meet expectations? Because I thought this was more like, here's new information for the LLM to be trained on versus looking at a model that's already been trained. And then does that make sense? I feel like I may have had this backwards. SPEAKER_160: Uh, so, uh, that's a great question, Alex. Um, so the way these models are trained Chamath Palihapitiya: is there's a step called pre-training, where you basically feed the model gobs and gobs of data. Yeah. The war in pre-training is mostly over. Like almost all the models are trained on the same subset of the internet. Now the war, the battlefield is post-training and there's a new field called mid-training, which is, I can get into that later. But with post-training, what you do is, uh, you have these human experts create question answer pairs where the, the, a software engineer might ask a question like, uh, Hey, how do you create an app that connects, uh, dog walkers to dogs? And, uh, can you write it in, uh, both, um, uh, Swift and, uh, Kotlin for Android, right? And now the system has to generate the app. So that's an example of a supervised fine-tuning dataset where you give it a prompt and a completion and the model learns how to, how to, um, how to auto, how to respond to that prompt, right? So you need experts in different fields. Now, the reason this has gotten harder is two years ago, uh, a very low skilled contractor was capable of creating SPEAKER_148: tokens that could have advanced the model. Now, because the floor has gone up, now you need PhDs from Stanford, Berkeley, MIT to figure out where the models break. You first have to break the model. You have to ask a question that literally stumps the model. And then the humans create good question pairs of questions and answers that you then feed into the model for fine-tuning. And then the model learns. SPEAKER_165: Ah, it's the question and answer pairs that generates the data for the model. Okay. Now that makes good SPEAKER_10: sense to me. So it sounds though, like what you've done is back to your Avengers Iron Man point, uh, just got together like a, like a super set of nerds. And basically you're applying the smartest humans to find the flaws in the most advanced LLMs. What happens, Jonathan, when we don't have PhDs that can ask questions that the model can't answer? Because it seems to me like we've raised the bar, Chamath Palihapitiya: but we're raising it towards the ceiling. Yeah. Yeah. I mean, I've oversimplified it a little bit. Like, um, so I would say like, uh, so one example I gave is, uh, you give question answer pairs. Uh, there's another type of, uh, data that you generate for reinforcement learning, where you, the humans are, uh, generating questions and verifiers. They're not generating the solution. They're generating a way to verify whether the solution is correct or not, which you can do with certain fields like coding, math, hard sciences. You can do that, right? Um, write a program to sort, numbers in Python. You can write test cases to verify whether the, the program that you wrote is SPEAKER_152: correct or not. But you're not telling them every single step along the way. You're just saying, I can verify that what you've done, the work you've done does generate an acceptable answer. SPEAKER_172: Correct. Correct. And, and, um, the version of this for enterprise is you create an RL gym, Chamath Palihapitiya: a reinforcement learning gym. I think this is one of the coolest things in computer science. Like just like a gym where humans go to train, this is like a gym where an agent trains. And in these gyms, you create basically clones of different websites, like a Doe Dash or an Uber Eats or, or a NetSuite or a Salesforce. You create this virtual environment with prompts or workflows and ways to verify whether the task is completed. Now, the task could be, hey, salesperson, why don't you research everything you need to about this prospect, and then update Salesforce in this way. And as long as, when you create this virtual environment, you create these prompts and you create these verifiers. Now, an agent tries out different combinations. The agent is basically trying to use the tools in an appropriate way to actually complete that task. And it's a cool way to learn whether the agent is learning through trial and error, how to execute a complex multi-step workflow. You're not teaching the agent, go to LinkedIn first, go to ZoomInfo first, and then check out Salesforce for prior conversations. You're not teaching it that, but it's learning. It's learning just based on feedback on, okay, if I accomplished this, I get a reward. And then it learns how to operate in a way that maximizes future rewards. It's really cool. To your question of what happens if we run out of human intelligence, right? Fortunately, I think we are still a significant distance away from that. We've run out of internet data, but there's lots of intelligence trapped in the minds of humans that has yet to be transferred from human minds to machine minds. I would say, Alex, even the way, like, I love your show. I love what you and Jason have done with this show. But when you're thinking about how to interview somebody who's on the show, what type of questions to ask, what type of prep to do, I guarantee that knowledge is not distilled SPEAKER_148: into the models from OpenAI or Anthropic or Meta or Google. The only way that's going to happen is if we hire somebody like an Alex or a Jason to work on Turing to use our tools. And we didn't get into this, but we have tools that make it easy to keep the quality of the data high. And we have AIs that will assist you in generating this data so that we can make this scale. Pun not intended. SPEAKER_10: No, I was going to let that one slide. We don't allow more than two dad jokes per episode. So we have to save them, if that makes sense. So here's my question. I understand that scale being subsumed SPEAKER_12: by Meta has been great for the Turing business. You told me before the show you guys are at nine SPEAKER_10: figures of revenue and profitable and growing. And honestly, hell yeah, I love that. But if I'm open AI, why wouldn't I try to replicate what Turing has built inside of my domain? Because if there's one thing that large AI model companies have today, it's a lot of access to capital. And so as we've seen from Meta trying to buy literally every human who's touched an AI before, there's a lot of money to play with here. So are you guys just so good that open AI doesn't want to try to replicate what you've built? Or is there something else that I'm missing in that it's good to have a third party versus an internal group be kind of telling you where you might want to work on things? Yeah. I mean, Chamath Palihapitiya: that's a great question, Alex. Firstly, to advance towards ASI, you need three things to move in parallel. Actually, you need four things. You need research and algorithms. You need compute, you need data, and you need the application layer. Those are the four things, right? Research, compute, data and applications. Now the labs are really good at research, right? And kudos to OpenAI, Anthropic, Meta, Google, Apple, all of these companies for advancing the frontier forward. So they're spending their energy there. We could flip this question to also ask, why don't they build their own compute? And in some areas, the labs are investing in it. I mean, we've heard of Google having TPUs and other companies trying to do that. But Nvidia clearly has the edge today. And there are some good companies like Grok, Cerebris, et cetera, that are also building custom chips. So compute is also advancing. And the third pillar is the data pillar, right? Now, generating this data is exceedingly complex. Again, three dimensions. It has to be complex. It has to be realistic, meaning it has to mirror how a real human would interact with the model. And third, it has to be diverse. It's very hard SPEAKER_148: for a lab to get all of this in-house. I'll give you an example. Like today at Turing, this month alone, we are hiring PhDs across the board in physics, chemistry, math, biology, at the level of granularity Chamath Palihapitiya: of somebody who's an expert in dark matter, somebody who's an expert in black holes, somebody who's an expert in molecular biology, right? Like these very niche fields. And you ideally need these humans part-time because the way these humans are good at the job of training the models is because they are also good at their day job, which keeps their skills sharp. So you kind of, and once you've ingested that knowledge, you might want to move to the next frontier and the next frontier and the next frontier. So you need a platform or a partner that can scale up very quickly to elite talent, manage the talent to like, make sure the talent is generating data part-time, the data is high quality. We have to build a ton of tools. Like we have this platform called Allen, where AI is used upstream of the human to minimize the work for the human. AI works alongside the human and AI does quality control. It SPEAKER_181: may sound a little dystopian. It's like, uh, you know, AI overseeing humans to improve AI. SPEAKER_10: I'm not afraid of our AI future personally. So that doesn't, that my, my P doom, I guess, is, is very low. Uh, but I want to go back to your black holes, dark matter point, because I'm a bit of a science fiction nerd and I'm also a bit of a space nerd. So those are topics that are near and dear to my heart. But if I spend a lot of time helping a particular LLM better understand why we think about dark matter, why we came up with the idea, you know, galaxies and spinning and not having enough matter, blah, blah, blah, blah, blah. Will that work to make the LLM smarter in that particular domain have spillover effects to other areas? Does it make it more intelligent writ large or more intelligent only in that specific domain? I'm not sure the answer Chamath Palihapitiya: here, so I figured I'd ask. Uh, yeah. Phenomenal question. Um, what we know for, what we know with relatively high confidence is that when the models get better at coding and math, they seem to get better in a wide variety of other tasks that have nothing to do with coding and math. Even it, something about coding seems to teach the models how to think in a more structured way, how to communicate with less ambiguity, how do you think step by step. So coding and math, there have been some experiments in how they have out of domain performance outside of just pure code generation. And coding and math is also interesting because sometimes when you ask complex questions, Alex, the sub steps involve being able to compute stuff or calculate stuff and pass the results forward. For example, Alex, if you had a question like, uh, Hey, what are some interesting themes in investing in AI? The answer to that might involve a model that knows how to write code to query pitch book and write some Python code to analyze the data and say, you know, AI and healthcare seem to be spiking. AI and retail seems to be going down. That requires the model knowing how to write some Python, uh, code to analyze the data. It knows, um, uh, math plot lib to plot the results and show you the results. So coding and math have broader, um, broader, um, applicability, but it's also true that this is one of the mysteries of, uh, these language models. The models seem to learn some representations about the world that seem to carry over to other areas. Uh, I can totally imagine you understanding human psychology. If the models were really good in human psychology, it could help the model write better website code that designs them, designs the website to be more persuasive, maybe like more conversion optimized, maybe makes it better for generating copy. So it may be, who knows? Maybe there is something to learn from studying the universe SPEAKER_148: that helps people in their day to day. It would, it could be like some indirect thing, but we don't SPEAKER_10: know for sure. So Jonathan, um, here's to a next great couple of years as the world gets, uh, bigger and faster and smarter. And, uh, just before I let you go, uh, where can people find Turing on the internet? And is there a particular job that you're having a hard time hiring for that you wanted to shout Chamath Palihapitiya: out to the world? Great. Uh, thank you, Alex. Uh, but if you're, um, if you're working on AGI research, um, or in human data, come talk to us like we are hiring across the board. Uh, and if you're building a frontier foundation model and you need data to make the model smarter at coding, reasoning, SPEAKER_174: STEM, et cetera, come talk to us. We're at turing.com and you can email me at, uh, Jonathan at turing.com. SPEAKER_10: Perfect. All right, Jonathan, thank you so much. We'll have you back on in another six months when AI is twice as smart and until then this is twist. Bye. All right. So here on twist, we have SPEAKER_191: talked ad nauseum about AI and the job market, mostly about how AI might impact the job market, SPEAKER_98: change who does what job change, what jobs are done. You get the idea. People have a lot of concerns that AI might take all the jobs. We'll have to see what happens today though. I want to talk to a company that wants to use AI to help people and not just get jobs, but get the right jobs and also to have companies find the right people, because as it turns out, and you know this, if you've done any hiring in your career, finding the right people pretty much is terrible, even with the modern tools we have today. So to help explain how this is going to work, please welcome to the show. It's Merkur CEO and co-founder Brendan foodie. Brendan. Hey, SPEAKER_194: how you doing? I'm doing great. It's awesome to be here and appreciate you having Alex. SPEAKER_12: It's my pleasure, man. I love doing these talking to people who are building what's actually next for SPEAKER_10: the economy literally never fails to give me a jolt, like a good espresso. So thank you. Now, one reason why I wanted to talk to your company is because you have this amazingly huge vision. You guys wrote that you founded the company because, and I quote, the labor market is the largest, most inefficient market in the world, which is a pretty big claim. So before we dive into how Merkur is going to approach this and how you're going to get into the market, tell me why you think that and what brought you to that as the problem you wanted to SPEAKER_199: solve. Yeah, it really comes down to this matching problem where when you were describing earlier, SPEAKER_200: the reason that it's painful is that when a candidate is applying to a job, they can only apply to a couple dozen jobs. And when a company is considering candidates in the market, they can only consider a fraction of a percent of the people available that are looking for work because they need to solve this matching problem manually. They need to manually review resumes, manually conduct interviews and manually decide who to hire. But when you're able to solve this matching problem at the cost of software, it makes way for this global unified labor market that every candidate applies to and every company hires from. And so that's the end vision of the company in the North Star that we work backwards from. Okay. So essentially, when I go out there to look for SPEAKER_10: jobs, I might look at my local area, I might look at jobs, companies that I've heard of, but no matter SPEAKER_12: how far I cast my net from my perspective, I'm still missing out on most of the gigs. And therefore, most of the companies aren't hearing from as broad a talent pool as they might. SPEAKER_203: Exactly. Yeah. SPEAKER_12: Okay. So in a world in which we're dealing with kind of the demise of remote work, isn't the labor market necessarily constrained by geography? SPEAKER_204: Well, not precisely. I don't think we're dealing with the demise of remote work. I just think that SPEAKER_200: in many ways, remote work might actually become even more prevalent. And part of this is that if you automate 90% of what it means to do remote knowledge work, that means that the bottleneck to productivity is the other 10%. And so for all the domains like software engineering, where demand is extremely elastic, we'll just produce 10 or 100 times more. And so I think that humans and the role that we play in knowledge work, both remotely as well as in person, is going to become amplified with these huge increases in productivity that are starting to happen. SPEAKER_12: So just gisting that down for folks, do you think that in time, remote work is not going to SPEAKER_10: go away? And therefore, there will be kind of a global talent pool. And therefore, if you want to access the best people, you're not only going to have to have a much larger pool, but you're also going SPEAKER_67: to look more broadly at the corporate level. This is certainly the case. And most importantly, SPEAKER_200: that I think it'll be centralized, right? Instead of all of these decentralized, fragmented job searches and companies that are looking to hire, I think it'll be one place that everyone goes, that every company hires from that can facilitate all of these job matches that people love and companies are SPEAKER_10: finding a lot of value in. Okay. And I think you want that one place, that central hub to be Merkur, your company. Yeah. And if you succeed, this means that recruiters, as we know them today, are going to have a hard time. Probably generic, un-AI enabled job boards like Indeed and similar SPEAKER_12: are all going to get hit. So that's the scale of what you guys are shooting for, essentially a revolution of how people find and secure jobs around the world. Exactly. That's fantastic. SPEAKER_200: Well, so one way of thinking about it is there's sort of these two ends of the spectrum. On one end of the spectrum, there's companies like LinkedIn or Indeed, their job boards, and they have this very broad distribution, but they only aggregate this very thin layer of the person's resume, right? And so they capture about one thousandth of the value chain associated with facilitating a hire. On the other hand of the spectrum, there's the services companies, the recruiting agencies, the staffing agencies that do all the manual work and heavy lifting to get their 30% of first year or whatever their fee structure is. Yeah. But no one has been able to combine the distribution of LinkedIn and Indeed with the value capture and value add of a recruiting and staffing firm, because previously it wasn't possible to automate services, right? But now that's all changing. Now it's becoming possible to automate everything that recruiters and staffing agencies would otherwise do in this unified way that has the same kind of distribution scale of a consumer platform SPEAKER_12: with hundreds of millions of people. All right. I wanted to start here because I want people to know where you're going to, but I now want to go the other direction and talk about what you're doing now, SPEAKER_10: because you guys have this really interesting, quote, public secret plan, if you will, of using a helping AI companies find the right talent that they need as essentially your wedge into the market. And to help explain this to folks, you guys want to learn how to place candidates very effectively in this one area where there's a short feedback loop, if you will. So you can learn pretty quickly SPEAKER_12: and then apply that more broadly over time. But I think you've been pigeonholed by some in the press is like helping people hire for AI when that does seem to be kind of a small fraction of the company's SPEAKER_10: vision. So what I want to know is how does it work today? I know you guys are working on one subset of the market, but walk me through from the candidate and the company side how Merkur actually does do SPEAKER_200: what you're describing. So a candidate will come to us looking for opportunities for work. They'll see a lot of the jobs available or they can apply to one of those or a talent pool more broadly. They'll upload their resume and they'll take an interview with the AI on our platform to then get evaluated to see what they're a good fit for, what they're interested in, what kinds of matches we might be able to create for them in facilitating huge volumes, tens of thousands of these matches with no human process and involvement. On the other side of the marketplace, there's companies that will give us these requests of saying they need 100 people in a particular domain and maybe a subset of software engineering or investment bankers or whatever the professional domain is. And we will go out to the supply side of our marketplace, find all the people that are a good fit for them and facilitate all of SPEAKER_12: those contract opportunities. Tell me more about the the AI interview. And I'll just say that I'm an AI bull. I think this stuff is really cool and going to work well. On the other hand, I do have some friends that are in less prestigious jobs who have had to deal with some AI onboarding, some AI screenings, SPEAKER_10: and they've been pretty negative about it. So I'm curious what you guys have built and how well it works and kind of what candidates think of that part of the process. SPEAKER_200: Yeah. So we built the first AI interviewer in March of 2023 when I was in my college dorm room. And initially, I remember it would hallucinate like like nothing else because it wasn't even GPT-4 at that point, right? It was check GPT had 10 second latency. And it's been extraordinary to just see this tailwind of model improvement, you know, lift all boats and make these applications possible. And so a good heuristic for it is emulating all the processes that a human would otherwise do, where similar to how a human would review resumes and conduct interviews, we automate all of the preparation for the interview and conducting the interview and evaluation of an interview with that similar format and heuristic. And while I think some candidates obviously prefer to talk to human because they think about it more as a selling process from the company rather than strictly a buying process, the overwhelming sentiment is that there's millions of people that apply to jobs and just get completely ghosted and don't even get the opportunity to interview, right? It'll be like less than 1% of people even get to talk with someone or show more than their resume for the most competitive jobs. And so globalizing that ability and not just like internationally, but also all across the U.S. for people to, you know, talk with the candidates and actually consider everyone for the opportunity has been really impactful. SPEAKER_96: Does the AI interviewer you guys have built in its current iteration, does it work with multiple languages? SPEAKER_12: Because I presume that if you're only doing English, you're still constraining yourself to a small portion of the overall possible job takers. SPEAKER_200: Yeah, we started out predominantly with English. Now we're starting to do others, working with a lot of our customers to try to make sure the models improve their multilingual capabilities so that the interviewers well set up for that. SPEAKER_12: And I presume that India is a market of choice for you guys, but are there any other hotspots around the world where you're seeing a lot of like talent saying, hey, we really want to be part of this SPEAKER_141: new marketplace that you're building? SPEAKER_200: Yeah, well, so actually the majority, the vast majority of our hires come from the U.S. now, over 60%. India is the second largest geography that we hire from, but also seeing lots of Eastern Europe, lots of South America. One thing to note is that the average pay rate in our marketplace is over $90 an hour. And so it's very different from most labor marketplaces in that way. It's a totally different league from what you would find with the crowdsourcing platforms like Scale and Surge that sort of built these legacy labor marketplaces with more in the range of $10 to $30 an hour pay rates. SPEAKER_12: But you guys are currently targeting a much more educated and rarefied employee pool, again, as your starting point, as your wedge. So I want to talk about how good your system is at finding and placing people because it sounds good in theory, but in practice, of course, that's what matters. So what's the right metric that you guys track in terms of finding the right person for the right job and then having both sides of that equation be happy? Is it repeat business? Is it successful, you know, completion of a contract? What's the KPI there? SPEAKER_200: Yeah, I think about it as a two-sided retention problem. And there's sort of leading indicators on each of those retention problems. So the supply side retention or applicant side retention is, of course, if they come back looking for work. And so we have models that predict what is the probability that they're going to be interested in particular job opportunity that we present them with based on all their prior experience so that we can ensure, you know, we're retaining candidates very well. And then on the demand side, it's how well are those candidates performing? Where we collect all of the performance reviews of who's doing well for what reasons and use that as our eval set, as our benchmarks to go internally on how do we predict based on the interviews, the kinds of people that are going to perform well, that are going to translate to, you know, a lot of value for our customers. And that's why our demand side retention is so large. So now we're working with six out of the mag seven, we have 1,605% net revenue retention on an annual basis. And so it's sort of SPEAKER_12: nuts. Wait, 1,605%, so it's 1,600%, so 16x net revenue retention. For folks out there who don't know why that's funny, mature software businesses will often turn in a net revenue retention number of 115% or plus 15. So the number that you put out there is humorously large and implies that customers are buying not just a little bit more of the services, but 16 times as much over a one-year SPEAKER_194: basis. Yeah, the two-side retention problem is working well. Yeah. No, two-side retention SPEAKER_10: solution. I don't think problem is the right word. Okay. So to me, AI companies that need to hire SPEAKER_12: experts is a very, it's a market that has a lot of reason to do well, to invest in new ideas, to try to find the right people, because they're in a massive race to build the best stuff and therefore capture the SPEAKER_10: market. How well do you think the Merkur model is going to translate when you go wider and you're dealing with people that might be less individually, resume-wise impressive and are more average folks applying for more average jobs? Because to me, difficult job, difficult person might actually be SPEAKER_146: easier to match than finding the right person from a more general pool for a more general gig. SPEAKER_200: Yeah. Well, we do hire a lot of people from general backgrounds as well. It's really, really broad across literally every industry and the economy. But maybe stepping back a little bit, the background of the company is initially we were automating the processes of hiring people for our friends. And then Scale.ai came to us and they used our platform to hire over a thousand people. And what we realized was that there was this huge transition in the market away from the crowdsourcing paradigm of low and medium skilled talent towards this medium and high skilled sourcing and vetting problem, not only with higher caliber people, but also people that work directly with the AI researchers to help them interpret evals and push the frontier of model capabilities. And that is one of the core reasons that the performance data we collect and all the things we learn and how we facilitate better matches is actually much more similar to a general work environment than most people realize on face. SPEAKER_12: When you say, Brendan, that you're collecting the evals, my impression of how this worked was I would come to you, resume, interview, job, and I would go work for someone else for some period of time. But if you're collecting the evals to ingest back into your process, does that mean that when I land a contract or a gig from Merkur that I'm actually working for you guys? SPEAKER_200: So technically, the work product that our contractors are producing ends up going to the customers. What we own is just the performance review on like, did this person do a good job on the project? SPEAKER_239: So the end customer tells you how the employee did and that's part of the overall arrangement. So the feedback does make it even if they're working for someone else. SPEAKER_200: Yeah, yeah. And one interesting thing is we facilitate the entire payment stack. So we learn from who's getting bonuses, who's getting raises, who's getting dismissed for what reasons, all as part of the data flywheel. SPEAKER_12: Okay, so people love to talk about the usefulness of data and how the more you have, the smarter you can be. So how much better have your systems become after ingesting increasing amounts of performance reviews, bonus information, contract retention and so forth, all the stuff we've talked about, how quickly does that filter back in and actually generate a material and measurable improvements in your ability to place candidates? SPEAKER_200: Yeah, pretty much on a weekly cadence. I mean, we went from where it's been most impactful and our focus has been is predicting people at the high end. And the reason is there's this dynamic on a project where if we're providing 100 people to work with a given customer, the top 10% of people are going to drive majority of the value. It's similar to like, if you have a team of 100 software engineers, probably the 10 core people are these like 10, this notion of 10X software engineers that are driving- Are you telling me that power laws exist? Power laws exist, right? And many, in many knowledge work verticals. But the implications of that from our standpoint are profound, because if we're able to build interviews and assessments and all this technology that can predict these power law outcomes, the amount of value that we drive for our customers and performance of those people and the SPEAKER_216: quality is extraordinary. Okay. So let's play devil's advocate here because I'm always on one hand, a capitalist technologist. And on the other hand, I'm a person who has friends who, SPEAKER_12: you know, have to make rent. So if you can help people find the 10X engineers and the 10X engineers only, like, let's say you can just say, look, here are the top five people and here's 45 others you can hire if you want. Aren't we going to end up with a labor market in which companies are able to hire the best and get more out of them and then they're going to need fewer total people? I'm just worried about folks who are like B students. Yeah. Well, so I sympathize with a concern, SPEAKER_200: but one thing I've come to appreciate is that people sometimes over index on the dimension of how exceptional someone is and under index on how relevant their experience is. And that if you match people with the right thing that they're extraordinary at, you can unlock these phenomenal SPEAKER_216: outcomes. All right. Now I want to talk about results because you guys raised, I think it was a SPEAKER_12: series B earlier this year, quite a large round, quite a large valuation. And I think TechCrunch said your revenue was somewhere in the realm of 75 million ARR. You guys have also talked about how you're growing, you know, 50% a month at one point in time. So one, since your last funny round, has growth stayed as hot as it had before? And then when do you think will be the right time SPEAKER_98: to use the wedge to expand your remit and try to take on a wider range of jobs and also longer SPEAKER_194: 10-year jobs? Yeah. So I'll answer the first part. The business has grown by an order of magnitude SPEAKER_200: since we received our term sheet for the series B. So the growth has been incredible. Happy investors, very happy investors. Well profitable the entire time. And so we have more cash in the bank than we've ever raised, which is unique for an AI company. And then to your second point of timing the expansion to new markets, the context for why we focus on the AI labs is we realize that there's more of a comparative advantage when we focus on hiring someone for five weeks versus five years. Like when we're hiring someone for five years, you want to get dinner with them, build a trust and a relationship. Five weeks, you want this fast, efficient AI interviews, automated process. And so we're leaning into that a lot now. In the wake of the scaled news with them no longer in the market. There's so much market pull that we're focused on capturing that. But what a lot of people don't realize is that throughout the duration of the business, we've still been doing lots of contract hiring outside of the market for AI labs and lots of full-time hiring for ourselves and our friends. We still have the customers from 2023 before we even started working with AI labs that are continuing to grow with us and so are starting to also ramp up investments and a lot of those other kinds of SPEAKER_216: hiring work. So it sounds a little bit like I asked a question that was more relevant maybe a year ago, SPEAKER_12: but instead you guys are expanding your work with AI labs and doing other things. So it sounds like SPEAKER_10: the wedge is wedging and you guys are already growing your remit. Okay. Now, before you go, two things. One, where can people find the company on the internet? And then two, is there a particular role that you're looking to hire for and want to shout it out into the broader world to see if the right SPEAKER_12: candidate is waiting for you? Which is ironic, I know, given the conversation. SPEAKER_200: Absolutely. People can go find us at Mercor.com, M-E-R-C-O-R.com. And we're hiring for a huge volume of roles, particularly lots of software engineers. So for any software engineers that are looking to join our team full-time, we're super eager to talk to you. For anyone looking to join our marketplace, whether it's part-time or full-time as a contractor, we pay exceptionally well. We have phenomenal satisfaction on the marketplace and retention as we talked about. And so would love the opportunity SPEAKER_12: to work with you. Do you use your own software as the place you source some of your own internal candidates? Of course. Absolutely. So you do like the taste of dog food. All right. Well, Brendan, thank you so much. We'll have you back on in, I don't know, probably six more months when you do eventually raise more money than I can ask you about that. But in the meantime, thank you. And we'll see SPEAKER_233: you soon. Yeah, for sure. See ya.