SPEAKER_00: Hey everybody, welcome back to Twist. This is Alex and we have a special interview for you today. I am an enormous fan of music. You may not know it, but I grew up playing classical and jazz trumpet throughout my youth and music has remained an absolute huge passion of mine throughout really my entire life. So when the AI revolution of the last couple of years came to the world of music, I was incredibly curious. Two companies have really caught our eye here on This Week in Startups, UDO, and of course its competitor, Suno. Today, we have David Ding, the co-founder and CEO SPEAKER_01: of UDO on the show to tell us what it's for, who's paying for it, and where AI-based music creation is SPEAKER_03: going. This Week in Startups is brought to you by .techdomains. Don't miss our Jam with Jcal contest. To apply and get more details, go to jamwithjcal.tech. Brought to you by .techdomains. LinkedIn ads. To redeem a $100 LinkedIn ad credit and launch your first campaign, go to LinkedIn.com slash This Week in Startups. And Brave. If you're building AI and search-based applications, train your models with the Brave Search API. Get started for free at brave.com slash Jason. SPEAKER_06: We're going to talk about AI and music. David, hi, how are you? And welcome to the show. SPEAKER_07: Hi, hello. I'm really excited to be here. So I want to start with some background stuff SPEAKER_00: because I know you were at DeepMind for a while, and UDO is a relatively young company. I think it was founded in 2023. So just give us, what was the moment in time in which you said, I have to leave where I am and go found this company because why? SPEAKER_11: Yeah, sure. So yeah, as you said, UDO was founded last year, November of 2023. And before that, I was a researcher at DeepMind. And throughout my entire childhood, I've always been interested in two things primarily. So one is technology, and the other is music. So as a kid, I always wanted to build computers that can simulate the way a human brain works. Maybe you can wire the neurons together and then try to model the brain. And then it turns out that when I went to college, this thing was starting to pick up traction. My first year of college, I took a machine learning course so that I can participate in this field. And my other passion growing up was music. So I played classical piano and played it for at least 10 years growing up before going to SPEAKER_13: college. And I always thought that it would be really, really cool if like a computer technology SPEAKER_11: could compose and make music. And so then fast forward to when I was working at DeepMind, generationally modeling really took off. You see technology like ChatGPT or DALI or Mid-Journey like Emerge that really revolutionized the way that computers can make art. And so at that point SPEAKER_13: in time, I was like, hmm, what happens if we apply the same technology that I've been learning how to build and apply to music to have a machine that can help people create music, ideate, and create songs. And so this is why we left DeepMind to create a company that produces a product to help SPEAKER_14: artists and songwriters turn their ideas into reality. So I want to go back to the point about SPEAKER_17: LLMs to image generation to music generation, because I mean, my day job is writing. So that's kind of what I know the best. And so to me, the idea of a large language model, taking in a lot of data SPEAKER_00: and then helping kind of do next word prediction, admittedly, it's more complicated than that, but I can really understand it. I kind of get how we can use LLMs to do image generation, but when we expand the work done to music, to me, I feel like I'm missing a link in how the technology actually functions. So without spilling any secret sauce, if you will, how does, you know, the AI models that I best understand end up creating tunes? Because it just seems to be like a real SPEAKER_13: stretch of what was possible, but it clearly works. Yeah. So similarly to how these large models learn how to produce images and texts, all models, they learn how to produce music by listening to lots of examples of music. So you listen to music, and it tries to synthesize the common elements across SPEAKER_11: music. So like a music theory, elements of music theory, like which chords follow which other chords, or how rhythm interacts with the overall structure of the song, as well as other elements. SPEAKER_13: What does it mean to be country music, which is rock music? Or how does a guitar string vibrate? And how does the sound of a piano echo and reverberate around the room? And finally, SPEAKER_11: how does this all interact with the recording technology? How do you turn this sound and turn it into stereo? And so this model, because it's trained on the final output music, it learns how to do everything. So from like the very fundamental music theory level, all the way to how the sound is recorded by the microphone. SPEAKER_00: Okay, so prepping for our chat today, I was playing around with you, Dio. By the way, I am now your most recent paying customer, shout out. And I decided to throw a curveball at your software. And I said, okay, look, I wanted to do a progressive metal song that sounds a little bit like Periphery, a band that I SPEAKER_25: love. But I'm like, look, let's do it in 6-8 time. Now, you're a classical trained pianist, I'm a SPEAKER_00: classical trained trumpet player. You and I know that when it comes to time signatures in the world of music, 6-8 is not very complicated, right? We're not doing like 11-8 or something crazy. You count sixes and then fives. This is pretty simple. And it kind of did it, but not perfectly. And I know this technology is still improving. So I'm not trying to be negative. But it is, is this going to a direction in which I could tell a service like you do like, I want to do a song, first half in 6-8, second half in 7-8. And I want to do a chord change from C to C major. And then like, how specific can we get? And then is that underpinned based on a very granular understanding of how music is put together? Or does the software better understand like, broader chunks of it versus like down to the SPEAKER_20: individual note level? Yeah, so this is definitely a direction that we do want to support, giving SPEAKER_13: users and musicians more ways of controlling the model like time signature, key, tempo, BPM, SPEAKER_11: or instrumentation, or even like dynamic levels, like, like start quiet, start swirling, and then SPEAKER_13: like, and then die down again. So this is something that we definitely want to support. A time signature is something that we do not support at the moment, because, because in our music, you know, when we train you all models, we do not teach it the concept of a time signature when we were annotating the data. However, a key signature, the key of the song is something that we do support. And this is SPEAKER_11: something that we did not support when we launched the model of version one, back in April, but it's something that we added in July, because we recognize that users want to be able to control the key. And so SPEAKER_13: then we annotated our data set to contain, oh, this is D major, this is C sharp minor. So that now when you go to SPEAKER_11: video, and if you specify a key, a minor, it will produce a song in that key. SPEAKER_00: Well, a minor is now the most famous key in I think all of music things to Mr. Kendrick Lamont. If you don't get that SPEAKER_32: reference, congrats for being offline for the last three months of music history. SPEAKER_36: Wow, this Jam with J Cal contest has been a blast. So far, I've had the opportunity to meet with four great founders from companies like CorePod, Ulama, Uptrends AI, and the Roam app, all because they all use dot tech domains. And we have room for one more. Do you want to come on the pod and tell me what you're building? Well, you only need two things to enter. You got to be a founder with under 2 million in funding. And you got to have one of those awesome dot tech domains. So head to Jam with J Cal dot tech, and tell me what you're building. And if you win, I will invite you on to this weekend startups. And you get to share your vision with me and the world. I'm working with dot tech domains because killer startups use them, you know, one x dot tech rabbit dot tech, so many others. And guess what? We use it too. That's right. That tech powers our founder Friday program. So tell me about your awesome dot tech domain and startup. Apply for the Jam with J Cal contest today at Jam with J Cal dot SPEAKER_17: tech. We're picking the final winner soon. Okay, so it sounds like what I did there was I asked Udo to do something that it doesn't do quite yet, which is probably why it got a little bit funky. But you said something interesting there, which is data annotation. And that I think is the thing that I was missing because it SPEAKER_00: sounds like you guys, um, the human label and like help it understand like, this is a rock drum beat in four four. So it does that create like a, a flag that then the software or the model can kind of go back to and like, and SPEAKER_43: like, point to and understand. SPEAKER_11: Yes, so by annotating the data in a training data set, you teach the model how to associate certain SPEAKER_13: descriptive words with musical elements. So then it sees three, four, the time signature, and it hears a song that's in three, four, and it knows, oh, three, four, it means like you have like, three, three beats. And then like the first beat is emphasized. When a user then asks the model to create three, four music, it can then like, hey, it's understanding of three, four and apply it to the composition of the song. It's kind of like when a human learns, um, if you never teach a human, oh, this is three, four, you can't ask the human, hey, create me three, four music. Uh, even if it can produce three, four music, it just doesn't know what three, four the words actually mean. SPEAKER_00: So it sounds like the data annotation then provides almost like a connective layer between music and the user's request and kind of helps natural language inputs translate to something the computer can understand as a, as a command prompt, essentially. SPEAKER_13: Yes, exactly. And we aim to improve our model by giving it more annotations to understand more elements of music so that, uh, the model can produce these elements, uh, upon command. SPEAKER_00: Okay. So I want to go back in time though, because I've been playing with UDO since, and this is a true story. Uh, one of my friends started sending us funny songs he made for us, uh, in the group chat and they were, they were whimsical things like Alex doesn't want to go to work tomorrow and like stuff like that. And I was like, okay, where is this coming from? And it was from you guys. And so I, I got a kind of early look at the software and I've made different songs and I've gotten to play with the new model some, but back in the beginning day, uh, when you were first getting like the, the point one version of this out one, how good or bad was it? And, and how, how easy was it to get from like proof of concept, if SPEAKER_32: you will, to something you were confident that people might want to actually use. SPEAKER_11: Yeah. So, uh, funny that you mentioned like the, uh, the first version of our model, uh, like the baby, uh, the SPEAKER_13: very, very baby version when we were still debugging our overall, uh, code base and training structure, we spent a couple of weeks trying to figure out, uh, why our model, uh, couldn't produce any lyrics. Uh, you, uh, you provide lyrics and the model just, uh, refuses to sing the lyrics, right? SPEAKER_11: And then, uh, we, uh, we spent a while looking at the model, like looking, analyzing like, uh, uh, different loss curves. And then, um, eventually we found, uh, the reason to be, uh, quite simple is that like when we were, uh, feeding the dataset to the model, there was some kind of bug that caused the lyrics to not appear. And so the model never saw the lyrics. And so therefore it couldn't SPEAKER_46: possibly know how to, uh, turn the lyrics into a song. SPEAKER_09: So essentially that they couldn't run the engine of lyrics because there were no words going in. SPEAKER_13: Exactly. Yeah. And so, uh, this really goes to show how, uh, the process is quite dependent. You had to like pay attention to detail and it's all about the, uh, the input data. And so we fixed the bug. And then after, after we fixed the bug, the model actually just kind of, uh, took off. SPEAKER_11: Uh, it, um, every week we saw improvements. The first week, it probably knows, um, like the broad SPEAKER_13: genres like rock versus jazz. As model training progressed, it started learning more, um, more specific keywords like energetic, what's hard rock, what is, um, uh, like smooth jazz. And also, the sound quality improved starting from something that sounds like very noisy to something that's SPEAKER_11: much more refined and more like what you would get from a studio. SPEAKER_00: Yeah, no, the, the actual fidelity is pretty good, um, in my experience. And one thing, as a, a fan of heavier music in general, there are certain heavy metal sub genres that depend a lot on orchestral additions that are mostly programmed. And so I'm familiar with like the current state of the art for studio music with, uh, added digital elements, if you will. And we're not that far off from this just with UDO's own creations offer. So that's very exciting, but it sounds like from the, the point of inception of the initial, like it works to go into market to 1.5 released in July is pretty quick. And that chart's going up and here's a model quality and fidelity and so forth. Do you think that trajectory continues for a long time? Or were there early winnings, SPEAKER_28: David, that let you improve faster than you might be able to now and in the future? SPEAKER_11: So obviously there is a point where you start from zero and then you get something. And so that's SPEAKER_13: the biggest, uh, delta. And as you observed, uh, our audio quality is actually, um, like pretty good. SPEAKER_11: Although there are still areas for improvement, which are, we are working towards, but the big focus going forward is additional controls for users. So giving people, um, like more SPEAKER_13: ways of controlling the music, like, uh, uh, maybe you want to provide, um, uh, like a guide, SPEAKER_11: like you had this, a lot like melodic line already, and you want the model to follow this melodic line SPEAKER_13: and, and add musical elements to it. Or maybe you have this like musical style, but you don't really know how to describe it using words. So how would you, um, take this musical style, synthesize it and, um, feed it as an example for the model to follow. And so we want to enable these additional controls because we recognize that music creation, the creator wants to have a lot of control over the SPEAKER_11: music because that it's their own creation, right? And so, um, uh, so that's the area that we really SPEAKER_00: want to focus on going forward. Okay. I want to, I want to do some demos in a little bit to show SPEAKER_17: people what we're talking about because you and I've used this of course a lot and they might not have, but one thing that I was thinking about is, is who this is for because I am an enormous music SPEAKER_00: fan. So to me, music is part of my day from kind of when I get up to when I go to bed, I'm either listening to audio books or music, right? And so to me, it's very personal, very important. And I know music theory and I love it. And it's, it's key to me, not everyone's like that. And so people have different music tastes, different consumption habits. And so I don't know, is, is UDO aimed for folks that want to create stuff for their own consumption? Is it more of a rough draft machine for artists as they explore new ideas? Um, is it a way to generate muzak for elevators? Uh, so I, SPEAKER_61: I guess kind of like, who is it, who do you think this is for now and in the future? SPEAKER_11: So we think that UDO is for, um, people who love music, people like, uh, like yourself, and also like, you know, uh, artists and songwriters who obviously love music as well. SPEAKER_13: We just want to create a tool to allow, to make music creation a lot easier than before, SPEAKER_11: kind of like other, uh, tools that have come, that came before in the past, like for example, like DAWs, sampling, drum machines. These are all innovations that, uh, turned, uh, something that was a SPEAKER_13: little bit harder before. And with the aid of new technology, uh, just making this creation process easier so that more people can participate in the creation process and that existing artists can leverage this to, um, try out ideas at a faster pace and, and come out with, uh, like music that incorporates these elements in creative ways that maybe even the creators of the technology never had SPEAKER_11: in mind. I think like one good example for this is auto-tune. Like when auto-tune came out, you know, a lot of people, uh, they had qualms about using it. It's like, oh, it's like cheapening SPEAKER_65: the experience. Qualms about using it. That's a polite way of saying it. Yes, but sorry, keep going. SPEAKER_13: Yeah. Like people were like, oh, this thing is like, uh, cheapening the, uh, the experience. SPEAKER_11: Like it allows like people who can't sing to, uh, to sing and that's a bad thing. But then like, you know, like, um, what really happened was like, you know, uh, like, um, really transformed, uh, the industry. Like people were using it and then people found ways of using it very, very creatively, like, you know, like bumping it up beyond the, like the spectrum, uh, and embracing SPEAKER_13: the auto-tune sound as, as a musical style. Right. And so, uh, we think that with these technologies, SPEAKER_11: it makes music creation easier and people will find like creative ways of using it. SPEAKER_67: Okay. So it sounds like for someone like me, a big music fan, I could use it to create fun things for SPEAKER_00: myself to listen to. If I was a musician, I can use it to expand ideas and give me new ideas, SPEAKER_17: but this doesn't replace, you know, I don't know my, my spouse's Spotify account at some point in time. SPEAKER_00: This is more like distinct acts of creation in the future versus passive consumption. SPEAKER_13: Exactly. So, um, I mean, uh, you might, you say that you play a trumpet, right? Uh, like your trumpet doesn't replace, uh, listening to like, uh, uh, like great trumpet players of the past, uh, on Spotify, right? Because you enjoy, uh, listening to music that other people create, but you also want the, SPEAKER_59: like the, the, the joy of creating music yourself. Yeah, no, I, I think that's right. And what I, SPEAKER_00: what I like about the idea behind taking modern AI techniques and applying them to music is it just allows a lot more people to do stuff. Uh, you know, David, like, uh, five years ago, people talked a lot about low code and no code. And there was this big chat about the democratization SPEAKER_61: of software development and that's kind of worked out, but I love the idea of more power to more SPEAKER_00: people. And this to me seems to fit into that. Now on, on the critical side, though, some musicians are worried that they're going to be replaced whole cloth or diminished in some way. I want to run my theory past you, which is that I don't think that's going to happen because the musicians that I love to listen to have their own very specific, sometimes experimental style that probably couldn't be replicated by even very intelligent models. So to me, this exists, if you will, side SPEAKER_19: by side with kind of how music is made today. I'm curious if that's your view as well. SPEAKER_11: Yeah. So we, uh, so that's definitely, uh, my view as well. Uh, we, we also, I believe that people SPEAKER_13: will continue making music the way, uh, they've always made music. And this is simply another tool in a toolkit that they can choose to use or they don't have to use it, but then it just, um, something additional, right? Like, uh, just because like electric guitar got invented, doesn't mean that the acoustic guitar got completely replaced, right? It just, just means that like SPEAKER_11: there's yet another instrument that you can add to your band. Yeah. Actually, I remember, um, SPEAKER_00: in my high school jazz band, we had a song, I think it was an old Buddy Rich tune and, um, it had a little bit of guitar by itself and our guitar player played electric. And one time he forgot to turn his guitar up. So we got to that part of the song and he played and no sound came out. And my, our director was like, well, you know, he was a trumpet player and he was making fun of the electric guitar for needing, you know, help essentially. And I was like, I don't know, that seems a little bit old fashioned, but this probably fits somewhere in there. Um, I do want to ask a quick question though, about, uh, where UDO will come up because you mentioned DAWs or digital audio workstations SPEAKER_17: earlier. Um, very much now the well-known kind of entity in the musical world. Does UDO ever become part of one of those, a plugin, an API that I can call, does it leave the website and end up somewhere SPEAKER_13: else? Oh, quite possibly. Like we think a lot of our, uh, power users, they use UDO to come up with ideas and then, uh, they, uh, download, uh, the individual stems, which is a feature that we SPEAKER_46: allowed, uh, we allow. So, uh, people can download stems and then load the stems up in their DAW for, SPEAKER_00: uh, for the post-processing. Uh, okay. So essentially they take the rod draft, bring the stems over and then you can do, okay. All right. That's pretty cool. Is, is it hard to do individual stems? Because that implies that the model is making a collection of tracks that are then mixed together. Is that how it's always been, or is that a new change to how the underlying model works? SPEAKER_13: So the underlying model always produces, uh, um, the fully mixed track. Um, but then, um, recently we, with version 1.5, we added the ability for users to download, uh, the individual stems, which are SPEAKER_16: separated post hoc from the mixture. Oh, post hoc. Interesting. Yeah. Oh, okay. So you, so it creates SPEAKER_00: something that's mixed and then you isolate. I would have thought this the other way around, but that's why we ask questions. Yeah. Okay. So, uh, before we talk about money, stems are the individual tracks inside of a song for, for example, bass or guitar or piano or whatever. Um, I just want to make sure that everyone listening understands stems. David, is that how you would define them as well? Yep. Okay, cool. So if you don't know what stems SPEAKER_89: are now you do. Okay. There are more than 50,000 venture-backed startups in the United States alone. This means marketing has to be perfectly targeted. You got a lot of competition out there, or you're just going to fade into the background and your money will go with it. All your ad spend will be for naught. You got to make sure you target the right prospects. So how are you going to do that, especially in a business to business context? Well, the answer is obviously LinkedIn ads where you can precisely reach the professionals who are likely to find your ad relevant. Just think about it. Wouldn't it be great to target your ads by the job title or the industry, the location that that company is, you know, and maybe even a very specific company, maybe you got a list of 20 lighthouse customers that you want to bear hug, that you want them to know about your product or services. LinkedIn ads is going to help you do that by building a relationship and driving results. LinkedIn is the environment where people are receptive to business. They're not there for food or politics or entertainment or music. They're there to do business. A billion members, 130 million of them are decision-makers and 10 million of them are C-level executives. So start converting your B2B audience into high quality leads today. Get $100 from your boy, Jay Cow, linkedin.com slash this week in startups to claim that credit. Again, linkedin.com slash this week in startups, no spaces, no dashes, terms and conditions. Why? Because they're giving you a SPEAKER_00: hundy. Okay. So UDO raised $10 million. That was earlier this year. And Jason Horowitz was in there, a number of artists, including the producer, Tay Keith, and love to see venture capital funds. Glad you guys raised some money. But my thought is this, I currently pay you $10 a month to use something like 1200 song creation credits. I look at that and I know how much AI costs to run. People talk a lot about that. I feel like I'm burning through your bank account. So is it as expensive as I imagine it is to run the model to create music? Because it sounds very compute intensive. SPEAKER_13: Yeah, obviously it's a balance for us. We want to make sure the price is set at a point where we allow people who are curious about the technology to try it out while being able to SPEAKER_11: make this process sustainable. So we chose a price in a way to basically allow for this, to be able to sustain usage while not being very expensive. And going down the road, we definitely SPEAKER_13: want to optimize our models, make them more efficient so that they can run at a cheaper cost, SPEAKER_11: because we want to maintain this commitment to users, but we also want to run a sustainable SPEAKER_00: business. Yeah. So we've seen this with, just to pick one example out there, OpenAI's GPT family of models. When 4.0 came out, it was one cost and then it's come down, I think, and we've seen that pretty frequently. Does that mean that you guys are able to extract a lot of efficiency from the underlying model and that this should get much cheaper to run over time? Or is there less low-hanging fruit SPEAKER_27: because it's doing music, which is just, to me, seems harder than doing text? SPEAKER_11: We think that there is a lot of room for improvement for sure. I wouldn't comment on whether or not it's SPEAKER_13: on the same scale as OpenAI. OpenAI obviously has entire teams of incredibly talented engineers working on this and we are a much smaller company, but we do believe that there are similar levels of efficiency SPEAKER_67: gains to be had. Okay. So essentially, yes, you know, Udeo is a smaller company. I think OpenAI has SPEAKER_00: over 1,700 people now, but with work, a similar-ish curve. Okay. That's actually a really good question SPEAKER_19: for me to ask. How big is the company today? What's your current staff size? So we currently have about SPEAKER_11: uh, 17 people. Uh, so we've, uh, grown quite a bit. How many people did you have before? When we, when we launched the company, uh, or when we launched, um, uh, our model back in April, we only had eight people. Eight. Oh God. Yeah. Um, and just because this is a startup show, SPEAKER_67: let's do some basics. Remote, hybrid, or in-office? Um, mostly, uh, mostly in-office, but some, uh, SPEAKER_00: some working remote. Okay. And, um, just thinking about staffing for the rest of the year, are you gonna keep hiring as aggressively as you have, or will that slow down that you've SPEAKER_20: more than doubled in size? Uh, we'll probably stay a little bit more steady. SPEAKER_00: Okay. So I know you guys raised, uh, from Andreessen, I mean, a venture firm that everyone watching this show knows, and you guys raised 10 milli. Is that enough money, David? Because some of your competitors have raised more and we are in the era right now of companies that use AI raising lots of money, let's say. So I'm just kind of curious why, why 10 million was the number and also, you know, how soon are you gonna be back on the show telling me about your shiny new round? SPEAKER_20: Yeah. So, uh, when we started, we raised 10 million because we want to be disciplined in how we, um, SPEAKER_13: spend the money. We believe that like, um, there's some amount of truth in the idea that scarcity SPEAKER_11: produces innovation. And, uh, and so, uh, or we try to be super efficient in a way that we use, uh, our capital to, uh, develop our models. SPEAKER_17: So tell me, tell me about that because, SPEAKER_00: you know, developing a model. I mean, people talk about how models are eventually gonna cost like a billion dollars to put together, but that's for a very general purpose model and so forth. So for, for you guys, how do you ensure that your capital expenditures on model creation and improvements are cash efficient? SPEAKER_13: Yeah. So one thing that, uh, we do is to try, is try to secure a cheapest compute power that's SPEAKER_11: available. Like the, the chip that's cheapest in terms of, um, floating point operations per second, SPEAKER_13: uh, versus, uh, dollars. And so, uh, we ended up like choosing Google clouds, uh, GPUs, which we identified as offering significant savings over other, uh, like chips, like, like NVIDIA GPUs. SPEAKER_57: Google's startup cloud program is one of the occasional sponsors of the show. So I just want SPEAKER_00: to point out that no, no one asked him to say that that was off the cuff, but there you go. So we're not, we're not being biased. Um, just to put it in your perspective for me though, cause I don't get to go to those negotiations. How much cheaper was, uh, GCP or UDO compared to competing providers? Was it a lot cheaper or was it more of a marginal differential? SPEAKER_11: Yeah. I'm not sure if I should comment on specific numbers, but, uh, it is. SPEAKER_57: Oh, you should. David, you definitely should. You should drop all the numbers you can right now. SPEAKER_11: Yeah. But it's, uh, definitely, uh, quite a bit, uh, cheaper. And, um, and so that's one factor. And the other factor is that we have quite a, quite a few, uh, really talented modeling, um, uh, research scientists, uh, among our co-founding, uh, among our co-founders. And, uh, SPEAKER_13: because they have a lot of experience training these really big alternative models, they know how to, uh, make maximal use of the available, uh, hardware, how to, uh, create really efficient SPEAKER_11: programs and how to like design architectures that can train efficiently. SPEAKER_25: So if you're doing that work though, cause that's, that's nitty gritty stuff. If you're SPEAKER_00: doing all that already, why not buy your own H 100s or equivalent and, and just run your own mini data center? It seems to me like if, if compute's gonna be such a core element of SPEAKER_61: what makes the, the digital brain that you use, why not own the, the neurons themselves? SPEAKER_13: I guess for us, uh, we, uh, as a startup, we didn't really want to deal with, uh, the logistics SPEAKER_11: of running our own data center. And we thought it would be simply to go with a cloud, uh, cloud option. SPEAKER_00: Do you see the company in, let's say five years from now, just looking down the road so far that I know we're making kind of almost like a joke here, but like, do you think that you'll still be on a major public cloud provider in five years? Or do you eventually off ramp when you have more, SPEAKER_17: more money and staff and so forth and do your own data crunching? SPEAKER_11: Yeah. It's hard to say, like the, the cost of, uh, computation on, on, on cloud has been, has been going down. Um, I think there's like, uh, some kind of new law, like one flaw or something, um, that supplanted, uh, Moore's law about the cost of GPUs, uh, over the years, the cost of a cloud SPEAKER_13: computing could, uh, go down like very significantly. And it's very hard to predict, uh, five years down the line, whether or not it will be more economical to buy your own chips or to lease them SPEAKER_00: from, from the cloud. Okay. So Huang's law, by the way, this was, of course, a reference to Jensen from NVIDIA, I presume. Yeah. Yeah. Okay. So if you know NVIDIA, you know, the guy Huang's law is, um, and I'm Wikipedia in this live, so this is not very lettered of me, but it's a general idea that as Moore's law predicted that the number of transistors would double about every two years. Huang's law is that GPUs will more than double their performance every two years. So it's essentially an acceleration or a faster version of Moore's law for GPUs. That speaks very well for you guys, because that means that the, your gross margins should improve over time, just naturally as, you know, chip companies make better chips. That's kind of cool. That's a tailwind for you as a CEO. SPEAKER_11: Yeah. It's definitely something that we're very excited about, uh, like cheaper compute, making even more powerful technologies, uh, possible. All right. Are you building the next SPEAKER_36: great AI product? Well, if you're doing that, you know how expensive all these APIs can be for model training data, obviously. And training AI is very expensive. That's a fact. We all know that. So you have to try Brave's new search API. Yes. I'm talking about Brave, the privacy browser that I use every day and on my mobile phone, Brave's browser has 65 million users. And that drives a lot of data into the Brave search engine, which is the only global scale independent search index outside of big tech. And that index is available to anyone with a Brave search API. So you're going to be able to use the Brave search API to power your chat bot or train models, inform answers to real-time queries, and serve images, web results, even rich tech snippets. The Brave search API features an easy to use intuitive data structure. So you're going to be able to get things done quickly. And its data is populated by real human interaction, not web crawlers. That's critical. And it's all done at a fraction of the cost of the major players free for up to 2000 queries per month. So you can try it on, play with it, really sort of brainstorming, and then plants are as little as $3 CPM. So here you go. If you're building next-gen AI apps or chat bots, you've got to try the Brave search API. Get started today at brave.com slash Jason. SPEAKER_00: On the public cloud front, I'm going to not ask about an individual provider, because I don't want you to get in trouble, but I have heard that there is a capacity crunch out there, that there's not enough total GPU-based compute for people that want it. Has UDO any issues getting the amount of SPEAKER_17: compute capacity that it needs at any point? SPEAKER_13: Yeah, I mean, it's always a balance where eventually we'll be able to get the compute, SPEAKER_14: but definitely at times, it just takes a while for different cloud providers to be able to find the chips that are available. SPEAKER_00: Okay. Now, I want to go from there to a demo so everyone can see the product that we're talking about from a compute perspective. So David, we drew straws before, and you're going to drive, because you told me that you have some new stuff to show off. So let's pull up UDO. If you're watching this on YouTube, you can see what we're doing live. If you're watching, listening to this on SPEAKER_17: Spotify or Apple podcasts, we will narrate as best we can, but we are going to do a little bit of testing around here to show off what we can pull off. So David, talk to me. What are you showing me? SPEAKER_20: Yeah. So I'm showing you the UDO create page. It's a dedicated, you can think of it like a creation SPEAKER_13: studio where you have a list of your recent creations on the right-hand side. And then on the left-hand side, you have a place where you can specify the type of music you want to create, SPEAKER_11: as well as any lyrics that you want to have. So maybe we can start very simple. Let's start with just creating, I don't know, like rock music. So here we can type rock. And then for simplicity's sake, I'll just create a song about New York. And so this is a feature that we launched just SPEAKER_13: yesterday, actually, where you can write, you can ask the language model to write lyrics for you before you submit the song and can even like tell it to, um, give it suggestions on what to do. Like for SPEAKER_11: example, let's say, uh, we want to make it a little bit shorter. Nice. Okay. So you can essentially SPEAKER_57: tell it to get more verbose or less verbose and do other things as well. I don't know. Uh, like SPEAKER_136: make sure to mention New York. Does it keep the last prompt in mind when you give it another SPEAKER_138: instruction? So is it still thinking, keep this short as you add the, make sure to mention New York SPEAKER_13: in the update box? Oh, uh, yes. So like, we actually have a prompt history that shows, uh, all the prompts that have accumulated so far. Um, so we aim for this to be, um, to improve upon our previous, uh, lyrics writing experience by giving people the aid of AI to help them, uh, come up with SPEAKER_11: ideas when, uh, they might have writer's block. They don't really know what to do, uh, for example, like, like me at this current moment. And so now like I had, now that I have like the genre and like SPEAKER_13: lyrics, I can hit create and that will, uh, queue up, uh, this creation. Yeah. And while, while that SPEAKER_00: goes, I just, I've been thinking about this because whenever I sit down to use UDO or a similar product, I tend to think in not genre terms, but in terms of, um, bands that I love and kind of how they approach the world. And, um, I'm kind of curious when you, when you're using UDO, do you tend to stick more towards like, like rock or do you get a hyper-specific, like make me a rock song with a, a touch of, I don't know, Tom Petty or something like that, because you can pull in different SPEAKER_11: influences. Yeah. So, uh, I would say that I usually, uh, just stick with, uh, genre information, SPEAKER_13: uh, but for users who, um, have a specific artist in mind, uh, we, uh, give, we provide functionality for a user to type in the name of the artist and we look up the style for the artist. So we don't actually put the name of the artist anywhere in the prompt because we don't want to create something that sounds exactly like the artist. But then for example, when you like type in like Led Zeppelin, it will like replace Led Zeppelin with, uh, the list of genres that he is, SPEAKER_11: that they are like, you know, like, uh, commonly, uh, associated with. So like, you know, like, uh, like maybe hard rock, uh, maybe, uh, male vocalist seventies, um, just things like that. SPEAKER_00: Okay. So if I put in, I mean, this is again, a niche genre, but like periphery, it's going to think progressive metal guitar, forward male vocals. So it'll essentially atomize an artist's name. And so essentially then artists become shorthands for genre and style. SPEAKER_20: That's correct. And, um, and we try to make sure that like the generated outputs, um, are definitely SPEAKER_13: influenced by the style of that artist, but it's not that exact style because we, that's one thing that we really do want to avoid dancing around the lawsuit here. And I'm trying to deliberately ask SPEAKER_17: questions you can't answer, but yeah, thank you for answering that. Can we play this? Let's, let's hear it. SPEAKER_100: Uh, yeah, sure. So this is, uh, the first, uh, example that came out. SPEAKER_149: It's better than something that I could, we could write. So shout out to that. SPEAKER_11: And then for example, uh, you can then, uh, like add additional descriptors to this. So maybe, um, not just raw, maybe you want to make it, uh, a minor, uh, and then, uh, you hit create again, um, and accuse it up. And so, uh, you, I'm curious about this though, because there is a, SPEAKER_17: there's a little bit of time that goes through to, uh, from when you click create to when UDO gives SPEAKER_00: you the song, which by the way, to me is, is, is no big deal. Um, but it does seem to be a variable SPEAKER_61: David. So what makes it a longer or a shorter calculation process on the UI side? SPEAKER_13: Yeah. So, uh, on the UI side, um, so we submit the request over to the server and the server, uh, reads on the prompt, like rock a minor and tries to figure out what to do with it before it, uh, sends it, uh, to the model. So we have some, uh, processing, uh, that goes on. We also have, we also run checks, um, for every song that gets submitted. We take the lyrics and we, uh, do a copyright check to make sure that, uh, the lyrics are not, are copyrighted lyrics. And this is something that's, uh, probably a little bit overly strict right now. There are a lot of, um, public domain songs or things that should not be, that are not really copyrighted that gets flagged, but we want to do on the side of caution rather than, um, uh, not flagging something that's actually SPEAKER_00: copyrighted. Okay. Now we have this new song, same idea, but now in A minor. Let's, uh, SPEAKER_165: let's listen to the new song, Chasing the Pulse. Yeah, let's see. I'm not sure if you have perfect SPEAKER_44: pitch or not, but like, uh, uh, it's a little bit hard for me to tell, but it's definitely a minor key. SPEAKER_167: David, I'm not gonna lie. Uh, I do not have perfect pitch. Indeed. If you had ever heard me SPEAKER_00: sing a bedtime song, you would think to yourself, that guy plays music? Cause it doesn't sound like it. SPEAKER_17: Um, so I can't tell if it's a minor or not. It did sound minor, but what, what hit me though, SPEAKER_00: is when the, in the course, or maybe it was the bridge, when the harmony voices came in, wasn't on the first note, they came in a little bit later, which feels very stylistic. And therefore, and I mean this in the best possible sense, like, like human, it felt like something that that would be a music editorial decision that a human can make to say, Hey, we're going to have lead singer. And then the harmony come in and delay. I don't know. I, I, I'm, I'm, I'm always a little bit torn with this between going, this is the coolest thing I've ever seen. And oh God, are, are, are, are humans going to lose just because, you know, I've sat in the guts of a symphony as we took Beethoven's fifth out of the studs and rebuilt it. And I, I don't want a future in which we lose that, but at the same time, I'm going to click these buttons a hundred times because it's a lot of fun. So maybe I'm part of my own problem, I suppose. SPEAKER_13: Yeah. Well, I, I think that, um, uh, people who play in bands, uh, they don't necessarily, uh, they won't stop just because, uh, there's this additional, uh, source of music. Uh, one of our co-founders actually, um, is involved in a band and he regularly tours the UK to, to perform with his band. And so, um, I think, uh, and he continues to enjoy this, right? Because it's just fun for humans to create music. And we just want to like give people, uh, more people the opportunity. Like he can create music with his band, but previously he couldn't create music in his bedroom, uh, or like SPEAKER_11: lying down, uh, like, uh, on his couch and then, uh, oh, I want a song. And like previously, what would you do? Like you can't even do anything. And so now, uh, this is possible. Yeah. And just because SPEAKER_00: I'm going to be an enormous brat because I can, um, court, one of our fine producers here at twist has given me a prompt. He would like us to try. So David, if you're up for it in zoom chat, there is a prompt entitled a jazzy neo-noir offbeat rap song about dinosaurs, which is evidence that court SPEAKER_17: is Gen X, but we'll leave that aside for now. Uh, but can we give, uh, can we give that a try? Chamath Palihapitiya: Yeah. Jazzy neo-noir offbeat rap song about dinosaurs. SPEAKER_00: Let's see what we get. Uh, the person who requested this song for everyone who's listening to this later on, um, was in a punk band once. So there you go. This is what a punk fan is going to put into the, the, the UDO, uh, generation process. All right. Uh, while this is waiting, David, one question I had written down just because, you know, I love this sort of thing. What's the craziest song that you guys have seen people come up with? Because everything that I've done thus far has been pretty standard, but I'm curious, has anyone like really blown your head off? Well, uh, one of the examples SPEAKER_13: of a song that took off pretty unexpectedly is a song called, uh, BBL Drizzy. It's, uh, uh, it's a sound that one of our users created. Uh, he himself is not a musician, but he is, uh, uh, a comedian, actually. And so he wrote like the funniest lyrics and he used, used UDO to turn this, uh, set of SPEAKER_11: lyrics into a song. And, uh, it ended up being, um, sampled from by Metro Boomin who created a beat, like part of like entire like Drake and, um, uh, Kendrick feud, uh, and challenged people to create, uh, like voices on top of this. And the funny thing is like Drake himself actually wrapped on top of it. Uh, it's, uh, it's, uh, pretty amazing watching it from the sidelines, uh, to see your tool, uh, be SPEAKER_142: genuinely immersed in pop culture. We'll, we'll come back to that in a minute, but I want to play SPEAKER_19: everyone this song. So here is the first sample clip of jazzy neo-noir offbeat rap song about dinosaurs. Hit it, David. SPEAKER_185: Rolling through the night for your storage war. Lost in the city lights. Geno feet at the floor. Brossaurus roof. T-Rex on a bowl. Pterodactyl fly rhythm, digging in my soul. Bones of the past. We dance like we're aceless. History breeds a jazzy night. Valid stages. Echoes in the alley. Shadows keep time. Jurassic jazz notes in the moon's climb. Dinosaurs in the urban globe. Rhythms of time of the ancient show. SPEAKER_188: Underneath the story flow. Where the city and wow collide. We go. SPEAKER_00: I mean, I'm not going to lie. That's not bad. That's not bad. Brontosaurus groove. T-Rex on a roll. Pterodactyl fly rhythm, digging in my soul. That is actually probably better than some stuff SPEAKER_07: that I listen to on Spotify currently. Yeah. Uh, our, uh, lyrics are very, uh, very peculiar. Uh, it's an artifact of, uh, language model. No, no, no, no. I, I meant all that. That SPEAKER_00: was not sarcasm. I mean, I, I never thought I'd see Brontosaurus T-Rex and Pterodactyl all within SPEAKER_57: one rhyming couplet essentially. Yeah. No, uh, again, I don't know if you answer this, but is the model that SPEAKER_17: writes the lyrics, the same model as what does the music generation, or are those two different models that then are, are brought together for this final product that we just heard? We use, uh, different SPEAKER_13: models. So, um, the, the model that creates the music is a proprietary model that we trained, uh, because there's nothing like that, uh, elsewhere. And, uh, and the model that writes the lyrics, we just use, uh, uh, GPT actually. Oh, simple enough. Yeah. Well, I mean, it works pretty well. SPEAKER_00: Okay. I want to talk about 1.5 a little bit, and then we'll wrap on virality. So 1.5 came out back in July, um, that brought key control improved, I think it was global language and then also audio SPEAKER_138: quality. So how has been the reaction to 1.5 and then what's next, uh, from UDO in the feature context? SPEAKER_14: Yeah, I think people were excited by the, uh, by the, by the changes, like for, um, a better global SPEAKER_13: languages. Um, uh, a lot of our, uh, Chinese speaking users remarked how the model suddenly became a lot better at producing lyrics that have, uh, Chinese in them. Our, um, uh, key control is definitely like, uh, some, a feature that people have wanted for a very long time. And people love the ability of like, uh, specifying the key and then modulating within the song. So you can, you know, like specify a key for the free section and we extend the section, you can specify a different key and you can kind of like, um, you know, uh, specify your own, uh, harmonic progression SPEAKER_00: throughout the course of the song. How long until that's like super visual? Like I can imagine myself like, um, let's say a song is three minutes and I'd like a, a line and I'm like, this chunk should be in a minor and be up tempo X, Y, Z. And then I want six measures of this, that like, does this become a SPEAKER_17: visual tool versus just something that I prompt, uh, at least in my experience with words? SPEAKER_11: Yeah. So this is something that, um, going forward, uh, we do want to make more visual over time. So we recognize, uh, there are some deficiencies in our current interface and make it a little bit SPEAKER_13: harder than necessary. And so we want to do like, um, uh, user research to figure out how best to, uh, craft the interface in a way that's intuitive for musicians. SPEAKER_00: Okay. So, cause musicians are already super familiar with editing software and so forth. So that kind of interface would be a second nature to them. Now I want to talk about virality because if we go back to when you guys announced your fundraise, I think Bloomberg reported and I have it somewhere in my notes here that you were seeing something like 10 songs created every minute or something like that. I forget the exact, um, pace, but how has the company been doing in, in usage terms in the last couple of months and how much bigger is it compared to that April, June, SPEAKER_13: timeframe? Yeah. So it was actually like 10, uh, songs every second, uh, not a minute, but, uh, people are, uh, people still like, uh, like super, uh, engaged with the entire process. We have a, like a very dedicated group of power users who, um, you know, go on discord all the time and they, SPEAKER_11: they share the songs that they have created. There's actually a bit of a collaborative flow as well, where people work on lyrics and songs together. And you see this because, um, in the final output, they will credit each other. Uh, uh, uh, they will say, oh, this song, uh, created with the help of this other user. And, um, and it's really fun to see people, uh, working on music, uh, in this way. SPEAKER_13: It's kind of like how people would jam together in the past, right? And while people still jam together, SPEAKER_11: but like now we have another way for people to jam together. And I think this is also part of what music is about, like bringing people together with a common passion. I agree with that entirely. I SPEAKER_138: was just thinking that, you know, you're right. Not everyone now has to go to a jam room, which SPEAKER_00: means that they don't have to get hearing loss. Like I did growing up in my ska band, which rest in peace, did not make it big and turn us all into multimillionaires. I'm sad to say. Yeah. Now on the virality point, you mentioned a discord, you mentioned power users. I learned about you guys from a friend, but I'm kind of curious is, is the product here inherently viral because you know, I was sent a song about me and my friends. I like to use it and I've been playing with it, showing it to people. Uh, and so I'm just kind of curious if that limits your sales and SPEAKER_17: marketing costs because people are almost taking your product to their own networks ambiently. SPEAKER_13: Yeah. I mean, most of our growth, uh, almost all of our growth is via, uh, completely organic channels SPEAKER_11: where, uh, people just like sharing, uh, amazing outputs that they have. And then people asking each other, how do you do that? And then like, uh, just spreading it this way. SPEAKER_17: So how has been, uh, like registered user growth at the company? Is it still SPEAKER_217: as quick as it was before in like percentage terms or gross number terms? How should I think about growth at the company essentially? SPEAKER_13: Yeah. I mean, obviously there was a, a very large, uh, initial spike, uh, when we launched, but we do, uh, still see like steady, uh, steady growth, uh, every single month. And, uh, we believe that like, once we, uh, launch, uh, newer versions of the model, SPEAKER_11: there'll be renewed excitement. And, um, as people find new ways of controlling the outputs and using the model for their own production needs. SPEAKER_59: Okay. All right. Well, I mean, I'm going to be watching with, with very close eyes because SPEAKER_00: I'm a user and now a customer, but one thing I've heard from VCs lately, I think Sarah Tavel from Benchmark wrote about this. And she said that a lot of the AI, the big AI model companies, the open AIs and so forth, um, are going to go kind of up stack in time. And so that startups, not yours, but some startups that do build products using well-known commercial models, for example, might eventually get supplanted by their model provider, essentially going up stack and taking their lunch. Um, are you at all worried as a company at one of the larger model companies, Amistral and Anthropic and open AI, I'm going, Hey, music is cool. We should do that too. And then kind of bulldozing into your market. SPEAKER_13: Yeah. So we think that, um, inevitably in the future, there will be more, um, companies who enter this music creation space. Uh, we believe that, um, music is sufficiently different from text. And there's a significant product element as well. People, uh, you want to have the right interfaces for people to interact with these, with these models. So like a chat bot, like, uh, kind of like, uh, like chat GPT is probably not the right interface for people, uh, who want to create music. SPEAKER_11: And so I think there's, uh, um, actually like an open question on like how to best produce this type of product. It's something that we are working towards. Uh, it's a, like, it's a tight coupling between the model and the product and like getting the right level of controls in the model so that you SPEAKER_13: can expose them to the user in an intuitive way in the product. Okay. And then just to wrap things up SPEAKER_00: here, David, I just want to ask you one more thing before I let you go, which is, you know, I think about UDO and Suno is the other company that I think people best know in your space. And so SPEAKER_61: I'm kind of curious, where do you see UDO today in comparison to Suno and, um, how many of their engineers are you currently trying to poach? So I think like we want to position ourselves, um, SPEAKER_13: as, uh, as allies to artists and our soundwriters and producers, we want to focus on like giving them the highest quality tools available. Um, so instead of focusing on like meme songs, uh, in particular, we want to focus on the really powerful creation tools, um, to help creatives, uh, like make music and make like high quality music that they're proud of. And maybe eventually, uh, they want to, uh, SPEAKER_138: incorporate in their other like music workflows. That was a very, very deft non-answer, but, um, SPEAKER_00: let me take another run at this. Uh, do you consider UDO's music model to be the best in the market SPEAKER_11: today? I would say so. Yes. It's the only model that, um, uh, that produces stereo music at, uh, like 44 kilohertz sampling rate. It's a lot higher, higher fidelity than any other music model that's out there. It has a better understanding of, um, genres and, uh, like almost any other music SPEAKER_59: model. Okay. I'll take that. I do want to have you back though in, I was going to say a year, but given SPEAKER_00: that you launched the product in April and it already feels like we've gone through two generations, probably sooner than that, because I'm curious to see how fast things improve, how competition evolves. And if you guys do decide to go back out into the market, because I, I think that given your traction early monetization and so forth, you should be able to raise more. Um, so it'll be very curious, but David, thank you so much for coming by twist. I really appreciate the information and the notes. And, um, thank you for making a new tool for me to play with because I absolutely love music. SPEAKER_11: Yeah. Thank you for hosting the podcast. It was really fun. SPEAKER_53: All right, everybody twist is back. We do live news. If you're not with us on YouTube, see you there. SPEAKER_00: We're also in every single podcast platform you can possibly find. And we are always trying to find the best and most interesting founders to explain the market as it is. This has been UDO, David Ding and Alex. Hey, we're out of here.