SPEAKER_00: your interviews with brian chesky i learned a lot from those episodes actually me too also i liked SPEAKER_01: his idea that like they were going to only ship the number of features he could keep in his brain and that his brain would be the maximum you know size of the canvas so if if i can't if one person can't keep all these changes in their brain let's put those changes into the next six month cycle i SPEAKER_06: thought that was pretty yeah awesome i actually borrowed a heuristic from there adapted it for our company which was if the person building the feature doesn't know how to write the code for it they're very good programmers but if they're finding it a hard time to break it down and actually SPEAKER_09: implement it then it's not worth shipping this week in startups is brought to you by crowdbotics SPEAKER_10: great ideas can change the world and crowdbotics is the fastest way to turn those ideas into code get a free scoping session for your next big app idea at crowdbotics.com twist vanta compliance and security shouldn't be a deal breaker for startups to win new business vanta makes it easy for companies to get a sock to report fast twist listeners can get one thousand dollars off for a limited time at vanta.com twist and open phone brings your team's business calls texts and contacts into one delightful app that works anywhere get 20 off your first six months at SPEAKER_01: openphone.com twist all right everybody welcome back to the program we have been having an amazing array of founders who are taking on the challenges of implementing ai in the real world and today will be no different uh we have our unvind shrin evas on the program he's the founder and ceo of perplexity SPEAKER_05: ai welcome to the program aravind do you have a nickname or you go by aravind aravind's good thank SPEAKER_01: you for having me here jason uh great to have you and um listen you there's a big battle going on between chat gpt and bard and you're right in the middle of that you are doing perplexity ai and you are trying to compete with these two giant you know software developers tell us a little bit about how that's going and how you see perplexity.ai which you can go check out right now you the interface will look familiar how do you plan on competing with them and how is that going SPEAKER_04: yeah so firstly we uh we started off a week after chat gpt came out we put it out and SPEAKER_06: there was a lot of difference at that time which is we were just a search bar and we give direct answers with citations uh whereas chat gpt was this entertaining hallucinatory bot that was not not just you know like correct many times but it was also equally wrong many times and its mistakes were also entertaining right so we focused a lot more on revamping search realizing that you know 10 years from now no one's going to be asking for 10 blue links you're going to ask for answers so we might as well start it today and the technology for that was ready both chat gpt and us were basically being powered by gpt 3.5 uh that was the fundamental breakthrough and then after that gpt4 even better than that so that's kind of how it started and we were seeing pretty different it's like oh you know if chat gpt lies or makes up makes up things there's this other side called perplexity you go there and like it's going to be this boring educated uncle kind of product but useful and you can trust it and that's kind of how we grew and then bard came out um i believe bard still hasn't solved the fake news problem completely um it does hallucinate and it doesn't actually have like real citations sometimes um um so we we are better than bard in that in the context of search but google's also rolling out this thing called magi or they call it search generative experience to the public but uh the wall street journal called it magi um so they're they're trying to do something pretty similar to us and so far the experience at least from what i've seen and myself used that and other people who use that is it's not very different from the way they used to extract text from the top link and put it space at the top it's not very different from what they've already done before um and they cannot afford to use really powerful models so like the search traffic that they have and if they actually want to really get it right but uh in the actual search bar itself outside of bard they're SPEAKER_32: going to lose a ton of revenue hmm yeah so they have two challenges there and so do have you done SPEAKER_01: a crawl of the web because you are giving citations and you have a language model behind this so tell us what is underneath the hood here because i have been saying hey if you're going to get a bunch of information and present it to me in a beautiful uh you know answer with bullet points and numbers uh like perplexity just did for me i asked it hey what are activities i should do with my seven year olds and uh they like cities and the outdoors and it gave me four popular uh destinations for cities and four popular destinations outside really good suggestions really tightly summarized and then at the bottom it said hey and then here are three citations uh trip advisor us news family vacationist and today the today show um so what's underneath the hood here yeah how is it generating the answer SPEAKER_31: yeah so lms are these great reasoning engines uh you throw you throw a lot of text at them and tell SPEAKER_06: them what to do with it and they'll do it for you um and then there's the other part that's great which is having a good index and a ranked ranked version of the index which is a traditional search engine and what we do where we come in is we combine the two together we say hey like elements are great we'll figure out what content to throw at them for a given query and we'll instruct them on how to actually uh take all the text that's thrown at them in the context of the query and get the needles from the haystack and present it in the right format to the user so they're doing more of the reasoning job they're not actually doing um pulling up actual facts that's been stored in the lm itself because some of them could be right some of them could be wrong the real actual facts are in your web pages so that's the content that we want to take and we have like our own index and also like we rely on other index providers and we collate from multiple different indexes multiple different crawls of the web and pull up the relevant links and then we ask the lm to do all the reasoning on top and then we give you the answer now the magic is that all this happens so fast we've put out the product in december and back then uh the latency used to be like five to six seconds per query in fact one of our investors daniel gross he used to joke uh to me saying you should call it submit a job and not submit a query it's that slow and now it's like almost as fast as google like you're hardly waiting um your the summary is like really generated really quickly and we still have like you know so much more room to improve there and i think at some point you're just going to take answers as the de facto search experience that's kind of what we want to bring together and our primary like superiority over the SPEAKER_39: existing products is the speed at which we deliver the really accurate well collated answers from so SPEAKER_01: many different sources and so but you are built off of today chat gpt4 correct yes we heavily use SPEAKER_06: chat gpt 3.5 and 4 and we also use a little bit of our own lms for many other things we have every question you ask on our site you see a few related questions that are being popped up right that's actually one of the favorite parts of the product for many of our users because they like asking more and that is sort of generated with our own lm for example so there are some parts of the product that we use our lms but i would say like most of the heavy lifting is being done with opening as lms right now SPEAKER_01: and so you added right now does that mean your plan on building your own um because it does seem like you're directly in competition with bing bing has the partnership with chat gp4 so it's almost like you're both using the same underlying technology correct they already have some scale so that that would be a difficult race there so how do you look at chat gpt's four's relationship SPEAKER_05: or open ai's relationship and microsoft's access to it i think we just need to win by building a superior SPEAKER_06: product there's just no other way and i believe so far we've done that we have not won against them but they still have a lot of distribution through windows devices um so a lot of people just go to edge and they can start using bing chat um but people have despite that we won against microsoft in the past like they google everyone went and searched for chrome as the first search query on like internet explorer to install it we all did that despite the friction they added um so that's one way there's only one way to win against the person who has much more distribution than you which is a superior product SPEAKER_31: now about using the same underlying technology it is the case today uh the reason is they have the best SPEAKER_06: models and there's still a lot of differentiation you can have and how you harness the power of these models these are so general purpose machines it's almost like you buy the engine from somewhere but you're building um a whole car with a lot of different parts and you can still build a better car um and if it is the case that open ai is just going to be the number one place by far and you want to give the best product to your users uh you don't need to use their model like there's no um like like you can say yeah i'm going to use my own model because i don't want to use someone else's but then if the search experience is pretty there compared to what you have with open ai you're not you're not going to get users and then you build your own modes of differentiation in other ways that just the person owning the llm cannot build as good competitive product as yours so if just the llm is the only reason this is working we don't have a chance but that's not the case here there's so many other things needed to be done to give you this experience where there's real-time facts being pulled up and presented in the right manner uh super fast reliable and like make the product engaging so all that also matters so for example with uh people have done comparisons between us and bing and um you know we have like much better accuracy in terms of how correctly we cite things a lot of academic research has been done there uh people spend on an average like two minutes more on our site than bing uh so that engagement is much better there so our bounce rates are much lower than bing so basically we only lack in one thing which is number of views on the site but that can only be addressed if we're given sufficient time to grow and like make people aware SPEAKER_59: of us all right we all know the one thing that separates great startups from the good ones is product velocity what does it mean product velocity fancy term right here you got your product and SPEAKER_61: your velocity speed the speed in which your product improves so can you ship updates can you release new features can you do bug fixes can you iterate on the interface can you solve problems for your customers and can you do it quickly because you're not alone you have competitors and your customers have choices they may fit solve their problems by writing their own custom code or they might use your solution this is what startups are about how fast can you get that product velocity going and so you know how do you supercharge it everybody says okay yeah we want to go faster but you got to go faster intelligently and crowdbotics is going to help you do that they're your cto as a service basically they provide you with the most optimal architecture to get your product to market as fast as possible you'll have access to an on-demand product manager and developer talent and they will help get your app into production 10 times faster than conventional development crowdbotics can work with your in-house dev team or you can just have them work independently and you own all the ip you own all the source code let the folks at crowdbotics supercharge your product velocity today no more waiting get a free build plan at crowdbotics.com twist that's a 4.99 value just for the twist listeners you get that for free that's c-r-o-w-d-b-o-t-i-c-s dot com twist for a free build plan how do you um get the citations if you SPEAKER_01: were asking this query i just did about like hey what cities should i take my seven years old to and then what outdoor locations how do you actually get the citations because chat gpt4 they they don't SPEAKER_04: provide citations um or do they they have this thing called a browser plugin um which is basically SPEAKER_06: powered by bing um but people hate that experience in the sense it's really slow and clunky uh yeah it is slow and clunky yeah and um so how do we do the citations we basically pull up the relevant SPEAKER_31: links to your query uh from a search index and then we combine that and tell the lms to write the answer we basically ask the lms to go read all those links and then pull up the relevant paragraphs SPEAKER_06: from each of those links and then make an answer out of whatever you thought was relevant but write down the answer as if an academic or a journalist would write it where each part of the answer has the corresponding citation like wikipedia you basically say hey like i want you to do the job of what a human does on wikipedia where when they're writing uh something about a new person or a new phenomenon or a new city this is basically going and like uh picking up a lot of web links about that um sifting through them and reading them and then coming back and writing an essay on it right so that whole human labor intelligence needed to do that is being automated now all of this happening in like seconds right that's yeah worth like hours of human labor and that's the value we are actually SPEAKER_01: adding to everybody got it and so you collect all those links give all those articles and then give the SPEAKER_69: the summary of them basically instruct the llm to like hey behave like a wikipedia person just just write SPEAKER_01: it like this so the core of this is prompt engineering and knowing how to prompt engineer um for different types of queries because different queries might require wikipedia editor SPEAKER_05: other ones might me need a more of a sensibility of a journalist and the llm knows the difference between SPEAKER_06: those things you need to make it know that's the skill there and and and you're right prompt engineering is a big part of it um but prompt it just because somebody might have your problem doesn't change much actually like prompts can leak uh so it's all about orchestrating the back end uh the making it work with the right sources too um so there you know there's a steve jobs movie with kate winslet in it where uh there's a scene between wasniak and jobs where wasniak's like i'm the guy writing all the code and i'm the i'm the code you don't write code you don't do design why do you why does everybody know you and not me and he says i play the orchestra so that's basically where anyone who aims to build a long-lasting company on top of llms though the thing you need to be really good at is playing the orchestra like having so many things work together reliably and efficiently and correctly and super fast so one of the pieces is searching the SPEAKER_01: webbing finding the right articles the next piece is knowing how to write the answer right what are the SPEAKER_31: other pieces here uh the relevant relevant parts from each article too like ah article has a ton of SPEAKER_06: content in it you only need a few for the query you asked making sure that you write the answer in the most accessible way initially we just started off with just putting text with citations then people are like hey i want neatly formatted answers i want markdown in it i want like code to be rendered in a specific way i want images in that i want like videos in it um i might want to customize it for kind of the domain i'm searching in like and then people keep asking for more and you learn more about the second part of google's mission right making it universally accessible and useful so that the first part is organized there was information the second part SPEAKER_01: is basically where lms are adding tremendous value now and how do you deal with specific verticals of data that are more siloed i see one of your co-founders or one of your founding team members were was from quora you of course have the reddit data set uh great for conversations you have twitter great for debates uh and funny one-liners and breaking news uh you have yelp you have google local you've got all these silos of data i asked hey what are some great greek restaurants did a pretty good job of telling me greek restaurants in the bay area and so how do you think about um those silos of data and are you intercepting searches and saying hey this search is about local businesses and restaurants this search is about something that the reddit data set would do better with how do you SPEAKER_06: think about that yeah so so uh the part about data like you know the access and things like that it's an ongoing debate and like i don't have like you know very strong opinions on what each person should do ideally if there's a need for us to pay any party for their data access we'll do it um as for how we do it like what links we know for you to use for which query we we do like take your query and figure out like which category it is and like try to use that information to um give you the right sources it's pretty hard actually google does a tremendous job at this SPEAKER_04: and um we are also doing something called focus searches where in the search bar instead of using SPEAKER_06: all of internet you can go and pick like academic or you can pick youtube you can pick reddit wikipedia SPEAKER_32: and you can just yeah yeah so you can there is a drop down called all yeah and i could just pick SPEAKER_01: youtube and then youtube you have access to the corpus of all the transcripts or just the metadata i guess SPEAKER_06: and titles for now we use metadata and titles but that's already amazing sometimes i can't find some SPEAKER_00: videos on youtube directly but these lms are so good at like doing the relevance ranking that's much SPEAKER_01: better than the youtube search algorithm um the language models do better than google's native search SPEAKER_06: algorithm wow sometimes not always got it most of the times it's equal but sometimes it's just really good at like these fine grain i was trying to find a video of like oh so there's a scene in this movie i want to find for watching for inspiration or something and then i couldn't find it on youtube and i come here and i get it uh it's very useful for reddit like i want to like learn about like you know the nothing phone like you know who's even using it who are those million people and then i don't have time to go to the subreddit nothing phone and like score over all these like links that's very useful there um people use it a lot for wikipedia like if they just want to focus on one thing like i was talking to the founder of wikipedia jimmy wales and he literally just asked for this feature like hey i just want to do search over wikipedia with an llm and i was like hey that's a SPEAKER_48: great idea yeah so i think i think they're building it now within wikipedia so um interestingly i did SPEAKER_01: a search uh for interviews uh with the ceo of airbnb mine didn't come up but other ones did but then it came up with i did ones from the past year and uh man that was kind of a bingo it kind of nailed it uh which is a kind of a nice feeling um i really think that's a creative idea and i can SPEAKER_79: see how what you're talking about is yeah got some um there is some point to this which is if you SPEAKER_01: narrow the scope or you build some interesting prompt engineering or narrowing uh and thoughtfulness you can get to a better answer so what's going to be your business model here you talked before SPEAKER_05: about how google is not going to make be able to make it work with advertising there's a group of people who believe that uh the chat interface will cannibalize their existing business so do you agree that the this chat gpt style interface or just the chat interface let's leave the gpt out of it um nobody owns the chat interface but is the chat interface anti-advertising or could advertising be SPEAKER_01: integrated into because on all in a lot of the i think three out of four besties thought hey advertising is not going to work and i thought i think advertising is gonna work great inside of this you have your citations but you could put right in embedded in the discussion you know all kinds of interesting things so if you were asking about places to travel with your kids and i'm disneyland and you didn't make it i could put in there hey and if you're thinking about outdoor stuff disneyland also has this adventure park and they do the safari and i could have like a really ai generated answer at the bottom so it gives me the correct answer or what it thinks is the correct answer but then it also gives an ad engines answer to it so am i right or my other three besties right SPEAKER_06: you decide i'm more with you here oh you are okay so firstly i think relevance can be even more targeted now than ever before i what what is the purpose of google it's just bringing two parties SPEAKER_31: together the advertiser and the consumer and they help you connect these two parties together with their query and link matching right at the end of the day the advertiser wants to get their content to the consumer or the content and llm can give you that needles in the haystack even better like it's even SPEAKER_06: more targeted honestly uh that if i were an advertiser i would just kind of focus on selling myself really well writing even better marketing copies with llms uh catered to the person i'm trying to sell to and we introduced this thing called ai profile on perplexity where you can just write about yourself uh yes i saw that and and that way you the the results are even more catered to you and then if you're an advertiser you can say i want to like target people who are of like having all these attributes in their profiles um and then uh the ranking will automatically take care of that so in some sense you're you're you're creating way more relevant and targeted ads than ever before i don't know if you use instagram but uh my experience in instagram is that the ads on instagram are SPEAKER_48: even more relevant than than um often that is the case here take a look at this can you see my screen SPEAKER_01: here's the query i did uh based on our little back and forth here will llms will the chat interface be uh accommodating to advertising well i put in here you're the ceo of disney parks pitch me on why i should take my seven-year-olds to one of your parks and so imagine this got appended to my previous search which hey what should i do with my seven-year-olds in a city or outdoors and i says oh thank you for considering one of our parks for seven years here are reasons we believe you'll have an unforgettable experience number one a place where everyone is welcome two more value and flexibility three SPEAKER_05: uh disability access service that's kind of weird uh number four new attractions and experiences that's really good uh memorable music that's good too actually uh park reservation system that's great uh we hope you'll consider this and then here it could have bookings and would you like to talk to an agent you have further questions and you could just hijack somebody's chat stream uh for your own purposes they could be thinking they want to go to europe for the summer and then you could sell them on going on a european disney cruise or something and i i think that this kind of um style of advertising where a company ceo starts a discussion with you in chat gpt uh and in a in a chat interface SPEAKER_96: uh is going to be magical yeah and and like you said you know you can give you an i can give you an SPEAKER_06: answer that's sort of neutral and unbiased and it's not targeted at you and i can also say by the way in case you were actually looking for something very much much to you and if you already shared that information with us fully transparent and you're in control uh we're not going to do it in a creepy way like facebook then um we should be able to give you the answer uh we should be able to help the advertiser sell to you uh sell to you even even better right so i think basically i'm going even more abstract first principle is thinking that uh it's not clear how you do it in the product and how you build a business model but at an abstract level the point of advertising is to reach the right person to sell to and SPEAKER_120: this can help you do that even better than the current system so therefore you should be able to SPEAKER_32: figure out something at a level below this if you're a sas or services company that stores customer data in the cloud then you need to be uh sock to compliant you knew that from a third party and you need that third party to close big deals and if you want to get compliant easier and faster you need to use vanta v-a-n-t-a vanta makes it so easy for you to get and renew your sock 2 on average van customers are sock 2 compliant in just two to four weeks prepare that to three to five months without vanta and vanta can save you hundreds of hours of manual work and up to 85 percent of compliance costs this is a total no-brainer and vanta does more than just sock 2 compliance they also automate up to 90 compliance for gdpr hipaa and more you can't afford to lose out our major customers we all know that listen it's a hard year last year was hard you can't lose those major customers because you don't have your compliance dialed in just work with vanta get your compliance automated and tight and tight is right lock down those big deals here's the best part vanta is going to give you a thousand dollars off that's 10 hundies get one thousand dollars off at vanta.com twist that's vanta.com twist for a SPEAKER_86: thousand dollars off your sock 2. so you've raised some money uh and you're currently trying to grow SPEAKER_01: the company tell me a little bit about uh what it's like to try to compete in this area for talent uh people are raising you've raised a lot of money but people have raised even more uh and there's a massive talent uh uh battle going on right now is it better to just hire great developers and have them learn uh because you're not building the fundamental model you're building something on top of it you know what's your strategy for talent here yeah so we we don't waste time trying to hire people SPEAKER_06: that same woman will be hiring anyway uh it's very it's you cannot compete they have way more cash way more like and they can give way less percent of the company because they have way bigger valuation so SPEAKER_31: what we do is go for these people who are still trying to get into ai very talented engineers who SPEAKER_06: haven't done ai before and want to be part of an amazing product that's growing and they want to feel the dopamine from shipping every week and want to see their stuff actually being put up and there there is a quite a lot of people there are quite a lot of people who are like that like who haven't done ai before very talented generalist programmers that's another thing that i look for which is are they generalists can they do back and can they do front and can they like strategize for the product can they do prompt engineering because all these are new skills like prompt engineering is not like a you know you don't need to but there's you cannot ask for years of experience there it's like a few month old skill uh so you just need to be somebody who's like pretty logical and like pretty good at like getting SPEAKER_132: things done you were at open ai for a while i was at open ai yeah yeah how long were you there and what SPEAKER_06: did you work on uh yeah i i worked on like diffusion models and the conversational models for like like chatbot like not not exactly like chat gpt but more like trying to get uh another modality into like conversations so that's kind of what i was focusing on um but the reason i started this company was because uh ever since i came to us for grad school uh in berkeley i was always interested in starting a company and um i was trying to look for people who were like me before who were like phd students who started a company and i could only find one example from the past that i really resonated with was larry and sergey so larry is my entrepreneurial hero like he he's the only reason i kind of wanted to do a company uh like in fact in a book he's written like he'd either do a be a professor or he would do a company and he would never work for anybody else i had more constraints in my like you know immigration and other stuff like that to have to like sort of work for a bit get some money and like learn more skills but that was sort of always there and it's not planned but it's just more like a coincidence happy coincidence that i'm working on search too um but yeah the work being at open ai was really helpful back then there was no chat gpt so i i didn't foresee the future where open ai is so successful but there was gpd 3.5 and was pretty good and like you know we knew like a lot of things were happening nobody knew that like if you put out these models in the chat ui the world would go SPEAKER_132: crazy that was very unknown so the fact that people are so used to the modality of uh chat because SPEAKER_01: they live in it all day long this was the breakout moment moment for ai because ai stuff had existed people were using it in the back ends to serve you up your for your page on tick tock or filling your search query or giving you a couple of words ahead in gmail and finish your sentence all that stuff SPEAKER_05: that was happening yeah but the it needed the interface to make it work yeah fascinating the SPEAKER_06: generality of the models was also amazing but it was all you if you remember opening i had a playground where you could go and enter a prompt and in green text you would see the completions but nobody cared about the average person in the world did not care about it and then no you put it into a chat ui and then the world goes crazy right huh makes you wonder if there's SPEAKER_01: another thing that you could do exactly we're all go even crazier and i got a thing uh and i'm interested in your thoughts on this were siri and alexa yeah just far too early they they had the ability to understand what you were saying yeah they just didn't have the ability to give you the right answer or any answer exactly i mean you could barely call you know you'd be like okay call my mom and it would be like calling mother teresa and you're like no no no no no no it's not what i want and just even getting it to play the right song took three tries now with chat gpt and and all these language models and bard and poe and what you're doing a perplexity it feels like talking to the computer who would work and i don't know why this doesn't exist yet it's going to happen it's going to happen yeah like if i have perplexity as running in the background on my phone in my ear pieces and i could just whisper to it and say hey hey perplexity what are some SPEAKER_05: greek restaurants near me uh that have uh lamb and that are over four stars and it just gave me the answer back and started talking to me and i could take out my phone yeah that would be so a magical and just using the the language models as your interface but using voice and having to talk SPEAKER_04: back to you would be incredible yeah it still doesn't exist in five years i think what's going to happen is we'll we'll talk uh we'll all wear glasses we'll talk and then we'll see the answer SPEAKER_06: render in in our glasses and then uh or it can speak back to us and we uh we can listen via the glasses or whatever uh okay why doesn't it exist today like as we speak uh i think you you can stitch together a demo uh with a speech recognition model and llm and then a text speech model right yeah um the latency wouldn't be enjoyable like um that it's mostly on the llm side not not even on the speed side uh you can make these uh asr and dts work pretty fast uh but if you had to wait for two SPEAKER_04: to three seconds it's a bit like talking to a socially awkward person like they would be like SPEAKER_06: staring at you for like two seconds and then giving you back the answer right yeah yeah so that's the experience you would get you it might not be very enjoyable like how you and i are talking right now SPEAKER_128: i think for that you need even smaller or even faster lms and ah so it's not it wouldn't have the SPEAKER_01: response time that people would find not annoying it would be it would quickly become annoying to have it giving those pauses yeah i find it quite charming now when my chat gpt interface like takes a second or bard is kind of skipping around and it stutters and then it plays and i'm like wait a second and then bard now just gives you the answer straight away boom yeah it doesn't do the typing but i've got i think the open ai apple app uh ios app has like kind of haptics in it yeah where it's like SPEAKER_165: typing i think it's part it's kind of a gimmick right it's not we chose not to do it uh but there SPEAKER_04: is this thing where you stream the output tokens uh token by token the reason we did that is because SPEAKER_06: you perceive the latency as lower ah like if i waited uh in my backend to generate the full answer and then display it like in the bard style you might just be like oh what the hell like i don't want to wait you know uh and then you just i just i'll bombard you with a huge paragraph it may not be as fun as like anticipating like you're reading along with the model generating tokens that's a different kind of ux um i like opening eyes choice here but we didn't do the haptics thing because i found it pretty annoying to use when we were beta testing it and so did the others in our company um so that's that that said you know like here's the thing with tts like you have to generate the full answer before feeding it into the texas speed system um if it's just going to read it word by word as the lmd goes the answer uh it's not going to get the tone of the sentence completely about saying things right so if there's an exclamation mark at the end SPEAKER_79: yes you started reading this yeah that's a fair point it's not going to know that i'm curious you also uh were a researcher at google uh in deep mind yeah before going to open ai and before launching SPEAKER_01: your own yeah a lot of these uh language models were based on you know seminal papers on tensors and whatever um and a lot of the code base was open source or open source ish i guess in facebook's it was leaked in the case of open ai the original models were open source how much overlap is there in the fundamental technology at this point and how much is different if we were to take SPEAKER_83: you know uh the top five language models how much shared dna do they actually have how different are they SPEAKER_04: at their cores so uh everything is a transformer uh which is the architecture built by google in 2017 um SPEAKER_06: and everything is generatively pre-trained with language models so all that is the same the difference comes to what data it's being trained on um where open ai puts in a lot of effort there compared to other organizations um the reason methods lama models were actually really good uh despite not being as big as open ai's models is because uh the researchers there put a lot of effort into SPEAKER_01: curating the right data ah well explain what that means to lay people here who are wondering what you're SPEAKER_23: talking about yeah um so how these uh intelligent uh language models are built is you have this giant neural SPEAKER_06: network and you download a lot of data from the internet uh terabytes of data and you make these neural networks predict the next word given the previous words you basically train them to be great autocomplete machines and by virtue of doing that they become really good at reasoning and things like that now uh that doesn't mean that if you just keep scrolling the web and scraping every page and creating the data set uh you're going to keep getting smarter and smarter uh in fact you get smarter by like not training on junk and actually training on good quality data um and and now like um it also turns out that if you train a lot on coding uh like github and other data sets uh you develop these reasoning capabilities to an even higher level than not training on coding it's kind of like thinking about like let's say you have a kid uh you send the kid to coding or math competitions uh even if they may not become the you know the you i am a medalist uh they might end up being great analytical and logical thinkers in their life and that might help them in their life so that's sort of what happens with these lms and so if you pay a lot of attention to what data they are trained on that helps you a lot in terms of what you can achieve achieve with them later so the base core iq of these models will be much higher if you put a lot more effort into like curating the training data more carefully and open ai was ahead of everybody else there um google has all the data in the world but they didn't pay enough attention to this um and uh now like people have caught up they understood you know this is where they need to pay attention on um as for like who's really ahead right now i think it's uh open ai like with gpt forward yeah yeah much far ahead who can likely catch up there's one more organization called anthropic sure and like they are the closest number two and uh both these organizations were more SPEAKER_128: or less the same people uh like the people who train gpd3 were the guys who and then started SPEAKER_187: anthropic leader are you still using your personal phone number at work at your startup in 2023 stop SPEAKER_190: such a common mistake founders make but open phone has totally rethought every detail of what a business SPEAKER_32: phone should look like in 2023 open phone makes it so easy to do this and so affordable that you have no excuse and you really don't want your team using their personal phones for business why well it could get creepy people start texting people on your team it could be that they leave your company and the sales person has all of these text threads going with all your clients and they bring them to your competitor do you want to deal with this nonsense you don't i can tell you open phone is amazing because we use it our sales team our ops teams we use it daily we also started using open phone for angel summit communications it's rated number one on g2 for customer satisfaction and let me tell you those g2 rankings those are dogged battles if you win that you really have to be the best twist listeners love open phone my sales team uses it ops team uses it customer support uses it and uh you SPEAKER_05: know what's great about it you can create a shared phone number like we did for the angel summit with multiple employees being able to field those calls and texts and keep it all sorted it's affordable at just 13 per user per month but twist users are gonna get 20 off that already ridiculously affordable price for six months at openphone.com twist and if you got an existing number open phone will port it over at no extra cost head to openphone.com twist to start your free trial and get 20 off SPEAKER_01: so when you look at the open source community they seem to be really moving fast now correct uh metas llama models were leaked leaked maybe or maybe leaked on purpose yeah uh you think that you think that SPEAKER_72: story is true that it was leaked on purpose to jumpstart the open source community i wouldn't be surprised but you know yeah it was accidentally leaked accidentally on purpose there's some parallels with the SPEAKER_01: covid leaks there i don't know yeah it was yeah it was a accidental leak but they might have leaked it SPEAKER_90: because yeah but this was actually good like it was good for the world that this anthropic got leaked i'm SPEAKER_06: sorry that llama got leaked yeah llama leaking was actually really good for the world it you know i think i think i think i think i think it gave more power to the rest of the world in terms of what SPEAKER_39: they can do with lms outside of open ai or google or anthropic so that's my question these open source SPEAKER_01: models you've got a lot of people working on them yeah and a lot of people are not happy with how closed open ai has become um even i've started referring to it as closed ai so if they're super closed and open tends to win if we're sitting here in five years who do you think wins open source or you know google and SPEAKER_202: open ai with closed models who do you think is going to win yeah it's you if you pattern match open will SPEAKER_04: open tends to win that's that's kind of correct but uh there's like a catch here which is the next SPEAKER_06: big wins are not necessarily going to come from whoever is going to continue to train more uh you need some algorithmic efficiencies to to make use of compute even better and you need really good researchers for that and the best researchers are sort of like nba players and like they're taken by these organizations uh who pay them millions of dollars a year and then if these guys who are building the tricks for making these models even better are in the close organizations then they'll always stay ahead of the open right so then and if and if these organizations stop publishing these techniques um and these guys sustain these organizations are paid to stay there forever um it's kind of like closing the walls um so the only way in which the open source world can catch up is like they're like amazing researchers who kind of like work in organizations that are actively open source models and i think right now there's only one big org that wants to do that which is meta and so as long as metas in the game i think there's a chance for open source to sort of stay there and like you know win in the long run every other organization doesn't want to publish anymore SPEAKER_01: that's a problem um nobody publishing except for meta except for meta uh and i guess that google feels like they made a mistake publishing all this stuff and giving them yeah i'm sure they do like SPEAKER_23: they they missed out on the whole revolution it's fascinating yeah and i didn't ask you about the SPEAKER_213: paid version what if i if i choose to pay what do i get so there's this thing called copilot uh that's SPEAKER_06: more like an interactive search companion ah uh that it does the equivalent of uh hundreds of search queries for you not just one so you can ask it really complex queries like go pull me all of jason's investments and all startup sites is done and like you know uh at what valuations is done like prepare a table for me and get it back to me uh if the information is there in public for example i could only find the valuation you invested in uber but not on robin hood ah so then it'll come back to me and give me the information or you can say like give me the year by year revenue of aws ever since its inception uh i want to track it and growth percentage year over year and it's going to come back to you with information so it's almost like you're having a researcher at your disposal oh SPEAKER_79: wow that's great and when you say it's a copilot is it something that lives in my system try and mac or SPEAKER_04: windows or a chrome no it's going it's on the browser the uh copilot is just meant to be like a SPEAKER_06: companion like that the world uses the word is a companion for search and it's going to help you plan travel by products um prepare meal plans according to your preferences and if you integrate your ai profile with it it's going to give you like much more detailed recommendations uh travel itineraries uh web research like i wanted to know a lot about like when did read off and start making money in linkedin like you know they they took a while to like start making revenue what was the hypothesis in with scaling all all these kind of things that you you're like not oh coding as well like write me a piece of code for pulling up all elon musk tweets to where he he tagged jeff bezos in it like you get the twitter api v2 code you can copy paste that and like go and execute it so it's it can read documentation pages so that way it's more factful than what code you get from charge 54. so all these kind of things it's very powerful so what we offer in the paid version is unlimited uses of that not full like technically unlimited so more like 300 queries a day which is practically unlimited for most people um and then everything else is free so the way we're thinking about it is the free version grows enough that we can do advertising there and the paid version is for power users who want to like use it for work or um very complex queries that they seek but uh the free users get like 25 queries a day even on the copilot version so you don't have to pay if you don't want to we SPEAKER_120: just want regular daily users to stop using google and use our product uh i will be one of them i'm just SPEAKER_01: signing up for the paid version as we wrap up the episode here uh you're hiring so uh where can people SPEAKER_06: learn more about what you're hiring for uh we're hiring for ios and android mainly right now so ios engineers if you want to come and help build our mobile experiences please join us that's the most SPEAKER_01: important time here uh and i think you can go to uh perplexity.ai slash about and you'll learn SPEAKER_72: more yeah all right we'll see you all next time on this week in startups bye bye on behalf of the SPEAKER_228: producers and the partnership team thank you for listening to episode 1770 we'd like to take one more time to thank our partners crowdbotics get a free scoping session for your next big app idea at crowdbotics.com slash twist vanta get a thousand dollars off your sock too at vanta.com slash twist and open phone get 20 off your first six months at open phone dot com slash twist if you're looking to become a partner of this week in startups you can email hannah at hannah at launch dot co that's hannah at launch dot co thanks for listening