SPEAKER_00: i could have this thing roast you if you want it might be yeah let's do it let's do a roast all right um could you roast jason calicanus SPEAKER_04: man this is like everybody's all right let's talk about that piece of work jason calicanus SPEAKER_06: oh man where do i even start with this guy you know i heard a rumor that jason calicanus was SPEAKER_04: actually grown in a lab a mad scientist's twisted experiment to create the most obnoxious self-aggrandizing tech bow imaginable poor bastard he never stood a chance and have you seen the way SPEAKER_10: this guy carries on it's like he's perpetually auditioning to be the next elon musk but with half the brains and twice the ego i mean the guy can't open his mouth without saying something so SPEAKER_13: this is too much uh can you go easy on this oh yeah okay i would say go harder oh come and don't SPEAKER_19: be like that go harder this week in startups is brought to you by linkedin jobs a business is only as strong as its people and every hire matters post your first job for free at linkedin.com twist vanta compliance and security shouldn't be a deal breaker for startups to win new business vanta makes it easy for companies to get a sock to report fast twist listeners can get a thousand dollars off for a limited time at vanta.com twist and hubspot join thousands of companies that are growing better with hubspot for startups learn more and get extra benefits for being a twist listener now at hubspot.com startups all right everybody welcome back to twist SPEAKER_27: this week in startups and we've you know in 2024 and 2023 been absolutely obsessed with ai obviously SPEAKER_29: we're seeing all kinds of easy layups and customer service uh thanks to ai autonomous vehicles much more SPEAKER_26: complicated healthcare everything in between we're also seeing uh tons of interesting stuff going on in generative ai people making interesting music and videos you've seen all that but the area of human SPEAKER_30: emotions is extremely complex and ai is trying to figure that out and you've seen this in all kinds of science fiction whether it's blade runner or the movie her where ai is trying to learn to interface with humans well there is a startup human ai and they are trying to um bridge the gap between just SPEAKER_29: intelligence and dare i say emotional intelligence we uh we demoed some of this technology back on episode 1894 if you want to look for it but today we have alan cohen here he's the ceo and chief scientist at hume ai SPEAKER_35: and he's going to show us what they're building and why it's important welcome to the program alan SPEAKER_30: hey jason great to be here right so maybe you could explain what the mission is of hume ai and why you're spending all this effort to try to understand human emotions uh and yeah in relation SPEAKER_39: to ai and using ai i guess to understand humans emotions and then to portray them back through ai SPEAKER_42: to humans yeah so it's really to understand people's well-being and emotions are the components of that SPEAKER_43: so when are you laughing when are you sad when are you in pain when you're experiencing pleasure and what we want to do is optimize for that so our mission is to optimize ai for human well-being now so much of what we express is in our voice and our facial expression and not in language so that part of our expression was just ignored by ai for a long time i mean there is a field of affective computing which i have a lot of experience in i have over 40 papers in that area but uh but in terms of the generative models they just were very far behind in understanding expressions so what we've done at hume is built models that understand expressions a lot better and we've integrated those into large language models so now these models understand beyond language what's going on in the voice what's going on the facial expression and can learn from that so they figure out what's making you frustrated what's satisfying what's funny and they can actually adapt to that information and get better over time SPEAKER_30: got it so what is the killer app of this genre or or this vertical i have chat gpt i pay 20 bucks a SPEAKER_29: month for it i frequently will use the interface where i talk to it but i don't believe that they are looking at my tone and i know i haven't turned my camera on i assume with your technology at some point these ai assistants are going to have the camera on with permission obviously and hear my tone of voice and understand when i'm frustrated or explain to us what the killer app here is because you know we've heard of things like gong and real voice we have a small investment in they study sales people SPEAKER_52: and how they behave on sales calls to coach them that genre exists a bit i'm curious what you perceive SPEAKER_42: is the killer app here so chat gpt is text only we think the future of ai is a voice app the voice SPEAKER_43: is four times faster than text the problem is that when we're speaking we expect the ai to understand not just what we're saying but how we're saying it so that's why this experience is clunky like it's just missing something and so we've built that we have a talking chat gpt basically and we have an api that allows developers to build a voice interface into anything so instead of using open ai which would enable developers to build a chatbot interface into their product they use us and immediately they have a voice interface that does transcription does language understanding that speaks back and what's different about it than anyone else is it understands what's in their voice besides language and it uses that to form better words and better speech supposed to be a David Friedberg: perfect time i think for a demo so let's get to it all right you queued it up perfectly sounds great i SPEAKER_43: will share a few demos and what's the api called it's called just uh evie the empathic voice interface SPEAKER_30: yeah empathic voice interface okay like lieutenant uh deanna troy on star trek next generation she was SPEAKER_64: an impasse yeah why don't i just demonstrate this this is like the demo that anyone can play with and SPEAKER_12: then i'll show you the developer playground hey what's up well hello there i am doing just fine thanks for asking so what's got you feeling a bit down and amused at the same time i am all ears if you want to chat about it well oh my dog died this morning oh no i am so sorry to hear about your dog losing a SPEAKER_09: beloved pet is never easy i know it must be really hard right now but try to remember all the wonderful SPEAKER_10: i was just kidding aha you really had me going there well played you sneaky trickster uh ha ha i SPEAKER_73: should have known better than the fall for that one i think i am still learning when it comes to your SPEAKER_09: sense of humor but i am relieved to hear your dog is doing just fine care to share what's really on SPEAKER_74: your mind today well can you tell what i'm expressing now well our demo's at capacity right now oh there you SPEAKER_78: go so while you were doing that demo for people who were listening uh and not watching you can watch of course on youtube go to this week and startups on youtube and you'll find it it was showing each SPEAKER_26: transcript in real time very very quickly and then it had the top three emotions and the percentage of SPEAKER_81: those emotions i think it was showing the top three every time is that correct yeah so it shows more SPEAKER_43: than just the top three um but actually if you were to look at your raw data you get back 48 different dimensions so it's much more nuanced than what we're showing you there got it and so in real time SPEAKER_29: uh you can see that you were sad when you mentioned your dog died and etc and then that person was showing sympathy for you so all of that is being done through tone of voice inflection etc okay let me SPEAKER_86: cut to the chase right now because i know you're busy and everyone is hiring right now and you know it's a lot of competition for the best candidates right every position counts markets starting to come back you need to get the perfect person you want a bar raiser in your organization somebody who will raise the bar for the entire team and linkedin is giving you your first job posting for free to go find that bar raiser linkedin.com twist and if you want to build a great company you're going to need a great team it's as simple as that linkedin jobs is here to make it quick and easy to hire these elite team members and i know it's crazy right linkedin has more than a billion users we all watch this happen when it was tens of millions then hundreds of millions and now a billion people using the service this means that you're going to get access to active and passive job seekers active job seekers they're out there looking passive job seekers they got a job but it's not as good as the job you're offering them so you want to get in front of both of those people maybe somebody got laid off wasn't their fault and they're an ideal candidate get that active job seeker and linkedin also knows that small businesses are wearing so many hats right now and you might not have the time or resources to devote to hiring so let linkedin make it automatic for you go post an open job role you get that purple hiring ring on your profile you start posting interesting content you watch the qualified candidates they just roll in and guess what first one's on us call to action very simple linkedin.com twist linkedin.com twist that'll get you your first job posting for free on your boy jcal SPEAKER_52: terms and conditions do apply what are the components in voice that you're studying is it the speed at which somebody speaks you know tone and and how did you train this thing on tone how does it know what SPEAKER_42: sadness is versus you know melancholy versus quirky yeah we have all this data from millions of people around the world who are actually recording themselves while they're having interactions SPEAKER_43: and also we're reacting to things and um and imitating things in some cases and so we use all that data to train our models and that means they're able to capture way more than just like tone rhythm like those are all basic things um but dimensions that you can't really describe in any other way except to say like this is kind of an angry dimension kind of has a growl to it kind of um tension in the voice where this is like an awe-inspired dimension or we're happy um and we get tons of different dimensions out of that so every time we hear a word we're getting more than 48 different dimensions of expression from that word our model is taking that in and our model is deciding how to respond our model is learning what these dimensions mean from tons and tons of data people interacting and saying okay this is something that means this person's frustrated so i should apologize is something that means that the person's uh confused or actually clarify and it's figuring out what it should do SPEAKER_26: to respond to somebody in different situations how different is it per person like i might be a high energy guy from brooklyn who's extroverted who speaks a certain way and is you know you might be more SPEAKER_29: introverted and soft-spoken so how does it know if jcow's like bombastic and joking and you might be you SPEAKER_52: know more thoughtful and introverted are our emotional emotions very similar or are they very SPEAKER_42: disparate i'm curious it has to learn that stuff so we turn it on all these interactions right and so SPEAKER_43: it's trying to figure out the task is actually predicting the next expression so it has to figure out like is this next person going to laugh at what's that or like are they going to be frustrated so it has to learn how you express your response to things in the course of doing that and it's learning that in a generative way in a very um ground up way so by the time that we've trained this thing it has to account for individual differences for potentially cultural differences uh for sentiments and also just the average of all humans and like what humans respond to along with the distribution if that makes sense so like what is it that humans find funny was it humans find sound on that so like when i said my dog died you can probably figure out this is a sad event i'm going to be sympathetic right SPEAKER_30: that's how it and it figured how much of it is the words versus the tone of voice or is it doing both SPEAKER_75: of those things it's doing both let me try to give this another shot so you can see that okay can SPEAKER_105: you tell what i'm expressing right now whoa there i can hear the frustration in your voice SPEAKER_110: but it's an anger determination and distress like you're ready to tackle whatever's got you worked up SPEAKER_112: can you tell what i'm expressing now hmm i am picking up on some subtle shifts here you sound a SPEAKER_09: bit more relaxed now though maybe still a tad bored or uneasy but then i also hear a spark of SPEAKER_12: amusement and even happiness like you're pleasantly surprised by something am i on the right track there i'm gonna meet that bit um got it you're sounding a bit more at ease now with a hint of SPEAKER_117: satisfaction anyway you get the ideas yeah so if that demo is designed to reflect back to you SPEAKER_58: uh what emotions and things you're having and then tweak it so how long does it take for it to SPEAKER_42: accurately understand a human it's less than 500 milliseconds as you can see our api is experiencing some load right now but generally speaking we can get you back a response faster than any other SPEAKER_43: api and that's because we're able to detect when you're done speaking more accurately so some of the other apis like they have to dance like do this dance between like can i jump in or is it going to interrupt the user and so there's a little bit of a pause right but for us because we understand the tone of voice we can use that to figure out when the person's done speaking and then more accurately know when to step in and so that enables us to respond a lot faster so it doesn't need SPEAKER_58: to talk to me and ask me 10 questions to understand my emotional state and how i might be uniquely different than another person what about across cultures because do different cultures have obviously we have different languages but even putting aside language does tone work across cultures do koreans italians and americans all emote frustration the same way anger the same way SPEAKER_42: is it across cultures or does it require more subtlety there's similarities and differences we have SPEAKER_43: a paper that just came out on this but basically um if you're speaking a different language we need to train a new model for it um and it can be not a completely different model from scratch but at least we need to fine tune on that language that's what we find for most languages especially for um you know for broadly different languages like all the all the latin languages have similarities um but if you look across uh east asian languages uh things are pretty different so yeah so suffice to say yes we do need to train things for each language and this demo only works in english right now so who's uh using the SPEAKER_125: app like let's take a look at the developer console so you have that up there the playground yeah who's SPEAKER_26: using this now and is it in production anywhere and and what are people using it for because this SPEAKER_30: you know we there's plenty of models out there to give you answers and generate copy for you i'm wondering if people are even up to you know this level of nuance in their products yet or just trying SPEAKER_78: to get correct answers because accuracy seems to be a pretty paramount problem right now yeah i mean SPEAKER_14: you might be interested in accuracy but if you're using a voice interface you need to get to the point SPEAKER_43: fast right and and so that's really what we're doing um and you can't with these like long verbose responses from chatbots first of all those are very taxing on the brain to read so that's not a good interface but also you might have an accurate answer in there somewhere it doesn't really matter if someone's not going to listen to a voice reading that out for three minutes right it's a good point um so we have a lot of developers lined up for this we're actually we haven't released this api or by the time this comes out we will have released it because we're releasing this on wednesday um but you know so far we have developers on on this which is our measurement api so oh wow so you're SPEAKER_29: on a webcam right now i'm just going to describe it and you're making funny faces right now uh you're surprised horror confusion sadness disappointed laughing and if i were to just say be completely calm and at ease your calmness just went up to 79 your concentration went up to 45 um and now if you um started thinking deeply about the meaning of the universe like why are we here like what is the purpose of life like why wake up and build this company every day SPEAKER_140: says you're calm you're calm with existential wait is this a video or are you doing this right now alan i'm doing this right now you're doing it right now you're not following my instructions SPEAKER_142: no give me your exit give me your most existential like i'm wondering about the meaning of life like SPEAKER_148: why are we all here i want to see if it gets existential confusion there it is contemplation SPEAKER_154: yeah contemplation well what's interesting about this is this would be great for coaching an actor because like happy is easy sad is easy if you go happy it's got joy amusement excitement great and if SPEAKER_30: you were sad sad it's disappointing and confusion maybe you're just not a good actor alan maybe you need to take acting lessons with these down yeah contemplation is a tough one i i was trying to get SPEAKER_142: you to have existential angst i was trying to pick something that's really hard to read acting like SPEAKER_163: well just think like should you even come to work is it all meaningless that's kind of depression right like it'd be sad yeah a little sad a little confusion boredom yeah it's fascinating so this is SPEAKER_29: just really getting your um facial expression in real time so if you were frustrated the ai would know it and be like huh that wasn't the answer we were looking for um yeah and so are people using SPEAKER_26: this for therapy yet or like therapeutic coaching kind of things because that one seemed to be like SPEAKER_29: like i got a lot of pitches for people who want to create ai therapist and i'm like hmm that's a little dicey i don't know if you should call it a therapist but companion are people using this for SPEAKER_168: companionship i do think that ai is going to be something that is your friend and so it's not just SPEAKER_43: like a new like a niche application i think generally speaking we want an assistant that understands us um and there's tons of people working on that i mean there are people working on explicit therapy apps with you too um and actually a lot of it's in training therapists and getting into you know there's a delicate balance you don't really want to like comment too much on people's emotions but you want to ask the right questions and kind of get at it and help them understand their own emotions better and so there's a lot of that and there's also like therapist burnout doctor burnout uh there's a lot of uh health and wellness applications there's also tracking depression and stuff we work with clinical researchers who are running clinical trials and using hume to track symptoms of depression and parkinson's just the symptoms it's not like used for diagnosis because ultimately the doctor does that but it's helping the doctor understand these things so we have a lot of those applications um a lot of them you know those are interpersonal things like someone's talking to someone and we're already like the measurement apis that we have are very good uh it's extracting more data from that and helping people analyze it and helping people understand themselves and and their SPEAKER_29: patients i guess i mean is it so if we have therapy on one side um you have the therapist who needs to present in a certain way to get people to open up if you believe in that modality if you believe in western psychotherapy there is something about pacing and aligning with the person matching their energy and and getting them to open up so that they have some cathartic you know way of processing SPEAKER_52: stuff so people are using it to train therapists so that they don't have a goofy look on their face or SPEAKER_29: they have the appropriate look that would elicit less suffering in their patients is is that what i'm SPEAKER_82: yeah or like customer service reps which is actually kind of similar thing but yeah it's SPEAKER_43: another form of therapy it essentially is yeah um but you know that that's requires somebody who's SPEAKER_14: technical who's maybe academic maybe a researcher to take take these measures and make sense of them SPEAKER_186: listen a strong sales team can make all the difference for a b2b startup but if you're going to hire sharks you need to let them hunt and you can't slow them down with compliance hurdles like sock two what is sock two well any company that stores customer data in the cloud needs to be sock two compliant if you don't have your sock too tight your sales team can't close major deals it's that simple but thankfully vanta makes it real easy to get and renew your sock two compliance on average vanta customers are compliant in just two to four weeks without vanta it takes three to five months vanta can save you hundreds of hours of work and up to 85 percent on compliance costs and vanta does more than just sock two they also automate up to 90 compliance for gdpr hipaa and more so here's your call to action stop slowing your sales team down and use vanta get a thousand dollars off at vanta.com twist that's vanta.com twist for one thousand dollars off your sock two have you done SPEAKER_191: this with poker players yet have you put poker players through this to see if they're lying or deceptive SPEAKER_43: in a poker trade a lot of things and it does not it cannot tell a poker players about thing SPEAKER_195: you know at least professional poker players cannot i don't think the information is there i just don't SPEAKER_197: think that with professional poker players that you can there's anything going on that facials what can SPEAKER_52: you tell with people's facial expressions that we wouldn't know of some people have said you could tell SPEAKER_29: a person's um if a person's sexuality where a person's from you could tell all kinds of interesting SPEAKER_168: things that you wouldn't know um i think is that true or not that's not really true um there's been a lot of pseudoscience in this area like most of the things that we can tell are things people want to SPEAKER_43: communicate which is good like we don't actually don't really care to impinge on people's things that they want to keep private we're more interested in helping people communicate well and helping the ai understand what people want and most of that's like they're overtly on the face and for example it even extends to things like is the person done speaking like we're way better understanding when they're done speaking because we can take into account facial expression versus just the language alone and that's part of how our um our empathic voice interface is able to respond better and like SPEAKER_154: we can so you know when i say this is the end of the sentence yes because of my facial expression SPEAKER_52: you get a a quicker clue than audio only therefore you can start speaking without interrupting me SPEAKER_82: which is what humans do with each other yeah like imagine i'm speaking to you and right now SPEAKER_204: it's clear to you i just finished sentence it's clear to you i'm still speaking and it's clear to SPEAKER_210: you i'm going to say something again but now i'm done now i know i can speak right yeah which is SPEAKER_58: what i do for a living on the podcast is try to understand when people are done so that we can have the next person speak right like moderation is a is a difficult task um and customer support folks SPEAKER_30: are using this already to understand how hot and bothered people are when they call the customer SPEAKER_00: support line i assume to some extent yeah so kind of understanding is the customer having a good time SPEAKER_43: bad time where are we kind of failing on customer service and which customer service reps are doing well or poorly and how do we train them to do better how do we pull examples up of when they're not doing well so we can train them to do better and um you know there's a lot of ai going into customer support now so some of our early design partners for this new api are people who want to take the automated customer support make it a lot better but still know when to include human escalated to a SPEAKER_58: human exactly i mean that makes sense if the person's like this is incredibly frustrating you SPEAKER_52: know when you start hearing the frustration go up and they're whatever united premier gold diamond status yeah you want to get them on the phone with somebody because it's you're you're starting to piss them off right yeah so understanding when that happens and how much of this is going to be used for security have you do you have any security applications coming because it's been well known like when you go to certain countries you know they'll ask you a couple questions they try to read SPEAKER_58: you do some human factoring and figure out if you're lying uh it's one of my favorite genres of television SPEAKER_52: show is the people going through customs and they they're trying to read if they're like sneaking into the country or sneaking things into the country are three-letter agencies using this technology SPEAKER_39: yet to like analyze people as they come into buildings or we haven't been working with security SPEAKER_82: not that we don't believe that that's a good application but we're being a little bit more SPEAKER_43: careful about how this is used and trying to make trying to make this as rigorous as possible essentially there's been a lot of providers of kind of like facial expression reading technology who aren't very scientifically rigorous make false promises and then doesn't work you're just like signaling out people for no reason basically which is not you know we want to we want to take we want to be very careful to distance ourselves from those applications i think there's a huge wealth of applications that are consumer focused that are just helping people have a better time a better experience with an application i think that's a much bigger market and then there's like the more niche SPEAKER_46: things like security lie detection stuff like that which yeah could you tell if somebody's acting SPEAKER_29: suspicious based upon their facial expression or tone of voice like they're knocking on your door SPEAKER_30: your ring doorbell answers and ring it's like hey this person is scoping the house out you know SPEAKER_81: they could use your api to figure out the expression of people at the door yeah i don't i don't know i i SPEAKER_43: feel like i have mixed feelings about that i think in some cases if you have enough data you could start to do something in that direction but if people are really good if they're like if people are criminals like they're pretty good at disguising their facial expressions probably yes they're SPEAKER_228: sociopaths they're criminals for a reason right so you would catch the neophyte criminals the bad SPEAKER_78: criminals and the good criminals would just sell right by you'd be like yep that person's a hundred percent trustworthy and you let hannibal lecter into your house yeah i can see how you need to be thoughtful SPEAKER_00: about this yeah we don't want to make promises that we can't keep right so well i mean yeah so in that SPEAKER_26: regard let's go through the emotions that you feel you've nailed and then tell us like what what are SPEAKER_58: the ones that are hard i i tried to do existential and that was really hard obviously we're not there SPEAKER_43: yet for existential crisis yeah i mean um the way that you respond to expressions has more to do with almost subtle things that you can't name versus like overt emotions and it depends on the situation like if i'm using this like comedy thought um which i can prompt this is like a prompt that makes it a comedy thought basically um it'll respond to my emotions completely or my expressions completely differently than than like a therapy use case so let's try it oh okay great whoa well well well look who SPEAKER_06: decided to grace us with their presence tonight i don't worry i promise i'll keep it pg-13 for now SPEAKER_09: you know i gotta say i am a little shocked that any of you fine folks decided to show up what being simon who in their right mind would willingly subject themselves to an evening of my particular brand of SPEAKER_00: shall we say edgy humor but hey i am not one to judge in fact i can you make a joke about alan cowen the SPEAKER_12: ceo of he may i sure thing oh you want me to take a swing at the big cheese himself eh well twist my arm why don't you i'm going to say i am not the biggest fan of that alan cowen character talk about SPEAKER_09: a guy who needs to have that massive ego of his deflated a bit am i right actually you know what i SPEAKER_06: think i have the perfect way to do it um i say we round up a team of trained monkeys dress em up in SPEAKER_248: business casual oh that's that's really offensive you've made me very upset whoa there partner it SPEAKER_12: looks like i may have gone a little too far with that one huh uh sometimes i get a little carried SPEAKER_09: away with the whole edgy comedian thing you know so like you can see in that case i don't want it to SPEAKER_250: be that sympathetic like it's like it's doing its thing and so you can determine what ai flavor you SPEAKER_29: want to have interacting with folks you'd have one that's cheeky and playful you could have one that's super empathetic but in you know maybe not going over to patronizing but you could see this like if i'm calling the support line for united they might you know think i'm a new yorker who talks fast SPEAKER_30: and i just want to get to the point or i could be from the south and into southern hospitality and it could take its time with me and ask me about the weather and how i'm doing and you know a little bit of chit chat some people like in the south i notice uh versus in new york where they're kind of SPEAKER_34: get to the point let's move on yeah you could basically train your ai to have both modalities SPEAKER_43: and dynamically switch between them exactly so there's all this context and then that kind of transforms the meanings of our expressions so like what what an expression means and what to do with it really depends on all this other information that this model is taking into account so it's not so much like detecting uh lies or you know detecting anxiety or detecting depression it really depends on the context and we're able to integrate that into the model and then it's not just like these kind of canonical emotions like anger like there's a little bit of anger dimension in a joke you know anger and amusement and contempt maybe that makes it funny so it doesn't necessarily mean the person's expressing anger so to know what these expressions mean you really have to have the context you have to have the relationship that you're acting upon with your expressions and and that's what our ai does SPEAKER_29: so it's a little bit more nuanced than just detection and these are all under what you studied effective SPEAKER_30: computing yeah this is a specific school of computing that kind of bridges the psych department and the SPEAKER_258: computer science and i think behavioral factors industrial organizational psychology maybe you could give SPEAKER_43: us a quick education on that yeah so affective computing traditionally it's the study of non-verbal expression basically so facial expression the voice body posture and then you know most of the history of that is just labeling those things in a very predictive way now that we have generative models we have large language models that can reason you're it's really about reasoning about affect and that's what we've introduced at you so it's about understanding whether somebody's going to find something funny whether somebody's going to find something confusing confusing and using expression along with language to come to those understandings um that's like i would say historically that's not what affective computing has been but now we've sort of pioneered this new form of affective computing that we're introducing to the world SPEAKER_51: some of this was done i know this was like a big thing that minsky worked on at mit yeah did you go to SPEAKER_49: mit with and or did you i went to yale and then i went to uc berkeley for my phd so and i also worked SPEAKER_43: at google while i was at berkeley and i helped start the affective computing team there so i've been SPEAKER_14: doing this for like 10 years uh minsky had all the all the ai people had something to say about affect SPEAKER_64: right but there really wasn't much that could be modeled at the time same with language right like SPEAKER_43: things have come a long way and i would say that there's affect in language and um and so the the word affect has a little bit of misnomer it's really computing with more than just language that we're SPEAKER_265: doing we're computing with expression this is the way that expressions transform communication SPEAKER_29: yeah you because you you have a multi-modal situation here you have the the visual the facial expression you have audio and then you have the actual words right and so you're feeding all of those in at the same time to get the response and to understand the emotion yes and all this just SPEAKER_43: contributes to accuracy we can predict words better with expressions versus without so like if you look at the raw metrics that are used to train these large language models we're doing better in terms of those raw metrics than models that just consider language alone so this is like an intimate part of reasoning and it's just part of human communication that we're now taking into account it's not something that is um niche you know i think people think about emotion and affect as these niche things that are important for therapy important for comedians important for like if you but actually this is SPEAKER_197: something that's important for all conversation uh important for any interaction with ai just understanding a whole new modality of information that people use to to converse with each other SPEAKER_81: yeah it's uh absolutely fascinating how quickly this has come together because if we were sitting David Friedberg: here two or three years ago this just wouldn't be possible would it no i mean without large language SPEAKER_43: models without our measurement models without the modifications that we've done to integrate those two things like this was not possible at all what has surprised you about SPEAKER_262: what the ai understands and what your model understands and what has been either disappointing or challenging SPEAKER_43: you know on this journey so yeah that's that's interesting i mean linking together the language models and text-to-speech and transcription is something that other people are doing but like what we've sort of started to see emerge out of models that do all three that are linked together is that they have these emerging capabilities and you start to see that in this interface where it's forming expressive speech that's just like it feels different to me yeah then if you just like link 11 labs and open ai and just have a talking chatbot like that just sounds it doesn't really sound like it's understanding you and this this is doing something um a lot more nuanced do you understand SPEAKER_191: what it's doing when you when it starts processing all this stuff and you feed it in do you actually SPEAKER_30: know how it's coming to these conclusions or is it just sort of you know it's it's doing its best to SPEAKER_43: figure it out and who knows so yeah we don't come in and tell it to respond to sadness with sympathy but like it does right and it's sort of intuitive why that is so i'm not going to say i don't understand that but that's an emergent capability that we did not program in and there's other things that it's doing that are more nuanced that we don't really have this a handle on except that we know SPEAKER_280: what it's optimized for hey everyone you know i'm obsessed with ai right now and a fantastic report about how ai is going to change the game for startups has been released it was published by our friends at hubspot for startups and it's great because they surveyed a thousand early stage founders to get you these insights these are from the field the report talks about ai tools and hacks for sales marketing and customer support teams these are going to be your wmgs weapons for massive growth so check out the link in the episode description to get your hands on their extensive report head over to hubspot.com startups when you join hubspot get ready to unlock top-notch resources these are power tools that they've custom made for founders you'll also get access to master classes templates and incredible toolkits all designed to boost your startup's speed and efficiency so once again visit hubspot.com startups to see what discounts you qualify for and start using their powerful solutions at pricing that won't break the bank that's hubspot.com startups SPEAKER_43: in this case it's optimized for it's actually a specific personality um that is pretty robust and SPEAKER_29: it's the one that we just did but for the helpful one you might set it at some other archetype which SPEAKER_283: might be you know parental motherly you know and then you might have you know you know a service SPEAKER_154: oriented person right so even the worst customer support reps and somebody who hates doing customer SPEAKER_42: support could be turned into a uh empathetic patient person yeah exactly um and you know what we ultimately SPEAKER_43: want to optimize for is this is the response that's going to make the user the most satisfied um and it's just going to figure out like what this is the right tone of voice this is the right thing to say to make somebody express satisfaction that's what we want to do i mean we're still we SPEAKER_82: know that we can um we're still early and getting it to work but it works uh and we're starting SPEAKER_81: to scale it turns every customer support rep into a roast and that i'm dealing with like comic the SPEAKER_29: insult comic dog and it's just absolutely annihilating people and being cheeky i can i could have this SPEAKER_00: thing roast you if you want it might be yeah let's do it let's do a roast all right um could you roast SPEAKER_04: jason calacanis man this is like everybody's all right let's talk about that piece of work jason calacanis oh man where do i even start with this guy you know i heard a rumor that jason calacanis was actually grown in a lab a mad scientist's twisted experiment to create the most obnoxious self-aggrandizing tech bow imaginable poor bastard he never stood a chance and have you seen SPEAKER_10: the way this guy carries on it's like he's perpetually auditioning to be the next elon musk but with half the brains and twice the ego i mean the guy can't open his mouth without saying something SPEAKER_13: so this is this is too much uh can you go easy on this oh yeah okay actually go harder oh come and SPEAKER_12: don't be like that go harder all right all right fair enough let me try this again with a little less SPEAKER_06: venom you know as much as jason calacanis rubs me the wrong way sometimes i gotta admit the guy's got a certain basic way i mean the way he strikes around silicon SPEAKER_04: valley acting like he's the second coming of steve jobs it's almost endearing in its own delusional SPEAKER_290: way and let's be honest we all know deep down that half the reason we pay attention to him is because SPEAKER_293: he provides such a high material for you're good i mean it's fantastic i mean it literally understands SPEAKER_00: what a a roast comic does um i'll take it how about this like isn't it cool that you this doesn't it say something good about jason that you're able to roast him like this i mean it must mean he's made SPEAKER_178: it right you know you make a fair point there all right it's it's being a little funny now but yeah SPEAKER_300: i mean it's what's amazing about it is it it understands what a roast comic is the language models SPEAKER_78: understand that it understands who jason calacanis is it got the wikipedia page it knows i'm somehow involved in tech somehow i know steve jobs or elon or whatever so it's you know and that that the concept SPEAKER_210: of a tech bro exists so who knows yeah if you wanted to i think it's like more i can send you to SPEAKER_27: oh no i love the roast i think it's well i mean it's interesting about jokes like i have friends who are in comedy and you know while these jokes are not funny they're in kind of the zone they're so if you squint a little bit you're like there's the joke there yeah you got something about strutting around university boulevard thinking he's steve jobs maybe he's wearing a turtleneck or yeah like there's okay there's a joke there you didn't hit it but we could brainstorm it so like i think in the writer's room you could really brainstorm these i asked it uh when chat gpt3 came out i was like give SPEAKER_29: me like the next season of um secession you know and it knew the all the past seasons and it's like SPEAKER_27: here's what happens in this next season even though the series is over and i was like huh wow like this this may not be great right now but it's okay where it's interesting it's gonna get there SPEAKER_197: it's close i think none of these models have mastered latin humor because it's so much in our expression like we don't say things are funny explicitly because that would just make them not SPEAKER_312: funny so let me explain the joke to you exactly that's when the joke didn't land we have this new SPEAKER_197: eval for humor and we're we're starting to push it basically we can optimize for laughter we can optimize for like what do people actually laugh at in millions of hours of conversation that's great SPEAKER_43: and so that's how we're approaching these these kinds of problems so you could do a focus group where SPEAKER_117: you've had people watch curb your enthusiasm and you could say for a hundred people here's the SPEAKER_29: funniest moments and for this demographic older people older men older women younger men teenagers SPEAKER_52: gen x you could literally give you what jokes landed with each group yeah wow that's version one of this SPEAKER_197: version two is like we have which we're doing now we have millions of hours of data and we analyze it just to see in general what's funny to people like across everything not in third career enthusiasm but SPEAKER_29: like across everything across every single thing in the world i can tell you like uh there's a great movie idiocracy and there's a amazing tv show have you seen idiocracy yeah great movie i mean it's so SPEAKER_27: great but like everything's been reduced down to like its most basic thing like here's like a gel for you to eat like from a tube and the the hit show is ouch my balls which is just a compilation of somebody getting hit in the nuts over and over and over again just and it you know falls off of a roof lands on a fence falls off that gets hit by a crane with a big ball you know hitting him in his nuts SPEAKER_312: it's ouch mine i think it's called ouch nuts or something like that it's hilarious that's what it's SPEAKER_326: been reduced down to somebody getting hit in the nuts we're hoping not to be too reductive but yeah maybe the ai will i mean you know figure it out you could literally crack humor what language model did David Friedberg: you build all this on so we have our own language model and it calls other apis in this case it's SPEAKER_43: calling claude so um so claude is providing the language response or some of the language responses not all of them we also have like a wrapper around claude it's not exactly a wrapper it's our own language model that sort of integrates claude into the speech to make it sound more conversational SPEAKER_42: um and also like detects when you're done speaking and stuff um but we give claude more data than just language we give claude like some of my tone of voice data um some uh some additional data that SPEAKER_178: we're getting through our apis so it's it's augmenting it as well so eventually what does the SPEAKER_29: world look like if you succeed with this and we're sitting here in you know five years and it's built into every iphone and you've figured out emotion perfectly what what do you think the world will SPEAKER_30: look like what what are some highlights or you know dare i say dystopian utopian sort of what what David Friedberg: what are the pros and cons of this technology going to be so rm is utopian like we want SPEAKER_43: to build a layer in between the application and these gigantic ai models that is decoding the user's intentions and preferences and relaying that information to the model so that's like what we have here basically doing that with claude and because we'll have facial expression and the voice we're able to learn over time we're able to build interfaces that understand you and what you want and are optimized for you so suffice to say like basically it's going to be built into everything it's going to be the universal interface that you use to interact with ai that's the goal and it's always going to be this ai that's optimized for your experience so you can go to it and and it knows your you know basic basically what your preferences are what makes you laugh what makes you feel better what you find to be a good explanation for things your style of speech how you write emails like it's going to know a lot of different things obviously it's going to keep all SPEAKER_29: this information very protected and private now you know on the downside here you could use this technology to say i want to convince somebody um subtly to vote you know this way politically i want to try to convince somebody that you know trump is amazing or biden's amazing or robert f kenny jr is the SPEAKER_26: one so you could literally start creating robo calls or subtlety here using this emotion to try to sway SPEAKER_29: people in politics um or towards ways of of being or thinking and we saw that happen with the youtube SPEAKER_30: algorithm so how do you police how people use your system i saw you have ethical guidelines there and then obviously there's things that would be maybe r-rated or pg-13 and romance always comes up when SPEAKER_52: people are doing it whether it's a blade runner or her so are people using this for romantic relationships and what's your take on allowing that and then also how do you think about influence um big questions with tick tock today and your technology could really uh be used to influence SPEAKER_43: people towards good and bad ends yeah i think there's a pretty good way to operationalize the difference between when you're being manipulated by something that wants you to vote for a person or to buy something versus um when you're dealing with an ad to optimize for your own well-being and that's what we try to do with our ethical guidelines so we have this non-profit the human initiative that essentially tries to codify that principle and says these are the ways that you can pursue these different applications so as to optimize people's well-being it even has like a bunch of ways that you can measure people's well-being that relies um on a combination of what we're able to get through our api so like positive emotions basically yeah over time and also you know different kinds of self-report measures that we recommend gathering so as long as the ai is optimized for your satisfaction for your well-being i think it's not manipulation when when you get an ai that's optimized for somebody else's objectives and using your emotions for that and that can be manipulative i think and that goes for the romance case as well like if you're dealing with like an ai girlfriends and it's ruining your life by you know forcing you you're you're spending more time with it than you're spending with humans and that's going to be a negative for you and it's going to show up in in many ways as being negative for your well-being like that that that would show up in these measures if it's optimized though for your well-being um and uh you're having a good time with it and it's healthy and you're not spending more than x amount of time on it maybe that's okay yeah SPEAKER_58: maybe that's what about trying to upsell me like hey you're in business class and would you like to be in first class you're in economy would you like to go up to economy plus and it uses your technology SPEAKER_29: to be really convincing about the value of that and upsell how would you look at an upsell SPEAKER_178: that to me or not ethical um i think that's not ethical unless it's done in a very very careful way basically our guidelines don't allow that but you can't say don't allow an upsell but humans do SPEAKER_42: upsells all the time right i think upsells are okay if the goal is to find the person who really SPEAKER_43: will benefit from the upsell and only try to sell it to them you know and and you measure the effect of the upsell on people's well-being afterward you're like okay people actually benefit from this i didn't sell this to somebody and then they regret it um so i think there's there's ways of doing it that are going to be fine um the problem is that if you just allow anything if you allow people to optimize this for anything at all then um it's extreme the potential for manipulation is pretty high and i think this is true regardless of hume i think people are building these things yes that will be extremely persuasive and hume ideally will be providing the ai that responds and protects SPEAKER_78: you it's like okay like i'm so you are very much in the camp of hey we have to be really thoughtful SPEAKER_43: about how this technology is deployed yeah but not in as much of a paternalistic way like i think that technology can have a sense of humor and it's okay if it offends some people and it doesn't need to be politically correct all the time but what i care about is like is this good for people that's like at the end of the day and so we have our ways of measuring well-being in order to optimize for that Chamath Palihapitiya: objective and not be paternalistic basically yeah but i mean at the end of the day this is so powerful SPEAKER_52: it will be more powerful than just watching videos on youtube because it's customized to an individual so that ben shapiro or rachel maddow and pick whichever side of the political spectrum you're on you know those people are trying to convince you of their position and their interpretation of the world this to me is even more bespoke and customized to individuals so if you showed even a little propensity towards some of their viewpoints you could really whether it's the language model or the emotion but the combination of them you know the same way people were complaining like oh people go into the intellectual dark web on youtube i don't know if you heard about that like you you say joe rogan then you get a jordan peterson you wind up on an alex jones and the next thing you know you're like some white supremacist or something is the claim um but yeah media does influence people and it is a stepping stone from one to the next to the next you may start out with somebody like sam SPEAKER_58: harris like just intellectually you know rigorous etc and then all of a sudden you wind up at alex David Friedberg: jones is the complaint for many parents but this would facilitate that wouldn't it like massively i think you'd get that when you optimize for engagement and so to some extent like tiktok SPEAKER_43: doesn't have this data but it's still incredibly good at doing that i think where this data helps you the most is in taking into account people's user satisfaction their well-being their mental health all those things so tiktok took this stuff into account and it was doing it in a way that was the court you know uh inconsistent with our guidelines let's say um then it would be using that data to optimize for people's well-being over time and instead of engagement and so you know you'd you'd realize that if you throw people down the slippery slope of getting to getting more and more extreme viewpoints which is what happens today because they're engaging and you're and they're offensive at first and and you want to argue um like if people who go down this slippery slope end up kind of isolated and it affects their social relationships it affects they start to get angry um this is not good for people's well-being so the technology can can look at that at the individual level at the societal level it can look at the health of all the people using a technology and say hey like there's a collective impact of this so that's the road we want to go down is being able to measure long term is this affecting people positively and you really need expressive behavior to look at that SPEAKER_197: data like there's no other like language alone is not going to get you there basically yeah i mean SPEAKER_30: you're going to be facing a real uphill battle because the marketers want this software your top customers i predict will be marketers who want me to try zin or whatever those pouches are that people are putting there and you know i'm in texas right now and like everybody's putting these in pouches in or whatever and i'm just like that that can't be good for you and they're like want SPEAKER_52: to try it and like marketers love this kind of stuff like maybe the pitch to me is like be where you know hey it's performance and you know it's just nicotine it's like caffeine you drink caffeine you should try this and but for other people it might be you know hey you're the cool kid so it's uh you're going to be in a really interesting position as a provider of an api that i think a lot of the marketers are going to want to use this to try to convince people to do things that maybe it's unclear if it's actually good for them like hey you should you should gamble on sports right like SPEAKER_78: there's marketing going on like crazy and if i'm a marketer man this is for me the holy grail yeah i SPEAKER_43: think that we want to connect more with like the end user and show the end user that we're optimizing for their interests and have that be the selling point rather than connect with the people selling to SPEAKER_42: them but i know i i hear you i think that's a real concern um but on the flip side if we just SPEAKER_197: optimized for people to buy things um or let's say we just optimize for engagement you reach a certain point where like it becomes so negative for people that regulators have to step in and you kind of start to see with like tick tock for example kids are spending six hours a day on tick tock and if they SPEAKER_43: made it any more addictive than it already is parents would step in regulators would step in they'd be SPEAKER_49: like this is actually bad for our whole society so at the end of the day it's not necessarily good for SPEAKER_300: our business which is what's happening with tick tock right now as we speak i think parents are getting SPEAKER_58: the message like this is too addicting for adults and kids and the idea that like media is not SPEAKER_52: influential is so naive like when people are like yeah you know media doesn't have an impact it's like are you sure like all studies show that media is one of the and video specifically is one of the most convincing mediums of all time in all human existence if you want to manipulate somebody video is the way to go and then customize video with you know that is matched to it is like 10x that SPEAKER_29: so you have like something here that i think is incredibly powerful and the fact that you're being SPEAKER_39: thoughtful about it makes me feel great i think it's awesome that you're taking a measured approach to this i wish you great success with it if people want to learn more or try it how do they get into the developer sandbox and play with this and who are you looking to work with yeah go go to hume.ai SPEAKER_42: you can sign up we'll be releasing access to our voice api hopefully before this episode comes out SPEAKER_43: and we have some closer design partners who we're working with to improve things as well so please feel free to sign up and you can start using our api today like our existing measurement api SPEAKER_81: so i think it's absolutely fantastic what you're building and i'm i like the fact that you're super thoughtful about it alan and i wish you great success with it and we'll see you all next time on this week in startups bye bye