Live from DevDay — the OpenAI Podcast Ep. 7

OpenAI · Beginner ·💻 AI-Assisted Coding ·9mo ago

Key Takeaways

The OpenAI Podcast Episode 7 explores how startups like Cursor, Abridge, SchoolAI, and Jam.dev are using OpenAI models to transform industries such as healthcare, education, and coding, with a focus on AI-powered tools and platforms for improving productivity, patient care, and software development.

Full Transcript

[Music] Welcome to the OpenAI podcast where we're live from OpenAI Dev Day. Here sitting with me from school AI is Caleb Hicks. Caleb, hello. >> Hi. Thanks for having me. This will be fun. >> So Caleb, you're working on tools for helping educators and helping people basically in the classroom understand progress of students. >> That's right. Yeah. So, first off, what was your reaction so far to Dev Day? >> Uh, a ton of fun. I think makes it uh a lot of things to be excited about that help us build, but also help students and teachers be more creative as well. So, that'll be fun. >> So, what have you been working on over the last year? What has changed with AI that's accelerated what you've been doing? >> Oo. Um, I think probably the biggest advancement over the last year for us, so we we uh put AI in students hands. That's the main uh the main thing that we focus on is safe managed uh AI that can act as kind of one-time personal tutors for students. >> Mh. >> And so probably the biggest change from open AI has been model progression. I think we get two advantages from that. One is uh significant leaps in intelligence >> uh and the other one is uh you know improvements in cost because we are working with uh an industry that isn't known for paying big dollars for software. Uh it's been important for us to be able to manage students using this in a cost-ffective way. So those have been the two areas that AI progression has helped from our work. has been a lot of uh orchestration which I'm sure we'll talk about a little bit just just getting different AI agents and models to work together uh for the best outputs for students in particular. So a couple of the releases we saw today, one was the agents SDK, you know, and you've talked about that. >> How much have one tools changed the ability of one to work faster into the scope of what you find is capable now? >> Yeah, I think we're seeing teams across industries work way faster and building better software because they've got kind of this always on expert hyper senior uh engineer next to them that they're pair programming with, right? And uh so we see that with our teams as well and that just allows us to to build better software faster uh and get it in the hands of teachers and students which is what we're here to do. >> What has been the biggest shift you've seen in talking to educators or people like that in regards to AI in general? >> Yeah, great question. So every teacher, school, and district is on a very similar journey. It starts with permission, right? Two and a half years ago, it was everyone under the sun was just banning AI altogether. Uh we've we've moved past that into productivity as teachers and and school leaders realizing, hey, this helps me in my job. >> Uh the leading schools are starting to get into that that really important spot, which is recognizing that every student has to know how to use this stuff. M >> if you're going into if you're graduating from high school and you're competing for colleges or jobs and you are don't know how to use AI yourself um you're at a severe disadvantage. So um most people are now orienting to like yeah we have to teach this. Uh we think there's a couple of special steps beyond those two uh where where you get to better support students. The AI tutor in every kid's pocket, right? But we think that is it's got to be classroom connected. It's got to know what you're doing in class and it's got to know um you know where you're trying to go and help you move that direction. Um, that last step we get really excited about is how you can put AI to work with teachers, uh, families at home, school leaders, kind of the system at large, uh, to really make school awesome for students. >> Could you tell me a little bit about the kind of stack you're using from both the teacher facing, student facing, backend or what you're working on in that regard? Yeah, from from the product side, we have kind of three different uh parts of the product that the that students, teachers, and school leaders use. So, there's uh kind of just a a pretty basic uh AI assistant, right? Um the the the GPT rapper as they say, right? Uh but but tuned to use cases in schools. I think a big thing that we felt like was really important is teachers should never have to become prompt engineers, >> right? >> Uh so >> you know I was a prompt engineer. >> Yes. Yeah. So so we do a lot of that uh kind of extra orchestration uh to we essentially enrich every prompt that that a teacher writes to get better output for them for their what they're teaching uh their grade level all that stuff. Um so there's there's that uh we call that dot. It's fun little blue animated character. Uh, and then we have tools. It's a form. You fill it out and it gives you an output, a lesson plan, adapted reading content, uh, things like that. Those we we like to call those the 101, like the the checkers level features that kind of table stakes that you've got to give to to teachers for them to move from can I even allow this to this is useful for me. Uh but the special part is when you start uh doing kind of these one-time guard railed safe managed AI tutors that the teacher can create >> give to their students students are interacting and then the teacher gets a real time dashboard of how the students are doing what they're doing with the AI and uh to make that very concrete uh the last 5 10 minutes of class a teacher may give what's called an exit ticket to their students Um, so they've got a recording of everything that they did during class that day and it loads up and says like, "Hey, how did the content go today?" And it does almost a what's called a a formative quiz. It it asks them questions and then it's coaching them on the an on on what they learned, what they want to learn, where they might go next, tease them up for whatever homework they might have. Um, and then it will do just kind of an emot kind of a social emotional check-in. and it will say like, hey, how like how was class? Um what feedback do you have for the teacher? Uh what are what are you looking for out of using this content? And rolling all of that up to the teacher so that they know how to better support those students in the future. Um a thing a lot of us don't recognize about teachers is they might be working with 300 students at a time. >> I had 42 desks in my classroom. >> So you started off as a teacher before all this. That informed a lot what you're doing. >> Yeah. So, I had 42 desks in my classroom and I was teaching seven or eight periods a day. >> Every day, teachers have to make this impossible choice. Do I work with uh the top 10% of students who love this? They get it. They want more of it. They'd stay after school if I let them. The 10% of students that are really struggling, maybe that's because they don't understand it. Maybe it's a learning disability. Maybe it's a problem at home. Maybe they got bullied in the last class. Or everyone else. And I care a lot about that. everyone else that middle 80% >> because I was one of those students and most of us were definitionally and we think about what we're building uh and what we've been able to build with with OpenAI and some of the tools that were announced today will be able to do even more of is uh give teachers almost a GPS for impact >> like these four students really need you today and you can jump in and support those four students in a way that you maybe wouldn't have even known that they had a concern >> that it brings up a very I think interesting point when I talk to developers often people confuse the tools for the product and what you have to do is you have to both understand the needs of the educator that comes from you working in the classroom and your peers working at school and understanding that and I think that's the thing you're able to bring to it where you look at this is a platform to build on top of it and that's something where you've identified all these areas that you can bring into to specialize it and make it very custom what you're doing. Yeah, I think that's something exciting we saw today in uh in the announcements was uh opportunities for uh people like me with subject matter expertise with that that really know a domain and we're able to build kind of with >> you all uh with OpenAI not just on top of OpenAI which is which is uh a really cool unlock for a lot of people. >> What are you most excited about that you saw today? >> The agent builder for sure. >> Agent builder. Yeah. Yeah. >> Yeah. That looks like uh it could be a lot of fun. And going back to when you know you you know how it was when we first started you had to just wire up a lot of that yourself and code tools certainly make that easier but just being able to drag and drop think something like file search the permission structure seemed really well thought out and particularly what you're doing in a classroom situation where you really have to have those safeguards in there. >> Yeah. uh we actually built our own and have for a number of years and I think we're excited to go and get our hands on on it and see what do we get to >> get rid of. Yeah. What do we get to uh make even better now that we have access to it ourselves? >> Um or or from you all. I think that's uh >> either way we're going to learn a lot from working with it and and that makes it easier for us to use but also again bring it into the hands of these people that aren't technical. They're not developers. They're not they're not thinking about, you know, which model they're using, but they just want to get their stuff done. >> Yeah, I think that's a great point in that the more you can have the people who understand who your customer is, in this case, students and teachers, understand what that problem is, and less time they're spending trying to wire things up and run, you know, fast API servers and stuff to do this. Seems like it can make the product get a lot better faster. >> Absolutely. >> What else have you been excited about you've seen? Um I think uh we've we've built some similar things and in our product we have this concept called powerups which are basically apps that you can use with dot our our character. I think we got to see some good patterns and examples of that today with the new apps that were that were announced. And I think one of the things we're really excited about is these partners that we've been developing uh that everyone is really solidified around MCP servers as a way to communicate with uh AI. And so um a thing that will definitely benefit us is OpenAI kind of drawing a line in the sand saying we're doubling down on this. Yeah. Uh and so now when we go to a partner that's already building for integration with Chad GBT, they can also bring that directly into school AI again for that safe managed guardrail experience >> um that uh that teachers and school leaders are looking for. >> Yeah. Yeah, one of the things I think it's going to be helpful to and you mentioned this before too is eval and the ability that as you run these systems and you run the models to be able to know how well they're performing because two or 3% may not seem like a lot but can make all the difference particularly with a student when you have five million students using uh your platform. Two to three% means a whole ton of issues every day. >> And it's one of those things it's often every company I talk to knows it's important. they want to do that, but to have the the opportunity and the time to spend building an eval suite to build something in there is just often kind of a thing they'll do in the future, but having that built into a system seems like a big >> we kind of saw it today, right? I don't have time to to do the eval. >> Exactly. Well, yeah. So, they did an eight minutes building an agent. Very impressive. Yeah. Take 10 and do the evals and that'll be fun. >> Yeah. That's it's kind of exciting too because for prototyping and spinning things out because I think a lot of what you're doing is probably going to be experimental and trying things with teachers to be able to sort of see if these things work >> and do you see that accelerating? >> Yeah. So we um one one of the things we do teachers when they're creating these like custom AI tutors they they do them lesson by lesson. It's like, hey, I'm teaching about the water cycle and I'm going to create this activity that's a uh starts out as a tutor and then turns into a game and then turns into a quiz. Uh just being able to preview that faster, but but almost having adaptive evals in the moment is one of the things that we've gotten into is >> uh really the meta prompting. How do we how do we not just tell the AI what to do, but tell it how to do the thing that it decides to do. >> Right. >> Right. Yeah. That's a it's a very interesting space to be in where the AI can help you build the prompt to tell it what to do, but knowing where you're going to start from has been super helpful. Um, where can we look forward to finding out more about School AI, what you're working on? >> Yeah. Uh, you can uh follow us on on X uh get school AAI. You can check out schoolai.com. Um, I think this is this is a fun time of year with back to school and seeing a ton of how how teachers and students are building >> um different uh different holiday stuff that gets really fun when when teachers are sharing with each other. So yeah, I would say uh on uh on X uh or Instagram or schooli.com. >> Last question. Any advice you give to developers about where to start with the tools what you've played with so far? Oh, I think um you can probably just start in the in like the GBT builder for most of the ideas that I hear is uh just start with the GBT builder and then expand from there. Um you've got uh you know in the developer portal you've got a ton of other tools to just start playing. Um and then I think the the agent builder that we saw today is going to be another fun one to start connecting the dots between all the different tools. >> Awesome. Thank you so much. >> Yeah, thank you Caleb. Hope you enjoy the rest of Dev Day. So, we're going to be talking to a few other developers about some of their different experiences and I'm always fascinated to see what it gets accelerated and how they're able to focus when a new tool comes out on really what is their core thing they're working on. So, up next we have Danny Grant who is with jam.dev which is an in browser tool for helping you evaluate your site and to figure out basically how to improve it. Is that correct? >> I'm so happy to be here and actually later today on stage we're going to announce a brand new tool. >> Oh boy. It lets any PM, designer, marketer Oh, I'm a little short. >> No, no, no. Micron's too tall. It's all good. >> I agree. Um, well, just like you just did, it helps anyone fix what's broken instantly without writing code. Um, it's called please fix. >> Okay. So, I would use just like my website, go into it and say just please fix or >> Yeah, I mean like yesterday before this tool was announced. Um, if you wanted to change like some copy on your site or like a like a button didn't look good or missed a hover state, you'd have to probably ask an engineer to do that for you. >> And the engineer, >> nobody wants to do that. No. >> Like, okay. And then they'd be like, can you make a ticket? You like, okay. You know, and then the the ticket gets prioritized and then like maybe they get to it. >> I worked for a company. I won't name it where we joked it was easier to release a model than put up a blog post. >> That's hilarious. >> Yeah, we won't name it. >> Um well, so so now as of today, what you can do is while you're looking at your site, you click on the please fix browser extension and then it lets you edit your site right there, like it's a Google doc, like it's a Figma. And then when you're ready and you like how it looks, you just click submit and it creates the PR for you. And it uses your design system. So engineers like the PRs. They're very clean. They're very tight. >> Very cool. What were you most excited about? about what you saw today. >> I mean, I think we just saw a new way to browse the web. Like I think OpenAI maybe just changed what we mean by the web, what we mean by a browser. Like if you think like web one is read, web two is read, write. Okay, there's web three. And then maybe this is like web four like read write think. I think we just saw a whole new way of experiencing the web where it's a lot less mechanical and it's a lot more stream of consciousness. So you talk about the apps inside chat GPT, right? >> It's freaking cool. >> Yeah, it is. It's it makes you think a lot about when you're building something, what the core functionality is, like what the purpose of it is. Any idea that if it's going to be presented inside chat GPT and so watching the demo where they interacted with with the Zillow website and actually having to have data about that and be able to drill down into it and it's it's exciting to think about the possibilities there. And so I could see for you all that seems like a pretty interesting area because people as they focus on usability, they're gonna want to make sure those things work really well. >> Yeah. Like now when you build your inside of chat GPT app, if you're the PM, if you're the designer and you want to tweak some things and make make the app look a little nicer, you can just use our browser extension, change it from the chat GPT interface and make the PR in GitHub. It's so cool. And I think, you know, these tiny tweaks, they often don't get prioritized today, but they should because the difference between a fine designed product and a well-designed product is that the well-designed product changes the world. Like the iPhone changes the world because it's usable. And there were actually attempts before that that I won't name that many people under the age of 30 have never heard of um because they just they didn't have that attention to detail. >> Have you seen your development process change with tools that have coming out, particularly AI tools? Yeah, I mean I like I'm biased. We've been using PleaseFix as we've been building PleaseFix and it like even today our pricing page changed dramatically because a PM could just go in and test a bunch of stuff without asking an engineer. I think the thing I'm most excited about is this idea that you can move as fast as your entire creative team altogether. Like no disruptions for engineers. It's like people can move without having to have bottlenecks on a few people. So it seems kind of cool because what you're talking about is the idea that people who aren't specifically engineer background are able to make those changes and has that been something you've seen basically within your company the adoption of tools like vibequ and stuff like this for people to contribute and be able to come up with ideas. >> Yeah, it's really cool. Um when we we talk to our users every day and some of the stories we hear are awesome. Like last week we talked to a user who's a firefighter and building software for firefighters. That's cool. um talked to a user a couple weeks ago who um grew up in the church system and is building software for churches has no software experience but is able to now make something that's very impactful to their community. I think that is awesome. I think we are about to see the Cambrian explosion of software the same way that with the web there's the Cambrian explosion of like news sources with you know Substack and Twitter. I I think this is about to happen for software and it's I think it's one of the best things for humanity. >> So how do things work at jam.dev dev as far as ideation, testing things with customers, releasing, and then figuring out what's best features and whatnot. >> The only thing we care about is does this deliver a wow for our users when they use it. And so that's what we're focused on. There's an emotional response we want people to have when they use our product because we work on the worst part of software development. Like I don't know anyone who is like today I want to fix some bugs, right? >> And so our job is just to make that whole part of the software experience a lot better. And so we're laser focused on that. We're just constantly talking to users. Every user that signs up from Jam hears from a co-founder. Every user who uses Jam hears from a PM. We are just in constant contact. >> When I worked at OpenAn, one of the most exciting things, and it sounds absolutely silly today was when we watched GPD3 spit out a React button. That was like, oh my gosh, I could do that. >> I remember that. I I don't remember it inside, but I remember from the outside. I remember seeing demos of that being like, the world just changed. That was that was the threshold for us being impressed was literally like four lines of code and to be able to oh that's really cool it can do that. What has been a big aha moment for you? >> I think look it's it's it's similar to what you're saying but when you think of what you just described coupled with what was just announced with the apps SDK. You can imagine that in the future yes humans are going to be building a lot of apps that are shown dynamically as they browse the web. But also I think that agents will dynamically build apps for you as you browse the web and that opens up a whole world of possibilities because you can have dynamic software right there. Like imagine inside of an organization right now if a PM wants to see how a product is doing they have to build a dashboard and it's a lot faster to do that today than even six months ago. >> But imagine the PM needs a dashboard and they're in chat GBT and chatgbt just gives it to them and no human had to do that. It means that there are two types of software. There's sort of long-term software that humans are going to work on. They're going to like really fine-tune, make it great, like Zillow, Canva, and then there's going to be sort of disposable one-time use software that an agent can just whip up. And that is just freaking cool. >> Yeah, it's it's a thing where, you know, as a developer, you're used to sometimes spinning up a tool one time to use that and then you're done with it. But it's another thing to think about that could be a new modality. You know, we get kind of fixated in sort of like the app store sort of idea. And I think that there's going to be, like you said, long-term stuff that you deploy inside a chat GPT, but then the ability for somebody just to spin up a thing they're only going to use once and that's the end of it. >> Yeah. >> And have you seen as far as things that sort of start off as hobby projects or things like that or maybe inside jam.dev, people said, "Hey, I have this idea. I wanted to solve for this thing, maybe even something from finance or comms or something that's turned out to be something useful." >> Yesterday we were um at the rehearsal for our 10x codebase session. It's four startups. We're all demoing and we're just sitting around kind of backstage talking and we're like, "Oh, how did you start your startup?" And almost everyone, three out of four, I'm the only one um where their startup started as an internal tool at their company that they needed and actually these startups pivoted to the internal tool because they found it so tremendously valuable to how they build software. >> What's you know how Slack got it start and other companies? It's a very interesting thing where the thing you spend the most time on is often the thing that's going to be the best product. And you know, one of the things they talked about OpenAI on here today was like how 70% of the codes coming from, you know, the PRs are coming from, you know, generated by codeex. And it's a very interesting thing to see when a company's using their own tool to build the tool, it seems to iterate much faster. >> It's actually so funny because if you listen to standard startup advice, like what would Paul Paul Graham tell you? It's hey, don't optimize the internal processes. Like just do things that don't scale. do them poorly. But actually, it's when you take a lot of care in your own processes that you can develop these products that can be their own companies and can help a lot of people. >> Okay. It might be that that is going to be the advantage is how quickly internally you make something because I think you're going to see a leveling effect with tools like the agent kit and everything else because it's going to be probably having a big technical bench isn't going to matter as much as having good product depth or understanding of your customer. >> Yeah. At the end of the day, a human has to use the software. And if a human has to use it, it has to be easy for the human to use it. And I do think design still today and maybe forever makes or breaks a difference. >> So you use jam.dev on the jam.dev site. >> Yeah. Yeah. And if you want to see what's brand new, go to jam.devfix. >> Okay. How often do you guys making updates on your own site? >> Um, right now we have an >> You did your pricing page. You did that. So yeah, >> too often. Is it become does that become sort of thing like well if it's so easy to change it you know >> we just added dark mode we added pricing page uh we do copy updates yeah now it's too easy >> is there going to be like a please fix wait you know >> oh please fix with a polite delay >> yeah or like maybe think about this like you know >> yeah we should add as a feature for engineers that if you made too many changes it like arbitrarily waits a few days and then asks for a change again >> do you really want this yeah >> what are you looking forward to what kind of tools would you like to see that don't exist yet. >> Um I well I was I was talking to the engineers sitting around me during the keynote >> and they were really excited about the optimizer and the evals as part of the agent kit >> and um and what they wanted to see was well like what if the eval could be written automatically for you using your own data and then use that to automatically optimize your prompt. So rather than people sort of sweating the details about these agents, what if the agents could improve themselves? And I think if we had that at Jim, like we could move a lot faster and we can make a lot more powerful software a lot sooner. >> Uh what advice do you have to founders or developers right now? >> It's it's never been a more fun time to build. I think we all just get to enjoy it. >> So how did you figure out that this your team wanted to work on with jam.dev? How did you decide this was the problem you wanted to solve? >> When uh my co-founder and I were product managers together at Cloudflare, we we worked on the fastest moving team in the company. It's this sort of skunk works team that would try to do big new things. It's the team that shipped Cloudflare workers um their cloud compute platform. It shipped 1.1.1.1 which now fields a trillion DNS queries a month. Um and it was just trying to move fast on this team that we realized a lot of the frustrations and bottlenecks come from just reproducing issues and there's no tooling out there and no user, no PM, no one in sales knows how to communicate to an engineer what the engineer needs to know. And we're like, we can't believe that there's no way to help the engineer get things done faster. We're spending all this like great brain time on communicating bugs and hopping in calls and screen sharing and and not enough on fixing them. And we thought that that that we can solve. >> It seems uh like a lot of things sort of in front of you as far as what you can do. How do you decide what you're going to do next? >> Okay. I think I think there's something that we undertalk about in like startup world which is if if your startup works up works out to your wildest dreams you're going to work on it for like 10 years >> and that's a really long time and so you better just love the problem because that makes it really really fun and so I think you work on the thing where if you get to wake up every day work on it and talk to users of the thing you're going to be pretty darn happy about that. >> That's pretty awesome. Uh how can people get started at jam.dev >> jam.dev dev/pleasfix and we'll please fix some code for you. >> Awesome. Thank you so much. >> Thank you. >> Uh so it's interesting to see what each kind of developer what they're looking forward to, what kind of excites them as far as that. A lot of that's based upon what they've been working on. I know a lot of devs have to build a lot of internal tooling and when you can solve for that with an SDK, it makes your life a lot easier because you can focus on the thing you want to work with. Up next we have Zach Lipton who is with a bridge. How you doing Zach? >> Not too bad. How you doing? >> Fantastic. So a bridge you're working on tools for helping the medical community people doing like transcription and that kind of area. Is that how you >> API platform for doctor patient conversations? Like if we were to like rewind to like the state of affairs, you know, uh post adoption of electronic health records but pre-ambient listening. um doctors were spending about two hours doing paperwork for every one hour of direct patient care. So it was this kind of like clerical uh burden crisis. Um and it was pulling it was a situation where technology was pulling doctors away from patients rather than bringing them closer. >> Um what we do is we provide this platform that helps kind of gives doctors superpowers, helps them with their paperwork, does all the note-taking in the background, preps them everything. that all these kind of documentation artifacts are ready for them the moment the visit's over in exactly the form they need so they can be fully present with the patient instead of spending all their time staring at the computer. >> What kind of metrics do you have so far as far as like time saved for doctors? >> Sure. Interesting story um and a difficult to track down in part because there's a time that doctors spend uh documenting during the day but the reality of the status quo before was that most doctors actually didn't finish their notes during the workday. So, they were home after hours logging into the EHR and doing what we called um pajama time. They're basically sitting there like pulled away from dinner with their families or logged in after hours, you know, sitting in sitting in bed finishing up their notes. Um our report, so we we we cobble together this information from a bunch of different sources, but we see doctors saving as much as an hour or more a day. Um we see, you know, doctors some doctors seeing like 10 15 patients in a day. we're talking about like you know 5 to 10 minutes often in note takingaking. So it's a it's a tremendous um >> kind of relief but even beyond the um actual time saves oftentimes there's an even larger sort of like perception of burden lift and that's because the doctor's worrying about less and able to focus on their patient. >> Yeah. I mean that's great because like you you know look at with an hour a day is either an hour they can spend with patients or just not have burnout and you know have more focus on that. You know that's we talk to before school.ai AI and looking at how they're basically able to help teachers spend more time in the classroom with students and that's the same problem you're dealing with here. How can doctors be there for students or student patients and also >> yeah well so we have a channel in Slack and it's what we call love stories and it's where we hear like uh kind of like all the feedback from the field from doctors and one of the one of the wild indicators that like we had landed on something big and this is relatively early um maybe like a couple months after we launched the first like enterprise pilot at a hospital system was when we started getting stories coming in and they weren't just talking about the clinical experience but they were talking about like uh I I spent and uh actually got to have dinner with my family every night this week for the first time in like 10 years or like a bridge is saving my marriage and that was that was not like where I was expecting things to land so quickly but uh you know that the problem is that big. >> What did you see today dev day that has you excited? >> Oh um so many things but maybe like two that are top of mind. Um one a lot of people have already talked about the agent developer kit >> and I think that's extremely exciting. There's this moment right now where I think everyone it's sort of a proto discipline. So everyone's been rolling their own tools in the hope of like kind of trying to figure out like what is the paradigm? How does this work? And there's so many things that have to work together. Um there's the context engineering, there's the prototyping, there's a sanenity checking, um there's evaluation, there's also down the road, you know, everything around production and monitoring. So, I'm really excited to see where these tools um ultimately go, but seeing OpenAI like take a strong position and put a a kind of comprehensive offering that brings together a lot of these things that you know people have been rolling their own orchestration tools, their own evaluation platforms, we certainly have um and seeing how much how much this is going to create a common platform and allow us to sort of lift off and focus more on the content um is something I'm super excited about. Um, and then in general, I've just been extremely excited and drawn a lot of inspiration from all the work that's gone on in terms of software developer tooling. And so, we are we are developing AI powered products for our customers, but we're also big consumers as uh, you know, productivity tools like codeex um, um, play a big role for us and just seeing how far we've come there. I mean, I remember um, so my background is academia. I was AI researcher. I am an researcher, a professor at Carnegie Melon and I did my PhD back in uh 2010s and I remember when Ilia had a paper, Iliach had a paper called learning to execute and it was just having like you know the idea was like code going in and like a model kind of anticipating output was like such a the fact that the model is doing anything at all in the space of code was kind of revolutionary. and to to see where we are now going from like maybe two years ago having our minds blown by just code completion and right now seeing these like larger like full codebased refactors taking place um you know it's it's kind of amazing to see the progress in the space and um I've been excited to to follow along. So when you work in medical it's a very high stakes area and so that's got to be something you think a lot about about how you deal with hallucination and also how you deal with customer concerns about these things >> 100%. >> So what do you look forward to in tools? What have you seen the biggest help in that area? >> Um that's an area where we've had to develop a lot of our own technology. Um >> you know what is a hallucination? It's kind of like like back in the old days of like developing simple classifiers. We had we had false negatives and false positives and now kind of like everything that's like if it's there and you don't want it, it's a hallucination. And the question is, well, what is a hallucination? Sometimes it's >> completely confabulated information that is asserting facts about the world that are not real. But in the context of medical note-taking, medical documentation, order placement, um it's there there's a kind of particular situated notion of what we really mean. It's something that's sort of unlicensed, you know, by the sort of surrounding context, the substantiating evidence, even if it might be true, you know, if uh if a doctor doesn't um if a kind of explanation of a disease like shows up in a generated patientf facing summary that the doctor never said. >> Yeah. you know, that's that's kind of like even if some of the information might be like either correct or plausible, like that's not within our so we have a kind of bespoke notion of what constitutes a hallucination. >> Um, but we able to like kind of >> often draw a lot of inspiration from what can like the frontier models already do out of the box. If we define our ontology of like these are the types of of errors we're concerned with, then we can go and say what is the ability of an out-of-the-box model given like each each documentation sort of sentence allocart to correctly designate them as belonging to the right category. And we find okay, we're already within the realm of like the model is able to judge even if even if it's not able to sort of never commit the crime in the first place, it's able to recognize when the crime's been committed. That gives us a sort of like proof of concept and now what we need to do is make it better, more accurate, cheaper, faster. So ultimately create our own special purpose models that are able to take in parallel every single sentence in all the generated documentation and surrounding artifacts and process for each one like sort of does it contain an error of of a unacceptable variety like of of what kind and then a kind of pipeline downstream for remediating. and we're able to do that um with about 97% recall at this point. >> So, do you have any advice for people who are trying to work on basically developing their own evals for hallucination or just sort of a good starting point? I >> I think it all starts with getting really crisp about what you really mean. And I think that's what we've seen before is that what is a hallucination for us is a little bit different from what is a hallucination for like a general um sort of like open world QA system. Mhm. So is there kind of like a uh boundary at which you keep expanding? For instance, you talk about medical say, okay, we feel maybe working in scribing right now is is an area that we can probably solve for and produce a pretty good product that's at or better than human level. Then do you look at like there's areas in which you would expand out to as you feel more confident? >> Absolutely. So for us, um, we've always bristled a little bit when, you know, like VCs put up this chart and they're like, "These guys over here are the the conversation agents and these guys here do are the coding startups and these guys are the scribing companies." And we we've never liked getting pigeonholed as a scribing company. That's because from the very outset >> um the central thesis wasn't just about scribbing. The central thesis was sort of >> um about medical conversations about this being this moment where um you know this is this is this magical spot. It's these 15 minutes that the patient waited maybe six months for are where the patient tells their entire story where the doctor goes through their entire reasoning process and within minutes after it's over the patient's forgotten 80% of what happened the doctor is like you know hours behind on their note-taking and so for us we've had this feeling that you know and we I've also like as an academic been working in actually applications in healthcare is like my kind of passion area for about a decade and I've been watching so many interesting machine learning ideas get developed as a proof of concept only to only to kind of sit on the floor and not get used. And so what we realized that the conversation was was this weigh in. It was this important arena where it was in some ways it's the most important moment in the entire experience of healthcare and sort of no one was providing value in that moment. And scribing we already knew was going to be a killer application because you know it was never going to scale. human scribing was never going to scale to all doctors. But those who could afford it were willing to pay tens of thousands of dollars per doctor per year to have a sort of like offshore scribe. And so that already kind of told us like there was this clamor for it. There was this need for it, but we could get in. But now that we're in um I I I'd view like take this broader view of like what what is like the entire picture rather than sort of being like we're we're quietly in the background just like minding our own business and then at the very end of the visit, boom, the magic happens. We don't want to go so far in the other direction that we become interruptive. But from a p perspective of just sort of like how do we support a doctor for the entirety of the visit from sort of their pre-charting experience, you know, before the patient even comes in the room through sort of every kind of cue or nudge that they might need during the visit to help them make the best possible decisions, help them uh tick all the boxes to make sure that like insurance is going to preapprove the particular test or treatment that they're going to do. So patient end doesn't end up wait wasting like uh you know a month waiting for care. Um, so we kind of like zoom out. we kind of see like this space of the conversation like the point of care is like now that we have ears in the visit and now that you have a sort of AI workforce to sort of do your bidding um what what are all these other jobs to be done that could be addressed in the moment and that includes everything from from from the sort of previsit experience through to real-time clinical decision support through all the kind of um anticipating and getting in front of all of the uh sort of like financial related ated documentation that needs to be done to ensure that that the doctor gets paid and that the patient gets their care in a timely fashion. >> Yeah. You know, on one end you have people working on, you know, the AI scientists and tools and trying to solve frontier problems. I've talked to other people who have told me that if you can have better intake in hospitals, you might be able to get rid of hepatitis that there's some very lowhanging fruits there. And is that something you've looked into or you sort of see a lot of opportunity there? >> Absolutely. >> Any particular area you'd like to see the tools get better sooner at? Oh. Um, so many. Uh, I'd say on a personal note, they're not very funny yet. >> Okay. >> So, I don't know if you've had this experience at Chat DVT that it's, uh, it's far better at like solving hard math problems than it is making a joke. >> Uh, Sora though is very good at jokes. I don't know if you, you know, tried that yet, but >> yeah. Um I think there's a tremendous amount of work that one still has to do to like um crisply define every single task for the model. And I think that like just how high models can come and I think you know like they're when it comes to like a crisply defined technical task they're they've gotten very high in the abstraction chain about breaking it down. Um but I think when you get outside the you know and start like addressing problems at like interacting with the system kind of like tackling the like the the the more like world problem you're you're discussing you find that like you kind of have to do all the driving and the system is a little bit more of a it's an information retrieval system. It is it is the world's knowledge at your fingertips but it is not kind of connecting dots at a more abstract level. And so I think you know in the in the coming you know months and years I'm excited to see uh to what extent does a model go from um more technical problem solver to uh a more independent interlocutor and the like normative side of problem solving. >> Going back how did you all decide this was the space you wanted to start with? >> Um you know I think we saw a few trends that were happening all at once. um at once like my research background was in deep learning um and any one given approach we kept running into a little plateaus here and there but if you zoom back and like look at the arc from 2012 through to you know maybe 2018 19 when the company was founded you saw there was there was a there was a a path of advances in speech recognition and path and advances in natural language processing that was you know preceding if anything accelerating And simultaneously um there was a crisis around physician burnout that maybe in 2018 wasn't the like number one burning priority on the minds of like CMIOS and like hospital system CFOs across the country but it was you know it might have been number five on their priorities and it was like rising up the charts and so we kind of knew there was this coming crisis of like doctors were burning out they were spending more and more time on documentation They were dropping out of med school. They were graduating med school with no intention to practice medicine, leaving to join to join tech companies, to join pharma, but to do anything but practice. Um, and so there was a turn problem or attention problem. And at the same time, like we we we kind of knew that there was like the the right family of tools were coming into fruition simultaneously. Um and you know for us uh I think me coming from um I was saying before this kind of academic perspective having you know a lot of us before had maybe operated in the machine learning for healthcare community a little bit on like >> what feels like a cool or important predictive problem but without maybe connecting all the dots when it came to like what were the priorities of the health system, what were the pain points of the health system, what were really like the choke points in maybe coming a little bit more from like what seems like an interesting clinical predictive problem. And I think at that moment we had a little bit of flash of insight that you know we didn't know if our timing would be right. Keep in mind that in 2018 like the typical like context length for a language model was maybe >> I don't know 256 words and these conversations are like 4 to 8,000 words like a median. Um but uh I think we just saw a lot of those convergence of like a few trends that that all spell that there was there there was this real opportunity to um to save time for doctors and ultimately hopefully you know um save money and also save lives. So in an area like medicine which is very high stakes what advice do you have for developers that are trying to build trust with their product because as you know probably better at anybody that there's been kind of a a road of broken promises of people who were very frustrated by you know oh this is going to do this this didn't work and when you come in with something that says hey it really works how do you win them over um I think it's uh it's it's it's a neverending I I think trust is trust is earned every single day. Trust is earned. I mean there's >> there's like the initial trust of like we've been talking about this vision for a long time and then we actually built the product and got it to work. But there's also the trust that's built through, you know, um, working with hospital systems is kind of like a hightouch like white glove enterprise. And I think you know through continued um delivery on everything from like our product commitments, our data security commitments, um you know the the kind of service we give people um the the continued fulfillment of of of every kind of promise and like continue to expand and serve not just you know initially maybe more primary care ambulatory now emergency inpatient nursing stakeholders. I think this this trust kind of acrus through this um continued delivery of everything from like the product through the experience of the the medical teams that are working with us in partnership over the course of now we're talking about you know many years. >> Awesome. Well, thank you very much. I appreciate it, Zach. And it's a bridge. People want to find out more. It sounds like a very exciting space to see where you guys are headed. >> Yeah. Thanks for having me. >> Enjoy the rest of Dev Day. >> See you. It's interesting to see where you have companies that are dealing with very high stake stuff like medical or education and a lot of it is the trust building. It's not a thing where you just pop out with a product and you say, "Hey, we're ready. We're done." You can see from the examples here between Danny and Zach and how they have to basically and um Caleb of basically just trying to sort of show the customers one step at a time, iterate on the product, improve it. And here to talk to us about tools for helping work on this is Lee Robinson from Curser. How you doing? I'm doing well. Thanks for having me. >> Lee, it is great to have you here. Uh, so cursor, um, I probably have about three cursor windows open right now on my computer that I'm thinking about right now. Uh, it has been an incredible product evolution for you all here. And just a little backstory, I was at open eye when we worked on the earliest version of codecs and code completions. >> Mh. >> And I kind of naively thought that, oh well, I just ask, you know, GPT3 at the 3.5 at this time or the codeex model or Da Vinci Code, whatever, and say just complete it and it's done and I got my code. I'm like, that's it. Code is solved with AI and we're done. >> Yep. Turned out that's not the case. Yeah, it turns out there's a lot that goes into it, especially from maybe simple text autocomplete, but to where we're at now with fully autonomous coding agents who can self-correct and fix their own errors and pull in information from the outside world. And it's wild how much better the coding with AI space has gotten just in the past year, I would say. >> And it's a tool very much what I appreciate is the fact that you guys are using cursor to make cursor better. >> Definitely. Yeah. One part of our culture that I think helps us produc

Original Description

The OpenAI Podcast is live for the first time. Host Andrew Mayne sits down with startups Cursor, Abridge, SchoolAI, and Jam.dev—each reimagining how AI can transform their industries. From healthcare and education to coding and collaboration, we explore how these builders are putting AI to work in the real world. Subscribe to the OpenAI Podcast on Spotify and Apple Podcasts
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from OpenAI · OpenAI · 0 of 60

← Previous Next →
1 Robots that Learn
Robots that Learn
OpenAI
2 Emergence of Grounded Compositional Language in Multi-Agent Populations
Emergence of Grounded Compositional Language in Multi-Agent Populations
OpenAI
3 OpenAI + Dota 2
OpenAI + Dota 2
OpenAI
4 Dendi vs. OpenAI at The International 2017
Dendi vs. OpenAI at The International 2017
OpenAI
5 Competitive Self-Play
Competitive Self-Play
OpenAI
6 Learning a Hierarchy
Learning a Hierarchy
OpenAI
7 Physical Spam Detection
Physical Spam Detection
OpenAI
8 Ingredients for Robotics Research
Ingredients for Robotics Research
OpenAI
9 OpenAI Five
OpenAI Five
OpenAI
10 OpenAI Five: Dota Gameplay
OpenAI Five: Dota Gameplay
OpenAI
11 Learning Dexterity
Learning Dexterity
OpenAI
12 Learning Dexterity: Uncut
Learning Dexterity: Uncut
OpenAI
13 OpenAI Five Benchmark: Post-Game Analysis
OpenAI Five Benchmark: Post-Game Analysis
OpenAI
14 Investigating Model Based RL for Continuous Control | Alex Botev | 2018 Summer Intern Open House
Investigating Model Based RL for Continuous Control | Alex Botev | 2018 Summer Intern Open House
OpenAI
15 Generative Modelling | Sadhika Malladi | 2018 Summer Intern Open House
Generative Modelling | Sadhika Malladi | 2018 Summer Intern Open House
OpenAI
16 A pathway to more efficient generative models | Will Grathwohl | 2018 Summer Intern Open House
A pathway to more efficient generative models | Will Grathwohl | 2018 Summer Intern Open House
OpenAI
17 Learning Dexterity | Alex Ray | 2018 Summer Intern Open House
Learning Dexterity | Alex Ray | 2018 Summer Intern Open House
OpenAI
18 Robust Vision-Based State Estimation | Hsiao-Yu 'Fish' Tung | 2018 Summer Intern Open House
Robust Vision-Based State Estimation | Hsiao-Yu 'Fish' Tung | 2018 Summer Intern Open House
OpenAI
19 Using Semantic Trees In Place of Sentences | Munashe Shumba | OpenAI Scholars Demo Day 2018
Using Semantic Trees In Place of Sentences | Munashe Shumba | OpenAI Scholars Demo Day 2018
OpenAI
20 Reinforcement Learning with Prediction-Based Rewards
Reinforcement Learning with Prediction-Based Rewards
OpenAI
21 OpenAI Spinning Up in Deep RL Workshop
OpenAI Spinning Up in Deep RL Workshop
OpenAI
22 Arena Announcement and Closing | OpenAI Five Finals (6/6)
Arena Announcement and Closing | OpenAI Five Finals (6/6)
OpenAI
23 Co-Op Match | OpenAI Five Finals (5/6)
Co-Op Match | OpenAI Five Finals (5/6)
OpenAI
24 OpenAI Five vs. OG, Game 2 | OpenAI Five Finals (4/6)
OpenAI Five vs. OG, Game 2 | OpenAI Five Finals (4/6)
OpenAI
25 OpenAI Five vs. OG, Game 1 | OpenAI Five Finals (3/6)
OpenAI Five vs. OG, Game 1 | OpenAI Five Finals (3/6)
OpenAI
26 Pre-Match Panel Discussion | OpenAI Five Finals (2/6)
Pre-Match Panel Discussion | OpenAI Five Finals (2/6)
OpenAI
27 Opening Keynote | OpenAI Five Finals (1/6)
Opening Keynote | OpenAI Five Finals (1/6)
OpenAI
28 OpenAI Robotics Symposium 2019
OpenAI Robotics Symposium 2019
OpenAI
29 OpenAI Scholars Demo Day 2019
OpenAI Scholars Demo Day 2019
OpenAI
30 Multi-Agent Hide and Seek
Multi-Agent Hide and Seek
OpenAI
31 Solving Rubik’s Cube with a Robot Hand: Uncut
Solving Rubik’s Cube with a Robot Hand: Uncut
OpenAI
32 Solving Rubik’s Cube with a Robot Hand: Perturbations
Solving Rubik’s Cube with a Robot Hand: Perturbations
OpenAI
33 Solving Rubik’s Cube with a Robot Hand
Solving Rubik’s Cube with a Robot Hand
OpenAI
34 Music Generation | Christine Payne | OpenAI Scholars Demo Day 2018
Music Generation | Christine Payne | OpenAI Scholars Demo Day 2018
OpenAI
35 Deephypebot | Nadja Rhodes | OpenAI Scholars Demo Day 2018
Deephypebot | Nadja Rhodes | OpenAI Scholars Demo Day 2018
OpenAI
36 Physics Net | Ifu Aniemeka | OpenAI Scholars Demo Day 2018
Physics Net | Ifu Aniemeka | OpenAI Scholars Demo Day 2018
OpenAI
37 Art Composition Attributes + CycleGAN | Holly Grimm | OpenAI Scholars Demo Day 2018
Art Composition Attributes + CycleGAN | Holly Grimm | OpenAI Scholars Demo Day 2018
OpenAI
38 Generating Emotional Landscapes | Hannah Davis | OpenAI Scholars Demo Day 2018
Generating Emotional Landscapes | Hannah Davis | OpenAI Scholars Demo Day 2018
OpenAI
39 Looking For Grammar In All The Right Places | Alethea Power | OpenAI Scholars Demo Day 2020
Looking For Grammar In All The Right Places | Alethea Power | OpenAI Scholars Demo Day 2020
OpenAI
40 Semantic Parsing English to GraphQL | Andre Carerra | OpenAI Scholars Demo Day 2020
Semantic Parsing English to GraphQL | Andre Carerra | OpenAI Scholars Demo Day 2020
OpenAI
41 Long term credit assignment with temporal reward transp… | Cathy Yeh | OpenAI Scholars Demo Day 2020
Long term credit assignment with temporal reward transp… | Cathy Yeh | OpenAI Scholars Demo Day 2020
OpenAI
42 Social learning in independent multi-agent reinfor… | Kamal N’dousse | OpenAI Scholars Demo Day 2020
Social learning in independent multi-agent reinfor… | Kamal N’dousse | OpenAI Scholars Demo Day 2020
OpenAI
43 Quantifying Interpretability of Models Trained on Coi… | Jorge Orbay | OpenAI Scholars Demo Day 2020
Quantifying Interpretability of Models Trained on Coi… | Jorge Orbay | OpenAI Scholars Demo Day 2020
OpenAI
44 Towards Epileptic Seizure Prediction with Deep Network | Kata Slama | OpenAI Scholars Demo Day 2020
Towards Epileptic Seizure Prediction with Deep Network | Kata Slama | OpenAI Scholars Demo Day 2020
OpenAI
45 Universal Adversarial Perturbations and Language M… | Pamela Mishkin | OpenAI Scholars Demo Day 2020
Universal Adversarial Perturbations and Language M… | Pamela Mishkin | OpenAI Scholars Demo Day 2020
OpenAI
46 Introductions by Sam Altman & Greg Brockman | OpenAI Scholars Demo Day 2020
Introductions by Sam Altman & Greg Brockman | OpenAI Scholars Demo Day 2020
OpenAI
47 Introduction by Sam Altman | OpenAI Scholars Demo Day 2021
Introduction by Sam Altman | OpenAI Scholars Demo Day 2021
OpenAI
48 Breaking Contrastive Models with the SET Card Game | Legg Yeung | OpenAI Scholars Demo Day 2021
Breaking Contrastive Models with the SET Card Game | Legg Yeung | OpenAI Scholars Demo Day 2021
OpenAI
49 Large Scale Reward Modeling | Jonathan Ward | OpenAI Scholars Demo Day 2021
Large Scale Reward Modeling | Jonathan Ward | OpenAI Scholars Demo Day 2021
OpenAI
50 Words to Bytes: Exploring Language Tokenizations | Sam Gbafa | OpenAI Scholars Demo Day 2021
Words to Bytes: Exploring Language Tokenizations | Sam Gbafa | OpenAI Scholars Demo Day 2021
OpenAI
51 Learning Multiple Modes of Behavior in a Continuous… | Tyna Eloundou | OpenAI Scholars Demo Day 2021
Learning Multiple Modes of Behavior in a Continuous… | Tyna Eloundou | OpenAI Scholars Demo Day 2021
OpenAI
52 Scaling Laws for Language Transfer Learning | Christina Kim | OpenAI Scholars Demo Day 2021
Scaling Laws for Language Transfer Learning | Christina Kim | OpenAI Scholars Demo Day 2021
OpenAI
53 Contrastive Language Encoding | Ellie Kitanidis | OpenAI Scholars Demo Day 2021
Contrastive Language Encoding | Ellie Kitanidis | OpenAI Scholars Demo Day 2021
OpenAI
54 Characterizing Test Time Compute on Graph Structur… | Kudzo Ahegbebu | OpenAI Scholars Demo Day 2021
Characterizing Test Time Compute on Graph Structur… | Kudzo Ahegbebu | OpenAI Scholars Demo Day 2021
OpenAI
55 Studying Scaling Laws for Transformer Architecture … | Shola Oyedele | OpenAI Scholars Demo Day 2021
Studying Scaling Laws for Transformer Architecture … | Shola Oyedele | OpenAI Scholars Demo Day 2021
OpenAI
56 Feedback Loops in Opinion Modeling | Danielle Ensign | OpenAI Scholars Demo Day 2021
Feedback Loops in Opinion Modeling | Danielle Ensign | OpenAI Scholars Demo Day 2021
OpenAI
57 Creating a Space Game with OpenAI Codex
Creating a Space Game with OpenAI Codex
OpenAI
58 “Hello World” with OpenAI Codex
“Hello World” with OpenAI Codex
OpenAI
59 Talking to Your Computer with OpenAI Codex
Talking to Your Computer with OpenAI Codex
OpenAI
60 Data Science with OpenAI Codex
Data Science with OpenAI Codex
OpenAI

The OpenAI Podcast Episode 7 explores the use of OpenAI models in various industries, including education and healthcare, and discusses the potential of AI-powered tools and platforms for improving productivity and patient care. The episode features interviews with startups like Cursor, Abridge, SchoolAI, and Jam.dev, and highlights the importance of trust, coding, and AI-powered tools in these industries.

Key Takeaways
  1. Use OpenAI models for education and healthcare
  2. Orchestrate AI agents for education and healthcare
  3. Create custom AI experiences
  4. Automate communication between teams
  5. Create internal tools and products
  6. Design AI-powered systems for education and healthcare
  7. Code with AI
  8. Create autonomous coding agents
💡 The use of AI-powered tools and platforms has the potential to revolutionize industries such as education and healthcare, and startups like Cursor, Abridge, SchoolAI, and Jam.dev are at the forefront of this revolution.

Related Reads

📰
OpenAI Just Bought Gitpod: The AI IDE Wars Are Officially On
OpenAI acquires Gitpod, signaling a shift towards cloud-based AI coding, and you can leverage this trend to enhance your development workflow
Dev.to AI
📰
Programming Assignments: A Complete Guide to Solving Coding Problems Faster and Smarter
Improve coding skills by learning strategies to solve programming assignments faster and smarter
Medium · JavaScript
📰
Will CAD Drafters Be Replaced by AI?
Learn how AI impacts CAD drafters and why it matters for the future of technical design
Medium · AI
📰
From "You Have a Bug" to "Here's the Root Cause" - Adding AI Code Analysis to My App Review Pipeline
Learn how to enhance your app review pipeline with AI code analysis to identify root causes of bugs and crashes
Dev.to · Ashish Mishra
Up next
How to Start Vibe Coding With Gemini AI: Beginners Tutorial
LoverFighterWriter
Watch →