Global AI Trends with Ben Lorica - #26

The TWIML AI Podcast with Sam Charrington · Beginner ·📐 ML Fundamentals ·9y ago

Key Takeaways

Discusses global AI trends and machine learning applications with Ben Lorica

Full Transcript

[Music] hey hello and welcome to another episode of twiml talk the podcast where I interview interesting people doing interesting things in machine learning and artificial intelligence I'm your host Sam charington so as you know a couple of weeks ago I announced the first anniversary of the podcast and to celebrate it our first anniversary listener appreciation contest I want to start this show by thanking everyone who participated in the contest you all have come out in Dr Groves over the past couple of weeks to shower the podcast with love and for that we are humbled and forever grateful overall we received nearly a 100 entries via our website and iTunes from listeners from 16 countries your stories have been awesome and we're most proud that we can help light that AI fire for you every week of course everyone who entered will receive a couple of twiml and AI stickers one for you and one for a friend those will be on their way soon but now the moment you've all been waiting for we have our winners for the first anniversary listener appreciation contest just a reminder our grand prize winner receives a bronze pass for the O'Reilly AI conference at the end of this month and our runnerup winner gets well hey Google what's the second prize the second prize in TL's first anniversary listener appreciation contest is a brand spanking new Google home like me all right without further Ado our second prize winner is ankor Patel Anker wrote in on the show notes page with this twiml congratulations on your first anniversary I have loved listening to the show I remember the news and machine learning from the first few months what a sheer amount of information but the more recent interviews format is excellent I particularly love the ones where you interview Folks at Ai and ml conferences very insightful indeed thanks for contributing to the broader ml Community well thank you anchor for being a loyal listener and we'll be in touch with you about that Google home and now our grand prize winner is Mason Grimshaw Mason also wrote in via the show notes page here's what he said the show is fantastic I started listening right when you switch to the interview format and I've definitely noticed your improvement as a journalist and the Improvement of the show great job I go to school in Boston so the show keeps me company on my 20-minute walk to school every morning and I couldn't ask for a more interesting companion I'm studying analytics and while I may not directly use AI in my profession I certainly use it as a hobby I really enjoyed Carlos guestrin when he came on to talk about the line paper and the Danny Lang discussion about video games well Mason we hope to see you in New York in a few weeks congratulations and thank you for being a loyal listener although there can only be one first prize winner we'd like to give everyone the opportunity to attend the O'Reilly AI conference with us with the code PC twiml that's PC wi ml all of our listeners will get 20% off of registration fees when purchasing passes for the conference please let us know if you're planning to attend the event we're looking forward to meeting up with twinl listeners there okay to continue with the O'Reilly AI theme this week I've got a very special guest I've invited my friend Ben Lura onto the show Ben is Chief data scientist for O'Reilly media and program director for strata data and the O'Reilly AI conference Ben has worked on analytics and machine learning in the finance and Retail Industries and serves as an adviser for nearly a dozen startups in his role at O'Reilly he's responsible for the content for seven major conferences around the world world this year in the show we discuss all of that touching on how Publishers can take advantage of machine learning and data mining how the role of data scientist is evolving and the emergence of the machine learning hey everyone I am on the line teolog data scien media and he's also responsible for content for both the strata data conference as well as the O'Reilly AI conference uh hey Ben how are you doing great great to be here Sam awesome awesome so uh every time I talk to you you are just getting off of a plane and it sounds like that is the case this time as well yeah so we had a outstanding uh Str data conference in London yeah uh basically many many people from uh uh many different parts of Europe came and spoke and attended and this year in particular we had uh concerted effort to focus on uh deep learning on the machine learning side but you know the the staple topics of strata were still remained popular particularly on architecting Big Data application okay awesome awesome well I thought we'd start this conversation by having you spend a little bit of time talking about your background and how you got uh started with data and um you know how you you know kind of your path leading up to O'Reilly sure so I I have a PhD in math and focused on uh nonlinear partial differential equations and uh towards the end of Graduate School uh I became interested in uh specific set of differential equations called stochastic differential equations which turn out to be at least um theoretically important for quantitative Finance so I I had some interest already in kind of the uh Industrial applications of of what I was doing but uh I definitely was on the academic track so after grad school I uh was an academic for five years but then at some point uh I decided uh I was actually much more interested in industry and uh uh at the risk of dating myself at least at the time when uh I I was contemplating moving to Industry there was no dat a science track so the the uh the exit the exit strategy was becoming a Quant and so that's what I did I did that for about two and a half years in a small hedge fund uh designing trading models risk management portfolio management and things like that and uh you know one of the first things I learned of course is the the stuff I thought uh was going to be super important uh the stochastic pdes while uh good to know not not exactly uh What uh uh I needed to do for my job so at that time at that time actually um the term machine learning I would say was kind of nent the people were still using the term data mining um and so that that's what I did at this I I I I applied basically statistical techniques in machine learning uh to financial time series but then at some point I realized my interests actually were much more on the tech side you know the programming and and and building software applications uh for uh analyzing these uh uh time series and so that and then I ended up um moving the Silicon Valley around actually believe it or not at the peak of the NASDAQ at the time which was was around March 2000 and and so then okay I was here I was here uh just in time for that first bust uh and then that first bust yep I was there around the same time yeah yeah yeah and so uh do you do you do you keep track of what's going on in the on the Quant quantitative side of the financial uh markets and how folks are how the technology has evolved since you worked in that space um I not not that closely I I still have friends and obviously uh my work leading these two big conferences that you describe Str data in the O'Reilly artificial intelligence conference brings me in contact with uh people working in finance so I I keep track of it that way um yeah so to some extent I do but I'm not immersed in the the latest uh techniques that they're using although I I would say that they're they are actually moving more towards our world you know of of using of using machine learning alternative data sources uh big data and uh some of these more uh uh bleeding edge uh machine learning techniques including deep learning so to the extent that they're actually showing up at the events I organize uh that's how I keep track mhm okay okay so I interrupted you you you uh oh yeah so then at some point got into the technical side of things I got into the technical side of things and um when I first moved to Tech I think uh you know this won't surprise a lot of your listeners but you know one of the big users of data in Tech are uh the marketing and sales people so that's how I kind of uh moved into Tech was basically took what I learned in uh in finance and started applying it in in marketing and sales applications so did a couple of startups joined a couple of startups that didn't really uh take off and then eventually at some point ended up at O'Reilly as a data scientist working on the types of data that we have which is a lot of it is sales data but a lot of unstructured and semi-structured text so I was uh I was definitely one of the first people doing a lot of these text Mining and uh uh machine machine learning for text from the early days I may have I may have been one of the first people to actually use this topic models LDA that David blle Andrew in and Mike Jordan wrote about uh for an IND for an industrial Consulting project so oh wow can you tell us a little bit more about that oh I mean so it's with a well-known tech company right so hired the as Ry to analyze our Tech sources unstructured Tech sources which include job posting all the job postings in the US okay and things like that and to just uh get them give them strategic advice so I thought I thought uh using uh LDA and topic models would uh provide some uh quantitative basis for the advice we were giving and from what I understand uh it did change the direction of this major major tech company tech company that everyone uh uh knows about uh so I guess one thing that I've been meaning to ask you for a while now actually is you know you've now got five strata conferences right and the two AI conferences and I think you're involved in all of those is that right yeah so I'm responsible for the program for all of those uh so basically everything the the we have at at these conferences we have two day trainings we have tutorials and we have sessions and then we have Keynotes so basically I I'm ultimately the person responsible for uh the the lineup for all these conferences okay and so my question then is do you actually have any time for doing data science at O'Reilly nowadays or are you uh how how could you possibly with all of these conferences you know this the you know I I would say that it's become it's become less and less right so I think uh I think as we kept adding more conferences and you know as you know many of these conferences are are spread out across geographic regions right so um so we have conferences in the us but we have conferences in Europe and Asia and so yeah I've I've I found myself with less and less time so to be to be honest so now I'm much more uh less of a practitioner more of basically uh Watcher from afar but I think that you know my background and my ability to read the original papers and talk to the researchers many of which whom I've known for many years I think I still have kind of some feel but maybe not as much of the hands on Hands-On field that I would like but I would say that uh the tradeoff for that though is I've gained a much more Global Perspective right so I I I don't know uh to what extent you've organized events but uh you know I mean uh you for us particularly since we organized events across the world you can't really just take your set of speakers from California and uh and take them to somewhere else right so you really have to know the local the local scene the local companies yeah uh the local communities and so uh to me I kind of uh found I found that very kind of rewarding right so that I I know I I know the the data scene in South e Asia in China in Europe and things like that so mhm well one thing that if we could maybe spend some time talking through a little bit of the kinds of problems that you tend to see at O'Reilly with regard to you know data science and machine learning in Ai and the types of approaches and techniques that you are using to solve them I think uh listeners would enjoy hearing a little bit of that detail I think at this uh like many companies uh this is not going to be a surprise uh O'Reilly is an older company it's not a startup so we have we do have uh uh many many different systems um some some some are kind of the new bleeding edge open source system some are older uh proprietary systems right and so actually to believe it or not one of our main problems is still to disate data integration you know I mean just CU we do have many many uh systems just uh getting the data all together in one place is it remains a challenge yeah and then uh beyond that I think uh luckily we do have a team dedicated to the data engineering part okay but it's it remains a a for us it remains a work in progress because uh we also keep adding adding systems that we're using you know cuz there's many many uh software as a service these days right so so different parts of the company starts using uh start using different things so that's one problem the other problem is still I think a lot of our a lot of our data is still unstructured right so I mean I guess there's some structure I mean so if you think about books there's some structured there right so uh but we also now I don't know to what extent you're following uh our safari platform uh which is increasingly relying on for example video right so uh okay um so so there's a lot of of uh uh data that we rely on that remains unstructured and one of the one of the challenges for us too as a kind of a media company that is uh building a learning platform for training is to take all of these uh many many data sources both structured and unstructured and organized um so it turns out actually that uh you know uh search and uh a a nice human curated taxonomy is still uh kind of do Remain the basic problems for companies like us right so let's say for example you wanted to learn something in uh in a new field of machine learning right so we may have thousands and thousands of sources because our safari platform doesn't just rely on our content it relies on our content Partners as well so we will have to organize you you can do a search so that's one way for you to probably navigate or Safar Safari platform but increasingly we we've we're finding that the people want kind of curated content right so how how do I learn about this new topic well we'll uh we'll use kind of the combination of humans and machines to organize a learning path for you or or a or a a resource center inside Safari so so I think that whole uh taxonomy creation uh and uh grappling with uh data integration and unstructured uh data uh those are our main challenges yeah and have you invested much in uh in personalization using you know some of the machine learning and and AI techniques that that folks are using for that kind of stuff I think uh I would say we're still in the early phases work in progress yeah that's a good point to be honest because uh people learn uh people learn differently right so each individual will learn differently from another and so to the extent that uh we're we are at least on our online division we are trying to build uh the best learning platform so a lot of that will increasingly have to rely on personalization but I would say we're still in the early early stages mhm and do you have a particular you know vision in mind for how you know how the technology is put to to use uh to enable um you know a certain kind of experience like do do you have a sense for what that experience you know looks and feels like and how it's different from you know what someone might experience today and then what the supporting pieces might need to be Yeah in our case because for example uh many Safari users uh use them through their companies let's say you work for a large company that has a a uh Subs subscription to Safari um right so it it will be kind of a combination of you us serving you content personalized to you but us also serving serving you content that kind of reflects what your particular organization is emphasizing or you know uh the rest of your team members are are learning about so um okay yeah and then the other thing Sam Actually I should mention is that Insight Safari we now have live online training on uh many of the topics that your listeners might be interested in including uh Big Data infrastructure and architecture uh machine learning um data science and uh increasingly uh we're finding actually there's a lot of demand for uh content that I would describe as uh much more non-technical you know so a lot of people are grappling with they read about a specific topic you know they may not uh need to implement it right away but they need to know at a high level what it is about and should they be should they be bringing that into their company and if so right what are what are some of the uh steps they should do to uh integrate the such and such technology or technique into their uh existing products mhh well so you talked a little bit about the kind of exposure you get um you know in terms in your roles at with the different conferences uh I'm interested in kind kind of taking your temperature on you know the various Trends you're seeing out there and what you're finding most interesting ah so good so I just gave a keynote about this in Israel yesterday oh really oh awesome um yeah yeah so uh I I would say uh on the machine learning side you know last year I kind of told people that this year I think deep learning will become a machine learning technique that uh people in the data science Community will will start using uhuh um so as you know deep learning is a lot more associated with the other conference that Iran which is the AI conference right uh where they're they're grappling with uh uh data from uh images and video and audio so computer vision and speech Technologies and things like that so what we're seeing is there's a lot of hunger in the data science Community uh to see if they can take uh deep learning and use it to replace some other uh existing machine learning technique that they're using so for example some people are looking at it in terms of recommender systems uh some people are looking at it to replace how they do search rankings and things like that so that's that's uh one Trend and uh so on the on the data science side the other thing we're kind of seeing in this might at this point be much more of a Bay Area thing is that the um people are starting to talk about a a role that kind of is a hybrid between uh the classic data scientist and the data engineer so uh a lot of people use the term machine learning engineer right um so what it is is basically someone who is maybe a little stronger on the software engineering side okay right so they write Co they write code with the express intention that this might or this will be deployed into production so it's not oneoff sloppy U and then they also tend to think much more holistically so uh if we're going to deploy this into production what is our logging infrastructure what is our AB testing infrastructure and so on and and then um yeah so then uh the the the emphasis is on production uh and less on prototypes okay it's interesting how the how the role um you know these roles just keep being redefined right I think a few years ago um the big conversation was that we saw we actually thought of the data scientist as this monolithic person right that needed to know how to do everything right they needed to be you know statistically Savvy uh you know know the math behind you know the analytics and the Machine learning stuff they needed to um you know understand how to get data from all these systems because you know their reality was that they spent you know 80% of their time or more just kind of shuffling data around uh and then we started to get the you know the rol started to split off a little bit um and you started to see you know data Engineers you know being thought of as separate from you know machine learning people and in some places you pair those you know those two with uh or I'm sorry data Engineers being you know separate from your data scientists and in some places You' pair those two with you know your real professional programers software Engineers you know now it sounds like you're saying you know we kind of came up with a new name and it's you know again it's this kind of unicorn that's supposed to know how to do everything am I reading it am I reading that right um I would say no I mean I slightly yeah but I think the original term of data scientist is was exactly what you describ which was the Unicorn this is not someone who's has a PHD in machine learning necessarily right so maybe the background here is much more of this in on the engineering side and then they learn uh enough machine enough machine learning to know uh how to uh basically we build machine learning enabled products got it got it and uh part of that is that uh you know the emphasis on production means also knowing uh what to do with these models once they hit production right so right how to how to tell when a model gets stale you know and and things like that and how when do we need to retrain it so but uh I would say the the background might be much more on the engineering side and then learn uh and machine learning to start uh uh being able to move much more fast MH right so and there there might still be data scientists in the in the organization to to kind of uh uh build the prototypes uh but increasingly I think uh at least the simpler uh machine learning things uh maybe this machine learning engineer can take on and in fact actually if if you talk to companies to uh use a lot of deep learning they even have a much more specific role which they call the Deep learning engineer right so okay so then uh other Trends so on the data oh so the other thing that I I I pointed I I've been pointing out to people is this notion that um we've had a lot of progress in machine learning right so you can just read read the online Publications and there's all sorts of papers being released lots of interesting developments right so and uh that's great right so but then uh I've been what I've been telling people is imagine a scenario where nothing happens in on the research front for the next 5 to 10 years or next five years right so I mean so my my uh my position is that there's still so much lwh hanging fruit in many companies including us or Riley yeah that you can just take what we know now uhhuh you going take what we know now and implement it implement it and we're we're we're going to be fine we're still we're still going to be fine right so like I said earlier s we're still grappling with data integration right right yeah and I don't think that that you guys are necessarily unique uh among you know Enterprises I think you know there are are some you know large sophisticated Enterprises that are kind of at the the front edge of this thing um but a lot of folks are really just in the stage of trying to figure out you know where it fits and how to best apply it and where they can extract the most value uh and and how to put together the the teams to be able to do it because of the you know the talent shortage as you're well aware yeah and then to be honest there's been a lot of progress so now we have to take a lot of those ideas and uh use them um um and to a large extent actually uh I think that uh while we have been fascinated with kind of horizontal platforms I think a lot of the interesting uh applications of these machine learning models will be in kind of verticalized applications right so and I I think we'll increasingly see companies uh uh specialize in in serving these these industries and interestingly actually there's a there's a uh intersection with the ai indust ai community in the sense in the sense that uh uh while we read a lot about the general AI uh actually the way the in DC community and the investors at least when you talk to them have been investing is they're they're investing in focused applications right so sure uh whatever that whatever that might be security drug Discovery and things like that um so the other thing that I I've been talking to people about is uh training data right so we sometimes forget that a lot of uh the development in deep learning really relied on the existence of large labeled training data sets right and to the extent that if you survey companies I think that Still Remains the bottleneck right so it's not the models it's not not uh delivering uh great models it's just coming up with uh training data right and and and there there's a lot of interesting things happening right so in the Deep Learning Community you have uh these generative models to uh mostly around gener generative adversarial networks right but then also in the data science Community there's interesting work for example by my friend Chris Ray at Stanford uh who had a system called Deep dive uh and now the next generation is snorkel where they basically uh are able to take uh uh noisier data sources so they start with less labeled data supplement it with noiser data sources and then they're able to build much more accurate models so I think there's a lot of uh we sometimes forget the importance of data uh and so I think there will there will be still a lot of of interesting uh research in how we get to training data sets much more efficiently and then uh on the machine learning side the other thing that I've been I've been talking a lot about recently is uh real time or live data so here I owe my inspiration to the rice lab the the successor to the amp lab so the amp lab as people know is the lab that originated Apache msos Apache spark in Alexia uh so the new lab rice stands for Real Time intelligent secure execution so what so live data is basically basically think about an agent interacting with an environment right so a user interacting with a website a robot navigating its environment self-driving car a player playing a a computer system player playing a computer game like a an Atari game or go so there you have an environment you have an a agent interacting with an environment so in the classic reinforcement learning sense you're trying to learn a policy right so given a state of the environment what action should what action should I take but you know if you actually take a step back and kind of look at the the flow of data in in in these types of applications the first part looks a little bit like what we've been dealing with in recent years so it might look like a streaming application uh so you have data ingestion and things like that uh stream processing and things like that but the machine learning part is slightly different right in the reinforcement learning since you're trying to learn this policy you have to run a lot of simulations you have heterogeneous computer graphs and if it's truly a live application you need to have have uh millisecond latency right so so then it turns out the uh existing Frameworks we have are not able to to do machine learning in these really live Dynamic environments yeah yeah so the classic tools that been that we've been using uh aren't able to to to do the machine learning you need in these environments which I think increasingly will be much more common right so um a as people as the tools get better the use cases will become much more clear so uh so the the people at Rice laab took a look and and kind of surveyed so what's out there and then they realized the you know computational framework for these types of applications don't exist right so so uh so then they ended up building building one called Ray which is out in Alpha in uh actually we're going to do a tutorial on on on this on reinforcement learning using Ray at our AI conference in San Francisco but the other thing I wanted to emphasize here is that uh the the the interesting thing to me is kind of machine learning in this live environment right so it turns out right now that people are using reinforcement learning but there's other techniques that might emerge mhm and the interesting the interesting thing about Ray is that you can use it for reinforcement learning or other approaches so for example the open AI folks recently published a paper where they use evolutionary algorithms right so to to to solve some some of the things that uh you would do with reinforcement learning but the interesting thing is that the r people uh took that paper and then basically they just implemented the this uh evolutionary algorithm in Ray no problem so like I said uh you know I think that I think the tools are are kind of a work in progress and that includes Ray and as the tools get better then uh people will start uh using these tools and then the use cases will be much more clear but basically think about uh any any kind of dynamic SE setting where you want to be able to take advantage of machine learning mhm and right now I would say sam that uh my other conference the AI conference is definitely much more uh aware and interested and focus on on these types of applications so for example we're already seeing that the reinforcement learning content and tutorials are very popular at that conference but I think as the tools get better maybe maybe we'll start seeing the data science Community start using them too right so for example you know look back through 3 years ago deep learning would have been inaccessible to the data science Community right we didn't have the tools like tensor tensor flow mxet big DL and P torch and all these things but as the tools got better people people start kicking the tires right so uh oh oh go ahead T I was just going to uh to drill into your comment about reinforcement learning um you know typically the literature uh in talking about reinforcement learning is looking at kind of an agent exploring and uh an environment uh and trying to maximize some policy so the offs sited example is like the an agent being trained to play Atari video games like breakout and and maximize their score um do you have a sense for how this translates into the real time you know streaming machine learning example that uh or even Industrial uh you know Enterprise type uh scenarios you know I think that uh uh the use cases still need to be worked out but you can imagine uh uh personalization on a website maybe right so where you're interacting uh with a website much like you're interacting with a game right um I mean on the AI side you can you can see applications in autonomous uh vehicles and drones sure um maybe in Inventory management if you really need real time Inventory management definitely the finance people might uh might be interested in this from Trading strategies or portfolio design and then uh resource allocation if you if you imagine the scenario where resource allocation uh with the live data becomes uh uh prevalent but Ro definitely robots are in any kind of robotic environment right so a robotic application is already using it but I think I I imagine uh uh a place a scenario where uh if the tools get better uh the people will probably find The Right Use cases mhm sure sure yeah one of the other things that I find interesting about the real time machine learning scenario and and let me know if you have come across this or have any thoughts on this but there seems to be uh there seems to be in in those kinds of environments a emerging of you know traditionally your training and inference are two totally different things and when we're talking about kind of real-time streaming data more and more I I see people wanting to do things like Active Learning where they're you know they're the the learning is real time in addition to the the inference which um you know folks more easily do real time now is that what you're seeing as well uh yes to some extent I mean I think that uh it the use cases that I'm interested in uh tend to still separate the training inference I mean the use cases that I'm much more familiar with tend to still separate the training in inference uh I think there's there are people who are kind of pushing the envelope towards much more of this online learning right scenario but then but then now we start getting into the scenario I just described right so we just kind of learning really with live data and uh uh and interacting with an environment right so uh where where you have where simulations and Explorations the ability to do those at large scale at very lowly see uhh come into play and those scenarios are really quite different yeah and actually this is a good segue to the last thing I wanted to emphasize which is uh uh compute right so uh you we're in the age of big models which is deep learning big data which is the training data and and live data and big compute right so as you mentioned um uh in in this in this scenario we need everything right so we need scale true put latency all of that with low power consumption so that I think there's a lot of interesting things happening there like uh what would what what is the future infrastructure for machine learning right so I I think those are still uh uh very active areas a lot of things happening uh at a rapid clip right so MH both from uh the GPU side the CPU side uh fbga and as6 and all of that right so um yeah so I think sometimes people forget that uh um to make all of this work you still need Hardware right um right and and the and hardware and Hardware you know there's a lot of trade-offs when you get to Hardware yeah yeah unfortunately I think for a lot of people working in the space you can't forget it enough right I think over time the you know the level of abstraction is going to have to raise where people yeah can just you know have the full flexibility to do the things that they want to do without having to think about you know how that how they configure their you know jobs to run on gpus or distributed or what have you um but still I think a lot of thought still has to happen more more thought than you know than it should be right yeah yeah and but I think you know I mean I think we're getting to the point where you've got these tools for hardware and software acceleration uhuh um and then the then then the software libraries so I think that uh for most practitioners the only time they'll think about it is when they look at their bill right nice uh at Le at least for most practitioners right so but then for the bleeding edge research who have to worry really at the low level like at the level of interconnects yeah and things like that yeah cuz they're trying to break or set the record for speech recognition a lot of that still matters but for regular people it's just the cost I think is uh what what they're going to end up that's how they're going to know what they're using so uh so the other thing that or or another thing that I often enjoy chatting with you about is the you know interesting startups that are you know doing interesting things you know both in Silicon Valley and around the world what uh any anything come to mind there or you know I'm particularly interested in you know ones that we don't hear about all the time that you know maybe in you know other other parts of the world you know there's a a bunch of startups in in China uh that I you know I'm not sure uh people here have heard of but just generally around uh applying uh deep learning to whatever whatever ver vertical right so manufacturing drones uh and to some extent do uh similar applications that we would see here but for that market right so for uh speech recognition and uh intelligent uh uh chat Bots and things like that uh but I would say I would say actually so if you look at the AI so the the the one country that I think is really interesting in terms of its excitement and Fascination for AI is China right so uh just just organizing a conference there and people just are dying for for content in in in this area so in terms of startups um I would say I don't know how you feel about it but I'm really interested in the startups that are much more focused yeah because uh you know I I think the whole if you're going to do a platform it's it's going to be tough to compete with these Cloud providers right so unless you have a platform that's focused on a vertical maybe right so uh but you know I mean Amazon Google uh Microsoft they have pretty impressive uh tools for doing uh for doing uh almost the end to end right so uh for building these applications but if you but if you are very very focused I think that's where you can ex Excel um so you might be a focus focused startup on drug Discovery or even actually you can take an area like computer vision right so I happen to advise a startup called matroid that's trying to be kind of a computer fion uh enabler for for many companies uh so then they can they can they can take kind of much more of the uh uh product approach in terms of how do companies use this easily uh to solve problems right so whatever whatever it might be summarizing surveillance cameras and and things like that I mean I guess you can build all of these things yourself using uh existing tools or the cloud but uh they already have a product that uh your analysts can use non programmers right so so I'm uh I'm uh quite interested in in in companies like that I've been trying to kind of get a as you mentioned earlier uh as to whether I pay attention mentioned the finance I've been recently trying to figure out what's happening in Finance on on some of these Technologies and I haven't uh I I I don't have a quick answer right so it seems like there's a mix of of hype and reality um but Finance is kind of also a peculiar industry in the sense that maybe the most interesting things are happening in companies who don't want to talk to you yeah yeah Finance can be like that so so so recently for example I tried to I had one of my editors I uh I introduced uh him to a bunch of companies right so in finance and uh to try to get a handle on what's happening and it's just hard the the people that we think are doing interesting things don't want to talk so yeah yeah H uh what else so I think uh I think like I said data is still going to be a important thing so I think people who have interesting data who are able to take public data make it usable right so I I guess uh Chris Ray and M C cafarella have this notion of dark data right so taking data that's very unstructured very unstructured and making and infusing it with struct so that you can use it for applications uh I think there's still a lot of there's still a lot of uh competitive advantage to people who have uh good data so what's next uh well the AI conference is next literally it's coming up at the end of June uh and I mentioned to you uh that this podcast is going to uh be published at the same time we're announcing a winner of a giveaway a ticket giveaway for the AI conference and one question that I had for you was you know I attended the first one it was a great event uh had lots of great conversations there heard great talks you what's going to be different about the second event so a few things one uh I the first event was a two-day event we did not have training or tutorials so at this event we will have both so for example we have two-day training on uh deep learning with tensorflow um but and and then a bunch of other trainings and uh one that stands out is a two-day training on on uh NLP with deep learning with my friend DP Brown oh wow and DP is also DP is also the organizer of uh uh a fake fake news Challenge and so the win ERS of which we'll present at the conference is this generating on the tutorial side detecting yeah yeah and on the on the tutorial side we have a bunch of interesting tutorials from uh reinforcement learning uh in this particular edition of the conference we're actually going to offer tutorials on a variety of deep learning Frameworks right so not just sensor flow we have uh big big DL and mxnet okay we're also trying not to be a deep we're we're trying to be the industry Gathering Place for AI where you can uh learn about uh many many different techniques in Ai and how to use it in your organization or your company okay so to that end we're also off we're also going to offer trainings in non-deep learning techniques like probabilistic programming uh um and what else so the other actually so reinforcement learning is a is a popular tutorial it's emerging as I mentioned earlier the other popular tutorials have are once aimed at the non-technical audience right so how do I bring how do I manage an AI project how do I bring uh AI back into my company right um and then on the Keynotes are going to be great right so we have uh David fuchi Dave fuchi who led the IBM team that one Jeopardy the quiz show okay so he hasn't spoken in many years but he has a new research outfit called Elemental cognition so he's going to give a keynote about what they're up to which is basically they're trying they're taking one of the grand challenges of AI natural language understanding and and basically uh TR trying to uh uh come up with a system that can uh do that well um and then besides Dave giving a talk one of his colleagues will do a 40-minute session Deep dive on what uh the technology and techniques Elemental cognition uh is doing to to correct natural language understanding okay uh Josh Josh stanan bomb of MIT uh you know I've long been fascinated by uh what they do so basically they're trying to develop uh techniques that that make that help machines learn and think like people so one I think one of the things that uh deep learning is great at is uh uh perception and large scale uh and and pattern recognition but it still it still relies on a lot of data and so Josh and his crew are trying to come up with alternative methods for maybe taking deep learning and infusing it with startup knowledge uh making it more much more efficient and much more similar to how people think okay and then uh I don't know if you followed recently but a group at Carnegie melon led by Thomas sand sandol won uh uh I don't I'm not a poker player but one of these poker tournaments uh where they beat out a bunch of uh human human top human players okay U so similar in you can think of it this achievement as basically almost at the scale of alao right but people don't aren't as aware of it so he's giving a keynote about this oh wow about how how they uh won uh the tournament it sounds like a great lineup and so yeah and just like the previous conference we have sessions on many of the techniques that the people are interested in but much more our focus is uh you know we we're also going to tr try to provide a track for people who are interested in uh how to bring these ideas and uh uh Technologies and methods back into their organizations and and uh Implement them into their products okay but uh we also try to mix it up I invited a bunch of my academic friends uh are going to be speaking at the conference about the really cool things that the industry people will find interesting and maybe kind of spark a conversation and and and see how uh we can uh you know be a true Gathering Place for industry uh interested in uh building AI products awesome it sounds like it's going to be a great time and I am certainly looking forward to it um and it'll be great to you know catch up with you in person once again cool yeah yeah yeah and uh at the risk of uh being uh putting another plugin but we also have an AI conference in in San Francisco in September absolutely absolutely and we're still we're still uh uh we're I would say 80% there as far as completing the lineup but it's already looking great and I'm sure you'll you'll be there too right oh of course yep looking forward to it um is a is there a cfp still open for that or has that been closed out that's been closed out for San Francisco okay all right great well Ben thanks so much for uh taking the time to be on the podcast it was wonderful having you on and again looking forward to seeing you in a few weeks thank you sir thanks bye-bye [Music] [Music] all right everyone that's our show for today we love love love hearing from listeners about the show you can leave your questions and comments over on the show notes page at twim ai.com talk sl26 where you'll find links to Ben and the various resources we mentioned in the show and as always our quote contest continues just drop us your favorite quote on the show notes page or your social media network of choice and we'll send you a laptop sticker once again thanks so much for listening and catch you next time

Original Description

This week I’ve invited my friend Ben Lorica onto the show. Ben is Chief Data Scientist for O’Reilly Media, and Program Director of Strata Data & the O'Reilly A.I. conference. Ben has worked on analytics and machine learning in the finance and retail industries, and serves as an advisor for nearly a dozen startups. In his role at O’Reilly he’s responsible for the content for 7 major conferences around the world each year. In the show we discuss all of that, touching on how publishers can take advantage of machine learning and data mining, how the role of “data scientist” is evolving and the emergence of the machine learning engineer, and a few of the hot technologies, trends and companies that he’s seeing arise around the world. The notes for this show can be found at twimlai.com/talk/26. Subscribe! iTunes ➙ https://itunes.apple.com/us/podcast/this-week-in-machine-learning/id1116303051?mt=2 Soundcloud ➙ https://soundcloud.com/twiml Google Play ➙ http://bit.ly/2lrWlJZ Stitcher ➙ http://www.stitcher.com/s?fid=92079&refid=stpr RSS ➙ https://twimlai.com/feed Lets Connect! Twimlai.com ➙ https://twimlai.com/contact Twitter ➙ https://twitter.com/twimlai Facebook ➙ https://Facebook.com/Twimlai Medium ➙ https://medium.com/this-week-in-machine-learning-ai
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from The TWIML AI Podcast with Sam Charrington · The TWIML AI Podcast with Sam Charrington · 30 of 60

1 Engineering Practical Machine Learning Systems with Xavier Amatriain - #3
Engineering Practical Machine Learning Systems with Xavier Amatriain - #3
The TWIML AI Podcast with Sam Charrington
2 How to Build Confidence as an ML Developer with Siraj Raval - #2
How to Build Confidence as an ML Developer with Siraj Raval - #2
The TWIML AI Podcast with Sam Charrington
3 Open Source Data Science Masters, Hybrid AI, Algorithmic Ethics & More with Clare Corthell - #1
Open Source Data Science Masters, Hybrid AI, Algorithmic Ethics & More with Clare Corthell - #1
The TWIML AI Podcast with Sam Charrington
4 Interactive AI, Plus Improving ML Education with Charles Isbell - #4
Interactive AI, Plus Improving ML Education with Charles Isbell - #4
The TWIML AI Podcast with Sam Charrington
5 Machine Learning for the Stars & Productizing AI with Joshua Bloom - #5
Machine Learning for the Stars & Productizing AI with Joshua Bloom - #5
The TWIML AI Podcast with Sam Charrington
6 Generating Labeled Training Data for Your ML/AI Models with Angie Hugeback - #6
Generating Labeled Training Data for Your ML/AI Models with Angie Hugeback - #6
The TWIML AI Podcast with Sam Charrington
7 Explaining the Predictions of Machine Learning Models with Carlos Guestrin - #7
Explaining the Predictions of Machine Learning Models with Carlos Guestrin - #7
The TWIML AI Podcast with Sam Charrington
8 Deep Learning: Modular in Theory, Inflexible in Practice with Diogo Almeida - #8
Deep Learning: Modular in Theory, Inflexible in Practice with Diogo Almeida - #8
The TWIML AI Podcast with Sam Charrington
9 Emotional AI: Teaching Computers Empathy with Pascale Fung - #9
Emotional AI: Teaching Computers Empathy with Pascale Fung - #9
The TWIML AI Podcast with Sam Charrington
10 Statistics vs Semantics for Natural Language Processing with Francisco Webber - #10
Statistics vs Semantics for Natural Language Processing with Francisco Webber - #10
The TWIML AI Podcast with Sam Charrington
11 Building AI Products with Hilary Mason - #11
Building AI Products with Hilary Mason - #11
The TWIML AI Podcast with Sam Charrington
12 Reprogramming the Human Genome with AI, w/ Brendan Frey - #12
Reprogramming the Human Genome with AI, w/ Brendan Frey - #12
The TWIML AI Podcast with Sam Charrington
13 Understanding Deep Neural Networks with Dr. James McCaffery - #13
Understanding Deep Neural Networks with Dr. James McCaffery - #13
The TWIML AI Podcast with Sam Charrington
14 Scaling Deep Learning: Systems Challenges & More with Shubho Sengupta - #14
Scaling Deep Learning: Systems Challenges & More with Shubho Sengupta - #14
The TWIML AI Podcast with Sam Charrington
15 Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - #15
Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - #15
The TWIML AI Podcast with Sam Charrington
16 Machine Learning in Cybersecurity with Evan Wright - #16
Machine Learning in Cybersecurity with Evan Wright - #16
The TWIML AI Podcast with Sam Charrington
17 Interactive Machine Learning Systems with Alekh Agarwal - #17
Interactive Machine Learning Systems with Alekh Agarwal - #17
The TWIML AI Podcast with Sam Charrington
18 Location-Based Intelligence for Smarter Marketing with Klustera - #18
Location-Based Intelligence for Smarter Marketing with Klustera - #18
The TWIML AI Podcast with Sam Charrington
19 AI-Powered Customer Support with HelloVera - #18
AI-Powered Customer Support with HelloVera - #18
The TWIML AI Podcast with Sam Charrington
20 Using AI to Simplify the Programming of Robots with Cambrian Intelligence - #18
Using AI to Simplify the Programming of Robots with Cambrian Intelligence - #18
The TWIML AI Podcast with Sam Charrington
21 Increasing Efficiency of Healthcare Insurance Billing with NLP, w/ Behold.ai - #18
Increasing Efficiency of Healthcare Insurance Billing with NLP, w/ Behold.ai - #18
The TWIML AI Podcast with Sam Charrington
22 Creating a Worldwide Financial Knowledge Graph with AlphaVertex - #18
Creating a Worldwide Financial Knowledge Graph with AlphaVertex - #18
The TWIML AI Podcast with Sam Charrington
23 From Particle Physics to Audio AI with Scott Stephenson - #19
From Particle Physics to Audio AI with Scott Stephenson - #19
The TWIML AI Podcast with Sam Charrington
24 Selling AI to the Enterprise with Kathryn Hume - #20
Selling AI to the Enterprise with Kathryn Hume - #20
The TWIML AI Podcast with Sam Charrington
25 Engineering the Future of AI with Ruchir Puri - #21
Engineering the Future of AI with Ruchir Puri - #21
The TWIML AI Podcast with Sam Charrington
26 Deep Neural Nets for Visual Recognition with Matt Zeiler - #22
Deep Neural Nets for Visual Recognition with Matt Zeiler - #22
The TWIML AI Podcast with Sam Charrington
27 Introducing Psycholinguistics into AI with Dominique Simmons- #23
Introducing Psycholinguistics into AI with Dominique Simmons- #23
The TWIML AI Podcast with Sam Charrington
28 Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - #24
Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - #24
The TWIML AI Podcast with Sam Charrington
29 Offensive vs Defensive Data Science with Deep Varma - #25
Offensive vs Defensive Data Science with Deep Varma - #25
The TWIML AI Podcast with Sam Charrington
Global AI Trends with Ben Lorica - #26
Global AI Trends with Ben Lorica - #26
The TWIML AI Podcast with Sam Charrington
31 Intelligent Autonomous Robots with Ilia Baranov - #27
Intelligent Autonomous Robots with Ilia Baranov - #27
The TWIML AI Podcast with Sam Charrington
32 Reinforcement Learning Deep Dive with Pieter Abbeel  - #28
Reinforcement Learning Deep Dive with Pieter Abbeel - #28
The TWIML AI Podcast with Sam Charrington
33 Robotic Perception and Control with Chelsea Finn  - #29
Robotic Perception and Control with Chelsea Finn - #29
The TWIML AI Podcast with Sam Charrington
34 Natural Language Understanding for Amazon Alexa with Zornitsa Kozareva - #30
Natural Language Understanding for Amazon Alexa with Zornitsa Kozareva - #30
The TWIML AI Podcast with Sam Charrington
35 The Power of Probabilistic Programming with Ben Vigoda - #33
The Power of Probabilistic Programming with Ben Vigoda - #33
The TWIML AI Podcast with Sam Charrington
36 Intel Nervana Update + Productizing AI Research with Naveen Rao and Hanlin Tang - #31
Intel Nervana Update + Productizing AI Research with Naveen Rao and Hanlin Tang - #31
The TWIML AI Podcast with Sam Charrington
37 Video Object Detection at Scale with Reza Zadeh - #34
Video Object Detection at Scale with Reza Zadeh - #34
The TWIML AI Podcast with Sam Charrington
38 Enhancing Customer Experiences with Emotional AI, w/ Rana el Kaliouby - #35
Enhancing Customer Experiences with Emotional AI, w/ Rana el Kaliouby - #35
The TWIML AI Podcast with Sam Charrington
39 Expressive AI-Generated Music With Google's Performance RNN with Doug Eck  - #32
Expressive AI-Generated Music With Google's Performance RNN with Doug Eck - #32
The TWIML AI Podcast with Sam Charrington
40 Smart Buildings & IoT with Yodit Stanton - #36
Smart Buildings & IoT with Yodit Stanton - #36
The TWIML AI Podcast with Sam Charrington
41 Deep Robotic Learning with Sergey Levine - #37
Deep Robotic Learning with Sergey Levine - #37
The TWIML AI Podcast with Sam Charrington
42 Deep Learning for Warehouse Operations with Calvin Seward - #38
Deep Learning for Warehouse Operations with Calvin Seward - #38
The TWIML AI Podcast with Sam Charrington
43 Cognitive Biases in Data Science with Drew Conway - #39
Cognitive Biases in Data Science with Drew Conway - #39
The TWIML AI Podcast with Sam Charrington
44 Data Pipelines at Zymergen with Airflow, w/ Erin Shellman - #41
Data Pipelines at Zymergen with Airflow, w/ Erin Shellman - #41
The TWIML AI Podcast with Sam Charrington
45 Web Scale Engineering for Machine Learning with Sharath Rao - #40
Web Scale Engineering for Machine Learning with Sharath Rao - #40
The TWIML AI Podcast with Sam Charrington
46 Marrying Physics-Based and Data-Driven ML Models with Josh Bloom - #42
Marrying Physics-Based and Data-Driven ML Models with Josh Bloom - #42
The TWIML AI Podcast with Sam Charrington
47 Machine Teaching for Better Machine Learning with Mark Hammond - #43
Machine Teaching for Better Machine Learning with Mark Hammond - #43
The TWIML AI Podcast with Sam Charrington
48 LSTMs, Plus a Deep Learning History Lesson with Jürgen Schmidhuber  - #44
LSTMs, Plus a Deep Learning History Lesson with Jürgen Schmidhuber - #44
The TWIML AI Podcast with Sam Charrington
49 Learning From Simulated & Unsupervised Images through Adversarial Training - TWiML Online Meetup
Learning From Simulated & Unsupervised Images through Adversarial Training - TWiML Online Meetup
The TWIML AI Podcast with Sam Charrington
50 Jennifer Prendki Interview - Agile Machine Learning - TWiML Talk #46
Jennifer Prendki Interview - Agile Machine Learning - TWiML Talk #46
The TWIML AI Podcast with Sam Charrington
51 Evolutionary Algorithms in Machine Learning with Risto Miikkulainen - #47
Evolutionary Algorithms in Machine Learning with Risto Miikkulainen - #47
The TWIML AI Podcast with Sam Charrington
52 Learning Long-Term Dependencies with Gradient Descent is Difficult - TWiML Online  Meetup
Learning Long-Term Dependencies with Gradient Descent is Difficult - TWiML Online Meetup
The TWIML AI Podcast with Sam Charrington
53 Word2Vec & Friends with Bruno Gonçalves -#48
Word2Vec & Friends with Bruno Gonçalves -#48
The TWIML AI Podcast with Sam Charrington
54 Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan  - #49
Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan - #49
The TWIML AI Podcast with Sam Charrington
55 Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
The TWIML AI Podcast with Sam Charrington
56 Intel Nervana DevCloud with Naveen Rao & Scott Apeland - #51
Intel Nervana DevCloud with Naveen Rao & Scott Apeland - #51
The TWIML AI Podcast with Sam Charrington
57 AI-Powered Conversational Interfaces with Paul Tepper - #52
AI-Powered Conversational Interfaces with Paul Tepper - #52
The TWIML AI Podcast with Sam Charrington
58 Topological Data Analysis with Gunnar Carlsson - #53
Topological Data Analysis with Gunnar Carlsson - #53
The TWIML AI Podcast with Sam Charrington
59 ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
The TWIML AI Podcast with Sam Charrington
60 Ray:A Distributed Computing Platform for Reinforcement Learning with Ion Stoica -#55
Ray:A Distributed Computing Platform for Reinforcement Learning with Ion Stoica -#55
The TWIML AI Podcast with Sam Charrington

Related Reads

📰
RQLM Part 3: The Instruction Pointer Was Already There
Learn about the concept of a synchronous quantum-hybrid machine and its implementation using a six-address quantum microcode sequencer
Medium · Machine Learning
📰
Do notebook ao deploy: por que tantos projetos de Machine Learning param no meio do caminho
Many Machine Learning projects fail to reach production, learn why and how to overcome this challenge
Medium · Machine Learning
📰
Do notebook ao deploy: por que tantos projetos de Machine Learning param no meio do caminho
Many Machine Learning projects fail to reach production, learn why and how to overcome these challenges
Medium · Data Science
📰
Adaptive Reasoning: Why the Smartest AI Models Are Learning to Think Less
Learn why overthinking can be a major issue in AI models and how adaptive reasoning can help mitigate this problem
Medium · Machine Learning
Up next
The Adam Optimizer is Just Momentum + RMSProp
DataMListic
Watch →