Solving any Machine Learning Problem | Approach and Steps Involved
Key Takeaways
The video explains the approach and steps involved in solving any machine learning problem, covering the fundamentals of machine learning and providing a practical guide for creating machine learning models. It discusses the importance of understanding the problem statement, data preprocessing, model selection, and evaluation metrics.
Full Transcript
foreign countries and good morning good evening for people who are from different parts of the world it's it's lovely to interact to a lot of people however I had a good introduction about the course beforehand about what it's going to be with the person who has connected to me uh Malik and talking about what are we going to do today sorry I have a little bad throat so I'm expecting all of you to please forgive me for that for the day it's it's like beyond control so what are we going to do today before I get into that um I'll take just half a minute to introduce me as raksha has did it very wonderfully uh keep in mind that the session is not about me the session is about the topic but uh what among that what about my experience is required for the session I'm just going to bring in those points so I am a corporate trainer I have spent most of my time uh training people I might look very young and actually I'm very young but I have spent a lot of time in in recent past working with people from Cisco Amazon QA simply learn and and logical companies so that that's what made me uh comfortable to teach on this topic here and my primary domain has been machine learning and data science uh why this information is required to you in just one second okay sorry yeah why this information is required to you is just to bring in a point that whatever I'm going to uh say today might go a little against your standard belief when I'm talking about those who are only students those who are working professionals May uh find my words very in line with their experience but those who are students you may find my words going little against your standard beliefs about machine learning or about data science but I want you to trust me or I should say Just note down these points and do your own research before you believe me on any of the point because whatever I'm going to bring in uh May contradict your educational knowledge and may give you a perspective from the industrial Viewpoint also uh I saw the poll poll result in the very start and I found there are people from different uh kind of experience there are people from six plus years of experience uh there are people from two to three years of experience honestly I don't even have six plus years of experience myself so it's glad to have an interaction with you don't take it like a training take it like a discussion and uh please do contribute from your side also in whatever way you can uh raksha just a quick question can people uh speak up by themselves or do they need any kind of commission from our site yeah actually uh they can interact with the chat only and okay perfect perfect so I will be keeping chart posted on my screen all the time so whatever question you have please come up on chat and whatever discussion whatever point you want to add on I just shouldn't just limit you to questions if you want to add on to any point please come up uh coming down to the topic like why this topic what this topic has to do with the state uh so I was thinking like what can I do in one hour what can I add on to people who may or may not know machine learning so I came up with this topic that there are a lot of students and machine learning being the buzzword around industry like you go to any coaching center tuition Center if they deal with it across the globe they must be dealing with python and machine learning which is making machine learning seems so easy but it is not like it may look very easy if you are looking from a college perspective but in the industrial scale it is not it is one of the Nightmare and and few steps in that are one of the nightmare in across the industry which is about data so I am going to talk a lot about um how do I being a developer or how do uh I being a Trader who interact with a lot of developers across industry from the small startups to all the way Amazon and Cisco what are the challenges this this world of machine learning Engineers face in the real world and if you're planning to be one of the machine learning or data science engineer you're most likely going to face the same challenges and I expect you to have your you know paper pencil or a notepad or some digital notepad or something be with you to keep with you and to type down or write down whatever we are discussing at least the bullet points so that you can think on that and research on that later in this entire hour I will not be teaching you anything I will not be coming up with a course kind of content that this is topic one and this is topic two no it's not possible but I will be giving you a lot of tangents to think upon I'll be giving you a lot of viewpoints a lot of perspectives which if you spend some time on they will really be going to help you out okay so I'll start my question with a very simple and a cliche uh query here that what is machine learning um I believe most of you have the idea even if you have taken a professional course or not you must be having an idea for this machine learning please put down your message on chat let's see what do you understand by the word machine learning everybody like do not worry about whether it is correct or incorrect what do you think is machine learning according to you please put out okay I got one more message it says it's by Chandra shekar it says learning from data another one is by predicting and analyzing using data analytics okay trying to predict okay sorry how the machine learned from its operation and does it iteratively give machine the capability to act and think like human okay perfect I'll give you a quick example which makes everybody on the same page about this question what is machine learning imagine a kit a very young kid and you know a kid which cannot even walk and and just struggling to crawl and have very little common sense no such belief systems Etc and the kid is crawling and on a table there's a cup of tea please cup of tea or coffee or anything which is hot and it is placed up there and the kid accidentally comes in and touch that cups touch uh that cup or the cup or the tea liquid gets spoiled onto to his or her hand okay what it is going to do to that kid the kid doesn't know that this is hot he doesn't know the meaning of the word hot the kid doesn't know the meaning of the word tea or coffee or liquid or gas or anything like that the kid just know one thing that this is a bad experience not in English but in terms of feeling the kid realized that this is a bad experience and the next time whenever that kid sees something like a cup placed on a table or any anywhere uh some sort of steam coming out of that the kid will restrain not to touch that object even that kid had no knowledge about temperature physics thermodynamics and all the concepts you know as an engineer the kids still understand that this is bad and this is good that's the level of intelligence that kid Is possessing at a young age of maybe three four five months okay or could be as young as one or two months and we as we grow older we keep on bringing up our Intelligence on the next step now where this intelligence is coming from uh how can a kid which is just two three four months uh old how can I get possess that knowledge because of a lot of genetic information placed in that kid's mind or neural system even before his or her birth if I say I am like 25 years old my brain could be having information millions of years old because of my parents and grandparents and great grandparents and so on that's why so many things comes naturally to us we learn walking in the age of one year uh and it's it's one of the very complex operation because of lot of genetical information placed in our plate make sense now when we talk about machines machine is just a piece of Hardware that's built from scratch has no genetical information has no brain has no neural network but if by the help of softwares we can help that machine to possess some sort of intelligence to decide what to do when and what not to do when is called as artificial intelligence there's a very clear difference you if you have worked on any programming language you may argue that hush this is the same thing that can be done with if else based programming right switch case or if else or else if all of them in every programming language yes but the differences in if else you are the one who decide what to do and machine does that you write that if x is greater than 5 do that else do something else okay but in case of machine learning you let the machine decide what decision to take and when that's the classical difference between a rule based programming and intelligent space program so given a quick brief about what is machine learning it is the ability it is one of the ability for machine to take decisions without hard coding those decisions if I give you one more quick example which makes which will make a lot of sense in your mind when you learn driving a car or a bike or a cycle or anything like recently I have learned couple of years back driving for how to drive a car and my elder brother was teaching me up uh wherever we were going on a road and I was driving that person my brother was guiding me with broader inputs like he was saying that okay there's a higher traffic you take a left and uh there's a higher traffic on the right side you go straight and then you go right Etc nobody was giving exact numbers like turn 80 degree right walk for five seconds and then turn 70 degree left and then put your speed below 25 kilometer per hour or anything like that but when you work with machines in the classical traditional environment you have to hard code everything you cannot instruct the machine that go a little slow there's no meaning of a little is the meaning of 20 kilometer per hour or 22 kilometer per hour or 15 kilometer per hour you need to hard code but with machine learning as Tesla and a lot of new companies are doing up that task you can bring up the form of decision making in the machine itself that's all about artificial intelligence and machine learning in a very general and simpler way okay uh I got a question or I'll take up the questions a little later let me have my points cleared up beforehand and whatever questions will be there I'll just keep taking it up because there's so many of them walking around the chat box and and it will be really difficult for me to take up in in like in synchronization okay so what are we going to talk about is how do you approach a problem and how do you design such kind of solution uh the very first thing the very very first thing is about understanding the requirement what is the requirement and understanding the possibility I'll give you an example my name is Harsh can you please tell me what is my mobile number you seem to be very intelligent forget about machines you as a human all of you must be like about 15 16 years of age I believe okay so few of you can be 25 30 35 years of age you are very intelligent you have done a lot of good work uh look up on phone book um I don't think so that's any kind of intelligence that's like a processing let me ask you another question my name is Harsh can you tell me what's my blood group I got the silence right that that's the answer that says that you cannot even I cannot right there are scenarios like this which are not solvable but if I say you that hey can you look at my uh face can you look at my video feed and can you predict what will be my age you can all of you will come down a number of 22 23 24 25 and and that's primarily correct right so you can do some kind of tasks considering it intelligent and you cannot do some kind of task considering it um like a fake guess uh rather just a way guess not not any form of intelligence so there are few things which have some sort of correlation this is called as correlation my facial image my video camera feed has something to do with my age because you are looking at my facial features you're looking at the wrinkles you're looking at the skin tones you're looking at a lot of facial features like we do and age is a derivative of that but my blood group is not a derivative of my name so you cannot predict that the very first thing we need to understand in approaching any machine learning or any intelligence based problem is is it solvable is it feasible does it have correlation if Amazon comes to me and say that hey we have this data in this so many people are coming in and they are purchasing and they are doing transactions I want to know next year what will be my total sales maybe there is some sense we can predict because based on past Trends based on current trends we can predict what will be the sales in next year but if Amazon comes to me and say that hey we are going to change our logo and looking at our data can you please tell me what color we are going to bring in not the question like what color we should bring if they are asking you what color we are going to bring in in logo it's it's a not correlated data or not correlated answer to predict form so the very first thing in a machine learning problem statement we do is we understand the requirement and we understand the need and we understand the feasibility uh whenever you learn machine learning uh nobody talks about that even I take a lot of machine learning courses for students and even I don't talk about that because they are limited hours courses 10 hours 15 hours 30 hours we talk about technology we don't talk about the business side effect but when you approach it in business you don't want to spend 30 40 50 hours working on a project and then realize that this is not even possible so the very first thing you should be doing is checking the feasibility checking the requirement the next thing you do is you collect the data now this point can go a little against your standard beliefs a lot of academic people a lot of students think like you will go to kaggle and you will download the data and then you will pass it to a DOT fit function and it will work it's not like that if you are planning to spend three months on a machine learning project trust me you will spend more than one and a half month only and only to collect the data two reasons for that one you need volume of data and two you need a very specific data I'll give you an example from India once again there's a company in Delhi called as chayos most of you might have heard of it if you are from the NCR it is in different cities also so chios is a tea brand as the name the word chai which is a Hindi word for tea so if you are not able to relate to it you can imagine cafe coffee day or Subway or McDonald or any of them basically the example is for trios but you can imagine any other country I was working with chayos and the tasks given to my team was to reduce their Dustbin ratio what is trust wind ratio these standard uh operating procedures their standard operating operating procedures ask them to throw something Beyond a 48 Hours of timestamp so if they make a sandwich today or just now if they make a sandwich at 12 noon on Saturday they are not supposed to use it after 12 noon of Monday okay as per trios different companies can have different timestamps but they are not supposed to use it even if they have 10 customers in front of them and they don't have any fresh sandwich they are still supposed to throw that so this is called as trustwin ratio that if they are making 100 sandwiches uh their number was 4.5 that 4.5 are going to dustbins so they wanted to reduce it of course the very first question I asked them was what's your problem statement that was clear uh that was kind of feasible also that we need to actually predict the demand and accordingly we need to produce the supply but the next question was do you have data and the answer was sure no that we do not have any form of data we do not have any CRM we do not have anything the only thing we have is our billing software and which is a third party software which does not allow us to store or get the data at a later stage that's a big full stop for me to proceed with the project and this is a real example by the way so we need to spend some three to four months uh wherein in the initial one week we built a tool a data collection tool we connected that to their Billing System and we just waited for three to four months for a lot of data to be connected every time they are making a sale our CRM was getting the data their Billing System was processing the bills and we were doing nothing we were just waiting for a lot of data to pile up in next like three to four months imagine like chai was doing some 500 transactions in a day we got 500 multiplied by 100 days of data approximately and that was a decently good number to start with so uh the very first thing again you should be focusing on after requirement Gathering is what's my source of data do I have data do I have to generate the data will I get the data in a structured format or will I get the data in a non-structured format is my data authenticated is my data actually correct or not I believe I'm clear so far that what exactly is the meaning of collection of data and why it's important a quick yes on the chat if you understand this okay anybody have any question perfect I got a lot of yes yes yes thank you so much okay so the next thing once you get the data then you should think of working with any of the tools you have learned like python R Matlab mat.lab or anything you know about that's where exactly you should be starting up with that too a lot of things before that whatever we have discussed these two steps should be on a paper pen or a PPT or some kind of managerial tool it's not on any coding tool once you get these two steps done then you should think of that I am a data scientist or a machine learning engineer I do my work using python or r or so I do my work in Python so what I will do in any of the scenario is I will once I got the data I will load my data into python I believe issue no machine learning you know the step you'll use pandas you load the data using read CSV in most of the cases or read Json or read txt or whatever way so I will not stress much upon the technical steps because they are like part of every single course you go through uh my stress will be on those steps which are not very clearly uh taught about so you will load your data once you load your data a lot of people just start doing technical things how do you identify the features to collect again that's about your requirement what are your requirements and in the initial stage you actually correct everything anything and everything that comes your way because you really don't know what could be relevant and what would not be the problem statements can also change at the latest okay you got this problem solved that's all our second problem then we cannot wait for three or four months again to collected it if you're collecting data we should collect like everything storage is very cheap like you get one GB of storage and hardly one rupee So when you buy for when you buy servers or large size smart businesses so that's very cheap you can just collect anything and everything that's coming your way so uh once you got your data and once you load it in Python do not feel eager to start doing something on your data do not feel eager to clean process okay uh you you have this function and that function and your knowledge and you have to apply no hold on firstly spend some time and understand that data firstly understand is that data correct or incorrect is that data relevant or relevant sometimes your correct data can also appeared I'll give you an example so uh let's say make my trip I believe all of you know those who are from outside India those who are from Africa UK region Thomas Cook is the next example so they are travel companies and these travel companies usually comes up with discount coupons and let's say you are the make my trip machine learning engineer and the task given to you is to find out how much percentage discount coupon should be launched should it be five percent ten percent fifteen percent or what of course you can launch a hundred percent discount coupon customers will be benefited but you will be in loss and you can launch a two percent discount coupon you will save your money but customer will not be benefited so you need to find an optimum number of number of percentage where the customers will also get the benefit and you will also be able to manage your Finance so in order to decide this you roll out a survey form and you ask people for where you are from what do you do what's your salary what's your net worth uh how frequently do you travel in which class like first class premium business class and which type of transport you transport mode you travel Etc and you get a lot of data wherein if I talk about in Indian rupees private salary of people annual salary of average annual salary of people are like five lakh four lakh six lakh 10 lakh and all of a sudden you see somebody's average salary is like 100 crores two reasons one that is accidental the person uh could be actually internal or could be intentionally wrong the person is just mocking with you like so many times we we see a form and we have no interest of our laptop and we feel random value uh could be it is accidental maybe the person wanted to write 100 000 which is one lakh or it could be like 10 lakhs 1 million but the person accidentally wrote like 100 crores uh or could be the person's actual salary is 100 crude what if Elon Musk is spelling your form but for any which reason for whatsoever reason because you are not launching a coupon code for Elon Musk or just resource or Mukesh Ambani you need to ignore that value you should not be considering that okay is this right is this wrong uh what if this hundred crore is the right value true that could be right but that person is 100 not your customers if a person is having 100 crores of annual salary and those who don't know what sounded crores it's like 1 000 Millions okay maybe like 1 billion yeah one billion uh Indian rupees so whatever the higher value is that person is not your Target customer you are not making your decision according to that person's salary because if you do if you do launch a coupon code for a person who is having 100 crores of salary you are completely deviating from your data collection requirement you are supposed to launch it for for the common people who may have a salary of two three four five lakh rupees per annum or maybe half a million uh Indian rupees per annum for those who are outside India okay so you should spend a lot of time finding your outliers in the data finding the types of data finding the patterns and data for more things like people don't understand there are four categories of data if you know or if you don't know uh I am typing it down on chat it's called as Noir where these are acronyms so n starts for nominal o stands for ordinal I stands for Interval R stands for range nominal and ordinal typically we classify data into two parts discrete and continuous can you also elaborate more about forms of data collection actually there are no static ways like there can be anything and everything that comes your way you may need to sit being a typewriter you need to type data at times you may need to build apis you may need to connect connect to pre-built apis you may need to download data there are a variety of ways like there are no static ways more information you have about it as much information you have about it that's going to help you if you are also a web developer it will be very easy for you to build apis if you are not a web developer you may need to hire somebody to build apis basically you need to have a lot of options in front of you while collecting it and case to case basis it will differ but very rarely you will get a CSV pre-built that you can take this and do the work whether it's a startup or a company like Cisco people firstly do not have the data and secondly people don't want to share the data so you need to generate your own data that's like the bad thing about data science okay coming back to the acronym part Noir so uh typically we divide our data into two categories which is discrete or categorical uh sorry discrete or continuous but we forget to divide it even further when your data is discrete it can still be divided into two parts one is called as nominal and one is called as artinal I'll give you an example I'll ask you a question that there are five names written let me pick five names from the attendance uh this of this webinar let's say it's Aditya a burner and Jerry and David and Chandra okay can you order them in such a way that it reflects something about their personality the only thing or the only way you can order is maybe alphabetically or maybe the time they have joined this meeting but you cannot actually order them in any such way which reflects some information right if somebody's name is Chandra and the initial word is C that does not reflect anything if somebody's name is Daniel that does not reflect anything you cannot put this logic that D comes after C so Daniel's height must be higher than chandra's height or anything like that this is like not making sense but if I give you another categorical data by the way this name is a categorical data if I give you another form of categorical data let's say I have uh three air conditioners in my house example what is a Five Star air conditioner one is a Four Star air conditioner one is a three star air conditioner can you imagine some sort of order in in your mind about them actually yes you can estimate the order of their price the five stars uh five star AC must be expensive than the four star and three star AC in a conventional way it's possible that the three star AC is an important one and it's much more expensive but in a general conventional way it's supposed to be cheaper the five-star AC is supposed to have higher lifetime the five star AC is supposed to consume lesser amount of electricity the five star is supposed to you know look better than a three-star AC again in a conventional way so these are few uh personality traits so to say about that data you can develop from the order that is mentioned so sometimes the order of the data reflects something and sometimes it doesn't if it reflects and I'm talking only about categorical data by the way right five star four stars research is categorical it's not continuous names are also categorical if the order of the data is reflecting something that is called as Cardinal data that's why the name ordinal coming from the order and if it is not reflecting something it's called as nominal data because most of the time these things are actually your names names email IDs pan numbers other card numbers password numbers Etc so they are called as nominal okay on the other side when you are into category sorry when you are into continuous site there are again two forms of data interval and ratio I'll give you an example for both of them [Music] my age is 25 years and my brother's age is 22 years okay value or let me say my age is 25 years and my brother's ages five years and I give you two options to pick for from first option is will you say I am five times Elder to my brother or will you say I am 20 years older to my brother pick the option A or B B lot of messages for B why because the equation five times is just valid for now next year it will change I will be 26 and he will be six now the five times doesn't hold valid but the equation of 20 years Elder is going to be fixed because the difference in our data is not of a ratio type it is of an interval type we have an interval gap of 20 years which will remain constant throughout the life but on the other side if I say I earned five lakh rupees this year and I have to pay 15 lakh rupees of tax that's 30 percent okay so not 15 lakh 1.5 lakh rupees of tax that's 30 and uh do you count the tax status as a ratio or do you count it as an interval will you if if your income increase one lakh by the next year will you increase your tax by one lakh or will you increase it as a percentage of one lakh that's where your ratio type which comes in so whatever type of data you have firstly divide them and you don't need to have a tool or a technology or a function for doing this it's your understanding when you have a data firstly you understand whether it is categorical or it is continuous if it is categorical understand if it is nominal or ordinal if it is continuous understand if it is continuous or sorry sorry there are a lot of tongue twister if it is continuous understand if it is in interval type or in the ratio type that will give you a lot of sense a lot of information about what to do with such kind of data okay moving ahead you need to check outliers you need to check the authenticity of the data uh what could be an example of authenticity let's say you are uh I'll give you a cliche example that's a data on kaggle and the data is about taxi fare prediction in New York most likely many of you might have touched upon that data earlier so the data and the task is about predicting the taxi fare in New York City and there are a lot of information like the pickup longitude latitude the drop of longitude latitude and how many passengers what was the time of the day what was the day was it Sunday Monday Tuesday Etc and so many other things now if you are doing an analysis of New York and you find a longitude latitude of India or of Dubai or of UK or of or any other city in USA that is purely invalid right now by looking at a number you will not realize that this is an invalid number but you need to put your domain knowledge your your overall understanding your brain ought to work as a data scientist to firstly study what are the longitude latitude range of New York City and then to filter your data that is there any longitude over and above this or below this is there any latitude above this or below this and if there are you need to discard those data or you need to rectify those dates because those data are not talking about New York could be accidentally they are wrong they were part of New York but somebody changed their longitude but could be they are like misplaced here okay but for whatever reason they are not valid data points so you have to clean your data you have to check whether my data is purely cleaned or not uh why divide it into two parts in the first phase I call it Eda I observe everything I don't perform anything I do a lot of read-only tasks like I plot graphs I plot or distributions I check the statistics I check the frequencies of everything occurring by check the minimum maximums out suppliers Etc and then I just note down things on my diary on my paper in the next this is called a city exploratory data analysis this is third step in my series first step is requirement Gathering second step is data collection third step is Eda the fourth step is I will act upon the the results I have received from Eda I will clean my data I will remove the outliers I will replace the values I will remove the missing values I will fill up those values etc etc okay I'll act upon curating those problems which I found out on Ada what could be the next step anybody the next step is very conventional I believe all of you know about that you will split your data into features and labels you'll separate your features with your labels okay you do a vertical split in which you say that how do you do that very simple there are drop functions for that or there are column selection operations for that one of any of them you can use it and the next step is you you do is you do a train test split again that's where you need to make a decision that what kind of mechanism are you going to pick for train test split are you going to do it in the conventional way like a random Trend test plate are you going to use k-fold are you going to use stratified K4 are you going to use splitwise k-fold are you going to use leave one out a full Etc there are so many options for that just give me a minute for turning off the video okay uh there are so many options available for that particular purpose and you need to decide on the basis of your understanding of two things one how well do you understand data and from where your answer will come from the third step that's PDA how well do you understand data one secondly how well do you understand each of the operations so whenever I start a machine learning course or somewhere in the middle I very frequently give this example suppose you have to go from from suppose you're a king or a very rich person and you have so many cars with you and you have to go from New Delhi to Bangalore those who don't know what is New Delhi to Bangalore like you can imagine any too far distant cities in your country okay for example New York to California or any other two countries in any country now which car will you pick it will depend on your understanding of two things first how well do you understand Terrain do you know what kind of terrain you are going to face from New Delhi to Bangalore I am from India I have worked in Bangalore I have worked in New Delhi and I have traveled a lot I know that it's a kind of plateau terrain it's most most of that is the urban area so I will pick a car which is more suited for urban areas okay for for developed areas I will not pick a an off-road car or I will not pick a very high high speed so sports car or anything like that and more likely pick a sedan or a hashbrown an SUV something the second thing is how well do you understand your car you know that it is a Terrain which is made up of lot of Urban Roads high quality roads but then do you understand about your do you understand things about your cars do you know which car is suited for long drives do you know which car is suited for uh cement roads and which car is suited for Dahmer roads and which car is suited for Desert roads and which car is suited for Icy areas and which car is suited for mountains those are the two things which you need to understand same here one question is how well do you understand your data are there any slides for the webinar actually no I I do not uh ever go inside whether it is Corporate Training or student training so first thing you need to understand is how well do you understand your data that's your terrain uh okay which terrain you are operating in is that a clean data balance data imbalanced data not clean data skew data or what kind of data it is and the second thing is how well do you understand your vehicles that's your machine learning algorithms which kind of split will you use which kind of regression model will you use which kind of classification model will you use so by an understanding of both of them in a deeper context you will be able to make up a good decision somebody just asked me for splits that which is the best split model something like that was a question I didn't see the exact wording um I'm still which would be a good approach for splitting this is the exact question so uh see there is no best approach so to say because imagine there is a best approach then then the other approaches must not be existing everybody will use the best of both the art in machine learning is to pick the option among the five or six available options pick the right option that's that's your task it's called as data science it's not called as you know a data technology unique it's science it's an art that's why people are getting paid very much higher because if you talk about the code it's just a 15 line of code which can predict things but it's the content in those 15 lines like you are picking what kind of model and why you're picking what kind of operation and why you're picking what kind of split and why that's why you're getting paid higher than other people okay moving down once you got your splits done you will pick a machine learning model how do you pick a machine learning model first way is your knowledge how well do you know about each of the uh each of the model how well do you know about each of them do you know what is the difference between logistic regression and svm do you know how k n operate do you know how uh decision trees up there do you know how gradient most operates Etc and uh again how many algorithms do you know suppose I am a middle class person and I have a single car I don't have options to pick from the only option I have is I have if I have to take my own car there's only one car and then I am a slightly rich person I have three cars and I have three options but if I am like Vijay mallya or Mukesh Ambani or Elon Musk I have like plenty of cards to pick from I can demand for what car I want if I want the sports car I have the best of the best sports car in my kit and if I want a limousine I have the best of the best Amazon available so here your money should be replaced by your knowledge if you are a beginner you may have two or three algorithms only to pick from if you are an intermediate person you have five algorithm if you are a person with mine level of experience you may have 20 algorithm if you are a person with higher experience you may have 100 algorithms to pick from or to customize from but there is one thing that can help you to pick among the given option it's your task to bring the option but there's something called as grid search cross validation or randomized search cross validation that can help you pick from the given option that given this data what is the best algorithm to use again please listen to my words given this data what is the best algorithm to use I have added the word given the state okay it's not like the global evest algorithm on every data there can be different algorithm which will work you have to figure out using grid search cross validation which is nothing but it will perform a lot of operations it will calculate every single combinations accuracy or loss and then tell you that this is the best performance performing model and you should pick this so grid search is your word to search for I'm giving you a lot of words to search forward for okay once you got your model uh finalized you need to tune your model you need to finalize what are the hyper parameters for that if you don't know what about if you don't know what hyper parameters you say model is equals to svm and model.fet that means you are using the default settings of svm there are so many things you can customize inside SCM and svm can give you 80 accuracy and the same svm if you customize a little can give you 95 percent of accuracy on the same data so there are parameters you need to tune there are a lot of configurations you need to make again grid search is going to help you out and your information is also going to help you out there once you get it done again a very cliche step model.fit you train your model and uh by the way like whatever we are discussing it's just not there for machine learning it is there for deep learning also it's the same process until data everything is same split is same finding the Deep learning model the same during the deploying model is same train the deploying model is also same okay and uh then you go forward for training the data once you train your data you have to validate your training results please listen to me you have to validate your training result training is the word search so you train your data it is like I'll give an example so let's say you are teaching a kid how to perform mathematical addition and you teach like two plus three is five one plus one is two and one plus four is also 5. and now you ask a kid a question that what is one plus one the kid may give you two answers either the kid will say two or the kid will say anything on a platform too this is called as training evaluation why it is training evaluation if the kid said anything apart from two it's 100 clear the kid have not learned but if the kid says to it's still not clear that they could have learned the kid may have crammed the things because you have told the kid one plus one is two and maybe the kid just had a friend the kid doesn't know how to calculate that the kid only knows that okay one plus one is two that means it is two the one plus two I don't know because that is never discussed okay this is uh where you will get a filter that okay whether my model has understood the training data or not if my model has not understood the training data I have to change things I have to go back to the very first step maybe my data collection is wrong maybe my requirement understanding is wrong maybe my idea is wrong maybe my processing is wrong maybe my split is wrong Etc or maybe my model is strong yes somebody said undercutting you know sitting no it's under fitting and overfitting yeah so uh but if training performance is good that does not give you the answer that your model is good it can be cramming which is also called as overfitting or it can be actually good training okay so you how will you test that you will ask a question which is never discussed if you assume that your kid has understood addition you ask him what is eight plus one if the kid answers you nine that means the kid has understood things if the kid doesn't answer you nine that means the kid has not understood okay uh virendra I have I will take your question or anything please give me some time I saw your hand keep your hand Rich I'll take your question so now how will you get this question eight plus one since it is addition you can curate and create any number of questions but when you are talking about data you have limited data with you that's a 500 records so that's why split was important you will firstly split your data into like 400 and 100 you'll trade on the 400 Parts 400 records and you will test on those 100 records that's how you validate on the test data once your validation fails you need to go back to the start and repeat everything but if your validation succeed that's again a point which most of the students do not do is to save your model you a lot of people do not write the code to save the model you need to serialize and create your model using uh packages like pickle job lab in machine learning sorry in deep learning tensorflow gives the package like dot save model and there are like various other ways but you need to save your model to keep it permanent you just can't keep it like a variable because that will be lost once you have got the training completed so you need to save your model and once the saving is done you need to deploy your model to whatever purpose you have built it for have you built it to make a website out of that for example Facebook is Facebook a website technically yes but is Facebook famous because of website no Facebook is famous because of their artificial intelligence it is famous for because of their high quality recommendation system Netflix and YouTube is famous because of their high quality recommendation system they are not only video hosting platform there are many like that it is famous for the recommendation they are giving you it is famous for analyzing their Pro your profile very well much better than any other platform so that's how their models are deployed and workings we need to deploy our model in the form of a mobile app web app game software API whatsoever for that again you need to take help from people who are working in that domain if you are deploying it to our website you need a web developer or you need to learn web development by yourself so if I quickly summarize all the steps one by one and then I'll take your question all those who have questions please raise your hands I'll take your question one by one I guess chat is not a feasible option to take questions from right now I'll give you audio access to ask the question so first step understand the requirement Second Step collect the data third step analyze your data perform the Ada fourth step act upon that Eda process your data clean your data transform your data do whatever is required fifth step split your features and labels sixth step split your data into training test using variety of options available next find your right model next tune that right model next train your model next check your training performance then check your testing performance if any of them is not working go back to the start and do things again with different combinations but if both of them are working save your model and employer model that marks the complete end-to-end pipeline of machine learning and that's how exactly an industrial project is being taken care of not like you would go to kaggle you download the data and you load it and you do model.fit and Tada it works no it doesn't go like that there are a lot of things you need to do before actually getting that data getting the data is like 50 of the task completed so yeah this is all about things from my side kaggle has toy data sets yes it is it is a very good platform by the way I'm not criticizing cattle what I'm saying is if you have to analyze Amazon you need to get data from Amazon you will not get the data from Target if you have to analyze Cisco you need to get the data from Cisco if you have to analyze Flipkart you need to get the data from Flipkart so not from Canada is like a general public data it's it's a mock data to practice it's a very good platform to practice I have practice from there all of you must have practiced from there everybody learns things from kaggle only but like again when you do things for production when you do things which makes money for companies you need to personalize data when do we use hyper parameterization I guess you are talking about hyper parameter tuning it is the same thing as grid search when I said tuning the model once you get your final model ready that okay I am going to use SVC or KNN or decision tree then what parameters of decision tree I will use what will be my max depth value what will be the mean life value what will be my n value Etc so that's how you do with grid search and randomized search if I know machine all those who have questions you can put down on chat I am looking at chat right now only at the new questions a lot of questions have skipped up I'm really sorry for that if I know machine learning how to start with deep learning like if you know machine learning and you have done three to four projects at least you can like uh keep going with deep learning you can take up a pore so you can follow YouTube you can follow Google there are so many things available um or you can like follow documentation of tensorflow and start learning from there how to know which algorithm fit for our data set it's again your knowledge about your car and the terrain that same example so uh there is no right or wrong answer to that everything works but but like you need to put your experience on work on that place uh grid search can help you a little on that how to check for data drift what do you mean by data drift do you mean the skew list or imbalance data yeah Ramesh please answer me this thing then I will answer your question further keshav asked if we if two different algorithms are giving similar result then how to make a choice again like I'm saying if you drive a car for 100 meter and you drive another car 400 meter and they come up with the same kind of performance it's your experience with those car which will help you choose which car to go forward like you cannot just have a global answer for this thing no algorithm is better than another algorithm I have seen logistic regression performing much better than any other complex algorithm even random random Forest I've seen logistic regression performing better than random Forest at variety of cases so it's not like logistic Recreation is a basic algorithm random Forest is a complex algorithm please scroll up I have a long question imbalance data okay so Ramesh there is something called as kurtosis to check for imbalance data that will help you out skewness and ketosis is what you can check forward for change of data with time or environment is equal to data drift I got it I got it that the cryptosis and sqls will help them out what are some good resources for guided ml projects using python to learn more about ML and how to apply them YouTube I guess is a very good platform there are like enough of channels about Python and machine learning on them we are on the platform of analytics with there so I believe analytics with yeah honestly I'm not promoting but honestly I have spent a lot of time on analytics with there while I was a student and I was learning even now sometimes I don't forward for articles and they're really good so this is a good platform any tips to find and work in Industry level project like uh it's like you cannot because you cannot get the data data is like a very secure thing nobody is ready to even if you go to Amazon and you say that hey I'm going to work for you for free will not give you data so it's difficult if you can actually find some connections and if you can find some startups that's the only way I guess but otherwise the traditional ways do something good in machine learning get into a job or an internship and get the access to data and work there that's like I guess a good way so I'll scroll up there was actually a very big question let me have a check yeah hi Harish actually I copy paste okay okay perfect one second uh pranav you have raised your hand have your query got solved please put down on chat if it is not how can we approach an industrial problem where our agenda is to predict whether our particular task is going to run into a problem or not the challenge here is that we have only six thousand of data to build the model plus the data we have uh okay I don't know who asked this question raksha has just posted it up it looks like a consultancy task there is a lot of information I need much more than this and moreover your message seems truncated it's it's given till all the data is categorical so there is a lot of things I guess it will take a lot of time I would suggest you reach out to me on LinkedIn I'll try to answer your question there better rather than spending a lot of time here and one more thing I'm putting down my LinkedIn here so all those who wants to connect you are more than welcome one ticket and uh harsh actually sorry to interrupt I have one question from my side sure so actually uh how far it is good to transform the target variable and if we do any transformation On Target then how to bring it back to its original form so it depends on how are you transforming are you transforming using the conventional method like label encoder or 14 quarter if that is the case then they themselves provide a reverse mechanism for for getting the result back like label encoding has their own reverse mechanism and one hot encoding also have their own reverse mechanism that the way with which you can reverse it but if in case you are not using the conventional method if you are designing your own function then you need to design the reverse mechanism also like the reverse map function for that so that's how you can get the answer in Reverse that's your second question sorry and your first question is how good it is just a minute okay so I'm really struggling from my cuff and gold okay so how good uh how good it is to transform the label is that your question right yeah your target variable how Target variable so there is no harm in doing that until you are preserving the nature of that and what do I mean by Nature so a lot of time when we try to convert the categorical values into numerical values we lose their sense of being nominal or ordinal we should treat the nominal data in ordinal data differently we shouldn't read just discrete data in a similar way discrete can be nominal can be ordinal and both of them should have a different treatment nominal should be transformed in such a way that after transformation it is also nominal same with ordinal should be transformed in such a way that after ordinal also it should be ordinal only if you are maintaining this then you are good to go but in my opinion you really do not need to transform labels at all because labels do not go as input labels are not used to find and predict the patterns they are just the mapping value and even if it is in text I guess lot of models can accept it yeah so like generally after generating predictions we need to bring it back to its original form yes suppose we did a lot of transformation or a square root or something like that like is it uh like is it not a wrong way to do no no that's absolutely not a wrong way uh one thing if you are preserving the nature then nothing is wrong but uh if you are losing out the nature then of course things will go wrong and another question like after prediction you need to change yes because your machine learning algorithm is designed to predict for the numerical value once you got your numerical value predicted pass that numerical value to a reverse mapping function and that will give you the reverse answer like suppose five means happen and 10 means elephant and two means banana and so on mm-hmm okay so you predict for a number let's say the number predicted out is 5 pass that number to a reverse map function and that will give you the elephant or apple or whatever is the mapping value okay is concerned about my health thank you so much and we can hopefully have more such sessions in the future thank you so much it was it is really kind of you but yeah I'm fine like it's just regular cuff and put okay uh any more questions if I skipped out please I am so sorry uh you can post it down here and I will answer it I guess we have five more minutes um I really hope you liked the session I really hope I was able to give you some value in one hour it was very very fast I have to keep it fast and continuously going because of so much of content to cover and I wish I Justified with the title itself yeah definitely and thank you so much on the behalf of analytics with as well I would like to thank you yeah you have delivered such a great session and I personally I was excited to attend this session because yeah it was very basic but it was uh like it's a great thing to understand how to approach a problem and she really answered like I have a few doubts regarding that and it was cleared in this session so thank you so much and I'm sure it was appropriate for lots of experience levels and hopefully we can conduct more sessions with you in the future questions let me answer that so it's by Shankar sahu it says any article you suggest for this relevant topic I guess there are not many articles about this topic that's why I chose this thing uh even if there are they are not very much promoted like medium articles or analytics with the articles there are not many articles on this I would suggest analytics you there to write something on that or I would write something like this on on some platform myself so and then I can share you even if there is maybe there are but I am not aware of any such article right now Google can help you write Ashok murias when do you decide for example model so my journey is like that whenever I develop I ask myself a question that is that a regression or classification data that's a fairly basic question if the answer is classification I decide whether this data has higher correlation or low correlation for example I'll give you a very simple example rather than a complex one you see Titanic data in Kagan very very common and simple data but none of those columns actually have higher correlation with the survived column none of those columns if you pass them to DF dot Corr function none of those columns actually have higher correlation with the survived column and there are a lot of production rate systems when you think the data is correlated but data is not so in those scenarios I start my journey with at least random forest and then I go forward with gradient boost and and add a boost and XT post and all of them if the data is correlated I don't really touch them I keep myself very simple with svm or logistic and so many I have like a flow which I follow and that's how I decide the algorithms would love to connect with you over LinkedIn sure enough I'll expect your connection request and I I would be happy to connect it was filled with very real examples so truly useful thank you so much and I hope this will add information to your knowledge and your value system and I wish you very good luck for whatever stage of life you are in your student or working professional or startup Enthusiast or businessman or employee intern whatsoever is your stage of life I wish you a good luck for that thank you so much and anything else you want to ask me you can I guess we are done for the for the session thank you raksha thank you so much for hosting it up and thank you all the participants who have taken out time on a festival season [Music]
Original Description
Machine learning is one of the most efficient and time saving tech domains of current times and hence every tech enthusiast must learn how to use it and how to create machine learning models for any problem statement.
In this DataHour, Harsh will explain in detail the approach and each step involved in solving any machine learning problem.
Do subscribe to Analytics Vidhya channel & get regular updates on videos:
Stay on top of your industry by interacting with us on our social channels:
Follow us on Instagram: https://www.instagram.com/analytics_vidhya/
Like us on Facebook: https://www.facebook.com/AnalyticsVidhya/
Follow us on Twitter: https://twitter.com/AnalyticsVidhya
Follow us on LinkedIn:https://www.linkedin.com/company/analytics-vidhya
Playlist
Uploads from Analytics Vidhya · Analytics Vidhya · 51 of 60
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
▶
52
53
54
55
56
57
58
59
60
The DataHour: Data Science in Retail
Analytics Vidhya
The DataHour: Anomaly detection using NLP and Predictive Modeling
Analytics Vidhya
The DataHour: Energy Data Science Project from Scratch
Analytics Vidhya
The DataHour: Explainable AI Need and Implementation
Analytics Vidhya
The DataHour: Google Cloud AI/ML
Analytics Vidhya
Prediction to Production in Machine Learning #machinelearning #prediction
Analytics Vidhya
Practical Applications of Data science in Ecommerce
Analytics Vidhya
How to tackle Overfitting?#machinelearning #overfitting
Analytics Vidhya
Building Data Pipelines on GCP #googlecloud #datapipelines #data
Analytics Vidhya
Hands-on with A/B Testing #abtesting #datascience
Analytics Vidhya
Efficient Implementations of Transformers #transformers #cnn #machinelearning
Analytics Vidhya
Modern Deep Learning Architecture #deeplearning #architecture #deeplearningtutorial
Analytics Vidhya
Key steps for Designing Artificial Neural Network (ANN) for Image classification #machinelearning
Analytics Vidhya
5 things you should know about Azure SQL #azure #sql #datahour #datascience
Analytics Vidhya
AI & ML in the Automotive Industry #machinelearning #ai
Analytics Vidhya
Building Machine Learning Models in BigQuery
Analytics Vidhya
NLP aspects in Telecommunication Industry
Analytics Vidhya
Practical Time Series Analysis
Analytics Vidhya
Fundamentals of Quantum Computing
Analytics Vidhya
A DAY IN THE LIFE of a Data Scientist (From waking up to working on algorithms)
Analytics Vidhya
Classification Machine Learning Model from Scratch
Analytics Vidhya
Knowledge Graph Solutions using Neo4j
Analytics Vidhya
Model Guesstimation (MLOps)
Analytics Vidhya
ETL Pipelines in Google Cloud Platform
Analytics Vidhya
Key steps for Designing Convolutional Neural Network(CNN) for Image Classification
Analytics Vidhya
Getting Started with AWS EC2 #amazon #aws
Analytics Vidhya
How to Use Azure NLP and Graph Databases for Intelligent Knowledge Mining
Analytics Vidhya
Certified AI & ML BlackBelt Plus Program #shorts
Analytics Vidhya
Visualizing Data using Python #machinelearning #visualization #python
Analytics Vidhya
DCNN for Machine RUL Prediction using Time-series Data #timeseries #machinelearning #datascience
Analytics Vidhya
M in ML stands for Math & Magic
Analytics Vidhya
An Unsupervised ML approach using Clustering
Analytics Vidhya
Customizing Large Language Models GPT3 for Real-life Use Cases #gpt3 #datascience
Analytics Vidhya
Model Parameters vs Hyperparameters - Techniques in ML Engineering #machinelearning
Analytics Vidhya
Practical MLOps #mlops #datascience
Analytics Vidhya
Data Engineering with Databricks #dataengineering #databricks
Analytics Vidhya
Multi-Objective Optimisation
Analytics Vidhya
When Airflow Meets Kubernetes
Analytics Vidhya
AI in Banking
Analytics Vidhya
Learn Convolutional Neural Network for Image Recognition
Analytics Vidhya
Extracting Value from Data
Analytics Vidhya
How to measure Marketing Channel Effectiveness
Analytics Vidhya
Transforming Lives | Data Science Immersive Bootcamp
Analytics Vidhya
Stock Market Analysis - AI driven approach
Analytics Vidhya
Become a Data Engineering Professional in 2022 | Future Trends + Skills Required
Analytics Vidhya
Ensemble Techniques in Machine Learning #machinelearning #ensemble #datascience
Analytics Vidhya
The Power of Visualization | Tableau Full Course | Analytics Vidhya
Analytics Vidhya
Demand for Data Engineers is on the Rise | Data Engineer | Analytics Vidhya
Analytics Vidhya
Data Visualization in Data Science | DataHour | Analytics Vidhya
Analytics Vidhya
Role of Optimization in Machine Learning & Deep Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Solving any Machine Learning Problem | Approach and Steps Involved
Analytics Vidhya
Topic Modeling Explained with Implementation | Using LDA in Python | DataHour by Arpendu Ganguly
Analytics Vidhya
Data Engineering in E-Commerce | The Best Case Study
Analytics Vidhya
Introduction to Classification using Azure Machine Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Introduction to Federated Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Diffusion Models for Generative Arts | DataHour | Analytics Vidhya
Analytics Vidhya
Master Google Analytics in 1 Hour | DataHour | Analytics Vidhya
Analytics Vidhya
Learn Hypothesis Testing | DataHour | Analytics Vidhya
Analytics Vidhya
A Practical Approach to Kaggle Competition | DataHour | Analytics Vidhya
Analytics Vidhya
Making AI work for Business | DataHour | Analytics Vidhya
Analytics Vidhya
More on: ML Maths Basics
View skill →Related Reads
📰
📰
📰
📰
Transfer Learning with MobileNetV2 in Keras: A Practical Guide
Medium · Data Science
From Data to Predictions: Understanding the Machine Learning Workflow
Medium · Data Science
The Inference Moat: Why Speculative Architectures Define the 2026 Competitive Edge
Medium · Machine Learning
How to Combine Optuna and K-Fold Cross-Validation
Medium · Machine Learning
🎓
Tutor Explanation
DeepCamp AI