Statistics You Must Know for Data Science ๐Ÿ“Š | Complete Beginner Guide

GeeksforGeeks ยท Beginner ยท๐Ÿ”ข Mathematical Foundations ยท12mo ago

Key Takeaways

Covers essential statistical concepts for data science, including mean, median, mode, and probability distributions

Full Transcript

I get uh some placement support after enrolling this course. >> Uh sorry uh can you be a little loud? I'm not able to hear you clearly. >> After enrolling this course, can I get some placement support? >> Um see I'm not the right person to answer this. Maybe you can get in touch with uh geeks for geeks. Somebody whom you would have contacted earlier. They will be the right people to guide you on this. >> Okay. Thank you. >> Yeah. Okay. Um so let's start. I'll start the session. So as I said this session is uh first of all good evening. Uh this session for the next one hour I'm going to talk about stats right now. uh stats is a huge vast uh subject and uh it's it's difficult to cover even if I have to just talk about statistic which is relevant for data science even if I have to cover this uh limited portion only 1 hour will not do justice to stats so what I am going to do is I'm going to talk about certain basic concepts I will give you an introduction and maybe a first level detailing for each of these concepts. Your job is to then go back and deep dive into each of these concepts to understand it better. So what I will do is I will try to make it as simple as interesting and give you maybe I'll take you from level zero to level one but then once this session is over maybe you can go back read some books or do uh some uh search on Google and then take your understanding from maybe level one to level two and then with more practice and uh more reading these concepts will slowly get better with time. Yeah. Uh everyone is with me so far? >> Yes sir. >> Okay. >> Yes sir. >> Yes ma'am. >> Okay. So let's get started. Um so basically uh I have structured the presentation in in uh in this format. Initially I will talk about what actually statistic is and why are we even uh studying stats. I I'm not sure what is the background of most of these students here but given that you are interested in data science I'm assuming most of you are engineers now as engineers we don't do stats right uh so why at this point we are looking at as uh stats as a subject I'll give you a little bit understanding of that then we deep dive into different types of statistics so there is descriptive stats there is inferial stats we look at what is the difference between these two and how Do these two statistics help us when we are doing a real uh data science project? Uh within descriptive stats, we will look at measures of central tendency and variation. And within inferial stats or inferial analytics, we will look at probability distributions, hypothesis testing, central limit theorem, correlation and coariance. Now these terms um don't get scared. It might look very heavy, very big and very new because most of you uh as I said if you are engineers you have not looked at these concepts earlier. It is okay. This is just an introductory class basically just giving you a little flavor. This is not meant it is not like after this 1 hour you will be able to you will be able to understand stats completely. That's not going to happen. This class is only to give you a little flavor so that then you can take up these concepts ahead and study yourself and of course once you join a proper class there will be more training provided. So this is an introductory class. I will try to make it as simple and as interesting as possible. Okay. Um so let's look at why stats or why for that matter data has become very important today. So you know when uh when we were when I was of your age were maybe about 20 years back most of the decisions that the companies uh companies like any company for that matter you can take right u if uh say for example HL or Reliance or or Infosys any company for that matter big or small or medium the decisions that they would take would be based more on intuition. Why? because the tools and the and the technologies that you require to capture data to analyze data and to gain insights from the data was not available then. So most of the decisions that were done were mostly subjective. It came from experience of senior people who had some intuition some understanding about how the business is performing but they did not have an end toend picture. But fast forward 20 years and if I look at the time today, we have all the tools and all the technologies that are helping the companies not only capture the data but also analyze this data and gain insights and then today data science and machine learning as a science has also evolved. So today most of the in fact all the decisions all the planning that companies make it actually happens based on the data and not on intuition and that is why today data science data analytics streams like this which never existed earlier right statistic as a science and stat statistician as a professional would used to exist but today you see a lot of data engineers data science data analytics data analyst professionals also today uh different roles you see right. So the reason why you see this is because today capturing data, storing data, analyzing data has become very simple, has become uh very easy and therefore most of the companies are now going through this route and therefore it has become very important that we look at data, we try to understand it, we make certain inferences from it, we analyze the data and then based on our understanding we go ahead and make certain decisions. Right? So this is about why the database decision making has become more of a norm than an exception. Now how do we go ahead and make decisions based on data. So that whole science is known as stats. So if I just have to define what is statistic, it is a science that deals with the collection of data, analysis, interpretation and presentation of empirical data. empirical data or you can just say data. Empirical is basically um you can say data that you collect from experiments but basically stats is the entire science from collecting data to analyzing interpreting and then you present it to your senior stakeholders who then use this data to make certain decisions. All of you with me so far? >> Yes ma'am. >> Yes ma'am. >> Yes ma'am. >> Yes ma'am. >> Yes ma'am. Okay, now let's go. Okay, now this is a question which >> Great. Thanks. Um, this is a question which I would want to ask you towards the end of this session. Maybe it is too premature to ask this but I would want to know after this session why do you think as a data scientist or as a data analyst or an aspiring data scientist data analyst why do we need to learn stats? So this is a question that I will ask all of you at the end of the session. For now let's go ahead. Now as I said stats there are two key branches which are relevant for the field of data science. First is descriptive statistics. So basically this is the branch of stats that deals with the past data something that has already happened right. So say for example um I am a company uh and I am I have been um uh I'm an organization which is almost 10 years old. So now for this 10 years I already have data in terms of what are my sales uh what are my cost what is the number of employees that I have right what are the profits that I am making so on and so forth. So this data is already existing. It is past data and descriptive stats what it does is that it uses this data to make certain insights. Right? So when I talk about again again let me give you another example. So say for example um I am geek for geeks and there are a lot of students uh I have the data of all the students that have been learning data science uh for the last 5 years. So I have their names the age the um you know what what are their educational uh qualifications so on and so forth. Okay. Uh I'm not sure uh how is this happening. I can see. Do you also see this or is it only I who can see certain marks on the >> Yeah, there is a cross. >> Yes ma'am. >> Yeah. Whoever has done it, can you please remove it? Yeah. Thank you so much. Thank you. >> Yeah. >> Hi. Say for example uh geek for geeks it has been in existence uh for the last 5 years. It has uh data for all the students that have taken various courses. So now geek forge geeks will have data for the student names where they are based out of what what are their uh educational qualifications. What is the number of what is the percentage of marks that they have got in standard 12 standard 10 what kind of jobs did they get? Do they have any professional experience before joining geek forge geeks. So these are various data points that geek for geeks will capture. Now once it has captured this data, it will go ahead and try to analyze what is the mean age of uh students who are coming for uh learning data science, right? Or it would look at what is the uh the the qualification of students who are coming to learn data science. So basically when you start looking into analyzing this data in terms and you summarize it right that is known as descriptive statistics. It is basically talking about it. It describes the data and how does it describes the data? It uses certain statistical measures like central tendency where we will look at these measures in the next slides but it will use measures of central tendency, measures of variation and come up with certain insights. So this part is descriptive stats. Basically it looks at the past data and captures what happened. The second branch of stats is inferial statistics. Now here what you're doing is you are uh this is basically or mostly used when you are conducting a certain experiment. Now when you do a certain experiment of course you would like uh the sample of the experiment to be as huge as possible. For example, there is a company uh let's try to understand this with an example. Let's say there is I I'm launching um I I have a company and I want to launch a certain soft drink in uh say tier one cities, right? I'm going to launch that. Now I want to know how much is the market demand for that product. So what will I do? I will uh look at certain set of tier one or tier 2 cities and I will look at whether the competitors already have launched. Then I would look at how much is the sales of those companies coming from um um these particular cities right what is their market share how much is the revenue that they are making so basically what I have done I have collected a certain sample and now I'm looking at the revenues of u my competitors in each of these cities so this has become a sample and using this sample now I'm going to conclude something or I'm going to make a statement saying whether this product is successful product or it is not a successful product. So when you instead of a population you look at a small sample and then you try to generalize about that sample that science is known as inferial statistics and I will give you more examples when we go further. Uh this inferial statistics requires a little bit more understanding and practice as compared to descriptive because descriptive more or less is your maybe fifth or sixth standard maths. Inferial is actually more about stats and less about maths where you collect a sample do an experiment on that sample and then you try to generalize that because my sample is be is saying me this possibly the population is also saying this. So you kind of generalize about a larger group by analyzing a smaller data set. And there's a lot of techniques that you use in inferial statistics like um hypothesis testing. You look at sampling, you look at analysis of variance which is another hypothesis test. So this is another part of stats which is inferial. It builds and test hypothesis from the data. Any questions so far? any um doubts? Are you clear with what descriptive stats does and what inferial stats does? >> Madam, I have one question. >> Sure. >> Yeah. >> So, uh is inferial statistics responsible for predictive analysis for AI for any models? >> Yes. Yes. >> Okay. >> Yes. Actually, both of them are. So what happens is when you are um uh when you are doing a real uh uh data science project right you will be given a lot of data there will be one dependent variable and there will be 100 200 maybe more than that independent variables. So what you will have to do is you will have to then analyze each column. Uh so when I say column means each variable. So you would look at how uh what is the mean value of of the variable right? How is it how much is the variation in that in that variable? Then you will look at how does the variation in that variable associated with the variation in the dependent variable. So this is where the descriptive statistics or descriptive analytics will play a role. When it comes to inferial stats based on the data you will create certain hypothesis right now you will have to test that hypothesis. What is hypothesis? Basically your data is saying that say X is related to Y. Right? Now you will have to do certain certain test to check whether your claim is correct or not. So you will do certain hypothesis testing there. For example, say you're looking at a salary data. Now you have salaries of men and salaries of women. What you see is that the average salary of men is more than the average salary of women. Now this is what you are seeing from the data. When you plot the data you can see the average salary of uh of men is say um uh 10 lakhs. The average salary of women is six lakhs. Now here once you have made that hypothesis you will have to do some testing which is known as hypothesis testing where you will do certain kind of test to conclude whether your claim is correct or not. That is where inferial statistics will play a role. So yes, as you are rightly said, it is both descriptive and influential statistics that will come into picture when you start working on real life um use cases. >> Thank you. >> Ma'am, I have a question. >> Mhm. >> Uh so we learned statistics and math for the first two years of engineering course. >> Uh do we need more of statistics if you want to go into data science? um see statistics um so I I'm the topics that have been given to you right it is uh so because what I was told is that there have been certain documents that they have shared with you uh which cover topics like uh certain distribution hypothesis testing correlation ANOVA all these topics are necessary especially when you are dealing with traditional data science I'm not talking about genai and But when you're talking about traditional data science where you're looking at algorithms like linear regression for example, right? Linear regression is actually a statistical technique. It's not a machine learning technique. So there when you look at individual p values, you look at um if you're doing an uh you would be looking at correlation, you would be conducting certain hypothesis, you would look at type one, type two error. So all of that becomes important. So yes, when you are starting out, a good amount of strong statistical knowledge is is required. But again, um if you understand the basic concepts and maybe to some extent a little uh in little detail about certain specific topics, that is sufficient. You do not need to be a master's in stats to actually be good in data science. >> Thank you so much, ma'am. Okay. Any more questions? Okay. Um I'll go to the next one. So yeah. So now let's look at descriptive stats. As you can see u there's there are two tables here. And because this is very easy I I'll take you through this quickly. Unless you have questions can you can uh you can stop me. So this first table is is where you have your student uh serial numbers and then you have the weights of these of the students. So this is basically a numerical column and in the second table you have your customer serial numbers and then you have their genders. Right? So now if I have to use descriptive stats or if I have to use descriptive analytics here, what is the first thing I will do? I will look at I'll calculate the mean of this column. Right? So mean is basically just your average. Um this is very simple just calculate the average of the weights of the students in the class. The second most important statistics after mean is the median weight. So what you do in median is basically you sort this column and you arrange it in ascending order and then uh the uh the the number that is that comes in in between which basically divides your data into two equal parts is your median value. Now here if say for example your number of observations are odd then it is simple to find out the middle value. If your number of observations are even then you will have to take the two values in between and take an average of that that will be your median value for the weight right and then you have mode which is basically the most frequently occurring value in the data. Now this is a numerical column. So here even if I say that mode is I don't know I've just randomly put numbers here but say yeah I can say mode is 58 it is okay it doesn't really give me any insight right it it's not telling me anything specific but here if I look at this column where my variable is categorical it is not numeric if I here the mode is female so say for example if I if I had a shop and I was just noting the gender of customers who are coming to me this is of importance to me right uh the mode is more important when you're looking at categorical variables the mean and medium becomes more important when you are looking at numerical uh variables so I have just done this um exercise here just to show you this so basically um this is your data this is basically just looking at these are all excel formulas where you can calculate the mean and medium. Now, um let's go one step ahead. Um this I'm sure all of you can do, right? And I've already answered the third question. For what kind of data is mode more relevant? Mode is more relevant for categorical data. So when you talk about measures of central tendency, mean is the most important uh statistic. In fact uh in hypothesis testing also uh most of the testing happens around on the mean. The mean of group A is different from mean of group B or mean of group A is equal to mean of group B so on and so forth. So while it is a very simple statistic to understand its use is very is is very high right. uh median is something that we use when you are drawing certain charts when you are say for example in when you visualize the data when you draw certain charts one of the charts that you will draw for numerical columns is um the box plot I don't have it ready here but you can of course Google and you would see that when you uh draw a box plot there will be a rectangular like structure which will show you where the mean is and where the median is right so it it basically will tell you what is the median value of salary across genders or maybe what is the median value of sales across cities across products so on and so forth. So median is also an important statistic for numerical columns. When it comes to categorical columns, mode becomes very very important. For example, you have a data set where you have been given cities, right? Cities is one of the column and there are 200 cities given to you, right? In that data set which has say about 500 observations, there are 200 different values of city. So do you think you would be using all the 200 values? >> What do you think? >> No ma'am. >> I can take the sample of of the values. >> Yeah. So what you would do is that you will you will not take a sample because you anyways have 500. Sorry, >> we can use the range. >> Yeah. So, um uh okay. So, let me just explain from where I am coming from. So, when I am uh and maybe this is a little uh mature because I'm directly jumping into a data science problem, but I'm just trying to link this to how uh the mode will matter. Right? So say for example you have a data set where there where you have 500 observations and there in 500 observations you have been given 200 cities each occurring at a different frequency right. So what you would do is you would basically take only the top 10 as different cities and the remaining you will just club it into others because when you are creating a model you do not really need you will you cannot use a categorical column as it is. you will have to do a uh certain kind of um encoding on that variable. So when you are encoding a variable what will happen? There will be a separate column for Delhi. There will be a separate column for Mumbai. There will be a separate column for all the 200 cities. Now this is going to become this is going to make your data very very huge right which is not which is not recommended. So what you would do is maybe you will just take the top 10 cities which are occurring at the highest frequency and for all the other cities where the frequency is very less you will just club it and create another column which is others. So this frequency is where mode comes into picture over here female gender it is three out of five the frequency is high. Now of course this is simple because you can only have two genders for but for example if I had a city column here and I had a lot of cities with few cities coming only once and few cities repeat getting repeated a lot of times then what would you do? You would keep the cities that are being repeated frequently and all the cities which are only coming once or maybe twice you would just club it and put it as others. So that you will only have say five columns. You would have Delhi, Mumbai, Chennai, Kolkata, Bangalore and then others. Right? So that is where mode plays an important role or that frequency plays an important role for categorical uh columns. If this is not making sense right now, it is it is absolutely okay. You just have to remember that mode is the most frequently occurring value and it is more important for columns which are categorical. >> Okay. Um I'll go to the next part of descriptive stats. Uh this is basically your um measures of variation. So we looked at uh the central tendency which is basically your mean, median, mode. Now for every data that you have while you have calculated the average of every column you also need to look at what is the overall variation in the data or in that particular variable. So I can say that the average weight of student is whatever say 50 50 or 53 but how much is the weights of other students varying from the mean average right that is calculated by or that is given by the variance. So basically it measures the variability in the data from the mean value. So if you see here let me use the laser. So this basically is what this is my observation. So for student one x i is 50 and say the mean value is 58. Say for example. So now I'm going to do what? I'm going to do 50 minus 58 and then I'm going to square it. Why am I squaring it? Because if I don't square it, what would happen? Can anybody guess? I'm going to >> sorry >> negative values will be present. >> Yeah. And these negative and positive values may cancel out each other. So this will not give me a complete picture of variation. That is the reason why I am squaring this. Right? So I square the difference and then I divide it by the number of observations in the data. This gives me the variance in the data set. Right? So when I say the variance of the data is x units. Basically I'm saying the variability in the data set from the mean value is x. Let me show this to you. I had done this to just give you a better idea. So now if you say if you look here is the excel visible guys. >> Yes ma'am. Yes ma'am. >> Great. So now you have a list of 18 students here and you have the values of the weights right 50 58 53 for uh so you can calculate the mean is basically just taking the average you get 52.7 now what I will do is for each of these observations I will calculate the difference so it is the difference between 50 and 52.7 so this is x minus the average and then I square it so when I square this all my values that I'm getting is is positive. Here it is highly possible that my positive and negative values can cancel each other if I don't square it. So what I do is I square all the values and I calculate the total sum. The sum is 1 6 1 0. So now what is the variance? The variance is the sum of the difference squared divided by the number of observations. So this is the sum the squared sum and this is the number of observations which is 89.4 four uh sorry which is the number of observations are 80. So when I do this calculation I get variance as 89.4. So this is the variance in my data. But the problem with this is and if you would have noticed what is the unit I'm getting here? It is the weight squared. Why? Because I have done square here. So the weight now that the the unit that I'm getting of the variance is different from the unit over here. Weight is in kgs. This I will get in kg squared. Now can I make any sense out of it? I maybe maybe make some sense out of it. But would that be intuitive? >> Would I be really be able to >> right? So then what do we do? We actually take the square of this variance. Right. So when we take the square of the variance now the unit is back I will get it in kgs. So I'm saying the standard deviation is 9.5. So when uh the standard deviation is actually the square root of the variance it does exactly what variance does. It gives you the dispersion of data from the mean. But the advantage of using standard deviation is that the unit of the standard deviation is the same as the unit of the variable that you are considering. So it becomes very easy to understand that okay the standard deviation here is basically 9.5 kgs. So the dispersion of data from the mean which is about 52.7. The standard deviation of this data is 9.5. So the data is varying from the mean by 9.5 units on an average. Is this clear? Is this concept clear? This has to be very clear because you will be using standard deviation extensively. >> Yes, >> ma'am. I have a doubt. >> Yeah, >> ma'am. This ma'am, here the variance is calculated for population mean. Right. But when you calculate the same for sample mean, we divide it by n minus one. Why nus? >> Right. Right. There's a there's a concept called bezels uh correction uh which is basically uh it talks about um um I'm I'm not able to recall uh it's it's basically taking into consideration the number of u in the degrees of independence right uh you you will have to look onto it I don't remember exactly what why they they do it but it is called bezels correction And they take into account the degrees of freedom. Sorry, not independence. The degrees of freedom reduces when you look at a sample instead of a population. And that is why you divide it by n minus one and not n. >> Okay. Any other questions? >> No ma'am. >> Okay. So, uh so yeah, you can actually uh look at why it is n minus one. As I said, it is more about using uh degrees of freedom. uh uh but just just look at it because I uh I don't remember on the top of my mind why it is n minus one but I do remember that it is to adjust the degrees of freedom when you're using a sample. >> Yeah. >> Yes ma'am. >> Okay. Um fine so this we already did. Why do we need standard deviation when we have variance? Uh one more thing I wanted to cover. uh not sure if you are covering in your documents that have been shared by geek for geeks but interquartile range is also a very important concept when you talk about uh u variation in the data right so interquartile range or commonly known as IQR it is a difference between the third quartile and the first quartile of the data now when you're talking about quartiles the first thing that you need to do is you need to sort this in ascending order so what am I saying when I say first quartile first quartile is basically the number of observations um in the 25% of the 25% of my data is less than a particular value that value is the first quartile so quartile you understand right it's 25 50 75% the the 1x4 of the data so when I say the first quartile I'm basically looking at 25% of the observation and I'm saying the 25% of my observations is less than this particular value. This is the first quartile. Similarly, the 50th uh the second quartile is basically your median which is 50 and then the third quartile is the 75% of the data. So I'm saying 75% of the data is less than a certain value that is the third quartile of the data. So what IQR does it it gives you the difference between the third quartile and the first quartile. Now why do you really in fact even need IQR? because IQR will basically tell you what is the variability in the middle 50% of the data right uh and the best use of IQR is to understand how many outliers do you have in your data which are these outliers what are their values so outlier is basically I'm not sure if you already know this but outlier basically you would have seen say for example here you see most of the numbers are in the range of 40 to 50 but there are certain outliers here say for example 75 this is way way beyond the mean value of 52.7 right it is very far from 52.7 similarly you see there's an underweight student with the weight 30 so this is also very far from the mean value now it is possible to have certain outliers it is natural to have certain outliers but we need to distinguish whether these outliers are because of some wrong calculation or wrong um surveying that you have done or is it it is just the um the nature of the data that these outliers are present and then based on whether it is case one or case two you decide whether you want to continue these you want to retain these outliers in your data or you would want to remove these outliers but the concept of IQR basically is in uh in the is identifying these outliers and you identify it by using these formulas so Q1 which is your first quartile minus 1.5 5 into IQR will give you the outliers which are in the lower range and Q3 + 1.5 into IQR will give you outliers in the range which are above 75. So let me show you something just to make this um little bit more intuitive. Uh have you heard or seen a box plot earlier? >> Yes ma'am. Yes ma'am. Yeah. >> Okay. Great. So now if you see here these values that you see so this these values that you see are basically the outliers and this is 1.5 lower. So I uh this is quarter 1 minus 1.5 IQR. This is quarter 3 plus 1.5 IQR. So the values that you see beyond these becomes your outliers. And therefore it is in this context that IQR becomes important. If you see this, this is basically the best maybe to show you how what IQR means. This is your 50% of the data. This is your lower quartile 25%, this is 50%, 50% is the median. This is the 75%. So when I say my upper quartile is say 75, which means that 75% of the observations are less than um um a particular figure. when I say uh the lower quartile or quarter lower quartile says 34 which means that 25% of the uh of the observations are less than that particular value and all the values which lie over here or over here is given by quartile 1 minus 1.5 IQR and quartile 3 + 1.5 IQR and therefore in this context IQR becomes uh an important concept to understand. >> Um, any questions here? >> Yes, ma'am. >> Yeah, tell me. >> Oh, ma'am, can you please explain how can we find first quarter and third quarter again? >> H Okay, so uh let's take this. Okay, so now what I have done is I have arranged this data in the ascending order and I have 18 observations, right? So uh 0.25 into 18 observations are 4.5. So uh what I'm saying is 25% of the data should be less than a particular figure. So if I look at the first four so 4.5 is the first 25% of 18 is 4.5. So I will look at the first four to five observations. The first quartile should be somewhere between 47 and 48. If you go here, right? And uh you you there's a formula here. You can see it is quartiles C3 to C20. This is my data range and I'm saying this is the first quartile. So it is giving me 48.3. If you see here, it has taken it has considered these values and it has calculated the quartile one which is 48.3. So I can say 25% of my observations is less than 48.3. Can I say this? This is 47 45 40 and 30. 25% observation means how many? Four. 4.5 or you can consider four. So four of the 18 observations are less than 48.3. This is my first quartile. Is this clear? >> Yes ma'am. 24 >> right similarly I will calculate the third quartile third quartile will be what 0.75 into 18 which is 13.5 so let's look at this how many are this this is 13 so I can say 75% of the observations which is basically 13 out of the 18 observations are less than 57 around what what do we get yes so we say we get the third quartile as 57.8 8. Is this clear? So this is how Yeah. So again you have this formula here. The the the the data range will stay will stay the same. You say third quartile is three. So this is your 57.8. Then the interquartile range is basically third quartile minus first quartile. So I just do a subtraction here. I get 9.5. Right? In fact, if I do quartile and then I select this data and I do two see this is giving me 52 which is the same as median. Why? Because the second quartile is again your 50% of the data. So 50% of the data lies or is less than 52. 50% of the data is more than 50 52 which is basically your median. So this is how you can calculate the first quartile, third quartile and the secondh quartile which is which is always your median. You can either calculate using this or you can calculate using the median formula. Clear? >> Yes ma'am. Yes ma'am. I have doubt ma'am. Uh >> yeah. >> Okay. >> How is IQ used to detect outliers? >> Okay. So uh what happens is uh if you see here this is your box plot and I'm referring to this diagram because only then you will be able to understand when I say this is this is my Q1. This is the Q1. This is the minimum value. This is Q this is the minimum value. This is Q1. This is median. This is the Q3 and this is max. Now when I'm saying it is Q1 minus 1.5 IQR basically I'm talking about values which will lie in this range then when I say Q3 + 1.5 into IQR I'm talking about values which will lie beyond the maximum value. So values which are in this range and values which are in this range right you will only have few values. Sometimes you can even have many values but mostly you would see there are outliers. Uh uh they they they will be between less to mean medium number of outliers. They will be lying in this range away from the median away from your quarter one away from the quartile Q3 values. So those are your um outliers. Let me see if I can find see over here you can see this is an outlier. Maybe when you start plotting or when you start looking at uh drawing that let's do this. Yeah. Do you see this? This chart these are your outliers. This area that you have is given by quarter 1 sorry Q1 minus 1.5 into IQR. This is this area is Q3 + 1.5 into IQR. any values that lie over here or here are your outliers and that is why IQR becomes important because it helps you to identify these outliers. The whole concept of box plot is actually showing uh what is the interquartile range, what is your uh 75th quartile, what is your 25th quartile and then are there values which are less than 25th quartiles or values more than the 75th quartile. Does it make sense now? Is it better? Do you understand it? >> Yes, ma'am. >> Yeah. >> Okay. So, fine. Uh, we have covered descriptive. Any questions on these two? They look simple but you will be using it extensively. Okay, great. So, now um let's go to another concept which is probability. Um how many of you feel comfortable with with the concept of probability? It's it's a tricky concept. Uh it again comes with Yeah. Right. Uh it it comes with practice. Uh it is initially when you start learning about what is a joint probability, what are independent events, what are what are mutually exclusive events. All of it sounds very confusing. uh I I I'll try to simplify it as much as I can. So first let's understand probability in a very simple term right now say for example um 100 students join uh geek for geeks in the month of uh July right or maybe 100 students joined geek of geeks last year in the month of July out of which 30 students could could get a job right so what is the probability of getting a job it would be 30 by 100 Right? Probability of getting a job is an event. Getting a job is an event. Right? So the number of observations in favor of the event it is 30 divided by the total number of observations which are 100. So there were 100 students who joined 30 got a job. So the probability of getting a job is 30 by 100. Right? If you look at you look at this example a company has thousand employees and every year 200 employees leave the job then what is the probability of attrition? Attrition means people leaving the company, right? What is the probability of attrition of an employee perom? So what is this? The the probability of attrition is the number of observations in favor which is 200 divided by the total number of observations which are 1,000. So it is 20%. >> This is clear. >> Yes ma'am. >> Yeah. Now let's look at certain rules of probability. These are very simple. You can let me know if you know the answers to this. You just this is basically common sense but it is also >> sorry. Okay. So let's go to the first exom. Exom means basically fundamental concepts of probability or you can say fundamental rules of probability. Right. The first is the probability of an event. What is the range of probability? What can be the maximum probability? What can be the minimum probability? >> 0 to one. >> 0 to 1 to one. >> 0 to one. Right. Very good. So yeah. So this is the first exam. The probability of an event E lies between 0 and 1. The probability of the universal set is >> first of all what is Yeah. Yeah. Tell me. >> One. One. >> One. Right. So for people uh who are still thinking universal is basically uh I put in all the events here for this for example say over here a company has 100 employees right uh and 200 employees leave the job so leaving the job is one event so either you are leaving the job or you are continuing in the job right so this is the total universe because these are the only two possible outcomes so the probability of person leaving the job is 20% %. So, and the probability of people continuing in the job is therefore 80%, because 800 employees would continue in the job. Now, when you add 20 and 80, what do you get? You get one, right? Similarly, you could also have a universal set which has more than two outcomes. But when you add up the individual probabilities, it will always add up to one. Therefore, we say the probability of the universal set is one. Yeah. Clear with this? >> Yes ma'am. >> Okay. Now let's go to the other one. This is very simple. For an event A, the probability of the complimentary event is >> 1 minus A. >> Great. Okay. Let's go to the other one. Yeah. The probability of an empty or an impossible event is >> zero. >> Zero. >> Zero. >> Right. The probability that either event A occurs or B occurs or both occur is this is a little >> B. >> Yeah, it is P A union B. It is the probability of event A plus the probability of event B minus the probability of A intersection B. Right? >> Um we will cover this uh in in the next in the next few slides also where we will start looking at individ different types of probability. But this is basically the concept that the probability that either event A occurs or B occurs or both of them occur is given by probability of A plus probability of B minus probability of A and B. The last one if A and B are mutually exclusive events the probability of A intersection B is >> probability of Ability. No, these are mutually exclusive means >> probability of A into probability of B. >> Okay. Um, can someone tell me what are mutually exclusive events? What do you mean by mutually exclusive events? >> Uh, common between both and event A is not event. >> The common element between A and B. >> Event A and event B. >> And the both events are not affected by each other's presence. >> When two events does not occur at the same time. Exactly. They cannot happen at the same time. So what is the probability that they are it is happening at the same time? It is zero because they cannot happen at the same time. Right? What is P A intersection B? It is P of A and B. So if A and B cannot happen together then the probability of A and A and B is zero. Correct? >> Yes ma'am. If I am if say for example and you know when you are learning probability the most common example that you would come across is either rolling a dice or uh taking a yeah or uh rolling a dice dice or a coin flipping a coin right or drawing a a certain card from the pack of cards. Now just take for an example you are rolling a dice. Can you get six and four at the same time? >> No ma'am. No ma'am. >> You can only get one. You can either get four or you can get six. So this event of getting four and an event of getting six when you are rolling the dice only one time is a mutually exclusive event. It can never happen. So the probability of this is zero. Clear? >> Yes. >> Yeah. Okay. Um let's go ahead. Okay. Now we will look at certain terms which again you will find a lot when you start looking at probability. Uh again once you start practicing this will you will start getting more comfortable. Let's first start with the simplest of it which is marginal probability. So this is the simplest because it's basically just looking at a probability of an event in isolation. I go uh to a shop and I uh buy a bread, right? So the probability of purchasing biscuits or probability of purchasing bread or probability of purchasing say chocolates. This is something which is happening in isolation. It has nothing to do with any events that have happened earlier. Right? So this is the unconditional probability. It does not depend on anything. it does not get influenced by anything any event that has occurred before it or it's occurring alongside it. Right? So marginal probability is simply a probability of an event occurring without considering any other events. For example uh say uh uh you are working in a bank and you are in the u in the uh say for example you are in the loan default team. You're basically looking at all the people who have defaulted. So the probability that a customer will default is the marginal probability. But now you have the data of all the customers who have made a default. You are analyzing this data. And now you want to calculate the probability. What is the probability that a person is a female or a person is married and has already committed a default? What is the probability of that? that becomes a conditional probability. So here the events you are calculating the event the probability of an event in isolation. Here you are looking at the probability of an event given another event has also occurred. So that is a conditional probability where you where uh given an event A has happened you're looking at a probability of event B. A is the customer has defaulted. B is the customer is a female. Right? So uh that that in such cases you calculate the conditional probability which is the probability of an event occurring given another event has also happened. The simplest example is I go to uh uh a provision store and I buy bread right that is a marginal probability. But once I have purchased bread, what is the probability of getting butter? That becomes a conditional probability. Right? Because when I'm just purchasing bread, it is okay. It is not linked to any events. But mostly when you buy bread, you will of course also go and buy butter, right? So conditional probability of purchasing the probability that a customer will add butter to his shopping basket given he has already added bread is a conditional probability. It is given by probability of a given b has happened equal to probability of a and b divided by probability of b where this the denominator of course has to be greater than zero. So forget about the formulas from a concept perspective. Is this clear? >> Yes ma'am. >> Yes ma'am. >> Okay. Um okay. Now let's look at joint probability. Right. How is it different from marginal and conditional? So joint probability represents the probability of two or more events occurring together. When you are drawing a card, right, from a pack of cards, the probability that it is red and it is a king, right? It is two events that are occurring together. The sequence of the events does not matter. So it is not like the probability of A given B has happened or probability of B given A has already happened. The sequence does not matter right. So that is where it is different from conditional probability. But it is different from marginal probability because here there are two or more events occurring together. When if it is only two events it is given by P of A intersection B. Now again within joint probability there can be independent events and there can be dependent events. For independent events the joint probability is probability of A into probability of B. For example I go and purchase bread and my friend also goes and purchase bread. Right? These are independent events. It has nothing my purchasing bread has nothing to do with my friend purchasing bread and vice versa. So these are independent events. for that the joint probability is the product of the individual probabilities. Yeah. Is this clear? >> Yes ma'am. >> Yes. >> But this will be different for dependent events. Why? For example, um again using the cards example, the probability of drawing uh say a king, right? When you have 52 cards will be different. But then once you draw a card and then you do not replace it, then the probability of again picking a card and having it as as a king will be different, right? Why? Because it is a dependent event. Say for example, you you pick a card first, you get a king, right? So now you only have one red king present. So the the probability of drawing a king again has reduced. So this is a dependency because it is dependent on the first event and because there is a dependency on the first event this becomes a conditional probability. It is given as P of A and B which is P of A conditioned to B has happened into the probability of B. Again don't um get confused in the formulas. Just understand the joint probability. What how what is joint probability in independent events? What is joint probability in dependent events? When you are talking about independent events, you are basically just looking at individual or marginal probabilities because it has no uh bearing on what others are doing or on any other events. They're independent. They don't have any kind of dependence. But for dependent events, the joint probability actually becomes a conditional probability. >> Okay. Ma'am. >> Yeah. Is it clear to everyone? >> Yes, ma'am. >> Okay. >> Yes. >> Great. Um >> yeah, we are already at six. So, um let me know. Do you want to continue or do you want to stop here? Is it has it become too much for you to take it? We can stop here otherwise we can proceed ahead. >> Continue on. Please continue. >> Okay. Great. So for people who want to drop Yeah. Yeah. I can't see people who raise hands. Uh yeah. Naven do you have a question? >> Yeah. Good evening ma'am. >> Yeah. Hi. >> Yeah ma'am. My question is uh is there any place where we can practice such uh statistics based based questions with their explained answers so that we can increase our understanding onto statistics that is required for that or is it simply we simply start with the uh projects so that we understand statistics how is it >> no no I I wouldn't recommend that you start with projects I would recommend you start with examples and first of all it's good to you may you can Basically look on Google you would see uh some documents where there would be certain examples given and I think geek for geeks would also be kind of having some documents where they are covering uh these concepts in more detail with the formulas and examples. So first thing is reach out to them if you have already enrolled in some course. Uh and if you haven't then uh you can look at YouTube, you can look at uh Google, there are lot of PDF books also available. You can search that and find a good number of examples that covers these concepts and that will make it more clear to you. >> Okay ma'am. Understood. Thank you. >> Uh ma'am, good evening. >> Yeah. Hi. >> Yeah. Hi ma'am. What are the days for the classes? >> Sorry. >> What are the days for the classes? >> Uh I am not able to hear you. >> Like what are the what are the fixed days for the classes? >> Oh okay. Okay. Uh again I am not the best person. I have just been asked to take a session on stats. So all these uh admin questions I think you should reach out uh to the admin guys uh in geek for geeks. They will be able to guide you better. Okay. Okay. >> Okay. Um maybe let's take another required to uh have a good understanding of stats and probability for this machine learning. >> See what happens is say for example um you are doing uh logistic regression. Do you know what logistic regression is? >> Uh not yet not yet. >> Okay. So logistic regression you are basically your dependent variable is a uh categorical variable. It's not a numeric variable. Right? So you're not using linear regression. You're using logistic regression. Now when you predict say for example uh the most again very specific use cases say a customer will default or not right uh either he will default or he will not default. But when you when the model is not going to give you 0 and one while it will give you 0 and one but what is what it is doing at the back end it is calculating the probability. So if you don't understand the concept of probabilities here or at least basic concept basic uh fundamentals of probability then when you start doing classification algorithms right for example random forest decision trees logistic regression which gives you output in the form of probabilities then you will not be able to kind of connect the dots or understand what the output is. While you have all the Python codes which will exactly give you how many zeros and how many ones in your test data set. Basically what happens at the back end is also something that you should understand. The model gives you a certain probability. It will give you say 0.6 or 0.3. Now if your benchmark probability is.5 anything more than.5 will be classified as one. Anything less than.5 will be classified as zero. However, if your benchmark probability is 08, anything less than 0.8 will be classified as zero. Anything more than8 will be classified as one. So there the concept of probability comes into picture and therefore little understanding of probability is required if not detailed is being honest with you. You don't really have to kind of deep dive into lot into probability but basic concepts. Yeah, I mean it's I mean enough uh we are enough with the basic like reading the things. Uh not much much is required about these probability and stats. >> So so whatever I have covered here at least just try to do that much. Try to understand the concept. Look at various examples. See if you can understand what do you what are independent events? What are dependent events? Not go and do 50 questions for each probability but just do about 7 to 10 so that the concept becomes clear. That much should suffice. >> Okay. Okay. Thank you. Thank you. >> Yeah. Any other questions? I saw a few hands that were raised. >> Okay. No questions then I'll move forward. >> Good evening ma'am. I have a one question. >> Yeah. Yeah. Tell me. >> Ma'am, currently I a final year student and currently my drives are aligned continuously days. >> So ma'am, basically I have just a question. So ma'am in interview focus on a probability and statics. continues with the ML and most of the focus >> is the most of >> yeah I'll tell you very frankly you require both because what happens is say there is an interviewer who is grilling you on linear regression now when you are going to explain linear regression a lot of statistics will come into play they will tell you what is the meaning of p value why will you reject a null hypothesis right all of this becomes is very important when you're looking at uh traditional machine learning. So a basic understanding I wouldn't say basic a little bit more than basic understanding of uh statistic is required specifically about p values. P value is a question which is asked every in every interview when you are starting your your career as a fresher in data science. This is the favorite question of all interviewers. They will ask you what is what do you mean by p value? What is type one error? What is type two error? What are the different kinds of distribution? Right? What you see on the slide, these are very very typical questions that are asked. So even if you know how a linear regression works or how a how classification algorithms work, stats, these basic questions of stats is something that you should be able to answer very confidently and then also link it with data science. Don't just talk about stats. You also have to make the interviewer explain how these algorithms use statistics. If you are able to do both then you have a good chance to clear. But for that you will need both knowledge of data science as well as um statistics at least basic knowledge is required. >> Okay. Thank you so much ma'am. Basically I practice the minimum two years of the stats. >> Mhm. >> Maybe it's good. M H okay >> previous interview I faced the I I prepare for the particular states but they did not ask anything particular statics >> then I think you are lucky because most of these students are while they are able to answer questions on algorithms they struggle with statistics and I think it is more about your own understanding than uh the interviews right because even later on as you progress in this field uh if you have a clear understanding of statistics you will be able to understand more complex algorithms also because at the back end most of these algorithms use statistics so learn it from that perspective don't learn it from a perspective that will I be able to answer questions in the interview of course that is required that is the ultimate aim you're doing this course so that we get a job but we are also doing this course so that we understand the basics right so do it from that perspective I think it should be okay but I I'm telling you because these are the questions that every interviewer asked. I'm surprised that nobody asked you questions on stats. It's it's quite surprising. >> Yes, ma'am. This is for me particular. >> But study it because it will help you is all I can say. It will help you get better understanding. See when you are doing decision tree or when you are doing random forest right they use these bagging and boosting uh approaches. Both bagging and boosting approaches are based on sampling. Now if you don't know what sampling is, how it is done, what is sampling with replacement, what is sampling without replacement, you will not be able to understand how decision trees and random forest are working, right? So that is why understanding these concepts is very critical. >> Yes ma'am. Thank you so much ma'am. >> Welcome. Okay. Um maybe what we can do um see I'm going to uh cover these are certain topics that I had kept uh if I go through all it would be a little late you can let me know what is something that you would like to understand I have a section on uh distributions then I am also talking about central limit theorem and then there's hypothesis testing and certain test that we would be doing and covariance and correlation So you let me know which is a particular section that you would like me to take and we can spend the next 10 15 minutes talking about this because if I cover the whole um whole deck it will take a little long uh and given it's a Sunday I'm sure most of you would want to uh complete the session at the earliest. So yeah let me know I'm okay to take up any >> I'll vote for the hypothesis please. Sorry, >> I'll vote for the hypothesis. >> Okay. Okay. Anyone else who wants to do something different? >> N square tests. >> Tests, right? Okay. >> N and K square. >> Yes. >> Okay. >> Basically, okay. So, let's do this. I have not covered K square. I have limited myself to basically the tests that are done on on numerical data. Kai square is basically done when you do a test on categorical but I can give you a a brief explanation of what Kaiquare does how do you use it in real use cases so let's do this let's go to um hypothesis testing this basically uh distributions um I would strongly recommend while there are a lot of distributions but you uh guys only need to focus on few of them so when you're talking about discrete distributions you should look at binomial, po and uniform. When you're looking at continuous distributions, you should look at normal, exponential and again uniform. But uh uniform distribution is both both discrete and continuous. But most of the time the data that you will get will always be either normal or exponential or it will be binomial or poison. Right? So focus on these um four and five distributions that should give you a good start. Um these are basically just examples. Central limit theorem is also important. If you are done with hypothesis I'll just quickly touch on central limit theorem. Again it's a favorite of interviewers. Um let's jump to hypothesis testing. Okay. So um I have what I have done is I have just added put these terms here and I have included a definition for each of the term. Don't bother to read these definitions because uh these are all textbook language. You will end up more confused. Just listen to me. I'll try to break these terms into uh in a simpler manner so that you understand the gist of it. Right? So first of all in hypothesis testing the first word is hypothesis. What do you mean by hypothesis? >> Assumption. >> Yeah. >> Assumption >> right >> right? You're making a certain assumption. You're making a certain statement or a certain claim. Now this assumption or claim or a statement can be made by anyone by you, me, any research agency, any educational organization, any client. uh these statements are known as hypothesis. If you remember some years back there was this one uh company com that that produ that manufactures comp plan right which is an energy drink and it it it said that if you have comp plan every day your average height would increase by say some x in right that is a claim that the company made right this is a hypothesis because we still don't know whether that claim is correct or not so such statements or exemptions or claims are known as hypothesis. Right? Now what is hypothesis testing? Now once I have made a claim, it is not a universal truth unless I test it. So the process of testing a hypothesis is known as hypothesis testing post which either the hypothesis will be accepted or it will be rejected. Yeah. So hypothesis testing is basically you perform certain test on the data right and then you conclude whether the hypothesis is correct or not based on the sample of data that you have collected. Now once you get uh a certain assumption right you have to make two hypothesis null hypothesis and alternative hypothesis. So basically what happens is a researcher or say the company right uh the company that is saying that if you drink comp plan your height will increase by say for example 5 in. So that is a claim that is made by the company that is the alternative hypothesis. Usually what we want uh a statement to be true. We want the claim to be true, right? That becomes a part of the alternative hypothesis. The complement of alternative hypothesis is null hypothesis. So null hypothesis basically is a very lazy hypothesis. Right? it will say there is no relation between drinking comp plan and the heights of uh children or say for example um again going back to the salary data set uh I have the salaries and I also have a gender now I'm just looking at the salaries and I'm making a claim that the average salary of men is greater than the average salary of women so this hypothesis I would want to prove that this claim is correct So this becomes my alternative hypothesis. While the null hypothesis again being a lazy hypothesis, it will say there is no relationship between the salaries of males and females. So it will say opposite to that of alternative hypothesis. So in this case my null hypothesis is the average salary of men is less than equal to the average salary of women. With me so far? Did you understand the example of null and alternative? Yes ma'am. >> Yes ma'am. >> Okay. >> Yes ma'am. >> Now comes P value. In a very very simple language P value is the evidence in support of null hypothesis. If you have a larger P value you will go ahead and accept the null hypothesis. If you have a lower P value you will go and reject the null hypothesis. Right? How do you decide what is a high value of P and what is a low value of P? For that we use alpha which is a significance level. Right? If if the alpha is if I take alpha as 05 and the p value is less than uh 05 I will reject the null hypothesis. If it is more than 0.05 I will accept the null hypothesis. Simply putting just remember P value is the evidence in support of null hypothesis. Lower P value indicates greater evidence against the null hypothesis. When you when you search for P value on Google, right, you will get a lot of statistical description of this term. It is not as complex as it sounds. when you start using p values in linear regression uh what will happen is say for example um you are looking you're you're doing a linear regression problem right where you are predicting say um the salaries of uh of of individuals based on their qualification based on their age how many years of experience they have and so on. So you will look at what is the um you you will look at each variable and you would look at the p value of each variable in connection with the dependent variable. When the p value is less than 0.05 or for that matter any level of alpha that you have selected what you are going to conclude this variable does not have a significant impact on the dependent variable. So a gender does not have any impact on the salary if the p value comes to be less than 0.05. This is where in practical you use the p value. But again as I said let's not go there. In the simplest form p value is the evidence in support of null hypothesis. All of you with me so far? >> Yes ma'am. >> Yeah ma'am. Great. >> Okay. So now let's look at how we actually do hypothesis testing. Right? These are the series of steps. The first thing is you write down the hypothesis either in your mind or on a piece of paper when you are solving problems that uh this is the hypothesis that is obvious from the data. Gender is playing a role on salary. Right? But so that is your hypothesis. Now basically hypothesis is defined using certain parameter. The average salary of men is greater than the average salary of women. We are talking about a certain parameter. We are talking about salary and we are talking about average salary. Right? So mostly hypothesis is about mean but it can also include proportion and standard deviation. However, in most of the use cases except for research, most of the use cases talk about mean as a parameter in hypothesis. Right? So, you describe your hypothesis. Then you define the null and alternative hypothesis. So for example um when I'm saying that uh based on certain exploratory analysis I have seen that the my assumption or the data what the data is telling me is that the salary of men is greater than the salary of women. So in here my alternate hypothesis is the salary of men is greater than the salary of women. The null hypothesis because it is a complement of alternative hypothesis becomes the salary of women is greater than equal to the salary of men. Right? For example over here if the claim is the average time spent by women on social media is more than men. Right? This is a claim. This is a hypothesis. What will be the alternate hypothesis? We want to the claim you would want to prove that the claim is true. Right? So that becomes your alternate alternative hypothesis. The opposite is null hypothesis. So what it that be? It would be that the meanantime spent by women on social media is less than equal to mean time spent by men on social media. Is it confusing or is it little clear? >> Clear man. >> Yeah. clear. >> Um, see these are certain rules that you need to keep in mind when you uh create your null and alternate hypothesis. A null hypothesis cannot have a not equal to sign. Even though it is a lazy hypothesis, if I am h if I have if my hypothesis is that the average height of students uh in class 4 is 4T, right? The null hypothesis will not say that it is not equal to 4T. the null hypothesis will still go ahead and say it is equal to 4 ft. Why? Because you cannot use not equal to sign in null hypothesis. Similarly, you cannot use equal to sign in alternate. So it can never be less than equal to or greater than equal to. It will always be either less than or greater than. Equal to sign only comes in null hypothesis. So null hypothesis can have an equal to sign or a greater than equal to sign or a less than equal to sign. These are certain rules that you keep in mind when you create a null and alternative hypothesis. Aligned >> so far. >> Yeah. Okay. So now I have defined my null hypothesis. I have defined alternative hypothesis. The next thing that I will do is I will calculate the test statistics. What is test statistics? I have a sample. I have a population. Right? So while I have collected say uh heights of 100 students in class 4, there will also be a population mean where it's much larger than the sample. Right? So my xbar is the average height of the students in the sample that I have. Mu that you see here is the average height of the students in the entire population. Right? So this this statistic will be xar minus mu divided by the standard deviation by roo<unk> n. This is a formula that you uh again this will differ for every u different kind of a distribution. When you're looking at a normal distribution this will be a zed statistic. When you're looking at a sample and a normal distribution this will be a t statistic. When you're looking at k square distribution, this will be a k square statistic. But basically, you calculate a test statistic which is the difference between the estimated value of the parameter and the hypothesis value. Right? Usually you will have you will have your xbar, you will have the mu of course standard deviation you can calculate or it would be given and n is always the sample size the number of observations. So you calculate your zed statistic. Right? Now once you have your zed statistic, you calculate the p value. But before you calculate the p value, you also have to decide what is the level of significance that I want to take. What should be the value of alpha? Alpha basically is the probability of making a type one error. Okay. Um I don't want to make it very confusing but just simply remember let's just look at type one error for today. Type two error you can go back and do your homework. Type one error is the conditional probability of rejecting the null hypothesis when it is true. You would want to have this probability as low as possible. So what happens is when you're doing certain medical experiments, right? Specifically in the medical field, you want to accept the null hypothesis, right? You don't want to reject it when it is true because the stakes are very high in in in the field of medical science, right? In in those experiments that are conducted in the field of of medicines, right? Your level of alpha is very very low. So basically depending on what is the context and what is the problem statement you will decide on the value of alpha or the significance level using the zed statistic you will calculate the value of p and if the p value is less than the significant significance level you will reject the null hypothesis. If it is more than the significance level you will accept the null hypothesis. So if I am um accepting or rejecting the null hypothesis, it will have an impact on what kind of variables am I going to take into my final model? If I'm just talking about how it is actually used in a data science scenario, clear questions, doubts, should I repeat the entire process again? Ma'am I have a question that >> ma'am here we are calculating z stat by this formula and from here we can calculate zcore as well by by the use of zcore table but how can we calculate uh p score over from here as because uh when I go for it in geek score we there have mentioned a critical score what is that critical score as well >> huh so this is basically there are certain functions in the excel and also in python which will help you calculate your critical zed value and the p value based on the zed statistics that you get. So you just need to Google it. Um nowadays nobody refers to those values which are given at the end of the book. You can just uh either do it in Excel or you can do it in Python and you will get the values of P because P is basically P value is basically given as uh in very layman's term when you plot a normal distribution curve. Right? P value will be the area under the curve. Once you get your zed or Z statistic, you will get a value of P based on whether it is a two sample test or a single sample test. If it is a two sample test then on both the sides of your curve you will get a certain p value which you will add up or if it is a single sample experiment that you are doing you will get a p value on one side of the on the curve. So basically to sum it up you can use both excel or you can also use uh python to calculate it. The functions you can easily google and get it if I am not wrong. Um how you calculate is basically u you do some uh it's it's a function called n o r ms i and inverse or something in excel you can use that also but it is you can easily calculate is what is the uh the summary >> okay ma'am thank here. >> Um okay. Um okay. What I would suggest is let's uh first of all try to get this these two steps correct. This is very simple, right? Because you already have these values. But if your null and alternate hypothesis are not correct, right? then even if you calculate the p value you will not be able to reject or accept the the right hypothesis. So first try to understand how do you create your null and alternate hypothesis and then you look into these steps. These are all mathematical steps. There is no logic or understanding here. Basically you calculate your test statistic then uh then you look at uh based on the zed statistic you look at your p values then you decide on any level of significance level and then based on that you accept or reject but the the critical thing is how you define your ho and ha these this is very important it is a little tricky also these are very simple examples but as you progress uh the examples would not be so straightforward so this is the key this is all maths this basically just formulas. >> Yeah, an tell me. >> Ma'am, I'm still not clear what's the difference between population mean and statistical mean and why are they not both the same thing? >> Uh, so you're not clear with Xbar and mu is what you're saying? >> Yes, ma'am. >> Okay. So see what happens is um your xbar will actually change uh it it will vary with what what what do you think the mean value of your sample will depend on >> the sample we take >> h it basically uh depends on your uh uh on the sample size right it mostly depends on first of all how unbiased your sample is, right? Then what is the uh the sample size or what is n? So your xbar will be different because you're looking at different samples. But your mu will also be different. Sorry, your mu will not change because there we are talking about a much much larger uh sample or we talking about the population in general. So what happens is I am making a claim about the population. So I'm saying that say the average uh income of families living in Bangalore is um is greater than 4200 for example. I'm making this then I create a random sample of uh 40,000 families and I I find that the disposable income for these 40,000 families is say 4,250 for example. So what happens? The xbar here is 4,250. This is the sample mean. But my mu is what? It is 4200 because it is based on a much wider sample size. Of course, we will never have the population. It is impossible. It is time consuming. It is very expensive to get the population mean. But there will always be certain estimate of the population mean. So that's why when you do Xar minus mu they will never be equal. They might be close there might be a difference in it but they will be different based on the kind of sample that you have the sample size the approach that you have taken to create that sample. So yeah, lots of >> Yeah, when you start looking at more uh examples, right, this concept of Xbar and mu will start um uh will start making more sense. >> Thank you very much. >> Yeah. Uh let's just quickly do this one more thing because I want to make sure that this is more clear. As I said, this is very important post. uh once we are done with this we can we can close the we can close the session. So uh I want you to now tell me let's look at these four tell me uh what would be the null hypothesis here. Average annual salary of data scientist is different for males and females. What do you think is the u is the null and the alternative hypothesis? Average annual data scientist is same for males and females. >> That is both equal. >> Salaries of both males and females are equal which is considered as non-hypothesis and we can consider the claim of that statement as alternative hypothesis. >> Right. Exactly. So I will say the annual salary of male data scientist is equal to the annual salary of female data scientist. That is the null hypothesis. Alternative hypothesis is it is not equal to. Right. Good. Now we say on average the professionals with MS in data science earns more than professionals with MS in engineering. What would be the null hypothesis? Only the null hypothesis here. >> So the null hypothesis will be the professional with MS in data science and MS in engineering will earn equally. >> Yeah. So the null hypothesis here will be the professionals with MS in data science is less than equal to what professionals with MS in engineering earn. Clear? >> Clear ma'am. Yes ma'am. >> Yeah ma'am. >> Okay. Let's just look at the last example. I think this has now started making sense to you. The average time taken by the passport office to process the passport application is within 30 days. The average time is >> 30 days. >> More than 30 days. Greater than. >> Yeah. So the null hypothesis it is greater than equal to 30 days. Yeah. And alternative hypothesis is it is less than 30 days. Good. So basically keep these rules in mind. This becomes simpler. Once you have the correct null hypothesis and alternative hypothesis, these steps are basically just just simple formulas and calculating the Z score, critical values and p values. This becomes simple. The key is this rest I think the others are self-explanatory or if it is just basically formulas. So yeah, I think uh we are um it's already 6:35. uh any specific questions on the topics that I have covered so far I am okay to take or anything just on your data science journey. Yeah, tell me >> ma'am, how much stats is required for becoming data analyst? >> Um, you should be very good in descriptive statistics. Uh, uh, that is definitely required. You will not be doing a lot of hypothesis testing. So, so yeah, uh, the a little difficult part you can let go. But from a data analyst point of view, uh, you should know how to analyze individual columns. So uh then you should also know how to plot them right because there are certain graphs that you can plot for numerical uh data there's certain graphs that you can plot for categorical data then how do you make insights from that data right what are the what is each chart giving you what message it is giving you that also becomes very important so based on that I think focus more on descriptive analytics focus more on correlation um that is that is good enough >> I mean inferial statistics is not required. >> It won't be required in so much depth. Uh inferial becomes more important when you are doing uh using models, right? Uh but uh that is when it becomes more relevant when you're working as a data analyst. Basically you will basically be looking at the data and trying to create certain insights which you will do you you will of course create certain hypothesis but uh I have not seen a lot of uh people using uh hypothesis testing and data analytics. You would basically just look at correlation or you would do a univariate and a biariate data analysis and per uh you know conclude your insights. But it is good to have see this is very basic. You can just you should know what a normal distribution is that is required even if you are go for a data analyst job. So you should know what different distributions are that is also required >> because if you are looking at a data say for example you are looking analyzing the customer call service data the number of calls that a customer care receives right most of the times it it follows uh a certain distribution known as poison distribution. Now if you don't know what is poison distribution then uh what kind of analysis can you do? You can't do much right? You should know what is the average rate of occurrence. How does it look like? How does the value of lambda change over a period of time? Does it change? Does it not change? So uh distributions are important. The only thing I think you might not require too much is basically deep dive into hypothesis testing. But I think all other parts that I have covered today will be important. Ma'am. >> Yeah. >> Thank you, ma'am. >> Thanks. Yeah, tell me. >> Uh, is it significant to remember all those sigma values, significance values? >> No, you don't really need to remember this. So, see based on your experiments, usually see the significance level is in nine out of 10 experiments. uh unless you are in a medical field or in a field where the stake is very high uh generally the p value is sorry the significance level is taken at 0.05 05 and you alone will not be deciding on this. This will usually be a group take. You cannot decide what kind of significance level I should take because when you are really performing experiments when you're actually doing it, right? It is more of a group or a team call than your own call because the cost of uh committing a tag one error is very high, right? So you of course you will have a say in it as you start progressing but it is mostly a team's call and uh as I said nine out of 10 times you use uh alpha 0.05 unless you are into a very specific kind of a research role. Okay ma'am. Okay. Thank you so much ma'am. I have >> I have one. >> Yeah. Yeah. Tell me. Uh >> you go ahead please. >> Okay. Thank you ma'am. So Python provides a libraries like sci and stats. So do we need to have a hands-on experience manual experience to calculate this test hypothesis testing or just Python libraries? >> Yeah. Yeah you can use Python libraries that is good. So only Python libraries is enough. Yes. Yes, you can use Python libraries. If when you are p uh you can use both Python and Excel, it is good to know, right? If you are just initially when you will start when you will begin your learning journey, it is good to see how things work in Excel. But slowly when you are on the job, you will you will have to use Python. Know both. But at the end, you will using more of it in Python only. Okay. Bur can go ahead. Ma'am, I have two questions slightly out of scope of today's session. Mhm. >> Number one, your perspective on data scientist versus AI researcher versus AI engineer. And second question is as a thirdyear computer engineer. We've been told to practice a lot of DSA questions because they consider that as holy grail of most of company interviews. >> So also your take on that because that takes a significant amount of time to practice. H I think third one is very important as a computer engineer and later on either as a data scientist or a data engineer data structures are very important. Uh spend as much time as possible because a lot of engineers uh even after completing engineering struggle with the concepts of data structures. So irrespective of whatever role you take in the future right and right now I think it is a little while you might feel that you want to go for X role or Y role uh you would use data structures in most of the roles so make that as a very strong foundation um your next question on whether as a data scientist or an AI engineer it's something that you will have to explore both of them are equally interesting depends on what uh What challenges you more? What makes you more excited? Both of them have equal opportunities. Uh what has happened lately is that most of the companies have now stopped making models because how much more models can you make, right? Um you might you you already have certain set of models. You try to tweak certain parameters. Uh try to make it a little bit more accurate. Uh but there's only a limit till which you can make it more your accuracy can increase. right? Uh most of the companies are now actually looking at uh how are actually looking at data engineering. So um it is your call as I said what makes you comfortable or excited about a role but yes uh data structures is something that will always be required and will always be in in in demand and you will always be tested on it as a computer engineer. >> It was a pleasure attending today's session. Thank you so much. >> Thank you. >> And as you said data set is very important. So in the field of data science also is it important to have a strong foundation of data structures or just uh basic questions are enough? M no see see I suggest you do it uh because as you start solving more uh complex problem with and very large data sets right um this data structures will will will help you a lot of time I've seen the teams uh struggling with uh deploying the model uh they're struggling with uh you know even basic data science appro approaches also while it might not require you to do uh use a lot of DSA but when you are cleaning the data and all of that you will require those concepts so please don't give it like a u uh give it importance give it priority this will come handy >> okay ma'am thank you thank you for your all these guidance >> most welcome any other questions >> the PDF link this slides link. >> Um I am not sure if I can share this. I will check with the admin guys if they are okay to share it. I will definitely share it. But what I understand is you guys have already been given certain notes on uh these topics that I have covered today. Is that true? >> Ma'am, there are only articles. So your DPT is very much useful. >> Okay. Sure. Okay. Sure. I I will share it. I will um I'm coordinating with someone called Vikas. I'm not sure if he's also the guy with whom you are coordinating. I will uh send this PPT to them to him and he can then uh forward it to all of you. >> Ma'am, you can share in the chat group. >> Uh I will have to check with Vikas first because uh they would want you to refer to Geek of Geeks articles, right? Uh if I share it, I'm not sure if there would be a conflict. Let me get uh these basic uh things in place first. I'm okay to share it and there's nothing there's no big deal in these slides but Vikas or the organization might have certain reservations. So I don't want to share it right now. Let me have a word and I will share it. It's it's I mean this is no u it's not a secret or it's not like very confidential stuff. I'm okay to share it. Let me get a go ahead from

Original Description

Statistics is the heart of data science โ€” it empowers you to make sense of data, draw conclusions, and build powerful models. In this video, weโ€™ll cover the most important statistical concepts every data science enthusiast must know, such as mean, median, mode, standard deviation, probability distributions, hypothesis testing, correlation, regression, and more. Whether youโ€™re a beginner or looking to revise your basics, this video simplifies complex ideas using real-world examples and use-cases in data analytics and machine learning. Get all important links here: ๐Ÿ”— Pre Register for Nation SkillUp - https://gfgcdn.com/tu/VJ5/ Visit website: https://geeksforgeeks.org/ Explore Premium LIVE, Online & Offline Courses (For maximum discount use code - GFGYT30) : https://geeksforgeeks.org/courses/ Solve POTD: https://www.geeksforgeeks.org/problem-of-the-day Ongoing contests, hackathons and events: https://www.geeksforgeeks.org/events Follow us for more fun, knowledge and resources, join us on our social handles: ๐Ÿ“ฑTake GeeksforGeeks everywhere in your pockets! Don't forget to download our official app: https://geeksforgeeksapp.page.link/gfg-app ๐Ÿ’ฌ X- https://x.com/geeksforgeeks ๐Ÿง‘โ€๐Ÿ’ผ LinkedIn- https://www.linkedin.com/company/geeksforgeeks ๐Ÿ“ท Instagram- https://www.instagram.com/geeks_for_geeks/?hl=en ๐Ÿ’Œ Telegram- https://t.me/s/geeksforgeeks_official ๐Ÿ“Œ Pinterest: https://in.pinterest.com/geeks_for_geeks/ Also, Subscribe if you haven't already! :) #StatisticsForDataScience #DataScience #MachineLearning #StatisticalAnalysis #Probability #DescriptiveStatistics #InferentialStatistics #DataAnalytics #HypothesisTesting #RegressionAnalysis #DataScienceBeginners #GeeksforGeeks #Learntocode #GfG
Watch on YouTube โ†— (saves to browser)
Sign in to unlock AI tutor explanation ยท โšก30

Playlist

Uploads from GeeksforGeeks ยท GeeksforGeeks ยท 0 of 60

โ† Previous Next โ†’
1 How I got into Walmart | Shailesh Sharma
How I got into Walmart | Shailesh Sharma
GeeksforGeeks
2 Upgrade yourself In 29 Days | GeeksforGeeks
Upgrade yourself In 29 Days | GeeksforGeeks
GeeksforGeeks
3 Learn AWS Fundamentals For Free
Learn AWS Fundamentals For Free
GeeksforGeeks
4 Conversation With Young Achievers | Meet the winners of Bi-Wizard Coding Contest | GeeksforGeeks
Conversation With Young Achievers | Meet the winners of Bi-Wizard Coding Contest | GeeksforGeeks
GeeksforGeeks
5 Meet The Winners Of Bi-Wizard Coding Contests | GeeksforGeeks
Meet The Winners Of Bi-Wizard Coding Contests | GeeksforGeeks
GeeksforGeeks
6 Interview Prep Strategies | PayPal
Interview Prep Strategies | PayPal
GeeksforGeeks
7 OLX Interview Preparation Strategies | Hukam Singh
OLX Interview Preparation Strategies | Hukam Singh
GeeksforGeeks
8 Meet Some More Winners Of Bi-Wizard Coding Contests | GeeksforGeeks
Meet Some More Winners Of Bi-Wizard Coding Contests | GeeksforGeeks
GeeksforGeeks
9 Live Mock DSA
Live Mock DSA
GeeksforGeeks
10 Microsoft Azure For Absolute Beginners
Microsoft Azure For Absolute Beginners
GeeksforGeeks
11 Python for Data Science | Data Science Master Bootcamp | Arpit Jain
Python for Data Science | Data Science Master Bootcamp | Arpit Jain
GeeksforGeeks
12 Getting Started with Data Analysis | Data Science Master Bootcamp | Ashish Jangra
Getting Started with Data Analysis | Data Science Master Bootcamp | Ashish Jangra
GeeksforGeeks
13 How to prepare theory subjects for SDE interviews | Geeks Summer Carnival 2022
How to prepare theory subjects for SDE interviews | Geeks Summer Carnival 2022
GeeksforGeeks
14 Get Your Tickets To The Geeks Summer Carnival | GeeksforGeeks
Get Your Tickets To The Geeks Summer Carnival | GeeksforGeeks
GeeksforGeeks
15 TED Talk Data Analysis Project | Data Science Master Bootcamp | Ashish Jangra
TED Talk Data Analysis Project | Data Science Master Bootcamp | Ashish Jangra
GeeksforGeeks
16 How I Secured AIR 9 in GATE'22 |  Tushar
How I Secured AIR 9 in GATE'22 | Tushar
GeeksforGeeks
17 Learn Java Backend Development | Geeks Summer Carnival | GeeksforGeeks
Learn Java Backend Development | Geeks Summer Carnival | GeeksforGeeks
GeeksforGeeks
18 How to Recognize which Data Structure to use in a question | Geeks Summer Carnival | GeeksforGeeks
How to Recognize which Data Structure to use in a question | Geeks Summer Carnival | GeeksforGeeks
GeeksforGeeks
19 Learn Data Structures and Algorithms | GeeksforGeeks
Learn Data Structures and Algorithms | GeeksforGeeks
GeeksforGeeks
20 Interview experience at Flipkart | GeeksforGeeks
Interview experience at Flipkart | GeeksforGeeks
GeeksforGeeks
21 Lets Prepare for GATE'23 the Right Way | Sakshi Singhal | GeekSummerCarnival
Lets Prepare for GATE'23 the Right Way | Sakshi Singhal | GeekSummerCarnival
GeeksforGeeks
22 Highest Paying Jobs in 2022 | Ishan Sharma | Geeks Summer Carnival 2022 | GeeksforGeeks
Highest Paying Jobs in 2022 | Ishan Sharma | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
23 Geeks Summer Carnival 2022 | 5th April- 11th April | GeeksforGeeks
Geeks Summer Carnival 2022 | 5th April- 11th April | GeeksforGeeks
GeeksforGeeks
24 Preparing for SDE interviews | Soham Mukherjee | Geeks Summer Carnival 2022 | GeeksforGeeks
Preparing for SDE interviews | Soham Mukherjee | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
25 Full Stack Development with React & Node | Utkarsh Malik | Geeks Summer Carnival | GeeksforGeeks
Full Stack Development with React & Node | Utkarsh Malik | Geeks Summer Carnival | GeeksforGeeks
GeeksforGeeks
26 Introduction to Open Source and Roadmap to GSOC 2022 | Geeks Summer Carnival 2022 | GeeksforGeeks
Introduction to Open Source and Roadmap to GSOC 2022 | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
27 Web Scraping in Action | Geeks Summer Carnival 2022 | GeeksforGeeks
Web Scraping in Action | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
28 Getting Hired at BITCS via GfG Job Portal | Get Hired With GeeksforGeeks
Getting Hired at BITCS via GfG Job Portal | Get Hired With GeeksforGeeks
GeeksforGeeks
29 How to build a faster landing Page | Geeks Summer Carnival 2022 | GeeksforGeeks
How to build a faster landing Page | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
30 Geeks Summer Carnival | 5th To 11th April, 2022 | GeeksforGeeks
Geeks Summer Carnival | 5th To 11th April, 2022 | GeeksforGeeks
GeeksforGeeks
31 How to get ideas for Startup | Geeks Summer Carnival 2022 | GeeksforGeeks
How to get ideas for Startup | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
32 Journey from Tier 3 to JusPay | GeeksforGeeks
Journey from Tier 3 to JusPay | GeeksforGeeks
GeeksforGeeks
33 Geeks Summer Carnival 2022 | GeeksforGeeks
Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
34 Dispelling Myths and Pre conceptions of Programming Languages
Dispelling Myths and Pre conceptions of Programming Languages
GeeksforGeeks
35 Must Do System Design Questions
Must Do System Design Questions
GeeksforGeeks
36 Understanding Sorting Techniques in an hour | Keerti Purswani | Geeks Summer Carnival
Understanding Sorting Techniques in an hour | Keerti Purswani | Geeks Summer Carnival
GeeksforGeeks
37 Get Hired at NEC | Job-A-Thon 8
Get Hired at NEC | Job-A-Thon 8
GeeksforGeeks
38 Journey from Tier 3 college to Microsoft | GeeksforGeeks
Journey from Tier 3 college to Microsoft | GeeksforGeeks
GeeksforGeeks
39 Get Hired with GeeksforGeeks at SuperK | Job A Thon 8
Get Hired with GeeksforGeeks at SuperK | Job A Thon 8
GeeksforGeeks
40 GeeksforGeeks: Redesigned
GeeksforGeeks: Redesigned
GeeksforGeeks
41 From Tier 3 to cracking multiple interviews | GeeksforGeeks
From Tier 3 to cracking multiple interviews | GeeksforGeeks
GeeksforGeeks
42 Live Mock DSA
Live Mock DSA
GeeksforGeeks
43 Youtube Data Analysis | Ashish Jangra | GeeksforGeeks
Youtube Data Analysis | Ashish Jangra | GeeksforGeeks
GeeksforGeeks
44 DSA Self-Paced Course Preview | Sandeep Jain | GeeksforGeeks
DSA Self-Paced Course Preview | Sandeep Jain | GeeksforGeeks
GeeksforGeeks
45 GATE Live Classes | Prepare for GATE CS 2023 | GeeksforGeeks
GATE Live Classes | Prepare for GATE CS 2023 | GeeksforGeeks
GeeksforGeeks
46 Journey from JIIT to Adobe
Journey from JIIT to Adobe
GeeksforGeeks
47 Life Is Unfair Ft. Shonty badmash | LIVE Discord Session | A GeeksforGeeks Exclusive
Life Is Unfair Ft. Shonty badmash | LIVE Discord Session | A GeeksforGeeks Exclusive
GeeksforGeeks
48 Interview Experience at Google | Tech Dose
Interview Experience at Google | Tech Dose
GeeksforGeeks
49 Live Mock DSA
Live Mock DSA
GeeksforGeeks
50 Interview Experience @ Amazon | GeeksforGeeks
Interview Experience @ Amazon | GeeksforGeeks
GeeksforGeeks
51 My journey through the tech world from India to US | Vidushi | GeeksforGeeks
My journey through the tech world from India to US | Vidushi | GeeksforGeeks
GeeksforGeeks
52 Complete Interview Preparation Course | GeeksforGeeks
Complete Interview Preparation Course | GeeksforGeeks
GeeksforGeeks
53 Live Mock DSA
Live Mock DSA
GeeksforGeeks
54 Getting Hired at FiftyFive Technologies | Job-a-thon 9.0
Getting Hired at FiftyFive Technologies | Job-a-thon 9.0
GeeksforGeeks
55 GFG Karlo, Ho Jayega | GeeksforGeeks ft. Khaleel Ahmed
GFG Karlo, Ho Jayega | GeeksforGeeks ft. Khaleel Ahmed
GeeksforGeeks
56 How I got job offers from 2 big companies : Arcesium & Microsoft | GeeksforGeeks
How I got job offers from 2 big companies : Arcesium & Microsoft | GeeksforGeeks
GeeksforGeeks
57 LINUX for Beginners | GFG x Itversity
LINUX for Beginners | GFG x Itversity
GeeksforGeeks
58 My interview experience at Walmart | GeeksforGeeks
My interview experience at Walmart | GeeksforGeeks
GeeksforGeeks
59 Get Hired at Speckyfox
Get Hired at Speckyfox
GeeksforGeeks
60 Live Mock DSA
Live Mock DSA
GeeksforGeeks

Related Reads

Up next
Solve Any Math Problem Step by Step โ€” Free (Type or Snap a Photo)
Zariga Tongy
Watch โ†’