Previously, we talked about z-scores and how they tell position on the Normal Curve - but how do we know if the data fits a Normal Distribution in the first place? A Normal Quantile Plot (also sometimes called a Normal Probability Plot) can give us an idea of Normality of the data. Put your data into L1, then go to Stat Plot and choose the very last graph. Don't forget to go to zoom - 9(stat) to get a good view of your graph!
We look for three characteristics of the data to assess Normality. 1) Does the data roughly form a straight line? Or do we see curvature at the ends? 2) Is the data centered about the x-axis - do we see roughly equal observations above and below? And finally, 3) Does a majority of the data appear to be in the middle of the graph rather than in the tails? If the answer to these 3 questions is "yes" then we can assume that our data appears roughly Normal.
Note that this is our ASSUMPTION based on our OBSERVATIONS. We didn't actually PROVE anything with our Normal Quantile Plot. Choose your language carefully when you interpret this.
Monday, September 16, 2013
Normal Curve and Z-Scores
The Normal Curve is otherwise known as the bell curve. It is a symmetric distribution centered about the mean and spread out by standard deviations. 68% of the data lies within the first standard deviation. 95% lies between the first two standard deviations. And 99.7% of the data falls within the first 3 standard deviations. Anything outside of the first three standard deviations is considered an outlier of the data (which means that 0.3% of the observations will be outliers).
To calculate position on the Normal curve, we can use the z-score formula, which is the sample mean minus the population mean, all over the standard deviation. Our answer (the z-score) can be interpreted as the number of standard deviations (positive or negative) above or below the mean.
To calculate position on the Normal curve, we can use the z-score formula, which is the sample mean minus the population mean, all over the standard deviation. Our answer (the z-score) can be interpreted as the number of standard deviations (positive or negative) above or below the mean.
Tuesday, September 10, 2013
Standard Deviation
The standard deviation is how far away each observation is, on average, from the mean of the data. We use standard deviation as a measure of spread about the mean, just as we use the interquartile range as a measure of spread about the median. Standard deviation, as paired with the mean, is typically used in symmetric or roughly symmetric data sets.
We can calculate the standard deviation using a long formula, or we can simply find it in the calculator when we do a 5-number summary (1-var statistics). Don't forget to always interpret it in the context of the problem!
We can calculate the standard deviation using a long formula, or we can simply find it in the calculator when we do a 5-number summary (1-var statistics). Don't forget to always interpret it in the context of the problem!
Monday, September 9, 2013
Guest Speaker from Fayetteville
Dr. Watson, a petrochemical engineer from the University of Arkansas, stopped by the library today to talk about a two-day trip opportunity. He also discussed ACT score and how important it is to the college application process. You need a 19 to get into the University of Arkansas unconditionally, and 18 to get in with remediation classes.
Remember, I do after school ACT prep from 3:40pm - 5:40pm in Mrs. Cain's room - 120. We cover all the subjects: Mrs. Cain does English and reading, I do math and science (aka how to read graphs properly). Come and join us - practice makes perfect!
Shout out to Nakeya and Antonia for coming after school today, and shout out to Trey and Megan for ZAP-ing their tests for a better grade.
Remember, I do after school ACT prep from 3:40pm - 5:40pm in Mrs. Cain's room - 120. We cover all the subjects: Mrs. Cain does English and reading, I do math and science (aka how to read graphs properly). Come and join us - practice makes perfect!
Shout out to Nakeya and Antonia for coming after school today, and shout out to Trey and Megan for ZAP-ing their tests for a better grade.
Sunday, September 8, 2013
Hints for Problem Set #2
Problem 1A: Think about what type of data is displayed in a histogram, and think about how many random variables that Mrs. Timmons is trying to work with. For part three, I'm asking what features she might see in the histogram IN GENERAL. I didn't provide you with one. But if she was looking at a histogram, what would she see? Think: CUSS....and go from there.
Problem 1C: Think about the shape of the histogram: given this specific shape, which estimate should be higher, the mean or the median? Check your notes from Friday, August 30th if you're stuck. I'm looking for an approximation here, there's no "right" answer as long as you justify it properly in part B.
Problem 1D: How can we measure center? There is, as we know from our notes from August 30th, more than one way....
Problem #5: Assume that you're taking the median of the 6 existing weights of luggage - that is, don't count x7 just yet. Otherwise, you're not going to be able to start the question.
Problem 1C: Think about the shape of the histogram: given this specific shape, which estimate should be higher, the mean or the median? Check your notes from Friday, August 30th if you're stuck. I'm looking for an approximation here, there's no "right" answer as long as you justify it properly in part B.
Problem 1D: How can we measure center? There is, as we know from our notes from August 30th, more than one way....
Problem #5: Assume that you're taking the median of the 6 existing weights of luggage - that is, don't count x7 just yet. Otherwise, you're not going to be able to start the question.
Thursday and Friday: Cities Survey/Experiment
Does giving a point of reference increase accuracy in estimation? Students worked in partners to explore this effect. Each group chose a city, then asked 15 different people (each): how many people do you think live in this city? OR, How many people do you think live in this city, given the 2010 population?
Most groups should have found that their data is more spread out without giving a reference point. When we throw in a reference point (otherwise known as an "anchor"), people then have grounds in which to base their estimation. Otherwise, they are relying solely on their prior knowledge of the city when making an estimation.
Students then turned in 7 questions about their findings. One key point to reinforce is that, when given symmetric data, use the mean as a measure of spread. When given skewed data, use the median as a measure of spread.
Shout out to Issac Wilburn and Josh Reynolds for working with remarkable intensity during the entire class period on Friday!
Most groups should have found that their data is more spread out without giving a reference point. When we throw in a reference point (otherwise known as an "anchor"), people then have grounds in which to base their estimation. Otherwise, they are relying solely on their prior knowledge of the city when making an estimation.
Students then turned in 7 questions about their findings. One key point to reinforce is that, when given symmetric data, use the mean as a measure of spread. When given skewed data, use the median as a measure of spread.
Shout out to Issac Wilburn and Josh Reynolds for working with remarkable intensity during the entire class period on Friday!
Wednesday, September 4, 2013
Boxplots and the 1.5 Outlier Test
Boxplots display univariate, quantitative data. The data is ordered and placed into quartiles. Quartiles 1 and 4 are represented by the whiskers, and quartiles 2 and 3 (called the Interquartiles, since they represent the middle 50% of the data) are represented by the boxes.
We can CUSS about the boxplot because they show center (the median), outliers (via the 1.5 outlier test), shape (look for symmetric versus skewed), and spread (look at the range or the interquartile range).
Problem Set 2 was given out in class today. Check tomorrow's blog post for hints for question 1.
We can CUSS about the boxplot because they show center (the median), outliers (via the 1.5 outlier test), shape (look for symmetric versus skewed), and spread (look at the range or the interquartile range).
Problem Set 2 was given out in class today. Check tomorrow's blog post for hints for question 1.
Subscribe to:
Posts (Atom)