Your FREE AP Statistics Practice Test 2026 – 230+ Q&A
Prepare with realistic, College Board-style AP Statistics questions — take a full multiple-choice practice test or drill one unit at a time.
How ready are you?
AP Statistics Test Prep
To find us again, just search Find us again. Search “Career Employer AP Statistics”
Lock each answer to reveal the explanation instantly as you go.
To find us again, just search Find us again. Search “Career Employer AP Statistics”
Practice by domain
More ways to study AP Statistics
Study GuideAP Statistics Study GuideThe most essential AP Statistics topics to master before exam day — with interactive chapter quizzes and flashcards built in.FlashcardsAP Statistics FlashcardsHundreds of AP Statistics flashcards in multiple study modes — flip, match, type, and quiz.
AP Statistics Practice Questions
Which measure of center is most resistant to the influence of extreme values in a data set?
The mean
The range
The median
The standard deviation
Correct answer: The median
The median is the most resistant measure of center because it depends only on the position of the middle value, not the magnitude of every observation. A single extreme value pulls the mean toward it but barely shifts the median. The range and standard deviation describe spread, not center, and both are strongly affected by extremes.
In a distribution that is skewed to the right, how does the mean typically compare to the median?
The mean is less than the median
The mean is greater than the median
The mean equals the median
The mean is always exactly twice the median
Correct answer: The mean is greater than the median
In a right-skewed distribution the mean is greater than the median because the long right tail contains large values that pull the mean upward while leaving the median near the bulk of the data. This relationship is reversed for left-skewed data, where the mean is pulled below the median. The mean and median coincide only in a symmetric distribution.
According to the empirical rule, approximately what percentage of values in a normal distribution fall within two standard deviations of the mean?
95%
68%
99.7%
50%
Correct answer: 95%
The empirical rule states that about 95% of values lie within two standard deviations of the mean in a normal distribution. About 68% fall within one standard deviation and about 99.7% within three. The value 50% corresponds to one side of the symmetric curve, not a standard-deviation interval.
A student scores 82 on a test with a mean of 74 and a standard deviation of 4. What is the student's z-score?
0.5
2.0
8.0
-2.0
Correct answer: 2.0
The z-score of 2.0 is found by subtracting the mean from the value and dividing by the standard deviation: (82 - 74) / 4 = 8 / 4 = 2.0. This tells us the score is two standard deviations above the mean. A positive result confirms the score is above average, ruling out the negative option.
Using the 1.5 x IQR rule, an observation is flagged as an outlier if it falls below which boundary?
The mean minus two standard deviations
Q1 minus 1.5 times the IQR
Q1 minus the IQR
The median minus 1.5 times the IQR
Correct answer: Q1 minus 1.5 times the IQR
The lower outlier boundary is Q1 minus 1.5 times the IQR; any value below it is considered an outlier. The upper fence is Q3 plus 1.5 times the IQR. The rule is built on quartiles and the interquartile range, not on the mean, standard deviation, or median.
A data set has Q1 = 30 and Q3 = 50. What is the interquartile range?
40
80
10
20
Correct answer: 20
The interquartile range is 20, calculated as Q3 minus Q1, or 50 - 30 = 20. The IQR measures the spread of the middle 50% of the data. Adding the quartiles (80) or averaging them (40) does not produce a measure of spread.
Which five values make up the five-number summary of a quantitative data set?
Mean, median, mode, range, standard deviation
Minimum, Q1, median, Q3, maximum
Minimum, mean, median, mode, maximum
Q1, Q2, Q3, mean, standard deviation
Correct answer: Minimum, Q1, median, Q3, maximum
The five-number summary consists of the minimum, first quartile, median, third quartile, and maximum. These five values describe center, spread, and the extremes of a distribution and are the basis for constructing a boxplot. The mean, mode, and standard deviation are not part of the five-number summary.
What does the length of the box in a boxplot directly represent?
The total range of the data
The standard deviation
The interquartile range
The mean of the data
Correct answer: The interquartile range
The length of the box in a boxplot represents the interquartile range, the distance from Q1 to Q3 spanning the middle 50% of the data. The whiskers, not the box, extend toward the extremes. The boxplot does not display the mean or standard deviation.
A symmetric, bell-shaped distribution of a continuous quantitative variable is best described as which type?
Normal distribution
Uniform distribution
Bimodal distribution
Skewed distribution
Correct answer: Normal distribution
A symmetric, bell-shaped distribution is described as a normal distribution, the foundation for z-scores and the empirical rule. A uniform distribution is flat, a bimodal distribution has two peaks, and a skewed distribution is asymmetric with a longer tail on one side.
The heights of adult women are approximately normal with a mean of 64 inches and a standard deviation of 2.5 inches. Between which two heights do about 68% of these women fall?
61.5 inches and 66.5 inches
59 inches and 69 inches
56.5 inches and 71.5 inches
62.5 inches and 65.5 inches
Correct answer: 61.5 inches and 66.5 inches
About 68% of the women fall between 61.5 and 66.5 inches, found by taking the mean plus or minus one standard deviation: 64 - 2.5 = 61.5 and 64 + 2.5 = 66.5. The empirical rule assigns 68% to one standard deviation. The interval 59 to 69 covers two standard deviations (about 95%).
If every value in a data set is increased by 10, what happens to the standard deviation?
It increases by 10
It increases by 10
It stays the same
It doubles
Correct answer: It stays the same
The standard deviation stays the same because adding a constant shifts every value by the same amount, leaving the distances between values unchanged. Standard deviation measures spread, which is unaffected by a uniform shift. Only the center, such as the mean, increases by 10.
A normally distributed exam has a mean of 500. Student A has a z-score of 1.5 and Student B has a z-score of -0.5. Which statement correctly compares their raw scores?
Student B scored higher than Student A
Student A scored higher than Student B
Both students scored exactly 500
The z-scores do not allow any comparison of their scores
Correct answer: Student A scored higher than Student B
Student A scored higher than Student B because a z-score measures how many standard deviations a value lies above or below the mean. Student A is 1.5 standard deviations above the mean while Student B is half a standard deviation below it. Since both come from the same distribution, the larger z-score corresponds to the higher raw score.
A real estate report states that the mean home price in a town is far higher than the median home price. What does this most strongly suggest about the distribution of home prices?
It is skewed to the left
It is perfectly symmetric
It is uniform
It is skewed to the right
Correct answer: It is skewed to the right
The distribution is most likely skewed to the right because the mean being much larger than the median indicates a long tail of expensive homes pulling the mean upward. In a left-skewed distribution the mean would fall below the median, and in a symmetric distribution the two measures would be roughly equal.
A data set of daily temperatures has Q1 = 58, Q3 = 70, and includes a reading of 92 degrees. Using the 1.5 x IQR rule, is 92 an outlier?
No, because it is below 88
Yes, because it exceeds 70
Yes, because it exceeds 88
No, because it is within two standard deviations
Correct answer: Yes, because it exceeds 88
The value 92 is an outlier because it exceeds the upper fence of 88. The IQR is 70 - 58 = 12, so the upper boundary is Q3 + 1.5 times IQR = 70 + 18 = 88. Since 92 is greater than 88, it is flagged as a high outlier. The 1.5 x IQR rule uses quartiles, not standard deviations.
Two boxplots compare test scores from two classes. Class X has a longer box than Class Y, but both have the same median. What can be concluded?
Class Y has greater variability in the middle 50% of scores
The two classes have identical distributions
Class X has greater variability in the middle 50% of scores
Class X has a higher typical score than Class Y
Correct answer: Class X has greater variability in the middle 50% of scores
Class X has greater variability in the middle 50% of scores because a longer box means a larger interquartile range, indicating more spread among the central half of the data. The equal medians mean the typical scores are comparable, so neither class is centered higher. A longer box reflects spread, not a higher center.
Which of the following describes the shape of a distribution whose histogram has a long tail extending toward the lower values on the left?
Skewed to the right
Symmetric
Uniform
Skewed to the left
Correct answer: Skewed to the left
A distribution with a long tail extending toward the lower values is skewed to the left, sometimes called negatively skewed. The direction of skew is named for the side where the tail stretches, not where the peak sits. A right-skewed distribution would have its tail toward the higher values, and a symmetric distribution would have no tail.
A teacher reports that a class data set has a mean of 75 and a standard deviation of 0. What must be true about the data?
Every value in the data set is 75
The values are evenly spread from 0 to 75
Half the values are above 75 and half below
The data set contains exactly one outlier
Correct answer: Every value in the data set is 75
Every value in the data set must equal 75 because a standard deviation of 0 means there is no spread at all and all observations are identical. Standard deviation measures how far values deviate from the mean, so zero deviation forces every value to match the mean exactly. Any variation would produce a positive standard deviation.
A scholarship requires applicants to score at least 1.0 standard deviation above the mean on a normally distributed admissions test. Using the empirical rule, approximately what percentage of test-takers qualify?
About 16%
About 32%
About 68%
About 5%
Correct answer: About 16%
About 16% of test-takers qualify because the empirical rule places about 68% of values within one standard deviation of the mean, leaving about 32% in the two tails combined. Since the distribution is symmetric, half of that 32%, or about 16%, lies in the upper tail beyond one standard deviation above the mean.
A data set's five-number summary is 12, 20, 25, 28, 60. Which feature of this summary most strongly signals a possible high outlier?
The median of 25 is close to Q3 of 28
The minimum of 12 is below the median
The first quartile of 20 is below the median
The maximum of 60 is far above Q3 of 28
Correct answer: The maximum of 60 is far above Q3 of 28
The maximum of 60 sitting far above Q3 of 28 most strongly signals a possible high outlier. The IQR is 28 - 20 = 8, so the upper fence is 28 + 1.5 times 8 = 40, and 60 exceeds it. The median being near Q3 and the quartiles falling on either side of the median are normal features, not outlier signals.
Which type of graph is used to display the relationship between two quantitative variables?
A scatterplot
A boxplot
A bar chart
A histogram
Correct answer: A scatterplot
A scatterplot is the correct display because it plots paired values of two quantitative variables as points, letting you assess the form, direction, and strength of their association. A boxplot and a histogram each display only a single quantitative variable, and a bar chart displays counts for a categorical variable.
The correlation coefficient r can take on values only within which range?
From 0 to 1
From 0 to 100
From -1 to 1
From -100 to 100
Correct answer: From -1 to 1
The correlation coefficient is bounded between -1 and 1 inclusive. A value near -1 or 1 indicates a strong linear relationship, while a value near 0 indicates a weak linear relationship; the sign shows direction. The ranges 0 to 1, 0 to 100, and -100 to 100 misstate these limits.
A least-squares regression line predicting weight from height has the equation y-hat = 50 + 4x, where x is height in inches and y is weight in pounds. How is the slope interpreted in context?
For each additional pound of weight, height increases by 4 inches
For each additional inch of height, predicted weight increases by 4 pounds
A person with zero height is predicted to weigh 4 pounds
For each additional inch of height, predicted weight increases by 50 pounds
Correct answer: For each additional inch of height, predicted weight increases by 4 pounds
The slope of 4 means that for each additional inch of height, the predicted weight increases by 4 pounds. The slope describes the predicted change in the response variable (weight) per one-unit increase in the explanatory variable (height), not the reverse. The value 50 is the y-intercept, the predicted weight when height is zero, not the slope.
A least-squares regression model reports a coefficient of determination of r2=0.81. What does this value mean?
The slope of the regression line is 0.81
About 81% of the data points fall exactly on the regression line
The correlation between the variables is exactly 0.81
About 81% of the variation in the response variable is explained by the linear model
Correct answer: About 81% of the variation in the response variable is explained by the linear model
An r2 of 0.81 means about 81% of the variation in the response variable is explained by the least-squares model relating it to the explanatory variable. It is the square of the correlation, so r would be about 0.9, not 0.81. It does not represent the slope and does not mean 81% of points lie on the line.
When a residual plot for a linear regression shows a clear curved (U-shaped) pattern rather than random scatter, what does this most strongly suggest?
The two variables have no association at all
A linear model is not appropriate for the data
The correlation coefficient must equal zero
The data contain no influential points
Correct answer: A linear model is not appropriate for the data
A curved pattern in the residual plot most strongly suggests that a linear model is not appropriate, because well-fitting linear models produce residuals scattered randomly around zero with no leftover pattern. A clear curve means systematic structure remains unexplained by the line. The pattern does not imply zero correlation or rule out influential points.
In a simple random sample of size n, what is true about every possible group of n individuals from the population?
Each possible group of n individuals has an equal chance of being selected
Only groups that represent every subgroup can be selected
Larger groups are more likely to be selected than smaller ones
Each individual is selected based on convenience to the researcher
Correct answer: Each possible group of n individuals has an equal chance of being selected
The defining property of a simple random sample is that every possible group of n individuals from the population has an equal chance of being chosen. This is stronger than merely giving each individual an equal chance. Selecting by convenience or forcing subgroup representation describes other, non-SRS methods.
A pollster divides voters into age groups and then takes a separate simple random sample from within each age group. Which sampling method is being used?
Convenience sampling
Cluster sampling
Stratified random sampling
Systematic sampling
Correct answer: Stratified random sampling
This is stratified random sampling because the population is split into homogeneous groups, called strata, and a random sample is drawn from within each one. The strata here are age groups, chosen because members within them are similar. Cluster sampling would instead sample whole groups, and convenience sampling would pick whoever is easiest to reach.
A researcher randomly selects several entire classrooms in a school and surveys every student in those chosen classrooms. Which sampling method does this describe?
Cluster sampling
Stratified random sampling
Simple random sampling
Convenience sampling
Correct answer: Cluster sampling
This describes cluster sampling because the population is divided into groups, called clusters, and entire randomly selected clusters are surveyed in full. Each classroom is a cluster meant to resemble the population as a whole. This differs from stratified sampling, where only some members are sampled from every group rather than all members from a few groups.
A student surveys friends in the cafeteria because they are easy to reach, then generalizes the results to the whole school. What is the primary weakness of this approach?
It requires too large a sample size
It guarantees an overestimate of the true value
It is a stratified sample that ignores clusters
It is a convenience sample that is likely biased
Correct answer: It is a convenience sample that is likely biased
The primary weakness is that this is a convenience sample, which is likely biased because the friends chosen may differ systematically from the rest of the school. Selecting whoever is easiest to reach gives no assurance the sample represents the population. The flaw is bias from non-random selection, not sample size, and the direction of any error cannot be guaranteed.
In a well-designed experiment, what is the main purpose of including a control group?
To increase the total number of subjects in the study
To provide a baseline for comparison against the treatment group
To guarantee that no confounding variables exist
To ensure the sample is selected at random
Correct answer: To provide a baseline for comparison against the treatment group
The control group's main purpose is to provide a baseline for comparison, so the effect of the treatment can be judged against subjects who did not receive it. Without this comparison, researchers cannot tell whether changes are due to the treatment. A control group does not enlarge the sample's reach or by itself eliminate every confounding variable.
Which study design allows researchers to draw a cause-and-effect conclusion between an explanatory variable and a response?
An observational study with a large sample
A survey using stratified random sampling
A randomized comparative experiment
A study based on voluntary response data
Correct answer: A randomized comparative experiment
A randomized comparative experiment is the design that supports a cause-and-effect conclusion, because randomly assigning treatments balances out other variables across groups. Observational studies and surveys, even with excellent sampling, can establish association but not causation because lurking variables are not controlled. Voluntary response data is especially prone to bias.
In an experiment testing a new fertilizer, the amount of sunlight different plots receive also varies and is tied to which fertilizer they got. Sunlight in this study is best described as what?
A control group
A response variable
A stratifying variable
A confounding variable
Correct answer: A confounding variable
Sunlight is a confounding variable because its effect on plant growth cannot be separated from the effect of the fertilizer, since the two vary together. This makes it impossible to attribute results to the fertilizer alone. It is not the response being measured, nor a control group, nor a basis for forming strata.
What is the most important benefit of using random assignment of treatments in an experiment?
It guarantees the sample looks exactly like the population
It tends to create groups that are similar except for the treatment
It removes the need to have a control group
It allows the results to be generalized to a larger population
Correct answer: It tends to create groups that are similar except for the treatment
Random assignment tends to create treatment groups that are roughly similar in all other variables, so any difference in outcomes can be attributed to the treatment itself. This is what permits causal conclusions. It does not make the sample mirror the population, that is the job of random sampling, which governs generalization instead.
A nutrition researcher records what 500 randomly selected adults already eat and tracks their health over time, without assigning any diets. The researcher finds that coffee drinkers have fewer headaches. What is the strongest valid conclusion?
There is an association between coffee drinking and fewer headaches
Drinking coffee causes fewer headaches
Random assignment removed all confounding variables
The study proves coffee should be recommended for headaches
Correct answer: There is an association between coffee drinking and fewer headaches
The strongest valid conclusion is that there is an association between coffee drinking and fewer headaches, because this is an observational study with no treatment assigned. Without random assignment, lurking variables such as overall lifestyle could explain the link, so causation cannot be claimed. Only a randomized experiment could support a cause-and-effect statement.
A factory wants to inspect bolts coming off an assembly line and selects every 25th bolt for testing. To best apply a simple random sample instead, the inspector should do which of the following?
Continue selecting every 25th bolt but start earlier in the day
Inspect only the bolts that are easiest to reach on the line
Inspect all bolts from the first production batch
Number all bolts produced and use a random process to choose which to inspect
Correct answer: Number all bolts produced and use a random process to choose which to inspect
To take a simple random sample, the inspector should number all bolts and use a random process so that every possible group of bolts is equally likely to be chosen. Selecting every 25th item is systematic sampling, not an SRS. Choosing the easiest-to-reach bolts is convenience sampling, and inspecting one batch in full is a single cluster.
A school administrator wants opinions from students and ensures that freshmen, sophomores, juniors, and seniors are each represented proportionally by randomly sampling within each grade. Why might this stratified approach be preferred over a single simple random sample?
It removes the need for any randomization
It guarantees a cause-and-effect conclusion
It ensures each grade level is adequately represented and can reduce variability
It is the same as taking a convenience sample of each grade
Correct answer: It ensures each grade level is adequately represented and can reduce variability
Stratifying by grade ensures each grade level is adequately represented and can reduce the variability of estimates when grades differ in opinion. A single random sample might by chance underrepresent a grade. The method still uses randomization within strata and, being a survey, cannot establish causation.
A medical team tests a drug by giving half the volunteers the real drug and half an identical-looking pill with no active ingredient, with no one knowing which they received. The purpose of the identical-looking inactive pill is best described as what?
A confounding variable introduced on purpose
A placebo serving as a control treatment
A stratifying factor for the volunteers
A method of cluster sampling the volunteers
Correct answer: A placebo serving as a control treatment
The inactive pill is a placebo serving as a control treatment, allowing a fair comparison by accounting for the psychological effect of simply taking a pill. The control group receiving it provides the baseline against which the real drug is judged. It is not a confounding variable, a basis for strata, or a sampling technique.
Two researchers study whether exercise lowers blood pressure. Researcher A randomly assigns volunteers to an exercise program or a no-exercise group. Researcher B simply compares people who already exercise to those who do not. Which statement correctly contrasts what each can conclude?
Both can conclude exercise causes lower blood pressure
Researcher B can claim causation while Researcher A can only show association
Neither can show any relationship between exercise and blood pressure
Researcher A can claim causation while Researcher B can only show association
Correct answer: Researcher A can claim causation while Researcher B can only show association
Researcher A can claim causation while Researcher B can only show association, because random assignment in Researcher A's experiment balances out lurking variables, isolating exercise as the cause. Researcher B's observational comparison leaves confounders, such as diet, uncontrolled. The presence or absence of random assignment is what separates a causal claim from a merely associational one.
A city surveys residents about a new park by mailing forms and using only the responses from people who chose to mail them back. Why is the resulting data likely to be misleading?
Voluntary response tends to overrepresent those with strong opinions
The sample is too large to analyze accurately
Mailing forms is a form of cluster sampling
Random assignment was not used to mail the forms
Correct answer: Voluntary response tends to overrepresent those with strong opinions
The data is likely misleading because voluntary response tends to overrepresent people with strong opinions, who are more motivated to reply. This produces a biased picture of the community's views rather than a representative one. The problem is self-selection bias, not sample size, and random assignment applies to experiments, not to fixing this survey flaw.
A random variable X represents the number of successes in a fixed number of independent trials, each with the same probability of success. Which type of probability distribution does X follow?
A geometric distribution
A uniform distribution
A binomial distribution
A normal distribution
Correct answer: A binomial distribution
X follows a binomial distribution because it counts the number of successes in a fixed number of independent trials, each having two outcomes and the same success probability. A geometric distribution instead counts trials until the first success, with no fixed number of trials. A uniform distribution assigns equal probability to outcomes, and a normal distribution is continuous rather than a count.
Two events A and B are described as mutually exclusive. What does this mean?
Knowing A occurred does not change the probability of B
Events A and B cannot both occur at the same time
Events A and B must both occur together
The probability of A equals the probability of B
Correct answer: Events A and B cannot both occur at the same time
Mutually exclusive means events A and B cannot both occur at the same time, so they share no common outcomes and their joint probability is zero. This is a statement about overlap, not about whether one event changes the other's probability, which describes independence. Mutual exclusivity does not require the two events to have equal probabilities.
What does the expected value of a discrete random variable represent?
The most likely single outcome of the variable
The largest value the variable can take
The long-run average value of the variable over many repetitions
The probability that the variable equals zero
Correct answer: The long-run average value of the variable over many repetitions
The expected value represents the long-run average value of the random variable if the random process were repeated many times. It is computed as the sum of each value times its probability and need not equal any single attainable outcome. It is not the most likely value, the maximum value, or a probability.
A geometric random variable counts which of the following?
The number of trials needed to get the first success
The number of successes in a fixed number of trials
The proportion of successes in a large sample
The total number of possible outcomes in an experiment
Correct answer: The number of trials needed to get the first success
A geometric random variable counts the number of independent trials needed to obtain the first success, where each trial has the same success probability. This contrasts with a binomial random variable, which counts successes within a fixed number of trials. It is a count of trials, not a proportion or a tally of all possible outcomes.
A fair six-sided die is rolled. Let event A be rolling an even number and event B be rolling a 5. Are events A and B mutually exclusive?
No, because 5 is an even number
Yes, because no even number is also a 5
No, because both events involve the same die
Yes, because the events have equal probability
Correct answer: Yes, because no even number is also a 5
Events A and B are mutually exclusive because no outcome belongs to both: 5 is odd, so it can never be one of the even numbers 2, 4, or 6. With no shared outcome, the two events cannot occur on the same roll. Sharing the same die or having any particular probabilities does not determine mutual exclusivity.
A multiple-choice quiz has 10 questions, each with 4 choices, and a student guesses every answer at random. What is the expected number of correct answers?
4
10
2.5
5
Correct answer: 2.5
The expected number correct is 2.5, found for a binomial setting as the number of trials times the success probability: 10 times 41 equals 2.5. Each question has a 41 chance of a correct guess, and expectation adds these across all 10 questions. The values 4, 5, and 10 ignore the correct multiplication of trials by probability.
Events A and B are independent, with P(A) = 0.5 and P(B) = 0.4. What is the probability that both A and B occur?
0.90
0.10
0.45
0.20
Correct answer: 0.20
The probability that both occur is 0.20 because for independent events the joint probability equals the product of the individual probabilities: 0.5 times 0.4 equals 0.20. The multiplication rule applies only because the events are independent. Adding the probabilities (0.90) or other combinations does not give the probability of both events happening.
In a class, 60% of students take Spanish and 30% of students take both Spanish and art. Given that a student takes Spanish, what is the probability that the student also takes art?
0.30
0.18
0.90
0.50
Correct answer: 0.50
The conditional probability is 0.50, found by dividing the probability of both events by the probability of the given event: 0.30 divided by 0.60 equals 0.50. Conditional probability restricts attention to the Spanish-takers and asks what fraction of them also take art. Simply reporting 0.30 ignores that the condition narrows the relevant group.
A basketball player makes 70% of her free throws, and each attempt is independent. What is the probability that she makes her first three free throws and misses the fourth?
0.7 times 0.7 times 0.7 times 0.3
0.7 times 0.3
0.7 times 4
0.3 times 0.3 times 0.3 times 0.7
Correct answer: 0.7 times 0.7 times 0.7 times 0.3
The probability is 0.7 times 0.7 times 0.7 times 0.3 because independent attempts let us multiply the probabilities of the specified outcomes: three makes at 0.7 each and one miss at 0.3. The multiplication rule for independent events applies to this exact sequence. Reversing the make and miss probabilities or adding terms does not match the described outcome.
A medical test is given to a population. Among people who actually have a disease, the test comes back positive 95% of the time. This 95% figure is best described as which kind of probability?
A conditional probability of testing positive given the disease is present
The probability that a randomly chosen person tests positive
A mutually exclusive probability
The expected value of a positive test
Correct answer: A conditional probability of testing positive given the disease is present
The 95% is a conditional probability of testing positive given that the disease is present, because it describes the chance of a positive result within the subgroup that actually has the disease. The condition restricts the population to diseased individuals before measuring the positive rate. It is not the overall positive rate, an expected value, or a statement about mutual exclusivity.
A salesperson has a 20% chance of closing each independent sales call. Let X be the number of calls up to and including the first successful close. What is the probability that the first success happens on exactly the third call?
0.2 times 0.2 times 0.2
3 times 0.2
0.8 times 0.8 times 0.8
0.8 times 0.8 times 0.2
Correct answer: 0.8 times 0.8 times 0.2
The probability is 0.8 times 0.8 times 0.2 because a geometric random variable requires the first two calls to be failures, each with probability 0.8, followed by a success with probability 0.2 on the third call. The independent trials let us multiply these in sequence. Using only success probabilities or multiplying by 3 misrepresents the required pattern of two failures then a success.
Suppose events A and B have P(A) = 0.4, P(B) = 0.5, and P(A and B) = 0.2. A student claims A and B are both mutually exclusive and independent. Which statement is correct?
They are mutually exclusive but not independent
They are independent but not mutually exclusive
They are both mutually exclusive and independent
They are neither mutually exclusive nor independent
Correct answer: They are independent but not mutually exclusive
The events are independent but not mutually exclusive. They are not mutually exclusive because P(A and B) equals 0.2, which is not zero, so the events can occur together. They are independent because the product P(A) times P(B), which is 0.4 times 0.5 equals 0.2, matches the actual joint probability. Mutually exclusive events with nonzero probabilities can never be independent.
A carnival game costs nothing to consider and pays $10 with probability 0.1, pays $2 with probability 0.3, and pays $0 with probability 0.6. What is the expected payout for one play?
$4.00
$1.60
$12.00
$2.00
Correct answer: $1.60
The expected payout is $1.60, computed by summing each payout times its probability: 10 times 0.1 plus 2 times 0.3 plus 0 times 0.6 equals 1 plus 0.6 plus 0, which is 1.60. Expected value weights every possible payout by how likely it is. Averaging the dollar amounts directly or adding them ignores the differing probabilities.
A quality inspector examines 8 independently produced parts, each defective with probability 0.05. Treating the count of defective parts as a random variable, why is a binomial model appropriate here rather than a geometric model?
Because there is a fixed number of trials and we are counting defectives among them
Because we are waiting for the first defective part to appear
Because the probability of a defect changes from part to part
Because the parts are not independent of one another
Correct answer: Because there is a fixed number of trials and we are counting defectives among them
A binomial model fits because there is a fixed number of trials, the 8 parts, and the random variable counts how many are defective among them. A geometric model would instead apply if we were counting parts until the first defective appeared, with no fixed total. The binomial conditions of a set number of independent trials with a constant success probability are all met here.
What does a sampling distribution describe?
The distribution of all individual values in a single sample
The distribution of values in the entire population
The distribution of a statistic computed from all possible samples of a given size
The distribution of the sample size across repeated studies
Correct answer: The distribution of a statistic computed from all possible samples of a given size
A sampling distribution is the distribution of a statistic computed from all possible samples of a given size. Each sample produces one value of the statistic (such as a sample mean or sample proportion), and collecting those values over all samples forms the sampling distribution. It is not the distribution of raw data within one sample, nor the distribution of the population itself, and it is built for a fixed sample size rather than describing how the sample size varies.
According to the central limit theorem, what happens to the shape of the sampling distribution of the sample mean as the sample size increases?
It becomes approximately normal regardless of the population's shape
It takes on the exact shape of the population distribution
It becomes increasingly skewed
It becomes uniform across all possible values
Correct answer: It becomes approximately normal regardless of the population's shape
The central limit theorem states the sampling distribution of the sample mean becomes approximately normal as the sample size increases, no matter the shape of the population distribution. This is why large samples allow normal-based inference even when the underlying population is skewed. It does not mirror the population shape, become more skewed, or turn uniform.
A population has mean 70 and standard deviation 12. For samples of size 36, what is the standard deviation of the sampling distribution of the sample mean?
12
6
0.33
2
Correct answer: 2
The standard deviation of the sampling distribution of the sample mean is 2, found by dividing the population standard deviation by n: 12 divided by 36, which is 12 divided by 6. The value 12 ignores the sample-size adjustment entirely, while the other values come from dividing incorrectly.
What is the mean of the sampling distribution of the sample proportion equal to?
The sample size times the population proportion
The true population proportion
The population proportion divided by the sample size
Always 0.5
Correct answer: The true population proportion
The mean of the sampling distribution of the sample proportion equals the true population proportion, which is what makes the sample proportion an unbiased estimator. The center of the sampling distribution sits at the population value it estimates. It is not scaled by the sample size and is not fixed at 0.5.
A sample proportion is described as an unbiased estimator of the population proportion. What does this mean about its sampling distribution?
Every sample produces the exact population proportion
The sampling distribution has no variability
The sampling distribution is always perfectly normal
The mean of the sampling distribution equals the population proportion
Correct answer: The mean of the sampling distribution equals the population proportion
Saying the sample proportion is unbiased means the mean of its sampling distribution equals the population proportion, so on average it neither overestimates nor underestimates the target value. Individual samples still vary, so not every sample hits the population value, the distribution still has spread, and normality depends on conditions rather than being automatic.
A population is strongly right-skewed. A researcher wants the sampling distribution of the sample mean to be approximately normal. Based on the central limit theorem, what should the researcher do?
Take a sufficiently large sample so the distribution of the mean is approximately normal
Take a very small sample to better match the skewed population
Transform every individual value into a z-score before sampling
Conclude that normal-based inference is impossible for skewed populations
Correct answer: Take a sufficiently large sample so the distribution of the mean is approximately normal
The researcher should take a sufficiently large sample so the distribution of the mean is approximately normal, which is exactly what the central limit theorem guarantees even for a skewed population. A small sample would leave the sampling distribution skewed, standardizing individual values does not change the population's shape, and normal-based inference is achievable precisely because larger samples produce an approximately normal sampling distribution.
A factory's bolt lengths have mean 50 mm and standard deviation 4 mm. As the sample size used to compute sample means increases from 16 to 64, what happens to the standard deviation of the sampling distribution of the sample mean?
It stays the same because the population standard deviation does not change
It is cut in half
It is reduced to one-fourth
It doubles
Correct answer: It is cut in half
The standard deviation of the sampling distribution is cut in half. It equals the population standard deviation divided by n; going from 16 to 64 multiplies the sample size by 4, and 4=2, so the standard deviation is divided by 2. It does change with sample size, and it does not shrink to one-fourth or grow.
In a large population, 40% of voters support a measure. For random samples of 200 voters, the sampling distribution of the sample proportion is centered at what value?
0.40
0.50
0.20
80
Correct answer: 0.40
The sampling distribution of the sample proportion is centered at 0.40, the true population proportion, because the sample proportion is an unbiased estimator. The value 0.50 is a common default that ignores the given proportion, 0.20 misuses the figures, and 80 is the expected count of supporters rather than a proportion.
Two sampling distributions of the sample mean are built from the same population, one using samples of size 25 and the other using samples of size 100. How do these two sampling distributions compare?
The size-100 distribution has more variability than the size-25 distribution
The size-100 distribution has less variability and is more tightly clustered around the population mean
Both distributions have identical variability because they share a population
The size-25 distribution is centered at a higher value than the size-100 distribution
Correct answer: The size-100 distribution has less variability and is more tightly clustered around the population mean
The size-100 distribution has less variability and is more tightly clustered around the population mean, because the standard deviation of the sampling distribution decreases as sample size increases. Both distributions are centered at the same population mean, so neither is higher than the other, and their variability is not identical since larger samples produce smaller spread.
A student claims that because a population of reaction times is heavily skewed, the sampling distribution of the sample mean for samples of size 80 must also be heavily skewed. Why is this reasoning flawed?
The sampling distribution always exactly copies the population's shape
Skewed populations cannot be sampled at all
The central limit theorem makes the sampling distribution of the mean approximately normal for a large sample like 80
The sampling distribution of the mean is always uniform regardless of sample size
Correct answer: The central limit theorem makes the sampling distribution of the mean approximately normal for a large sample like 80
The reasoning is flawed because the central limit theorem makes the sampling distribution of the mean approximately normal for a large sample like 80, even when the population is skewed. The sampling distribution does not copy the population's shape at large sample sizes, skewed populations can certainly be sampled, and the sampling distribution is not uniform.
When setting up a significance test, what is the role of the null hypothesis?
It states a claim of no effect or no difference that the test attempts to find evidence against
It states the effect the researcher hopes to prove is true
It specifies the sample proportion observed in the data
It guarantees that any difference found is due to chance
Correct answer: It states a claim of no effect or no difference that the test attempts to find evidence against
The null hypothesis states a claim of no effect or no difference, and the test gathers evidence against it. It typically expresses that a population proportion equals a specific value, such as p = 0.5. The alternative hypothesis, not the null, states what the researcher suspects is true, and the null is about the population parameter rather than the observed sample proportion.
In a one-proportion significance test, what does the p-value measure?
The probability that the null hypothesis is true
The probability that the alternative hypothesis is false
The proportion of the population that supports the claim
The probability of observing a sample result at least as extreme as the one obtained, assuming the null hypothesis is true
Correct answer: The probability of observing a sample result at least as extreme as the one obtained, assuming the null hypothesis is true
The p-value measures the probability of getting a sample result at least as extreme as the observed one, calculated under the assumption that the null hypothesis is true. A small p-value means such a result would be unlikely if the null were correct, providing evidence against it. The p-value is not the probability that the null is true, nor a population proportion.
A confidence interval for a population proportion is constructed as the sample proportion plus or minus a margin of error. What two components are multiplied together to form that margin of error?
The sample size and the population proportion
The point estimate and the confidence level
A critical value and the standard error of the sample proportion
The p-value and the significance level
Correct answer: A critical value and the standard error of the sample proportion
The margin of error is the product of a critical value (a z* from the standard normal distribution for the chosen confidence level) and the standard error of the sample proportion. Together they set how far the interval extends on each side of the estimate. The margin of error does not multiply the sample size by the proportion, nor does it involve the p-value.
In the context of a significance test, what is a Type I error?
Rejecting the null hypothesis when it is actually true
Failing to reject the null hypothesis when it is actually false
Choosing too small a sample size for the test
Calculating the standard error incorrectly
Correct answer: Rejecting the null hypothesis when it is actually true
A Type I error is rejecting the null hypothesis when it is in fact true, meaning the test concludes there is an effect when there is none. Its probability equals the significance level alpha. Failing to reject a false null is instead a Type II error, and a Type I error is about an incorrect conclusion rather than a sample-size or calculation mistake.
What does the power of a significance test represent?
The probability of making a Type I error
The probability that the null hypothesis is true
The width of the confidence interval
The probability of correctly rejecting a false null hypothesis
Correct answer: The probability of correctly rejecting a false null hypothesis
The power of a test is the probability of correctly rejecting the null hypothesis when it is actually false, that is, the chance of detecting a real effect. Higher power means the test is better at finding true differences. It is not the probability of a Type I error, which is alpha, and it is unrelated to the width of a confidence interval.
A researcher tests whether more than half of voters support a measure. Which set of hypotheses correctly states this one-sided test?
H0: p = 0.5 versus Ha: p > 0.5
H0: p > 0.5 versus Ha: p = 0.5
H0: p-hat = 0.5 versus Ha: p-hat > 0.5
H0: p = 0.5 versus Ha: p < 0.5
Correct answer: H0: p = 0.5 versus Ha: p > 0.5
The correct hypotheses are H0: p = 0.5 versus Ha: p > 0.5, because the claim of more than half support is a one-sided alternative in the greater-than direction. Hypotheses are written about the population proportion p, not the sample proportion p-hat. The null always carries the equality, and the less-than alternative would test the opposite direction.
In a sample of 400 adults, 240 said they exercise weekly. What is the sample proportion used as the point estimate when building a confidence interval for the population proportion?
0.40
0.60
240
0.024
Correct answer: 0.60
The point estimate is the sample proportion 0.60, found by dividing the number of successes by the sample size: 240 divided by 400 equals 0.60. This value serves as the center of the confidence interval for the true population proportion. The count 240 is not a proportion, and 0.40 reverses the success and failure groups.
A 95% confidence interval for the proportion of residents who recycle is found to be 0.58 to 0.66. Which statement is the correct interpretation of this interval?
There is a 95% probability that the true proportion falls between 0.58 and 0.66
95% of residents recycle at a rate between 0.58 and 0.66
95% of all possible samples will produce a sample proportion between 0.58 and 0.66
We are 95% confident that the true proportion of residents who recycle is between 0.58 and 0.66
Correct answer: We are 95% confident that the true proportion of residents who recycle is between 0.58 and 0.66
The correct interpretation is that we are 95% confident the true population proportion of residents who recycle lies between 0.58 and 0.66. The 95% refers to the long-run capture rate of the method, not a probability that this particular fixed interval contains the parameter. The interval describes the population proportion, not the behavior of individual residents or sample proportions.
A test of H0: p = 0.30 against Ha: p > 0.30 yields a p-value of 0.02 at a significance level of 0.05. What is the correct conclusion?
Reject the null hypothesis; there is convincing evidence the proportion exceeds 0.30
Fail to reject the null hypothesis; there is no evidence the proportion exceeds 0.30
Accept the null hypothesis as proven true
The result is inconclusive because the p-value is positive
Correct answer: Reject the null hypothesis; there is convincing evidence the proportion exceeds 0.30
Since the p-value of 0.02 is less than the significance level of 0.05, we reject the null hypothesis and conclude there is convincing evidence the proportion exceeds 0.30. A p-value below alpha signals a result too unlikely to attribute to chance under the null. We never accept the null as proven, and a small p-value is decisive rather than inconclusive.
A study reports a p-value of 0.46 when testing whether a coin is unfair. What is the best interpretation of this p-value in context?
There is a 46% chance the coin is fair
If the coin were fair, a result at least as extreme as the observed one would occur about 46% of the time, so there is no evidence the coin is unfair
The coin is unfair 46% of the time it is flipped
The probability of a Type I error in this test is 0.46
Correct answer: If the coin were fair, a result at least as extreme as the observed one would occur about 46% of the time, so there is no evidence the coin is unfair
The correct interpretation is that, assuming the coin is fair, a result at least as extreme as the observed one would happen about 46% of the time, which is quite common and gives no evidence the coin is unfair. A p-value is a conditional probability about the data, not the probability that the null hypothesis is true and not a fixed Type I error rate.
A polling firm wants to estimate the proportion of likely voters favoring a candidate but finds its margin of error is too large. Which change would reduce the margin of error?
Decreasing the sample size
Increasing the sample size
Raising the confidence level from 95% to 99%
Switching from a proportion to a count
Correct answer: Increasing the sample size
Increasing the sample size reduces the margin of error because the standard error of the sample proportion shrinks as the sample size grows, narrowing the interval. Decreasing the sample size would widen it, and raising the confidence level enlarges the critical value, which also widens the margin. Reporting a count instead of a proportion does not address the margin of error.
A medical researcher compares the proportion of patients improving under a new drug to the proportion improving under a placebo, using data from two independent groups. Which inference procedure is most appropriate?
A one-proportion z-interval
A test of a single population mean
A two-proportion z-test
A goodness-of-fit comparison of one group to a fixed value
Correct answer: A two-proportion z-test
A two-proportion z-test is most appropriate because the researcher is comparing categorical success proportions from two independent groups to see whether they differ. A one-proportion procedure handles only a single group, and a mean-based test applies to quantitative data rather than proportions. The setup involves two observed proportions, which is exactly what the two-proportion z-test addresses.
A drug company sets a very small significance level so it rarely concludes a useless drug works, but as a consequence its test often fails to detect drugs that truly do work. This tradeoff is best described by which relationship?
Lowering the chance of a Type I error tends to increase the chance of a Type II error and lower power
Lowering the chance of a Type I error also lowers the chance of a Type II error
Type I and Type II errors are unrelated to the significance level
Reducing the significance level increases the power of the test
Correct answer: Lowering the chance of a Type I error tends to increase the chance of a Type II error and lower power
Lowering the significance level reduces the chance of a Type I error but tends to increase the chance of a Type II error, which lowers the power to detect a real effect. Making it harder to reject the null protects against false alarms at the cost of missing true effects. The two error types trade off rather than moving together, and shrinking alpha does not raise power.
Before performing a one-proportion z-test, a student checks that the expected number of successes and the expected number of failures are each at least 10. Why is verifying this condition important?
It guarantees the sample is a simple random sample
It ensures the sampling distribution of the sample proportion is approximately normal so the z-procedure is valid
It proves the null hypothesis is true
It eliminates the possibility of a Type I error
Correct answer: It ensures the sampling distribution of the sample proportion is approximately normal so the z-procedure is valid
Checking that the expected successes and failures are each at least 10 ensures the sampling distribution of the sample proportion is approximately normal, which is what makes the z-based test valid. This large-counts condition justifies using the normal model for the proportion. It does not address whether the data were randomly sampled, prove the null, or remove the chance of a Type I error.
A survey records each respondent's favorite music genre. How is this variable best classified?
Quantitative and discrete
Categorical
Quantitative and continuous
A parameter
Correct answer: Categorical
Favorite music genre is a categorical variable because it places each respondent into a group or label rather than recording a numerical measurement. Quantitative variables, whether discrete or continuous, take number values that can be added or averaged. A parameter is a numerical summary of a population, not a type of variable.
Which graphical display is appropriate for showing the distribution of a single categorical variable?
A histogram
A scatterplot
A boxplot
A bar chart
Correct answer: A bar chart
A bar chart is appropriate for a single categorical variable because it shows the count or proportion in each category with separated bars. A histogram and a boxplot display quantitative data, and a scatterplot shows the relationship between two quantitative variables. The gaps between bars distinguish a bar chart from a histogram.
What key visual feature distinguishes a histogram from a bar chart?
A histogram uses pie slices instead of bars
A histogram's bars touch because the horizontal axis is a continuous numerical scale
A histogram always has exactly five bars
A histogram can only display categorical data
Correct answer: A histogram's bars touch because the horizontal axis is a continuous numerical scale
A histogram's bars touch because its horizontal axis represents a continuous numerical scale divided into intervals, so adjacent bins share a boundary. A bar chart leaves gaps between bars because its categories are distinct labels. Histograms display quantitative data, the opposite of the categorical-only claim.
A relative frequency table reports the proportion of students in each grade level. The relative frequencies in such a table should sum to what value?
The number of categories
1
The total number of students
The largest single frequency
Correct answer: 1
The relative frequencies must sum to 1 because each relative frequency is a part of the whole expressed as a proportion, and the parts together make up the entire data set. If expressed as percentages they would total 100%. The raw counts, not the relative frequencies, sum to the total number of students.
In a stemplot (stem-and-leaf plot) of the value 47, where the stems are tens, what is the leaf?
47
4
7
0
Correct answer: 7
The leaf is 7 because in a stemplot the stem holds the leading digit or digits and the leaf holds the final digit. For 47 with stems representing tens, the stem is 4 and the leaf is 7. The full number 47 is not a single leaf, and 0 is not a digit of this value.
Which display preserves the individual data values while still showing the overall shape of a small quantitative data set?
A pie chart
A stemplot
A bar chart
A relative frequency table of categories
Correct answer: A stemplot
A stemplot preserves the individual data values while revealing the shape because every observation appears as a leaf attached to its stem. Histograms show shape but group values into bins, hiding the exact numbers. Pie charts and category bar charts are for categorical data and do not display quantitative values.
When describing the distribution of a quantitative variable, which four features should be addressed?
Shape, center, variability, and unusual features such as outliers
Color, size, label, and title
Mean, median, mode, and range only
Slope, intercept, correlation, and residuals
Correct answer: Shape, center, variability, and unusual features such as outliers
A complete description of a quantitative distribution addresses shape, center, variability (spread), and unusual features such as outliers or gaps. Reporting only the mean, median, mode, and range omits shape and unusual features. Slope, intercept, and correlation describe relationships between two variables, not a single distribution.
A dotplot of quiz scores shows two clearly separated clusters of dots with a gap between them. What does this most directly indicate about the distribution's shape?
It is uniform
It is roughly symmetric and bell-shaped
It appears bimodal
It has no center
Correct answer: It appears bimodal
The distribution appears bimodal because two distinct clusters of dots correspond to two peaks in the data. A uniform distribution would show roughly equal heights across the range, and a bell-shaped distribution would have a single central peak. Every distribution has a center, so that option is incorrect.
The 30th percentile of a distribution of incomes is $42,000. What does this value mean?
About 30% of incomes are at or below $42,000
Exactly 30 people earn $42,000
About 30% of incomes are above $42,000
The average income is $42,000
Correct answer: About 30% of incomes are at or below $42,000
It means about 30% of the incomes are at or below $42,000, because a percentile gives the percentage of observations falling at or below a given value. It is not a count of people, and it is the lower 30%, not the upper portion. A percentile describes position, not the average.
A value at the 90th percentile of a test-score distribution can best be described as which of the following?
A value that 90% of scores exceed
A value at or below which about 90% of scores fall
A value equal to 90 points
A value exactly nine standard deviations above the mean
Correct answer: A value at or below which about 90% of scores fall
The 90th percentile is the value at or below which about 90% of the scores fall, marking a relatively high position in the distribution. It is not a value most scores exceed, nor is it necessarily 90 points. Percentiles measure relative standing, not standard deviations.
A z-score is computed for a data value and turns out to be negative. What does this indicate?
The value is below the mean
The value is an outlier
The value is above the mean
A calculation error was made because z-scores cannot be negative
Correct answer: The value is below the mean
A negative z-score indicates the value lies below the mean, since the z-score is the value minus the mean divided by the standard deviation. Values below the mean produce a negative numerator. Z-scores can certainly be negative, and a negative z-score does not by itself mark an outlier.
Two students take different exams. One has a z-score of 1.2 on her exam and the other has a z-score of 0.8 on his. Why does comparing z-scores allow a fair comparison?
Z-scores convert each value to a common scale of standard deviations from its own mean
Z-scores convert all scores to a 100-point scale
Z-scores ignore the standard deviation
Z-scores can only be compared if the exams are identical
Correct answer: Z-scores convert each value to a common scale of standard deviations from its own mean
Comparing z-scores is fair because each z-score expresses a value as the number of standard deviations from its own distribution's mean, standardizing different scales. The student with a z-score of 1.2 stands higher relative to her group. Z-scores depend on the standard deviation and do not require identical exams or a 100-point scale.
A density curve is used to model a distribution. What must be true of the total area under any density curve?
It equals the number of data values
It equals the mean
It can be any positive number
It equals 1
Correct answer: It equals 1
The total area under any density curve equals 1 because the curve represents the entire distribution as proportions, and all proportions together make up the whole. Areas under regions of the curve correspond to proportions of data. The area is fixed at 1 regardless of the number of data values or the mean.
On a density curve that is skewed to the right, how do the locations of the mean and median compare?
The median lies to the right of the mean
The mean lies to the right of the median
They are always at the exact same point
The mean lies at the peak of the curve
Correct answer: The mean lies to the right of the median
On a right-skewed density curve the mean lies to the right of the median because the long right tail pulls the mean toward the larger values while the median stays near the bulk of the area. The mean and median coincide only on a symmetric curve. The mean is not generally located at the peak.
In a perfectly symmetric, single-peaked density curve, where are the mean and median located?
The mean is left of the median
Both lie at the center of the curve
The median is at the peak but the mean is in the tail
Both lie at the far right of the curve
Correct answer: Both lie at the center of the curve
In a symmetric, single-peaked density curve both the mean and the median lie at the center, which is also the peak, because the curve balances evenly on either side. Skew is what separates the mean from the median. Neither measure is pulled into a tail when the curve is symmetric.
Approximately what percentage of observations in a normal distribution fall within one standard deviation of the mean according to the empirical rule?
About 50%
About 99.7%
About 68%
About 95%
Correct answer: About 68%
About 68% of observations fall within one standard deviation of the mean under the empirical rule. About 95% lie within two standard deviations and about 99.7% within three. The 50% figure corresponds to the median splitting the distribution in half, not a one-standard-deviation interval.
Battery lifetimes are approximately normal with a mean of 40 hours and a standard deviation of 5 hours. Within which interval do about 99.7% of the lifetimes fall?
35 to 45 hours
25 to 55 hours
30 to 50 hours
20 to 60 hours
Correct answer: 25 to 55 hours
About 99.7% of lifetimes fall between 25 and 55 hours, found by taking the mean plus or minus three standard deviations: 40 minus 3 times 5 is 25, and 40 plus 3 times 5 is 55. The empirical rule assigns 99.7% to three standard deviations. The interval 30 to 50 covers two standard deviations (about 95%).
Which measure of variability uses every value in the data set in its calculation?
The range
The interquartile range
The standard deviation
The median
Correct answer: The standard deviation
The standard deviation uses every value in the data set because it is based on each observation's squared deviation from the mean. The range uses only the maximum and minimum, the interquartile range uses only the quartiles, and the median is a measure of center, not variability.
A data set of house ages has a minimum of 2 and a maximum of 88. What is the range?
88
45
86
90
Correct answer: 86
The range is 86, found by subtracting the minimum from the maximum: 88 minus 2 equals 86. The range measures the total spread between the extreme values. Using the maximum alone (88) or adding the values (90) does not give the range.
Why is the standard deviation, rather than the variance, often preferred when describing the spread of a single quantitative variable?
The variance is always negative
The standard deviation ignores outliers entirely
The variance cannot be calculated for real data
The standard deviation is in the same units as the original data
Correct answer: The standard deviation is in the same units as the original data
The standard deviation is often preferred because it is expressed in the same units as the original data, making it easier to interpret, while the variance is in squared units. The variance is never negative and can certainly be computed. Neither measure ignores outliers; both are affected by extreme values.
If every value in a data set is multiplied by 3, what happens to the median?
It stays the same
It increases by 3
It is divided by 3
It is multiplied by 3
Correct answer: It is multiplied by 3
The median is multiplied by 3 because multiplying every value by a constant scales all measures of center and spread by that same constant. The middle value, like all others, becomes three times larger. Adding a constant would instead shift the median, but here the values are multiplied.
Adding a constant of 5 to every value in a data set has which effect on the measures of the distribution?
Both the mean and the standard deviation increase by 5
The mean increases by 5 but the standard deviation is unchanged
The mean is unchanged but the standard deviation increases by 5
Neither the mean nor the standard deviation changes
Correct answer: The mean increases by 5 but the standard deviation is unchanged
Adding 5 to every value increases the mean by 5 while leaving the standard deviation unchanged, because a uniform shift moves the center but does not alter the distances between values. Measures of center respond to added constants, but measures of spread such as standard deviation do not.
A frequency table shows 40 freshmen, 30 sophomores, 20 juniors, and 10 seniors. What is the relative frequency of juniors?
0.20
0.10
0.30
20
Correct answer: 0.20
The relative frequency of juniors is 0.20, found by dividing the 20 juniors by the total of 100 students. Relative frequency expresses a category's count as a proportion of the whole. The value 20 is the raw count, not a proportion, and 0.10 corresponds to seniors.
When is the median generally a more appropriate measure of center than the mean?
When the distribution is strongly skewed or has outliers
When the distribution is perfectly symmetric
When every value is identical
When the data are categorical
Correct answer: When the distribution is strongly skewed or has outliers
The median is generally more appropriate when the distribution is strongly skewed or contains outliers, because it resists the pull of extreme values that distort the mean. In a symmetric distribution the mean and median agree, so either works. Measures of center do not apply to categorical data.
A cumulative relative frequency graph (ogive) reaches a height of 0.75 at the value x = 80. What does this tell you?
Exactly 75 values equal 80
About 75% of the data values are above 80
The value 80 is the mean of the data
About 75% of the data values are at or below 80
Correct answer: About 75% of the data values are at or below 80
It tells you that about 75% of the data values are at or below 80, because a cumulative relative frequency graph plots the running proportion of data up to each value. The height of 0.75 marks 80 as roughly the 75th percentile, or third quartile. It does not represent a count or the mean.
Which of the following is a quantitative variable?
Eye color of survey respondents
Number of text messages sent per day
Type of pet owned
Preferred brand of cereal
Correct answer: Number of text messages sent per day
The number of text messages sent per day is quantitative because it is a numerical count that can be averaged and ordered meaningfully. Eye color, type of pet, and preferred cereal brand are all categorical, since they sort respondents into groups rather than measuring a quantity.
A histogram of reaction times is described as right-skewed. Which statement about its tail is correct?
The tail extends toward the larger values on the right
The tail extends toward the smaller values on the left
There is no tail in a skewed distribution
The tail is always exactly symmetric
Correct answer: The tail extends toward the larger values on the right
In a right-skewed histogram the tail extends toward the larger values on the right, which is where the distribution is named for. Most observations cluster on the lower end with a few high values stretching the tail. A left-skewed distribution would have its tail toward the smaller values.
A pie chart and a bar chart both summarize the same categorical variable. What is one advantage a bar chart has over a pie chart?
A bar chart can display quantitative intervals
A bar chart makes it easier to compare the sizes of categories directly
A bar chart shows the exact individual data values
A bar chart requires the categories to sum to 100%
Correct answer: A bar chart makes it easier to compare the sizes of categories directly
A bar chart makes it easier to compare category sizes directly because the heights of bars are simpler to judge than the angles of pie slices. Both displays are for categorical data and do not show individual quantitative values. Neither requires the categories to total 100%, though relative frequencies do.
A data set of ages includes the values 5, 6, 6, 7, 8, and 95. Which measure of center would be most misleading for describing a typical age, and why?
The median, because it ignores the value 95
The mean, because the value 95 inflates it far above the typical age
The mode, because it equals 6
The minimum, because it equals 5
Correct answer: The mean, because the value 95 inflates it far above the typical age
The mean is most misleading here because the extreme value 95 pulls it far above the cluster of younger ages, so it no longer represents a typical age. The median resists that single outlier and better reflects the center. The minimum and mode are not measures of typical center in this context.
A standardized assessment is normally distributed. A score with a z-score of 0 corresponds to which value?
The minimum score
The mean score
The maximum score
A score one standard deviation above average
Correct answer: The mean score
A z-score of 0 corresponds to the mean score, because the z-score formula subtracts the mean and a result of zero occurs only when the value equals the mean. It marks the center of the distribution, not an extreme. A z-score of 1, not 0, would be one standard deviation above the mean.
A distribution of values is described as roughly uniform. Which dotplot or histogram pattern matches this description?
A single tall peak in the center that tapers off on both sides
One long tail stretching to the right
Bars of approximately equal height across the range of values
Two separate clusters with a gap between them
Correct answer: Bars of approximately equal height across the range of values
A roughly uniform distribution shows bars of approximately equal height across the range of values, meaning each interval contains about the same number of observations. A single central peak describes a bell shape, one long tail describes a skewed shape, and two clusters describe a bimodal shape.
A manager arranges all 800 employees in a list ordered by hire date, randomly picks one of the first 10 names, and then selects every 10th employee after that. Which sampling method is this?
Systematic random sampling
Stratified random sampling
Cluster sampling
Simple random sampling
Correct answer: Systematic random sampling
This is systematic random sampling because a random starting point is chosen and then every kth member of an ordered list is selected. Choosing every 10th employee after a random start defines the systematic pattern. It is not an SRS, since not every group of employees is equally likely, and it does not form strata or sample whole clusters.
A researcher wants to estimate the average commute time of city workers but only has a list of downtown office buildings, not a list of all workers. What is the main reason this list could cause undercoverage?
The list will make the sample too large to handle
Workers outside downtown offices have no chance of being selected
Random assignment cannot be applied to the list
It forces the study to become an experiment
Correct answer: Workers outside downtown offices have no chance of being selected
The answer is that workers outside downtown offices have no chance of being selected, which is the definition of undercoverage. A sampling frame that omits part of the population leaves those members unreachable, biasing the estimate. The flaw is the incomplete frame, not sample size, randomization, or study type.
In a phone survey, interviewers can only reach people who answer their landline during business hours. People who work during those hours and never pick up are systematically excluded. This is an example of what type of error?
Sampling variability
Confounding
Nonresponse bias
Voluntary response bias
Correct answer: Nonresponse bias
This describes nonresponse bias, since selected individuals who cannot be reached or do not answer differ systematically from those who do. The working people excluded may have different commute or schedule patterns, skewing results. It is not random sampling variability, and no one is volunteering or being assigned a treatment here.
A pollster asks, 'Don't you agree that the wasteful new tax should be repealed?' Why is this question problematic?
It is too short to gather useful information
It requires a random sample to be valid
The wording is leading and can bias responses toward one answer
It introduces a confounding variable into the survey
Correct answer: The wording is leading and can bias responses toward one answer
The answer is that the wording is leading and can bias responses, because describing the tax as 'wasteful' pushes respondents toward agreeing it should be repealed. Loaded language is a recognized source of response bias in surveys. The issue is question phrasing, not length, sampling method, or confounding, which applies to experiments.
An agricultural scientist divides a field into four plots, applies the same fertilizer to all four, and measures yield. A colleague notes the study cannot detect whether the fertilizer helps. What essential element of an experiment is missing?
A response variable
Random sampling of plots
A larger field
A comparison group receiving a different or no treatment
Correct answer: A comparison group receiving a different or no treatment
The missing element is a comparison group receiving a different or no treatment, since without one there is no baseline to judge the fertilizer's effect. Comparison is a core principle of experimental design. A response variable (yield) is present, and random sampling and field size do not address the lack of comparison.
In an experiment, subjects are assigned to treatments, and neither the subjects nor the people measuring outcomes know who received which treatment. This design feature is called what?
Double-blind
Stratification
Replication
Blocking
Correct answer: Double-blind
The answer is double-blind, because both the subjects and those assessing the response are unaware of the treatment assignments. This guards against bias from expectations on either side. It differs from stratification and blocking, which group subjects, and from replication, which refers to using many experimental units.
A drug study gives one group the new medication and another group a pill with no active ingredient. Several patients in the no-medication group report feeling better anyway. This improvement among those receiving the inactive pill is best described as what?
A confounding variable
Sampling bias
The placebo effect
Replication
Correct answer: The placebo effect
This is the placebo effect, the phenomenon where subjects respond favorably simply because they believe they are being treated. It is exactly why studies include a placebo control. It is not confounding, not a sampling flaw, and not replication, which refers to applying treatments to many units.
An exercise study expects that men and women may respond differently to a training plan. Researchers separate subjects into a male block and a female block, then randomly assign treatments within each block. What is the purpose of this blocking?
To allow the results to generalize to the whole country
To turn the study into an observational study
To guarantee no placebo effect occurs
To reduce variability by accounting for a known source of differences
Correct answer: To reduce variability by accounting for a known source of differences
The purpose of blocking is to reduce variability by accounting for a known source of differences, here sex. Comparing treatments within similar blocks isolates the treatment effect from sex-related variation. Blocking does not control generalization, change the study to observational, or address the placebo effect.
A researcher conducts a matched-pairs experiment in which each subject receives both treatments in a random order at different times. What is the main advantage of this matched-pairs design?
Each subject serves as their own control, reducing person-to-person variability
It eliminates the need for any randomization
It makes the study an observational study
It guarantees the sample represents the population
Correct answer: Each subject serves as their own control, reducing person-to-person variability
The advantage is that each subject serves as their own control, reducing person-to-person variability so differences are attributed to the treatments. Randomizing the order of treatments is still required. The design remains an experiment and does not govern how well the sample represents the population.
A study of 60 plants applies a growth hormone to 30 plants and none to the other 30. Why is applying the treatment to 30 plants rather than just 1 important?
It makes the experiment double-blind
It removes the need for a control group
Replication across many units allows real effects to be distinguished from chance variation
It converts the experiment into a census
Correct answer: Replication across many units allows real effects to be distinguished from chance variation
The answer is that replication across many units lets researchers separate a genuine treatment effect from random plant-to-plant variation. A single plant could differ for unrelated reasons. Using many units does not create blinding, eliminate the control group, or make the study a census.
A polling firm correctly takes a simple random sample of 1,000 adults but discovers afterward that wealthy respondents were far more willing to complete the long survey than others. The conclusions could still be biased mainly because of what?
The sample size was too small
Random assignment was not used
Nonresponse that differs by income group
The use of stratified sampling
Correct answer: Nonresponse that differs by income group
The bias stems from nonresponse that differs by income group, since wealthier respondents completed the survey at higher rates, skewing results. A good random selection cannot fix who chooses not to respond. The problem is not sample size, random assignment, or stratification, which was not used.
Which of the following best describes the difference between the population and a sample in a statistical study?
The population is always smaller than the sample
The sample includes everyone, while the population is only those who respond
They are identical in every well-designed study
The population is the entire group of interest, while the sample is the subset actually examined
Correct answer: The population is the entire group of interest, while the sample is the subset actually examined
The answer is that the population is the entire group of interest while the sample is the subset actually examined. Researchers study the sample to draw conclusions about the larger population. The population is not smaller than the sample, the sample is not everyone, and the two are not identical.
A survey on personal income asks respondents to state their salary to a live interviewer face to face. Many people overstate their income. This tendency to give answers seen as more favorable is best classified as what?
Response bias
Undercoverage
Sampling variability
Confounding
Correct answer: Response bias
This is response bias, where respondents give inaccurate answers, here inflating income to appear more favorable. The face-to-face setting encourages socially desirable responses. It is not undercoverage, which concerns excluded groups, nor sampling variability or confounding, which belong to experiments.
A scientist wants to know if a tutoring program improves test scores and randomly assigns students to receive tutoring or not. Students who knew they were being tutored tried harder simply because of the attention. To reduce this, the tutoring and control activities should be made as similar as possible. The unwanted effect being reduced is best described as what?
Undercoverage of the student population
A lurking effect from subjects' awareness of treatment, similar to a placebo effect
Sampling bias in the selection of students
Lack of replication in the experiment
Correct answer: A lurking effect from subjects' awareness of treatment, similar to a placebo effect
The answer is a lurking effect from subjects' awareness, similar to a placebo effect, where extra effort comes from knowing one is in the special group rather than from tutoring itself. Making activities comparable controls for it. This is not undercoverage, sampling bias, or a replication issue.
Researchers randomly select 50 schools from a district and then survey all teachers within those selected schools. This combination is best described as what?
A simple random sample of teachers
Stratified sampling, with schools as strata
A systematic sample of teachers
Cluster sampling, with schools as clusters
Correct answer: Cluster sampling, with schools as clusters
The answer is cluster sampling with schools as clusters, because whole schools are randomly chosen and every teacher within them is surveyed. Sampling entire selected groups in full is the hallmark of cluster sampling. It is not an SRS or systematic sample of individual teachers, nor stratified sampling, which samples within every group.
An experiment finds a real, statistically convincing difference between two treatments, but all subjects were unpaid volunteers who responded to an online ad. What limitation does this place on the conclusion?
No cause-and-effect conclusion can be drawn at all
Random assignment was clearly not used
The results may not generalize beyond people like the volunteers
The study must be an observational study
Correct answer: The results may not generalize beyond people like the volunteers
The limitation is that results may not generalize beyond people like the volunteers, since they were not a random sample of any broader population. Random assignment still supports a causal claim for these subjects, but generalization requires random selection. The study is a valid experiment, just limited in scope of inference.
A study randomly selects participants from the entire population and randomly assigns each to one of two diets. Which two conclusions are most appropriately supported by this design?
Causation, and generalization to the population
Only association, and no generalization
Generalization only, but no causal claim
Neither causation nor generalization
Correct answer: Causation, and generalization to the population
The answer is causation and generalization to the population. Random assignment supports a cause-and-effect conclusion, while random selection from the population supports generalizing the result. Having both random selection and random assignment is what permits both inferences at once.
A factory inspector wants every item in a shipment to have a known, equal chance of selection and uses a random number generator to pick item ID numbers. If two items happen to share the exact same ID due to a labeling error, what problem does this create for the sampling process?
It makes the study an experiment
It introduces the placebo effect
It guarantees nonresponse bias
It creates undercoverage because identical IDs cannot both be distinguished and selected fairly
Correct answer: It creates undercoverage because identical IDs cannot both be distinguished and selected fairly
The answer is that it creates undercoverage, since two items sharing one ID cannot be told apart, so at least one lacks a fair, distinct chance of selection. A proper frame requires every unit to be uniquely identifiable. This is not an experiment, a placebo issue, or nonresponse.
A health magazine reports that readers who eat breakfast tend to weigh less, based on a survey of its subscribers, and concludes that eating breakfast causes weight loss. What is the most accurate criticism?
The survey used too large a sample
Random assignment was used incorrectly
This is observational data, so a lurking variable could explain the link
Blocking should have been applied to the subscribers
Correct answer: This is observational data, so a lurking variable could explain the link
The accurate criticism is that this is observational data, so a lurking variable such as overall health habits could explain the association. Without random assignment of who eats breakfast, causation cannot be claimed. The issue is not sample size, and random assignment and blocking were never part of a survey.
In designing a survey, why is it generally better to use a chance-based method to select respondents rather than letting the researcher choose who seems representative?
Chance selection always produces a larger sample
Chance selection guarantees zero sampling error
Chance selection allows a causal conclusion
Chance selection avoids the researcher's conscious or unconscious bias in choosing
Correct answer: Chance selection avoids the researcher's conscious or unconscious bias in choosing
The answer is that chance selection avoids the researcher's conscious or unconscious bias in choosing who participates. Letting a person pick 'representative' individuals invites systematic favoring of certain types. Chance methods do not guarantee larger samples, eliminate sampling error, or, in a survey, enable causal claims.
A treatment group and a control group in a well-run experiment differ in their average response. Before concluding the treatment caused the difference, what must researchers also consider?
Whether the difference is larger than what random assignment alone could plausibly produce
Whether the sample was a census of the population
Whether voluntary response was used
Whether the population was divided into strata
Correct answer: Whether the difference is larger than what random assignment alone could plausibly produce
The answer is whether the difference is larger than random assignment alone could plausibly produce, since chance creates some difference even with no real effect. Only a difference too large to be explained by chance signals a treatment effect. Census, voluntary response, and strata are unrelated to this judgment.
A company emails a feedback link to all customers and analyzes only the responses it receives. Compared with this approach, why would randomly selecting a subset of customers and actively contacting them give more trustworthy results?
It would produce a census of all customers
It would turn the study into an experiment
It avoids voluntary response bias by not relying on who chooses to reply
It removes the need for a sampling frame
Correct answer: It avoids voluntary response bias by not relying on who chooses to reply
The answer is that random selection with active follow-up avoids voluntary response bias, which arises when only self-motivated customers reply. Reaching a chosen subset reduces the dominance of strong opinions. It does not create a census, become an experiment, or remove the need for a sampling frame.
An experimenter gives every subject the new energy drink and then asks if they feel more alert; most say yes. Why can the experimenter not conclude the drink increases alertness?
The sample was selected without replacement
Stratified sampling was not used
There is no comparison group, so the placebo effect and other causes cannot be ruled out
The response variable was measured incorrectly
Correct answer: There is no comparison group, so the placebo effect and other causes cannot be ruled out
The answer is that with no comparison group, the placebo effect and other explanations cannot be ruled out, so the reported alertness may not come from the drink. A control group is needed to isolate the effect. The problem is the missing comparison, not replacement, stratification, or measurement of the response.
A researcher selects a sample by writing each member of the population on an identical slip, mixing the slips thoroughly in a bowl, and drawing slips without looking. Which sampling method does this physical procedure carry out?
Simple random sampling
Stratified random sampling
Cluster sampling
Convenience sampling
Correct answer: Simple random sampling
The answer is simple random sampling, because identical slips drawn blindly give every possible group of members an equal chance of selection. The bowl-and-slips method is a classic way to physically implement an SRS. It does not form strata, sample whole clusters, or pick whoever is easiest, which would be convenience sampling.
A spinner is divided into regions so that the probability of landing on red is 0.25, on blue is 0.35, and on green is some unknown value, with no other outcomes possible. What must the probability of green be?
0.60
0.30
0.40
0.50
Correct answer: 0.40
The probability of green must be 0.40 because the probabilities of all possible outcomes in a sample space must sum to 1: 1 minus 0.25 minus 0.35 equals 0.40. Since red, blue, and green are the only outcomes, their probabilities are exhausted by this total. Values such as 0.30 or 0.60 would make the probabilities fail to add to 1.
For any event A in a sample space, what is the relationship between the probability of A and the probability of its complement, not A?
They are always equal to each other
They sum to 1
They multiply to 1
Their difference is always 0.5
Correct answer: They sum to 1
The probability of A and the probability of its complement sum to 1 because every outcome either is in A or is not in A, covering the entire sample space exactly once. This is the complement rule, written as P(not A) equals 1 minus P(A). The two probabilities are not generally equal, and they do not multiply to 1.
Events A and B are not mutually exclusive. According to the general addition rule, how is the probability that A or B occurs calculated?
P(A) plus P(B)
P(A) plus P(B) minus P(A and B)
P(A) times P(B)
P(A) minus P(B)
Correct answer: P(A) plus P(B) minus P(A and B)
The probability of A or B equals P(A) plus P(B) minus P(A and B) because simply adding the two probabilities counts the overlap region twice, so the joint probability must be subtracted once. This general addition rule applies whether or not the events overlap. Just adding P(A) and P(B) is correct only when the events are mutually exclusive.
A bag contains 4 red and 6 blue marbles. A marble is drawn, its color recorded, and it is replaced before a second draw. What is the probability both marbles drawn are red?
0.40
0.12
0.08
0.16
Correct answer: 0.16
The probability both are red is 0.16 because drawing with replacement keeps the trials independent, so multiply the probabilities: 0.4 times 0.4 equals 0.16. Replacement restores the bag to 4 red out of 10 for the second draw. Failing to multiply, or treating the draws as dependent, would give a different value.
A bag contains 4 red and 6 blue marbles. Two marbles are drawn without replacement. What is the probability that both are red?
0.1333
0.16
0.40
0.20
Correct answer: 0.1333
The probability both are red is about 0.1333 because without replacement the second draw depends on the first: 104 times 93 equals 9012, which is approximately 0.1333. After removing one red marble, only 3 red remain among 9 marbles. Using 0.16 would wrongly treat the draws as independent with replacement.
The probability distribution of a discrete random variable lists each possible value with its probability. Which condition must this list of probabilities satisfy?
Each probability must be greater than 0.5
The probabilities must all be equal
Each probability is between 0 and 1, and they sum to 1
The probabilities must sum to the number of outcomes
Correct answer: Each probability is between 0 and 1, and they sum to 1
A valid probability distribution requires each probability to be between 0 and 1 and the probabilities to sum to 1, since these are the basic rules every probability must obey. The values need not be equal, and there is no requirement that any single probability exceed 0.5. Summing to the number of outcomes would violate the rule that total probability equals 1.
A discrete random variable X takes the value 0 with probability 0.5, the value 1 with probability 0.3, and the value 2 with probability 0.2. What is the expected value of X?
1.0
1.5
0.5
0.7
Correct answer: 0.7
The expected value is 0.7, computed as the sum of each value times its probability: 0 times 0.5 plus 1 times 0.3 plus 2 times 0.2 equals 0 plus 0.3 plus 0.4, which is 0.7. Expected value weights each outcome by its probability rather than averaging the values evenly. The simple average of 0, 1, and 2 would mistakenly give 1.0.
When two events are independent, which statement about conditional probability is true?
The probability of A given B equals zero
The probability of A given B equals the probability of B
The probability of A given B equals the probability of A
The probability of A given B is always larger than the probability of A
Correct answer: The probability of A given B equals the probability of A
For independent events, the probability of A given B equals the probability of A, because knowing that B occurred provides no information that changes A's probability. This is the defining feature of independence. A conditional probability of zero would describe mutually exclusive events, not independent ones.
A random variable that can take any value within an interval, such as the exact height of a randomly chosen adult, is best described as which type of variable?
A discrete random variable
A continuous random variable
A categorical variable
A binomial random variable
Correct answer: A continuous random variable
Exact height is a continuous random variable because it can take any value within an interval rather than only separated, countable values. Discrete random variables, by contrast, take a countable set of values such as whole-number counts. Height is numeric rather than categorical, and it is not restricted to a binomial count of successes.
Let X be a random variable with mean 10. A new variable is defined as Y equals 3 times X plus 5. What is the mean of Y?
15
30
35
18
Correct answer: 35
The mean of Y is 35 because the linear transformation rule gives the mean of 3X plus 5 as 3 times the mean of X plus 5: 3 times 10 plus 5 equals 30 plus 5, which is 35. Both the multiplier and the added constant affect the mean. Forgetting to add the constant 5 would wrongly yield 30.
Let X be a random variable with standard deviation 4. A new variable is defined as Y equals 2 times X plus 7. What is the standard deviation of Y?
8
15
11
4
Correct answer: 8
The standard deviation of Y is 8 because multiplying a random variable by a constant multiplies the standard deviation by the absolute value of that constant, while adding a constant does not change spread: 2 times 4 equals 8. The added 7 shifts all values equally and leaves variability unchanged. Adding 7 to the standard deviation would be an error.
Two independent random variables have means of 12 and 5 and variances of 9 and 16, respectively. What is the variance of their sum?
5
625
25
12.5
Correct answer: 25
The variance of the sum is 25 because for independent random variables variances add: 9 plus 16 equals 25. This addition of variances holds for both sums and differences of independent variables. Standard deviations, unlike variances, cannot simply be added, so combining them directly would be incorrect.
Two independent random variables have means 12 and 5 and variances 9 and 16. What is the variance of their difference, the first minus the second?
7
25
Minus 7
5
Correct answer: 25
The variance of the difference is 25 because variances of independent random variables add even when the variables are subtracted: 9 plus 16 equals 25. Subtraction does not subtract variances, since variability accumulates regardless of the sign. Reporting 7 by subtracting the variances would be a common mistake.
In a binomial setting with 20 trials and probability of success 0.3 on each trial, what is the standard deviation of the number of successes?
6
4.2
About 1.45
About 2.05
Correct answer: About 2.05
The standard deviation is about 2.05 because the binomial standard deviation is np(1−p): 20×0.3×0.7, which is 4.2, approximately 2.05. The value 4.2 is the variance, not the standard deviation. The mean, 6, is np and is unrelated to spread.
A fair coin is flipped repeatedly until the first head appears. On average, how many flips are expected before the first head, given a success probability of 0.5 on each flip?
1
2
0.5
4
Correct answer: 2
The expected number of flips is 2 because the mean of a geometric random variable is 1 divided by the success probability: 1 divided by 0.5 equals 2. On average, it takes the reciprocal of the per-trial success probability to reach the first success. Using 0.5 itself confuses the probability with the expected waiting time.
A survey finds that 70% of households own a pet, and among pet-owning households 40% own a dog. What is the probability that a randomly chosen household both owns a pet and owns a dog?
0.40
0.70
0.28
1.10
Correct answer: 0.28
The probability is 0.28 because the general multiplication rule multiplies the probability of the first event by the conditional probability of the second given the first: 0.70 times 0.40 equals 0.28. The 40% applies only within the pet-owning group, so it must be scaled by the 70%. Adding the percentages would incorrectly exceed both individual values.
A study reports P(A) = 0.6, P(B) = 0.5, and P(A and B) = 0.2. What is the probability that A or B occurs?
0.9
1.1
0.7
0.3
Correct answer: 0.9
The probability of A or B is 0.9 because the general addition rule subtracts the overlap once: 0.6 plus 0.5 minus 0.2 equals 0.9. Without subtracting the joint probability, the overlap would be double-counted, giving an impossible value above 1. The result stays within the valid 0-to-1 range only after the subtraction.
A simulation is used to estimate the probability that at least two people in a group of 30 share a birthday. Why might a simulation be a reasonable approach for this problem?
Because the exact theoretical probability is impossible to define
Because simulation can approximate a probability that is tedious to compute exactly
Because simulation always gives the exact answer
Because probability rules do not apply to birthdays
Correct answer: Because simulation can approximate a probability that is tedious to compute exactly
Simulation is reasonable because it can approximate a probability that is tedious to compute exactly, repeatedly modeling the random process and recording the proportion of times the event occurs. The theoretical probability does exist and follows probability rules, but the calculation is laborious. Simulations give estimates that improve with more trials, not guaranteed exact answers.
A teacher claims that whether a student passed an exam is independent of whether the student attended a review session. Which equality, if true, would confirm this independence?
P(passed and attended) equals zero
P(passed given attended) equals P(passed)
P(passed) equals P(attended)
P(passed or attended) equals 1
Correct answer: P(passed given attended) equals P(passed)
Independence is confirmed if P(passed given attended) equals P(passed), because that shows attending the review session does not change the probability of passing. Equal marginal probabilities or a joint probability of zero do not establish independence; a zero joint probability would instead indicate mutually exclusive events. Comparing the conditional probability to the unconditional one is the correct test.
A discrete random variable representing the number of cars passing a checkpoint in one minute is most naturally measured on what kind of scale?
Any real number in an interval
Categories with no order
Whole-number counts
Percentages between 0 and 100
Correct answer: Whole-number counts
The number of cars is measured as whole-number counts because you can have 0, 1, 2, or more cars but never a fraction of a car, which is the hallmark of a discrete random variable. Continuous variables, by contrast, take any real value in an interval. Counts are numeric and ordered, not unordered categories or forced percentages.
Two fair dice are rolled. What is the probability that the sum of the two dice equals 7?
61
121
91
367
Correct answer: 61
The probability the sum is 7 is 61 because 6 of the 36 equally likely outcomes produce a sum of 7, and 6 divided by 36 equals 61. The favorable pairs are (1,6), (2,5), (3,4), (4,3), (5,2), and (6,1). Counting fewer favorable outcomes, such as treating order as irrelevant, would understate the probability.
A weighted die is constructed so that the probability of rolling a 6 is 0.4 and the probabilities of rolling 1 through 5 are equal. What is the probability of rolling a 3?
0.20
0.10
0.40
0.12
Correct answer: 0.12
The probability of rolling a 3 is 0.12 because the remaining 0.6 of total probability, after the 0.4 for a 6, is split equally among the five other faces: 0.6 divided by 5 equals 0.12. Total probability across all faces must equal 1. Assuming a fair value of about 0.167 would ignore the weighting of the die.
A binomial random variable counts successes in 50 independent trials with success probability 0.6. What is the mean number of successes?
30
20
12
25
Correct answer: 30
The mean number of successes is 30 because the binomial mean equals the number of trials times the success probability: 50 times 0.6 equals 30. This expected count reflects the long-run average successes across many repetitions of the 50 trials. Using the failure probability of 0.4 would wrongly give 20.
In a game, a player wins $5 with probability 0.2 and loses $2 with probability 0.8. Over many plays, what is the expected gain or loss per play?
A loss of $0.60
A gain of $0.60
A gain of $3.00
A loss of $2.00
Correct answer: A loss of $0.60
The expected result is a loss of $0.60 per play because expected value sums each outcome times its probability: 5 times 0.2 plus (negative 2) times 0.8 equals 1 minus 1.6, which is negative 0.60. A negative expected value indicates a long-run average loss. Ignoring the loss term, or its sign, would misstate the result as a gain.
What is the difference between a parameter and a statistic in the context of sampling distributions?
A parameter describes a sample, while a statistic describes a population
A parameter describes a population, while a statistic describes a sample
A parameter and a statistic are two names for the same quantity
A parameter is always larger in value than the corresponding statistic
Correct answer: A parameter describes a population, while a statistic describes a sample
A parameter describes a population, while a statistic describes a sample. A parameter is a fixed but usually unknown numerical summary of the whole population, and a statistic is computed from sample data to estimate it. The roles are not reversed, the two terms are not interchangeable, and there is no rule forcing a parameter to be larger than a statistic.
An estimator is described as biased. What does bias refer to in a sampling distribution?
How spread out the values of the statistic are across samples
A systematic tendency for the statistic to over- or under-estimate the parameter
The chance that any single sample is collected improperly
The number of outliers present in one particular sample
Correct answer: A systematic tendency for the statistic to over- or under-estimate the parameter
Bias is a systematic tendency for the statistic to over- or under-estimate the parameter, meaning the center of the sampling distribution does not sit at the true parameter value. Spread is described by variability, not bias; how a single sample is collected is a separate sampling-method issue; and the count of outliers in one sample is unrelated to the long-run centering of the estimator.
Two estimators of the same parameter are both unbiased, but estimator A has a smaller standard deviation of its sampling distribution than estimator B. Which estimator is generally preferred and why?
Estimator B, because more variability gives more information
Estimator A, because lower variability means its values cluster more closely around the parameter
Neither, because unbiased estimators are always equally good
Estimator B, because higher variability reduces bias
Correct answer: Estimator A, because lower variability means its values cluster more closely around the parameter
Estimator A is preferred because lower variability means its values cluster more closely around the parameter, making individual estimates more reliable. More variability does not give more useful information about a fixed parameter, unbiased estimators are not automatically equal in quality, and variability does not affect bias since both estimators are already unbiased.
When sampling without replacement, which condition must be met so that the standard deviation formula for a statistic remains approximately valid?
The sample size must be at least 10% of the population
The sample size must be no more than 10% of the population
The population must be exactly twice the sample size
The population must be normally distributed
Correct answer: The sample size must be no more than 10% of the population
The sample size must be no more than 10% of the population, the so-called 10% condition, so that sampling without replacement does not appreciably change the standard deviation of the statistic. Requiring the sample to be at least 10% would violate this condition, a fixed population-to-sample ratio of two is not the rule, and normality of the population is a separate consideration from the 10% condition.
For the sampling distribution of a sample proportion to be approximately normal, which condition is checked using np and n(1 - p)?
The independence condition
The randomness condition
The large counts condition
The 10% condition
Correct answer: The large counts condition
The large counts condition is checked using np and n(1 - p), both of which should be at least 10 for the sampling distribution of the sample proportion to be approximately normal. The independence and 10% conditions concern whether observations affect one another, the randomness condition concerns how the sample was selected, and none of those use the np and n(1 - p) calculations.
A population proportion is p = 0.3 and a random sample of size n = 50 is taken. What is the standard deviation of the sampling distribution of the sample proportion?
Approximately 0.065
Approximately 0.30
Approximately 0.21
Approximately 0.0042
Correct answer: Approximately 0.065
The standard deviation of the sampling distribution of the sample proportion is approximately 0.065, found by taking np(1−p), which is 50(0.3)(0.7). The value 0.30 is the proportion itself, 0.21 is p(1 - p) without dividing by n or taking the root, and 0.0042 is the variance before taking the square root.
Why must the population be normal (or the sample size large) for the sampling distribution of the sample mean to be normal?
Because only normal populations have a defined mean
Because small samples from non-normal populations may leave the sampling distribution non-normal
Because the sample mean is undefined for skewed data
Because randomness alone guarantees normality at any sample size
Correct answer: Because small samples from non-normal populations may leave the sampling distribution non-normal
Normality is required because small samples from non-normal populations may leave the sampling distribution non-normal; only a normal population or a large sample (via the central limit theorem) ensures the sampling distribution of the mean is approximately normal. Non-normal populations still have means, the sample mean is defined for skewed data, and randomness by itself does not produce normality at small sample sizes.
A population is exactly normal with mean 100 and standard deviation 15. For samples of size 9, what is the shape of the sampling distribution of the sample mean?
Skewed, because the sample size is small
Exactly normal, because the population is normal
Uniform, because all sample means are equally likely
Unknown, because the central limit theorem does not apply
Correct answer: Exactly normal, because the population is normal
The sampling distribution of the sample mean is exactly normal because the population is normal, and that holds for any sample size, including a small one like 9. A small sample does not introduce skew when the population itself is normal, sample means are not all equally likely, and the central limit theorem is not even needed since normality is inherited directly from a normal population.
What does the standard deviation of a sampling distribution measure?
The spread of individual data values within one sample
How much the statistic typically varies from sample to sample
The distance between the population mean and the population median
The number of samples needed to estimate the parameter
Correct answer: How much the statistic typically varies from sample to sample
The standard deviation of a sampling distribution measures how much the statistic typically varies from sample to sample, capturing the sampling variability of an estimate. It is not the spread of raw values within a single sample, not a gap between population center measures, and not a count of how many samples are required.
When constructing a sampling distribution for the difference between two sample proportions, what is the mean of that sampling distribution?
The sum of the two population proportions
The difference between the two population proportions
Always zero
The product of the two population proportions
Correct answer: The difference between the two population proportions
The mean of the sampling distribution of the difference between two sample proportions is the difference between the two population proportions, because each sample proportion is unbiased for its own population value. It is not their sum or product, and it is only zero in the special case where the two population proportions happen to be equal.
Two independent samples are taken to study the difference of two sample means. How is the standard deviation of the sampling distribution of the difference computed from the two individual sampling-distribution standard deviations?
By adding the two standard deviations directly
By taking σ12+σ22, the root of the sum of the two variances
By subtracting the smaller standard deviation from the larger
By averaging the two standard deviations
Correct answer: By taking σ12+σ22, the root of the sum of the two variances
The standard deviation of the sampling distribution of the difference is found by taking σ12+σ22, because variances of independent quantities add. Standard deviations themselves cannot be added or subtracted directly, and averaging them does not correctly combine independent variability.
A researcher uses a computer to repeatedly draw many random samples and record the sample mean of each. What is this process most directly used to approximate?
The population distribution of individual values
The sampling distribution of the sample mean
The exact value of the population mean
The bias of the sampling method
Correct answer: The sampling distribution of the sample mean
Repeatedly drawing many random samples and recording each sample mean is used to approximate the sampling distribution of the sample mean, since the collected means form a simulated version of that distribution. It does not reproduce the population distribution of individual values, does not pin down the exact population mean, and is not a measure of the sampling method's bias.
A simulated sampling distribution of sample means is centered very close to the true population mean but has visibly large spread. What does this suggest about the estimator?
It is biased but has low variability
It is approximately unbiased but has high variability
It is both biased and low in variability
It is neither unbiased nor variable
Correct answer: It is approximately unbiased but has high variability
Being centered near the true population mean indicates the estimator is approximately unbiased, while the large spread indicates high variability, so the estimator is approximately unbiased but has high variability. It is not biased since the center matches the parameter, and the visible spread rules out describing it as low in variability.
A population proportion is 0.6, and samples of size 100 are drawn. The large counts condition is checked with np = 60 and n(1 - p) = 40. What does this confirm?
The sampling distribution of the sample proportion can be treated as approximately normal
The population proportion must actually be 0.5
The samples were collected without randomness
The 10% condition is automatically satisfied
Correct answer: The sampling distribution of the sample proportion can be treated as approximately normal
Because np = 60 and n(1 - p) = 40 are both at least 10, the large counts condition is met, confirming the sampling distribution of the sample proportion can be treated as approximately normal. This calculation says nothing about forcing p to be 0.5, does not address how the samples were collected, and does not by itself verify the separate 10% condition.
Why does increasing the sample size reduce the standard deviation of a sampling distribution but not its bias?
Because bias depends on the centering of the estimator, which sample size does not change, while spread shrinks as n grows
Because larger samples always remove all bias automatically
Because sample size affects only the population mean, not the statistic
Because bias and standard deviation are the same quantity
Correct answer: Because bias depends on the centering of the estimator, which sample size does not change, while spread shrinks as n grows
Increasing sample size reduces spread because the standard deviation of a sampling distribution shrinks as n grows, but bias depends on the centering of the estimator, which sample size does not change. Larger samples do not automatically remove bias, sample size does not alter the fixed population mean, and bias and standard deviation are distinct concepts rather than the same quantity.
A one-proportion z-interval requires the standard error of the sample proportion. Which expression gives that standard error when the sample proportion is p-hat and the sample size is n?
N times p-hat times one minus p-hat
p-hat times one minus p-hat, divided by n, with no square root
np^
np^(1−p^)
Correct answer: np^(1−p^)
The standard error for a confidence interval is np^(1−p^). The sample proportion is used because the true proportion is unknown when estimating. Omitting the square root leaves the variance rather than the standard error, dropping the failure term ignores variability from non-successes, and multiplying by n instead of dividing reverses how sample size affects spread.
When conducting a one-proportion significance test, the standard error in the test statistic is computed differently than in a confidence interval. Which value is used inside the standard error for the test?
The sample proportion p-hat
The hypothesized proportion p0 from the null hypothesis
The midpoint of p-hat and p0
The margin of error
Correct answer: The hypothesized proportion p0 from the null hypothesis
For a significance test the standard error uses the hypothesized proportion p0, because the test assumes the null hypothesis is true and builds the sampling distribution around p0. A confidence interval instead uses p-hat since it makes no null assumption. Averaging the two values or substituting the margin of error are not part of the correct test-statistic formula.
A one-proportion z test gives a test statistic of z=2.5 for the alternative Ha: p > p0. Which statement best describes what this z-score represents?
The sample proportion equals 2.5
The probability of the result is 2.5
The observed proportion is 2.5 standard errors above the hypothesized proportion
The margin of error is 2.5
Correct answer: The observed proportion is 2.5 standard errors above the hypothesized proportion
A test statistic of z=2.5 means the observed sample proportion lies 2.5 standard errors above the hypothesized proportion under the null model. The z-score standardizes the distance between p-hat and p0. It is not a proportion itself, not a probability such as a p-value, and not the margin of error, which belongs to interval estimation rather than the test statistic.
For the same data set, how does a 99% confidence interval for a proportion compare to a 90% confidence interval?
The 99% interval is centered at a different sample proportion
The 99% interval is narrower because higher confidence reduces uncertainty
The two intervals have identical widths
The 99% interval is wider because a higher confidence level uses a larger critical value
Correct answer: The 99% interval is wider because a higher confidence level uses a larger critical value
The 99% interval is wider because raising the confidence level increases the critical value z*, which enlarges the margin of error. Greater confidence requires capturing the parameter more often, so the interval must stretch farther. The interval does not become narrower with higher confidence, the widths differ, and both intervals share the same center at the sample proportion.
A researcher wants the margin of error for a proportion to be no more than 0.03 at 95% confidence and uses 0.5 for the planning proportion. Why is 0.5 chosen for this calculation?
It makes the math simplest to perform by hand
It produces the largest possible standard error, giving a conservative sample size
It is always the true population proportion
It minimizes the required sample size
Correct answer: It produces the largest possible standard error, giving a conservative sample size
Using 0.5 is chosen because the product p times one minus p is maximized at p equal to 0.5, which yields the largest standard error and therefore the most conservative, or largest, required sample size. This guards against underestimating the needed sample when the true proportion is unknown. It is not selected for arithmetic ease, is not assumed to be the true value, and it maximizes rather than minimizes the sample size.
A 95% confidence interval for the proportion of voters favoring a tax increase runs from 0.42 to 0.49. Based on this interval, what can be concluded about the claim that a majority favor the increase?
There is convincing evidence that a majority favor the increase
The interval gives no information about the majority claim
There is convincing evidence that a majority do not favor the increase, since the interval lies entirely below 0.50
A majority is plausible because 0.50 is close to the interval
Correct answer: There is convincing evidence that a majority do not favor the increase, since the interval lies entirely below 0.50
Because the entire interval from 0.42 to 0.49 falls below 0.50, there is convincing evidence that a majority do not favor the increase. A confidence interval can be used to assess a claim about a value: since 0.50 is not a plausible value, the majority claim is not supported. The interval clearly addresses the claim, and 0.50 being nearby does not make it plausible when it sits outside the interval.
In a two-proportion significance test of H0: p1 = p2, the data from both groups are combined to estimate a single proportion. What is this quantity called and why is it used?
The combined proportion, used to widen the confidence interval
The pooled proportion, used because the null assumes the two population proportions are equal
The average proportion, used to compute the margin of error
The critical proportion, used to set the significance level
Correct answer: The pooled proportion, used because the null assumes the two population proportions are equal
It is called the pooled proportion, and it is used because the null hypothesis assumes the two population proportions are equal, so all successes and failures are combined to estimate that common value. This pooled estimate goes into the standard error of the test. It is not used to widen an interval or set alpha, and it is more than a simple average since it weights by group sizes.
Which of the following is the correct point estimate reported in a two-proportion confidence interval comparing groups 1 and 2?
The pooled proportion combining both groups
The ratio of the two sample proportions
The sum of the two sample proportions
The difference between the two sample proportions, p-hat 1 minus p-hat 2
Correct answer: The difference between the two sample proportions, p-hat 1 minus p-hat 2
The point estimate for a two-proportion confidence interval is the difference between the two sample proportions, p-hat 1 minus p-hat 2, because the interval estimates the difference in population proportions. The pooled proportion is used in a significance test, not in the interval estimate. A sum or ratio of proportions is not the parameter being estimated when comparing two groups by a difference.
A two-proportion z-interval for p1 minus p2 is computed as -0.02 to 0.11. What does the inclusion of 0 in this interval indicate?
The two population proportions are definitely equal
Group 1 has a higher proportion than group 2
There is no plausible difference, so it is plausible the two proportions are equal
The interval was computed incorrectly
Correct answer: There is no plausible difference, so it is plausible the two proportions are equal
Because 0 lies within the interval from -0.02 to 0.11, it is plausible that the two population proportions are equal, since a difference of zero is among the plausible values. This does not prove the proportions are exactly equal, only that no difference is plausible. It does not establish that group 1 is higher, and the presence of 0 is a normal, correct outcome, not an error.
Before running a two-proportion z-test, which condition must be checked regarding the counts in each group?
Each group must have at least 10 successes and at least 10 failures
Only the total combined sample size must exceed 30
Only one of the two groups needs at least 10 successes
The two sample proportions must be exactly equal
Correct answer: Each group must have at least 10 successes and at least 10 failures
The large-counts condition requires that each group separately have at least 10 successes and at least 10 failures, so the sampling distribution of each proportion is approximately normal. Checking only the combined total or just one group fails to validate the normal approximation for both. The sample proportions need not be equal; in fact the test exists precisely to compare them.
A significance test for a proportion uses data from a sample collected without randomization. How does this affect the validity of the conclusions?
It has no effect as long as the sample is large
It undermines the ability to generalize results, since the randomness condition is not met
It guarantees a Type I error will occur
It automatically makes the p-value smaller
Correct answer: It undermines the ability to generalize results, since the randomness condition is not met
A lack of randomization undermines generalizing the conclusions because the randomness condition, required for valid inference, is not satisfied. Inference procedures assume data come from a random sample or randomized experiment. Large sample size does not repair selection bias, the missing randomness does not guarantee a specific error, and it does not predictably change the p-value's size.
A test yields a p-value of 0.08. Using a significance level of 0.10, what decision is made, and using 0.05, what decision is made?
Reject the null at 0.10 but fail to reject at 0.05
Fail to reject at both 0.10 and 0.05
Reject the null at both 0.10 and 0.05
Reject the null at 0.05 but fail to reject at 0.10
Correct answer: Reject the null at 0.10 but fail to reject at 0.05
With a p-value of 0.08, the null is rejected at the 0.10 level because 0.08 is less than 0.10, but it is not rejected at the 0.05 level because 0.08 is greater than 0.05. The decision depends on comparing the p-value to the chosen significance level. The same p-value can lead to different conclusions at different alpha levels, which is why both comparisons matter.
How is the significance level alpha related to a Type I error in a significance test?
Alpha equals the probability of a Type II error
Alpha equals the probability of a Type I error when the null is true
Alpha equals the power of the test
Alpha equals the p-value of the test
Correct answer: Alpha equals the probability of a Type I error when the null is true
The significance level alpha is the probability of making a Type I error, that is, rejecting a true null hypothesis. By choosing alpha, the researcher sets the acceptable false-positive rate before collecting data. Alpha is not the probability of a Type II error, not the power, and not the p-value, which is computed from the observed data rather than chosen in advance.
All else held constant, which change would increase the power of a one-proportion significance test?
Decreasing the sample size
Lowering the significance level alpha
Increasing the sample size
Making the true proportion closer to the hypothesized value
Correct answer: Increasing the sample size
Increasing the sample size raises power because a larger sample reduces variability and makes it easier to detect a true difference from the hypothesized value. Decreasing sample size lowers power, lowering alpha makes rejection harder and reduces power, and a true proportion closer to the null value makes the effect harder to detect. Only a larger sample reliably increases power here.
A medical team concludes a treatment is effective when in reality it has no effect. In terms of patient consequences, which type of error has occurred?
A Type II error, leading to a missed effective treatment
A Type I error, leading to use of an ineffective treatment
No error, because the conclusion favored action
A sampling error unrelated to hypothesis testing
Correct answer: A Type I error, leading to use of an ineffective treatment
Concluding the treatment works when it truly does not is a Type I error, since a true null of no effect was wrongly rejected. The practical consequence is using a treatment that provides no benefit. A Type II error would be failing to detect a real effect, which is the opposite situation, and the mistake is a genuine inference error, not just sampling noise.
A confidence interval and a two-sided significance test at the same confidence level can lead to consistent conclusions. If a 95% confidence interval for p does not contain the value p0, what would a two-sided test of H0: p = p0 at alpha = 0.05 conclude?
Fail to reject the null hypothesis
The two methods cannot be compared
Reject the null hypothesis
Accept the null hypothesis as true
Correct answer: Reject the null hypothesis
If a 95% confidence interval excludes p0, then a two-sided test at the 0.05 level rejects the null hypothesis, because p0 is not a plausible value for the parameter. The duality between intervals and two-sided tests means values outside the interval would be rejected. The methods are directly comparable, and one never accepts the null as proven true.
A school claims at least 80% of graduates are employed within a year. To test this skeptically, which alternative hypothesis matches an investigator doubting the claim?
Ha: p > 0.80
Ha: p = 0.80
Ha: p < 0.80
Ha: p-hat < 0.80
Correct answer: Ha: p < 0.80
An investigator doubting the claim of at least 80% employment would use Ha: p < 0.80, testing whether the true proportion falls short of the claimed level. The alternative reflects the suspicion that employment is lower than stated. A greater-than alternative tests the opposite direction, the alternative never contains an equality, and hypotheses are written about the parameter p rather than the statistic p-hat.
Two news outlets report the same poll. Outlet A states a margin of error of plus or minus 2 percentage points and outlet B states plus or minus 4 percentage points for the same proportion. Assuming equal sample sizes, what most likely explains the difference?
Outlet B used a higher confidence level than outlet A
Outlet A used a higher confidence level than outlet B
Outlet A used a larger sample
Outlet B reported the point estimate incorrectly
Correct answer: Outlet B used a higher confidence level than outlet A
With equal sample sizes, outlet B's larger margin of error most likely results from using a higher confidence level, which increases the critical value and widens the margin. A higher confidence level for outlet A would shrink, not widen, its margin relative to B. Equal sample sizes rule out a sample-size explanation, and a misreported point estimate would not change the margin of error.
Why must inference for a proportion be based on a categorical variable rather than a quantitative one?
Quantitative variables cannot be sampled randomly
Proportions are computed by averaging numerical measurements
Categorical variables always produce larger samples
Proportions summarize the share of observations falling into a category
Correct answer: Proportions summarize the share of observations falling into a category
Inference for a proportion uses a categorical variable because a proportion represents the share of observations falling into a particular category, such as success versus failure. Quantitative data are summarized by means, not proportions, so averaging numerical measurements describes a mean instead. Sample size and randomization apply to both variable types, so those reasons do not distinguish proportions.
A study finds a statistically significant difference between two proportions but the confidence interval for the difference is very close to 0, such as 0.001 to 0.012. What does this illustrate?
Statistical significance always means a large practical effect
The proportions are exactly equal
The test must have been performed incorrectly
A result can be statistically significant yet small in practical importance
Correct answer: A result can be statistically significant yet small in practical importance
This illustrates that a result can be statistically significant while the actual difference is small in practical importance, since a narrow interval near 0 shows the effect, though real, is tiny. Large samples can detect minuscule differences. Significance does not guarantee a large effect, the test is not necessarily flawed, and an interval excluding 0 means the proportions are not exactly equal.
When the conditions for a one-proportion z-test are met, which distribution is used to find the p-value?
The standard normal distribution
The t-distribution with n minus 1 degrees of freedom
The χ2 distribution
The binomial distribution with parameter p-hat
Correct answer: The standard normal distribution
The p-value for a one-proportion z-test is found using the standard normal distribution, because the standardized test statistic follows an approximately normal model when the large-counts condition holds. The t-distribution applies to inference about means with unknown standard deviation, not proportions. The χ2 distribution is used for categorical tables, and the exact binomial is not the basis of the z-procedure.
A two-sided test of H0: p = 0.4 produces a sample proportion below 0.4. How is the two-sided p-value related to the one-sided tail area in this case?
It is double the one-sided tail area
It equals the one-sided tail area
It is half the one-sided tail area
It is unrelated to the tail area
Correct answer: It is double the one-sided tail area
For a two-sided test, the p-value is double the one-sided tail area because both tails of the distribution count as results at least as extreme as the observed one. Since the alternative allows deviations in either direction, the area in the matching tail is doubled. It is not equal to or half of a single tail, and it is directly built from the tail area, not unrelated.
A researcher increases the sample size in a one-proportion confidence interval while keeping the confidence level and sample proportion the same. What happens to the width of the interval?
It increases
It becomes zero
It stays exactly the same
It decreases
Correct answer: It decreases
Increasing the sample size decreases the width of the interval because a larger n shrinks the standard error, reducing the margin of error on each side. With the confidence level and sample proportion unchanged, only the standard error changes. The width does not increase or remain fixed with a larger sample, and it approaches but never reaches zero for any finite sample.
In reporting the conclusion of a one-proportion significance test, why is it incorrect to say the test proves the null hypothesis is true when failing to reject it?
Failing to reject means there is not enough evidence against the null, not proof it is true
Failing to reject always means a Type I error occurred
The null hypothesis can never be written about a proportion
Proving the null requires a larger significance level
Correct answer: Failing to reject means there is not enough evidence against the null, not proof it is true
Failing to reject the null means only that there is insufficient evidence against it, not that it is proven true, because the test can never confirm a parameter exactly. A non-significant result is consistent with the null but also with small effects the test could not detect. It does not imply a Type I error, the null is routinely about a proportion, and no alpha level can prove the null.
In a one-sample t interval for a population mean, what does the margin of error represent?
The difference between the largest and smallest values in the sample
The product of the critical t value and the standard error of the mean
The hypothesized value of the population mean
The probability that the interval misses the true mean
Correct answer: The product of the critical t value and the standard error of the mean
The margin of error in a one-sample t interval is the product of the critical t value and the standard error of the mean, and it is added to and subtracted from the sample mean to form the interval. It is not the range of the data, not the hypothesized mean, and not a probability, since those quantities describe other features of the data or procedure.
A 95% confidence interval for a population mean is reported as 48 plus or minus 6. What is the point estimate of the population mean?
6
48
42
54
Correct answer: 48
The point estimate is 48, because in an interval written as a center plus or minus a margin of error, the center value is the sample mean that serves as the point estimate. The value 6 is the margin of error, while 42 and 54 are the lower and upper endpoints of the interval rather than the point estimate.
A teacher wants to estimate the mean score on a final exam with a smaller margin of error while keeping the same confidence level. Which action will accomplish this?
Increase the sample size
Decrease the sample size
Raise the confidence level
Use a one-sided interval interpretation
Correct answer: Increase the sample size
Increasing the sample size will shrink the margin of error while holding the confidence level constant, because a larger sample lowers the standard error of the mean. Decreasing the sample size would enlarge the margin of error, raising the confidence level would also widen it, and switching to a one-sided interpretation does not address the goal of a smaller two-sided margin.
For a one-sample t procedure with a sample of size 18, how many degrees of freedom are used to find the critical value?
17
18
19
9
Correct answer: 17
The degrees of freedom equal 17, found by subtracting one from the sample size of 18, since a one-sample t procedure uses n minus 1 degrees of freedom. The value 18 ignores the subtraction, 19 adds instead of subtracts, and 9 incorrectly halves the sample size.
When the conservative approach is used to find degrees of freedom for a two-sample t procedure by hand, which value is typically chosen?
The sum of the two sample sizes
The smaller of the two sample sizes minus one
The larger of the two sample sizes minus one
The product of the two sample sizes
Correct answer: The smaller of the two sample sizes minus one
The conservative approach uses the smaller of the two sample sizes minus one for the degrees of freedom, which produces a larger critical value and a wider, safer interval. Summing the sample sizes overstates the degrees of freedom, using the larger sample size is not conservative, and multiplying the sample sizes does not correspond to any degrees of freedom rule.
A matched-pairs t test is most appropriate when the data consist of which of the following?
Two independent random samples from two separate populations
Counts of successes and failures in two groups
A single categorical variable measured once
Pairs of related observations, with the analysis performed on the differences within each pair
Correct answer: Pairs of related observations, with the analysis performed on the differences within each pair
A matched-pairs t test is appropriate when the data consist of pairs of related observations, and the test is carried out on the differences within each pair. Two independent samples call for a two-sample procedure instead, a single categorical variable is not quantitative, and counts of successes and failures call for proportion-based methods.
Which set of hypotheses is appropriate for a paired t test where the mean difference is denoted mu sub d?
H0: p = 0 versus Ha: p is not equal to 0
H0: sigma = 0 versus Ha: sigma is not equal to 0
H0: x-bar = 0 versus Ha: x-bar is not equal to 0
H0: mu sub d = 0 versus Ha: mu sub d is not equal to 0
Correct answer: H0: mu sub d = 0 versus Ha: mu sub d is not equal to 0
The appropriate hypotheses are H0: mu sub d = 0 versus Ha: mu sub d is not equal to 0, because a paired t test concerns the mean of the population of differences. Using p refers to a proportion, using x-bar incorrectly states the hypothesis about a statistic rather than a parameter, and using sigma refers to a standard deviation rather than a mean difference.
In a two-sample t test for the difference of two means, what is the null hypothesis usually stated as?
One population mean is twice the other
The two population means differ by exactly 10
The two sample means are equal
The two population means are equal
Correct answer: The two population means are equal
The null hypothesis for a two-sample t test is usually that the two population means are equal, equivalently that their difference is zero. The hypothesis is about population parameters, not the observed sample means, and it does not by default specify a difference of 10 or a doubling relationship unless a problem explicitly states such a claim.
A researcher calculates a one-sample t statistic of 2.40 for a test of H0: mu = 50 against Ha: mu greater than 50. What does the positive sign of the test statistic indicate?
The sample mean is below the hypothesized value of 50
The sample size is large
The population standard deviation is positive
The sample mean is above the hypothesized value of 50
Correct answer: The sample mean is above the hypothesized value of 50
A positive t statistic indicates the sample mean is above the hypothesized value of 50, because the numerator is the sample mean minus the hypothesized mean. A sample mean below 50 would yield a negative statistic, while the sign of the statistic is unrelated to whether the standard deviation is positive or to the size of the sample.
Which of the following is the correct interpretation of a p-value in a one-sample t test for a mean?
The probability that the null hypothesis is true
The probability that the alternative hypothesis is true
The probability of getting a sample result at least as extreme as the one observed, assuming the null hypothesis is true
The proportion of the population that lies within the confidence interval
Correct answer: The probability of getting a sample result at least as extreme as the one observed, assuming the null hypothesis is true
The p-value is the probability of obtaining a sample result at least as extreme as the one observed, assuming the null hypothesis is true. It is not the probability that either hypothesis is true, since hypotheses are not assigned probabilities in this framework, and it is not a statement about the share of the population inside an interval.
A 99% confidence interval for the mean weight of a product is (245, 255) grams. A quality inspector claims the mean is 250 grams. Based on this interval, what should the inspector conclude?
The claim is plausible because 250 is inside the interval
The claim is contradicted because 250 is outside the interval
The interval proves the mean is exactly 250 grams
The confidence level must be lowered to test the claim
Correct answer: The claim is plausible because 250 is inside the interval
The claim is plausible because 250 grams lies inside the interval from 245 to 255, and any value within a confidence interval is consistent with the data at that confidence level. The value 250 is not outside the interval, the interval cannot prove an exact value, and there is no need to change the confidence level to assess this claim.
A study reports a 95% confidence interval for the difference in mean reaction times (treatment minus control) as (-0.8, 0.5) seconds. What does this interval suggest about the difference in means?
There is no convincing evidence of a difference, since the interval includes zero
There is convincing evidence the treatment mean is lower
There is convincing evidence the treatment mean is higher
The two means are exactly equal
Correct answer: There is no convincing evidence of a difference, since the interval includes zero
Because the interval from -0.8 to 0.5 includes zero, there is no convincing evidence of a difference between the two means at the 95% confidence level. An interval that contains zero allows for the possibility of no difference, so it neither establishes that one mean is higher or lower nor proves the means are exactly equal.
In a one-sample t test, a Type I error occurs when which of the following happens?
The null hypothesis is rejected when it is actually true
The null hypothesis is not rejected when it is actually false
The sample mean equals the population mean
The confidence level is set too high
Correct answer: The null hypothesis is rejected when it is actually true
A Type I error occurs when the null hypothesis is rejected even though it is actually true, producing a false positive conclusion about the mean. Failing to reject a false null is a Type II error, the sample mean matching the population mean is not an error, and the choice of confidence level is a design decision rather than an error.
For a fixed effect size in a test about a population mean, increasing the significance level alpha generally has what effect on the power of the test?
It decreases the power
It has no effect on the power
It increases the power
It makes the power exactly equal to alpha
Correct answer: It increases the power
Increasing the significance level alpha generally increases the power of a test about a mean, because a larger rejection region makes it easier to reject a false null hypothesis. Raising alpha does not decrease power, it does have an effect, and power is not set equal to alpha, since power measures the chance of correctly rejecting a false null.
A nutrition researcher computes a one-sample t interval but the sample of size 12 shows a strong skew with a clear outlier on a graph. What is the main concern for the inference?
The interval will be too narrow to be useful
The degrees of freedom will be negative
The normality condition may be violated, making the t interval unreliable for this small sample
The sample mean cannot be computed
Correct answer: The normality condition may be violated, making the t interval unreliable for this small sample
The main concern is that the normality condition may be violated, making the t interval unreliable, because a small sample of size 12 with strong skew and an outlier suggests the population may not be normal. The skew does not force the interval to be too narrow, degrees of freedom cannot be negative, and the sample mean can always be computed from the data.
A two-sample t interval for the difference in mean heights of two plant varieties (A minus B) is (1.2, 3.6) centimeters. What does this interval indicate?
Variety A and variety B have equal mean heights
Variety A has a higher mean height than variety B
Variety B has a higher mean height than variety A
No comparison of the means can be made
Correct answer: Variety A has a higher mean height than variety B
The interval from 1.2 to 3.6 centimeters indicates variety A has a higher mean height than variety B, because the entire interval for A minus B lies above zero. Since the interval does not include zero, the means are not plausibly equal, and a positive interval rules out variety B being taller and clearly allows a comparison.
An experimenter measures the cholesterol of 30 patients before and after a medication and runs a paired t test. How many degrees of freedom does this test use?
60
58
29
30
Correct answer: 29
The paired t test uses 29 degrees of freedom, found by subtracting one from the number of pairs, which is 30 pairs minus one. A paired analysis reduces to a one-sample test on the 30 differences, so the count is not based on 60 total measurements, and the values 60 and 58 incorrectly treat the data as two independent groups.
Why is a paired design often preferred over two independent samples when the same subjects can be measured under two conditions?
It eliminates the need to check any conditions for inference
It always produces a larger p-value
It reduces variability by controlling for differences between subjects
It removes the need to compute a sample mean
Correct answer: It reduces variability by controlling for differences between subjects
A paired design is often preferred because it reduces variability by controlling for differences between subjects, since each subject serves as their own comparison. It does not remove the need to check inference conditions, it does not guarantee a larger p-value, and computing a mean of the differences is still required.
A one-sample t test gives a p-value of 0.12 at a significance level of 0.05. What is the correct conclusion?
Reject the null hypothesis; there is convincing evidence about the mean
The test should be repeated until the p-value is below 0.05
Accept the alternative hypothesis as proven
Fail to reject the null hypothesis; there is not convincing evidence about the mean
Correct answer: Fail to reject the null hypothesis; there is not convincing evidence about the mean
Because the p-value of 0.12 exceeds the significance level of 0.05, we fail to reject the null hypothesis and conclude there is not convincing evidence about the mean. A p-value above alpha does not justify rejecting the null, we never accept the alternative as proven, and repeating a test just to chase significance is not a valid procedure.
What is the standard error of the difference of two sample means built from in a two-sample t procedure?
The sum of the two sample standard deviations
n1s12+n2s22
The difference of the two sample means
The larger sample standard deviation divided by the total sample size
Correct answer: n1s12+n2s22
The standard error of the difference of two sample means is n1s12+n2s22, because the variances of independent samples add. Adding the standard deviations directly is incorrect, the difference of the means is the estimate rather than its standard error, and using only the larger standard deviation ignores the second group.
A sample of 16 measurements has a mean of 25 and a sample standard deviation of 4. What is the standard error of the mean for a one-sample t procedure?
4
0.25
1
16
Correct answer: 1
The standard error of the mean is 1, found by dividing the sample standard deviation of 4 by 16, which is 4 divided by 4. The value 4 is the standard deviation itself, 0.25 divides incorrectly by the sample size, and 16 is the sample size rather than a standard error.
In hypothesis testing for a mean, what does the significance level alpha represent?
The probability of a Type II error
The probability of correctly failing to reject a true null
The power of the test
The maximum probability of a Type I error that the researcher is willing to accept
Correct answer: The maximum probability of a Type I error that the researcher is willing to accept
The significance level alpha represents the maximum probability of a Type I error that the researcher is willing to accept, setting the threshold for rejecting a true null hypothesis. It is not the probability of a Type II error, not the chance of correctly retaining a true null, and not the power of the test, which measures detecting a false null.
A study comparing the mean income of two cities uses independent random samples. Which condition specifically addresses whether one sample's values could influence the other's?
The randomness condition
The symmetry condition
The large counts condition
The independence condition, including independence between the two groups
Correct answer: The independence condition, including independence between the two groups
The independence condition, including independence between the two groups, addresses whether the values in one sample could influence the values in the other, and it must hold for a valid two-sample t procedure. The randomness condition concerns how samples are selected, the large counts condition applies to proportions, and there is no separate symmetry condition for this procedure.
A 90% confidence interval for a population mean is (62, 78). What is the margin of error for this interval?
8
16
70
62
Correct answer: 8
The margin of error is 8, found by taking half the width of the interval, since the width from 62 to 78 is 16 and the margin is half of that. The full width 16 is twice the margin, 70 is the center of the interval, and 62 is the lower endpoint rather than the margin.
When the population standard deviation is unknown and a sample standard deviation is used instead, why is the t distribution preferred over the normal distribution for inference about a mean?
Because the t distribution has a smaller spread than the normal distribution
Because the t distribution accounts for the extra uncertainty from estimating the standard deviation
Because the t distribution is only defined for large samples
Because the t distribution requires the population to be uniform
Correct answer: Because the t distribution accounts for the extra uncertainty from estimating the standard deviation
The t distribution is preferred because it accounts for the extra uncertainty introduced by estimating the standard deviation from the sample, giving it heavier tails than the normal distribution. The t distribution has a larger spread, not a smaller one, it is defined for all sample sizes through its degrees of freedom, and it does not require a uniform population.
As the degrees of freedom for a t distribution increase, what happens to its shape?
It becomes more skewed to the right
It becomes flatter with heavier tails
It approaches the standard normal distribution
It becomes a uniform distribution
Correct answer: It approaches the standard normal distribution
As the degrees of freedom increase, the t distribution approaches the standard normal distribution, because the extra uncertainty from estimating the standard deviation shrinks with larger samples. The t distribution is symmetric rather than skewed, its tails grow lighter rather than heavier as degrees of freedom rise, and it never becomes uniform.
A car company tests whether a new tire design changes braking distance compared with the old design using the same 25 cars under both designs. Which procedure is appropriate?
A two-sample t test treating the two designs as independent groups
A matched-pairs t test on the differences in braking distance
A test for a single population proportion
A two-proportion z test
Correct answer: A matched-pairs t test on the differences in braking distance
A matched-pairs t test on the differences in braking distance is appropriate because the same 25 cars are measured under both tire designs, creating paired data. Treating the designs as independent groups ignores the natural pairing, and proportion tests do not apply to the quantitative braking-distance measurements.
A one-sample t interval for a mean assumes the data come from a random sample. What goes wrong with the inference if the sample was actually a voluntary response sample?
Bias may make the interval fail to capture the true population mean reliably
The standard error becomes negative
The degrees of freedom must be recalculated
The confidence level automatically rises to 100%
Correct answer: Bias may make the interval fail to capture the true population mean reliably
If the data come from a voluntary response sample, bias may make the interval fail to capture the true population mean reliably, because the sample may not represent the population. Degrees of freedom depend only on sample size, standard errors are never negative, and a flawed sampling method does not raise the confidence level to 100%.
A two-sample t test produces a test statistic of t=−3.1 for comparing mean test scores (group 1 minus group 2). What does the negative value indicate?
The mean of group 1 is less than the mean of group 2
The mean of group 1 is greater than the mean of group 2
The two sample means are equal
The standard error is negative
Correct answer: The mean of group 1 is less than the mean of group 2
A negative test statistic indicates the mean of group 1 is less than the mean of group 2, because the numerator computes the group 1 mean minus the group 2 mean. A positive difference would give a positive statistic, equal means would give a statistic near zero, and a standard error is always positive.
A 95% confidence interval and a two-sided significance test at alpha = 0.05 are used on the same data for a mean. If the hypothesized value lies outside the 95% interval, what is the corresponding test conclusion?
Fail to reject the null hypothesis
Reject the null hypothesis
The two methods give contradictory results
No conclusion is possible without the p-value
Correct answer: Reject the null hypothesis
If the hypothesized value lies outside the 95% confidence interval, the corresponding two-sided test at alpha = 0.05 will reject the null hypothesis, because the interval and the test are linked. A value outside the interval is not consistent with the data, so we do not fail to reject, the methods agree rather than contradict, and the interval alone supports the conclusion.
A researcher wants to detect a smaller true difference between two population means without changing alpha. Which change increases the power of the two-sample t test?
Decreasing both sample sizes
Increasing the variability within each group
Increasing both sample sizes
Reducing the true difference even further
Correct answer: Increasing both sample sizes
Increasing both sample sizes increases the power of the two-sample t test, because larger samples reduce the standard error and make a true difference easier to detect. Decreasing sample sizes lowers power, more within-group variability lowers power, and a smaller true difference is harder, not easier, to detect.
A confidence interval for a mean is computed at the 80% confidence level. Compared with a 95% interval from the same data, the 80% interval will be which of the following?
Wider, because lower confidence requires a larger critical value
Centered at a higher value
Identical, because the data are the same
Narrower, because lower confidence requires a smaller critical value
Correct answer: Narrower, because lower confidence requires a smaller critical value
The 80% interval will be narrower because a lower confidence level uses a smaller critical t value, producing a smaller margin of error. Lower confidence does not require a larger critical value, the intervals differ despite using the same data, and both intervals share the same center at the sample mean.
A one-sample t test is conducted with Ha: mu less than 30 (a left-tailed test). For which kind of sample mean will the p-value be small?
A sample mean much greater than 30
A sample mean exactly equal to 30
A sample mean much less than 30
Any sample mean, regardless of its value
Correct answer: A sample mean much less than 30
For a left-tailed test with Ha: mu less than 30, the p-value will be small when the sample mean is much less than 30, since that result provides strong evidence in the direction of the alternative. A sample mean above 30 points away from the alternative, a mean equal to 30 gives a large p-value, and the p-value clearly depends on the sample mean.
What is the correct way to interpret a 95% confidence level for a method that produces intervals for a mean?
In repeated sampling, about 95% of the intervals constructed this way will capture the true mean
Each individual interval has a 95% chance of being correct after it is computed
95% of the population values fall inside any one interval
The sample mean equals the population mean 95% of the time
Correct answer: In repeated sampling, about 95% of the intervals constructed this way will capture the true mean
The 95% confidence level means that in repeated sampling, about 95% of the intervals constructed by this method will capture the true mean. It is not a probability statement about a single completed interval, it does not describe the share of population values inside one interval, and it does not claim the sample mean equals the population mean most of the time.
A two-sample t test requires that the two samples be independent. Which of the following study designs satisfies this requirement?
Measuring the same people twice under two conditions
Measuring one group at two different time points
Pairing each subject with a similar partner before measuring
Randomly assigning different subjects to two separate treatment groups
Correct answer: Randomly assigning different subjects to two separate treatment groups
Randomly assigning different subjects to two separate treatment groups satisfies the independence requirement for a two-sample t test, because the two groups consist of distinct, unrelated individuals. Measuring the same people twice, pairing subjects, or measuring one group at two times all create dependent data better suited to a paired procedure.
A confidence interval for the difference of two means is found to be entirely negative, such as (-5.0, -1.0). What does this indicate about the two population means (computed as mean 1 minus mean 2)?
Mean 1 is plausibly larger than mean 2
Mean 1 is plausibly smaller than mean 2
The two means are plausibly equal
The difference is plausibly zero
Correct answer: Mean 1 is plausibly smaller than mean 2
An interval that is entirely negative, such as -5.0 to -1.0, indicates mean 1 is plausibly smaller than mean 2, because the difference mean 1 minus mean 2 is below zero throughout. A positive difference would mean the first mean is larger, and since zero is not in the interval the means are not plausibly equal nor is the difference plausibly zero.
In a paired t procedure, what quantity is actually analyzed to test for a mean difference?
The two original groups of measurements treated separately
The set of within-pair differences
The ratio of the two measurements in each pair
The combined list of all measurements pooled together
Correct answer: The set of within-pair differences
A paired t procedure analyzes the set of within-pair differences, reducing the problem to a one-sample t analysis on those differences. It does not treat the two groups separately, it does not use the ratio of paired measurements, and it does not simply pool all measurements together, since pooling would discard the pairing structure.
A researcher reduces the significance level for a test about a mean from 0.05 to 0.01 to be more cautious. Holding everything else fixed, what is one consequence of this change?
The probability of a Type I error decreases but the probability of a Type II error increases
The probability of a Type I error increases
Both error probabilities decrease at once
The power of the test increases
Correct answer: The probability of a Type I error decreases but the probability of a Type II error increases
Lowering the significance level from 0.05 to 0.01 decreases the probability of a Type I error but increases the probability of a Type II error, because a stricter threshold makes the null harder to reject. The Type I error probability does not increase, both error probabilities cannot drop at once from this single change, and the power decreases rather than increases.
To find us again, just search “Career Employer AP Statistics”
According to the empirical rule, approximately what percentage of values in a normal distribution fall within two standard deviations of the mean?
Pick an answer to see the explanation
Click Start Test above to launch a full-length AP Statistics multiple-choice practice test weighted exactly like the real exam, or drill a single unit — from Exploring One-Variable Data to Inference for Slopes. Every question includes a clear explanation so you learn the reasoning, not just the answer.
The AP Statistics exam is a college-level assessment that measures your ability to collect, explore, and draw conclusions from data using statistical reasoning across nine units.
It is administered by the College Board and is given once a year in May, with most students taking it after a year-long AP Statistics course.[1] A strong score can earn you college credit or advanced placement.
These practice questions follow the published AP Statistics course and exam description, mirroring the content and pacing of the real multiple-choice section so you can build readiness across every unit.[2] To round out your prep, pair these with our free study guide, flashcards, and cheat sheet.
Dates, fees, and policies change — always verify the current details at collegeboard.org before you register.
AP Statistics is one of the 17 AP exams — explore all our AP practice tests to compare and prep across the whole family.
AP Statistics at a Glance
AP Statistics at a glance
Detail
AP Statistics
Questions
40 multiple-choice + 6 free-response (practice covers the 40 MCQ)
Format
Hybrid digital exam (Bluebook MCQ + handwritten free response)
Time limit
About 3 hours total (1 hr 30 min MCQ + 1 hr 30 min FRQ)
Scoring / Result
Scored 1-5; a 3 or higher is generally passing and may earn college credit
Calculator
Graphing calculator permitted throughout; built-in Desmos in Bluebook for 2026
Administered by
College Board (given once a year in May)
Eligibility
Open to any student; no formal prerequisite, but a year-long course is recommended
Cost / Fee
Approximately 99intheU.S.(about129 internationally); verify at collegeboard.org
Retakes
Offered only once a year in May — you can only retake it the next year
What Is on the AP Statistics Exam?
The AP Statistics exam has two equally weighted sections: 40 multiple-choice questions (50% of the score) and 6 free-response questions (50% of the score), the last of which is a longer investigative task. The content is organized into nine units, from Exploring One-Variable Data through Inference for Slopes.[1]
Each unit carries an official weighting on the multiple-choice section, with Exploring One-Variable Data the most heavily tested. Our full practice test mirrors these proportions:
AP Statistics weighting by unit
Unit 1: Exploring One-Variable Data20% · 15-23%
Unit 3: Collecting Data15% · 12-15%
Unit 4: Probability & Random Variables15% · 10-20%
Unit 6: Inference for Proportions15% · 12-15%
Unit 7: Inference for Means15% · 10-18%
Unit 5: Sampling Distributions10% · 7-12%
Unit 2: Exploring Two-Variable Data5% · 5-7%
Unit 8: Inference for Chi-Square3% · 2-5%
Unit 9: Inference for Slopes3% · 2-5%
Practice Questions by Unit
Use Start Test for a full weighted AP Statistics simulation, or open the hub and pick a single unit to drill your weak area. After each full exam, your results show a per-unit breakdown so you know exactly where to focus — most students need the most reps on Exploring One-Variable Data, Collecting Data, and the inference units.
Who Is Eligible to Take the AP Statistics Exam?
The AP Statistics exam is open to any student — there is no formal prerequisite and you do not have to take an AP course to sit for the exam.[6]
That said, the exam covers a full year of college-level statistics, so most successful examinees have completed an AP Statistics course or equivalent coursework in data analysis, probability, sampling, and inference. A solid second-year algebra background is recommended.
If your school does not offer AP, you can usually arrange to test at a nearby school that administers AP exams. Contact that school’s AP coordinator early, because seats and ordering deadlines fill well before May.
How Do You Register for the AP Statistics Exam?
You register for the AP Statistics exam through your school’s AP coordinator, not directly with the College Board. In My AP, you indicate that you plan to test, and the coordinator orders your exam.[6]
The standard exam fee is approximately $99 at schools in the U.S., U.S. territories, Canada, and DoDEA schools, and about $129 internationally. Your AP coordinator collects any fees you owe.[4]
The final ordering deadline for full-year courses is typically in mid-November, and a late order fee applies after that. Verify the current fee and deadlines at collegeboard.org, as they change each year.
If your school does not offer AP, contact a participating school’s AP coordinator to arrange testing — do this in the fall, well ahead of the spring deadlines.
How Is the AP Statistics Exam Scored?
AP Statistics is scored on a scale of 1 to 5, where 5 means extremely well qualified, 3 means qualified, and 1 means no recommendation.[5]
The multiple-choice and free-response sections each count for 50% of your composite score, which is converted to the final 1-5 scale. There is no penalty for wrong answers, so you should answer every multiple-choice question.
A score of 3 or higher is generally considered passing, and the College Board and ACE recommend that colleges grant credit for a 3 or above. Each college sets its own policy, so check the AP credit requirements at your target schools.
How Hard Is the AP Statistics Exam?
AP Statistics is considered challenging because it tests statistical reasoning and clear communication across nine units rather than simple computation.[2] The difficulty is choosing the right procedure, checking its conditions, and justifying your conclusion in context.
The multiple-choice section pairs concept questions with data-based items — reading graphs, output, and tables quickly matters as much as knowing formulas, even with a graphing calculator in hand.
The free-response section then asks you to design studies, run inference, and defend claims with evidence, ending in a longer investigative task. Communicating reasoning clearly is what separates strong scores from average ones.
1-5
Score scale
3+ generally passing
40
Multiple-choice Qs
50% of the score
9
Units tested
One-Variable Data weighted most
The takeaway: drill until you’re consistently scoring at or above your target college credit threshold on full-length, unit-weighted practice — especially the heavily tested units — before exam day in May.
What to Expect on Exam Day
AP Statistics is a hybrid digital exam: you answer the multiple-choice questions and view the free-response prompts in the Bluebook testing app, then handwrite your free-response answers in a paper booklet.[1]
You work through 40 multiple-choice questions in the first 90 minutes, take a short break, then complete 6 free-response questions in the final 90 minutes. A graphing calculator is permitted throughout — and an approved Desmos calculator is built into Bluebook for 2026 — and a formula sheet and statistical tables are provided.
Bring an acceptable photo ID if required by your school, arrive early, and leave phones and personal items as instructed. Having simulated the full multiple-choice timing with practice tests makes the pacing feel routine.
How to Use This AP Statistics Practice Test
Recreate exam conditions. Take the full multiple-choice test timed, with your calculator.[2]
Diagnose, then drill. Use a full simulation to find weak units, then drill them.
Prioritize the heavy units. Exploring One-Variable Data and the inference units move your score most.
Learn the why. Read every explanation — reasoning beats memorizing.
Answer everything. There’s no guessing penalty, so never leave a question blank.
Why the AP Statistics Exam Matters
A strong AP Statistics score is one of the clearest ways to earn college credit, skip an introductory course, and strengthen your college applications — it gives admissions officers and colleges an objective measure of college-level readiness.[5] Because the exam is offered only once a year, every rep counts: you can’t retake it until the following May. These free AP Statistics practice tests are the most efficient way to walk in ready the first time.
Conclusion
Performing well on the AP Statistics exam comes down to statistical reasoning across nine units, sharp data analysis, and the ability to justify conclusions under timed conditions. Use this free AP Statistics practice test to find your weak units, drill them to mastery, and pair it with our free study guide, flashcards, and cheat sheet to walk in confident on test day.
AP Statistics Practice Test FAQ
The AP Statistics exam is a college-level assessment administered by the College Board that measures your ability to collect, analyze, and draw conclusions from data. It is intended for high school students who want to earn college credit or advanced placement by demonstrating mastery of introductory statistics. Scoring well can let you skip an equivalent introductory statistics course in college.
The AP Statistics exam is about 3 hours long and has two equally weighted sections. Section I is 40 multiple-choice questions in 1 hour 30 minutes (50% of the score), and Section II is 6 free-response questions in 1 hour 30 minutes (50% of the score), including one longer investigative task. Our practice test focuses on the 40-question multiple-choice section, weighted to the official unit breakdown.
AP exams are scored on a scale of 1 to 5, where 5 means extremely well qualified and 1 means no recommendation. The multiple-choice and free-response sections each count for 50% of your composite score, which is then converted to the 1-5 scale. A score of 3 or higher is generally considered passing and is recommended for college credit.
A score of 3 or higher is generally treated as passing, and the College Board and ACE recommend that colleges award credit for scores of 3 and above. However, each college sets its own credit policy, so some competitive programs require a 4 or 5. Check the AP credit policy of your target schools to know the score you need.
The AP Statistics exam fee is approximately $99 per exam at schools in the U.S., U.S. territories, Canada, and DoDEA schools, and about $129 elsewhere (verify the current fee at collegeboard.org, since fees change). You register through your school's AP coordinator, not directly with the College Board. If your school does not offer AP, you can arrange to test at a nearby participating school.
Yes. A graphing calculator with statistics capabilities is allowed on both the multiple-choice and free-response sections of the AP Statistics exam, and you should bring one you know well. For the 2026 exam, an approved Desmos graphing calculator is also built into the Bluebook testing app. The exam also provides a formula sheet and statistical tables.
The AP Statistics exam covers nine units: Exploring One-Variable Data, Exploring Two-Variable Data, Collecting Data, Probability and Random Variables, Sampling Distributions, Inference for Proportions, Inference for Means, Inference for Chi-Square, and Inference for Slopes. Exploring One-Variable Data is the most heavily weighted unit. Our practice test mirrors these official unit proportions so your practice matches the real exam.
Because AP Statistics rewards reasoning with data across nine units, the most effective preparation is repeated, unit-weighted multiple-choice practice under timed conditions, paired with free-response practice that asks you to justify conclusions. Read every explanation to learn the underlying reasoning, not just the answer. Reinforce weak units between sessions with a study guide, flashcards, and a cheat sheet.
Career Employer is the ultimate resource to help you get started working the job of your dreams. We cover topics from general career information, career searching, exam preparation with free study materials, career interviewing, and becoming successful in your career of choice.
Here at Career Employer, we focus a lot on providing factually accurate information that is always up to date. We strive to provide correct information using strict editorial processes, article editing, and fact-checking for all of the information found on our website. We only utilize trustworthy and relevant resources. To find out more, make sure to read our full editorial process page here.