Mock Interview STAT

Author

Cecilia Regueira

Top 10 Statistics Interview Questions

Note

💡 Mock

This document presents the top 10 statistics questions and answers that most frequently come up in job interviews. These questions are based on interviews I have personally participated in, either as a candidate or as an interviewer when assessing new hires. The goal is to highlight the statistical concepts that employers commonly expect candidates to understand.

1. What is the difference between descriptive and inferential statistics?

✅ Click to reveal answer

Descriptive statistics summarize and describe the main features of a dataset (e.g., mean, median, standard deviation, plots).
Inferential statistics use a sample to make conclusions or predictions about a larger population (e.g., confidence intervals, hypothesis tests).


2. What is the difference between correlation and causation?

✅ Click to reveal answer

Correlation measures the strength and direction of a relationship between two variables.
Causation means that changes in one variable directly cause changes in another.
Correlation alone does not imply causation; establishing causation typically requires controlled experiments, causal models, or strong assumptions.


3. What is a p-value and how should it be interpreted?

✅ Click to reveal answer

A p-value is the probability of observing results as extreme as (or more extreme than) the data, assuming the null hypothesis is true.
A small p-value suggests the data are unlikely under the null hypothesis, but it does not measure the probability that the null hypothesis is true.


4. What is the difference between Type I and Type II errors?

✅ Click to reveal answer
  • Type I error: Rejecting a true null hypothesis (false positive).
  • Type II error: Failing to reject a false null hypothesis (false negative).
    Which error is more costly depends on context (e.g., medical testing vs. fraud detection).

5. What assumptions are required for linear regression?

✅ Click to reveal answer

Key assumptions include:

  • Linearity

  • Independence of errors

  • Homoscedasticity (constant variance)

  • Normality of errors (for inference)

These can be checked using residual plots, QQ plots, and diagnostic tests.


6. What is the Central Limit Theorem and why is it important?

✅ Click to reveal answer

The Central Limit Theorem (CLT) states that the sampling distribution of the sample mean approaches a normal distribution as sample size increases, regardless of the population distribution.
It enables statistical inference using normal-based methods, even when the data are not normally distributed.


7. What is a confidence interval and how do you interpret it?

✅ Click to reveal answer

A confidence interval (CI) provides a range of plausible values for an unknown population parameter.
A 95% CI means that if we repeated the sampling process many times, about 95% of those intervals would contain the true parameter.


8. What is overfitting and how can it be prevented?

✅ Click to reveal answer

Overfitting occurs when a model captures noise instead of the underlying pattern, performing well on training data but poorly on new data.
It can be prevented using techniques like cross-validation, regularization, simpler models, and early stopping.


9. How would you handle missing data in a dataset?

✅ Click to reveal answer

Common approaches include: - Deleting missing observations (when missingness is minimal)
- Mean/median imputation
- Model-based or multiple imputation
The choice depends on the amount and mechanism of missingness (MCAR, MAR, MNAR).


10. How do you evaluate the performance of a statistical or predictive model?

✅ Click to reveal answer

Model performance can be evaluated using: - Regression metrics: RMSE, MAE, R²
- Classification metrics: accuracy, precision, recall, F1-score, AUC
Evaluation should be done on validation or test data to assess generalization.


11. What happens to a confidence interval when the sample size doubles?

✅ Click to reveal answer

When the sample size doubles, the width of the confidence interval decreases, because the standard error scales with (1 / ).
Doubling the sample size reduces the margin of error by approximately 29%, making the estimate more precise.


Bonus


Question 1: Statistics Foundations

You are analyzing ride data for Uber.

  • The probability that a user takes a ride in the morning rush hour is 10%.
  • The probability that a user takes a ride in the evening rush hour is 20%.
  • The probability that a user takes an evening rush hour ride given that they took a morning rush hour ride is 50%.

If we observe that a user took an evening rush hour ride, what is the probability that they also took a morning rush hour ride?

✅ Click to reveal answer

Let:

MMM = event that the user took a morning rush hour ride EEE = event that the user took an evening rush hour ride

Given:

  • \(P(M)=0.10\)

  • \(P(E)=0.20\)

  • \(P(E∣M)=0.50\)

Step 1: Use Bayes’ Theorem

  • \(P(M \mid E)=\frac{P(E \mid M) * P(M)}{P(E)}​\)

  • \(P(M \mid E) = \frac{0.5 * 0.10}{0.20} = 0.25\)

This means that given a user took an evening rush hour ride, there is a 25% probability that they also took a morning rush hour ride.

Question 2: Experiment Design & Power Analysis

When designing an experiment, given a fixed sample size, significance level, and desired power, we compute a Minimum Detectable Effect (MDE).

Suppose you run the experiment, collect the data, and observe an effect size that is smaller than the MDE, yet the result is statistically significant.

How is this possible? Shouldn’t effects smaller than the MDE be non‑significant?

✅ Click to reveal answer

This situation is possible and not a contradiction. The key reason is that the Minimum Detectable Effect (MDE) is defined in terms of statistical power, not statistical significance.

  1. What the MDE Actually Means?

The MDE is the smallest effect size that your experiment is designed to detect with a specified probability (power), usually 80% or 90%, if that effect is truly present. In other words:

If the true effect equals the MDE, you have (for example) an 80% chance of obtaining a statistically significant result.

It does not mean:

Effects smaller than the MDE can never be significant

  1. Statistical Significance Is Random Hypothesis testing involves sampling variability. Even if the true effect is smaller than the MDE:

Random variation may produce a larger observed effect The standard error may be smaller than expected The test statistic may cross the significance threshold

As a result, you can still obtain a statistically significant p‑value.

  1. Power Is About Probability, Not Certainty If the true effect is below the MDE:

The probability of detection is less than the target power But that probability is not zero

So significance can happen — just less frequently.

  1. Relationship Between Effect Size, Power, and Significance
Concept What it Represents
Statistical significance A single realization: did this sample cross the significance threshold?
Power The long-run probability of detecting a true effect when it exists
Minimum Detectable Effect (MDE) The smallest effect size that can be detected with a specified power
Observed effect A random estimate of the effect size obtained from sampled data

A statistically significant result answers:

“Is this result unlikely under the null hypothesis?”

The MDE answers:

“How likely am I to detect this effect if it is truly this large?”

These are related but not the same question.

Question 3: Causation

Amazon Prime members place an average of 10 orders per month, while non‑Prime users place 2 orders per month.

Does Amazon Prime membership cause users to place more orders? Explain your reasoning.

✅ Click to reveal answer

There is selection bias.

Prime members self‑select into the program, and their higher number of orders may reflect pre‑existing shopping behavior rather than a causal effect of Prime membership.

To establish causality, we would need a randomized experiment or a credible causal identification strategy.