Q & A for DS
Modelling Questions
1. Whats is the role od a data scientist in an organization?
✅ Click to read more
A data scientist bridges the gap between raw data and actionable business insights. They design experiments, build predictible model and communicate findings to non technical stakeholders to drive data informed decisions.
2. Explain the bias-variance trade-off
✅ Click to read more
Bias measures how far the model predictions are from the true values (underfitting risk), while variance measures sensibility to fluctuations in training data (overfitting risk). A polymonial degree-1 model on curved data has a high bias, a degree 20 model on noisy data has high variance. The sweet spot minimizes total error.
3. Whats cross validation and why should you use k fold instead of a single train-test split?
✅ Click to read more
Cross-validation repeatedly splits data into training and validation folds to estimate generalization error k-fold CV uses every sample for both training and validation across k rounds, giving a lower-variance performance estimate and making better use of limited data than a single split.
4.Walks through the key steps of data processing before modelling
✅ Click to read more
The pipeline typically covers:
- data collection & auditing
- handling missing values
- outline detection
- features encoding for categorical
- scaling/normalization
- feature engineering
- train-test splitting.
Each step reduces noise and ensures the model receives clean, comparable input.
5. Compare L1(Lasso) and L2 (Ridge) regularization - when would you choose each?
✅ Click to read more
L1 adds the sum of the absolute coefficient values as a penalty, shrinking irrelevant coefficients to exactly zero and producing sparse models - ideal for feature selection. L2 adds the sum of squared coefficients, distributing shrinkage across all features without zeroing them - preferred when all features may be relevant. Elastic combines both.
6. What strategies exist for handling missing data, and what are the trade-offs?
✅ Click to read more
Options include: - (1) listwise deletion -simple but loses information - (2) mean/ median/ mode imputation - fast but ignores uncertainty - (3) K- nearest neighbour imputation - captures local patterns - (4) multiple imputation - gold standard, preserves variability - (5) model- based imputation using algorithms such as a Random Forest.
The best choice depends on the missing data mechanism (MCAR, MAR or MNAR)
7. How does K-means clustering work, and what are the mains limitations?
✅ Click to read more
K means iteratively assigns each point to its nearest centroid, then recomputes centroids until convergence, minimizing within-cluster sum squares. Limitations: requires the number of clusters K to be specified in advance, assumes spherical clusters of similar size, is sensitive to initialization and ouliers, and struggles with non-convexed shapes.
8. What metrics would you use to evaluate a classification model, and when does accurancy mislead you?
✅ Click to read more
Key metrics: precision, recall, F1, score, ROC-AUC, and log loss, accuracy misleads when datasets are imbalanced - a model predicting the majority class for every observation can achieve 99%, accuracy on 99:1 data set while being useless. In such case F1 score or PR-AUC are more informative.
9. Interpret a confusion matrix and explain precision and recall are derived from it
✅ Click to read more
A confusion matrix cross-tabulates actual vs predicted classes. True Positives (TP) are correct positive predictions: False Negatives (FN) are missed positives, True Negatives (TN) are corrected negative predictions. Precisions = TP/ (TP +FP), Recall = TP / (TP + FN). Precisions answers how many flagged cases are real?, recall answers: ” how many real cases did we catch?”
10. What is feature scaling important, and which algorithms are most sensitive to it?
✅ Click to read more
Feature scaling ensures all features contribute equally to distance and gradient computations. Algorithms sensitive to scale includes K-nearest neighbors (uses Euclidean distance), SVM (maximizes margin in feature space), neural networks (gradient descent converges faster), and PCA (variance-based). Tree-based models such as Random Forest and XGBoost are scale-invariant.
11. How a Random Forest difer from a single Decision Tree, and what makes it more robust?
✅ Click to read more
A single decision tree is prone to over fitting because it memorises training data . Random Forest builds many trees on boostramp samples of the data at each split considers only a random subset of features (bagging + features randomness). The ensemble averages predictions, dramatically reducing variance without increasing bias.
12. Explain the difference between bagging and boosting
✅ Click to read more
Bagging (e.g.Random Forest) trains models independently in parallel on random data subsets and average results - primarily reducing variance. Boosting (e.g. XGBoosr, AdaBoost) trains models sequentially, with each learner focusing on the errors of its predecessor - primary reducing bias. Bagging is less prone to over fitting, boosting often achieves higher accuracy, but needs carefully tuning.
13. What is gradient descent, and how to do batch, mini-batch and stochastic variants differ
✅ Click to read more
Gradient descent iteratively updates model parameters by moving in the direction of the steepest loss decreases. Batch GD uses the full data set per update - stable but slow for large data. Stochastic GD updates after each sample - fast but noisy. Mini-batch GD (the standard in the deep learning) uses a small random batches, balancing speed and stability.
14. What is multicollinearity, how do you detect it, and what can you do about it?
✅ Click to read more
Multicollinearity occurs when predictor variables are highly correlated, inflating coefficient standard errors and making model interpretations unstable.
Detection methods include calculating Variance Inflation Factor (VIF > 10 indicates a problem), examining correlation heatmaps, and checking condition numbers.
To address it: remove one of the correlated features, combine them, apply PCA to create orthogonal components, or use regularised models like Ridge regression that handle collinearity effectively.
15. Explain the intuition behind Support Vector Machines (SVM)
✅ Click to read more
SVM aims to find the optimal hyperplane that separates classes by maximising the margin—the distance between the boundary and the nearest data points (support vectors). These support vectors define the decision boundary.
For non-linearly separable data, SVM uses kernel functions (e.g., RBF, polynomial) to map inputs into a higher-dimensional space where linear separation becomes possible. SVMs are effective in high-dimensional spaces and work well with smaller datasets.
16. How does the Naive Bayes classifier work, and why is the ‘naive’ assumption problematic?
✅ Click to read more
Naive Bayes applies Bayes’ theorem assuming that features are conditionally independent given the class label. It calculates posterior probabilities for each class and predicts the class with the highest probability.
The “naive” assumption is often unrealistic because features are usually correlated (e.g., words in text). Despite this, Naive Bayes performs well in practice, especially for high-dimensional problems like text classification, due to its simplicity and efficiency.
17. How do you handle outliers in a dataset?
✅ Click to read more
First, determine whether outliers are genuine data points or errors.
For genuine outliers: use robust models (e.g., tree-based) or robust loss functions (e.g., Huber).
For errors: remove them or cap values using winsorization.
Transformations such as log or square-root can reduce skewness. Always apply any treatment using training data only to avoid data leakage.
18. State the Central Limit Theorem and explain its practical significance in data science
✅ Click to read more
The Central Limit Theorem states that the sampling distribution of the sample mean approaches a normal distribution as sample size increases, regardless of the population distribution (assuming finite variance).
Practically, it justifies the use of normal-based statistical tests (z-tests, t-tests), supports confidence interval estimation, and explains why many real-world aggregated measurements tend to follow a normal distribution.
19. What is Principal Component Analysis (PCA) and when should you apply it?
✅ Click to read more
PCA is a dimensionality reduction technique that projects data onto orthogonal components capturing maximum variance. These components are linear combinations of the original features.
It is useful when features are highly correlated, when reducing dimensionality for visualisation (2D/3D), or when lowering computational cost. A key drawback is reduced interpretability since components are not directly tied to original features.
20. Explain ensemble learning and give examples of when it outperforms individual models
✅ Click to read more
Ensemble learning combines predictions from multiple base models to improve accuracy and robustness. It performs best when models have complementary strengths, when data is noisy (reducing variance via bagging), or when the problem is complex (reducing bias via boosting).
Common methods include Random Forest (bagging), Gradient Boosting (XGBoost, LightGBM), and stacking. Ensembles often dominate in real-world competitions and structured data problems.
21. What is XGBoost and what makes it more effective than traditional Gradient Boosting?
✅ Click to read more
XGBoost extends gradient boosting with several optimization: second-order Taylor approximation of the loss, built-in L1/L2 regularization, column and row sub sampling, efficient parallel tree construction, and native handling of missing values.
These enhancements make it faster, more regularized, and highly effective on tabular data sets.
22. How do you select the optimal number of clusters K in K-Means?
✅ Click to read more
Common approaches include: (1) Elbow method — plot inertia vs K and identify the point where improvement slows; (2) Silhouette score — measures cohesion and separation (higher is better); (3) Gap statistic — compares clustering performance to a reference null distribution.
Final choice should be validated with domain knowledge and interpretability.
23. Explain the purpose and mechanics of LSTM networks
✅ Click to read more
LSTMs are a type of recurrent neural network designed to capture long-term dependencies in sequential data. They address the vanishing gradient problem using gating mechanisms: forget gate removes irrelevant information, input gate controls what new information is stored, and output gate determines what is passed forward.
They are widely used for time series, text, and sequential signal modelling.
24. What is the difference between RMSE and MAE, and when would you prefer each?
✅ Click to read more
RMSE (Root Mean Squared Error) squares errors before averaging, penalising large errors more heavily. MAE (Mean Absolute Error) averages absolute differences, treating all errors equally.
Use RMSE when large errors are particularly undesirable (e.g., forecasting), and MAE when robustness to outliers is important. RMSE is smooth and easier to optimise, while MAE is more interpretable.
25. How do you handle imbalanced datasets in a classification problem?
✅ Click to read more
Strategies include: (1) Resampling — oversample minority class (e.g., SMOTE) or undersample majority; (2) Class weighting — penalise minority class errors more; (3) Threshold tuning — optimise decision threshold for F1 or recall; (4) Algorithm choice — use models robust to imbalance (e.g., XGBoost with scale_pos_weight).
Evaluation should use metrics like F1, PR-AUC, or MCC instead of accuracy.
26. What is the purpose of a recommendation system, and what are the main approaches?
✅ Click to read more
Recommendation systems predict user preferences to surface relevant items.
Main approaches: (1) Collaborative filtering — finds similar users/items or learns latent factors; (2) Content-based filtering — recommends based on item features and user history; (3) Hybrid methods — combine both to improve performance and handle cold-start problems.
27. Explain the vanishing gradient problem in deep networks and how modern architectures address it
✅ Click to read more
In deep networks, gradients shrink exponentially during back propagation, especially with sigmoid/tanh activation, preventing early layers from learning effectively.
Solutions include: (1) ReLU activation — reduces saturation; (2) Batch Normalization — stabilizes activation; (3) Residual connections (ResNet) — allow gradients to flow directly; (4) Proper weight initialization (He, Xavier).
28. What is overfitting, and what is your toolkit for combating it?
✅ Click to read more
Overfitting occurs when a model memorises training data and fails to generalise.
Mitigation techniques include: (1) More training data; (2) Cross-validation; (3) Regularization (L1/L2, dropout); (4) Early stopping based on validation loss; (5) Reducing model complexity; (6) Data augmentation.
29. Describe the concept of backpropagation in neural networks
✅ Click to read more
Back propagation computes the gradient of the loss with respect to each model weight using the chain rule of calculus, propagating errors backward from the output layer to the input layer.
The forward pass generates predictions, the backward pass calculates gradients, and an optimizer (e.g., SGD, Adam) updates weights in the direction that minimizes loss. This process repeats across many iterations (epochs).
30. What is the purpose of t-SNE, and how does it differ from PCA for visualization?
✅ Click to read more
t-SNE is a non-linear dimensional reduction technique designed for visualizing high-dimensional data in 2D or 3D. It preserves local structure, meaning nearby points in high dimensions stay close in the projection.
Unlike PCA (which preserves global variance through linear projections), t-SNE reveals clusters and manifold structure but is computationally expensive, non-deterministic, and does not preserve global distances between clusters.
31. What is feature importance in tree-based models, and how is it measured?
✅ Click to read more
Feature importance measures how much each feature contributes to model predictions. In tree-based models, it is typically computed as the total reduction in impurity (Gini or entropy) from splits involving that feature, weighted by sample counts.
Permutation importance is a model-agnostic alternative that measures performance drop when a feature is randomly shuffled, making it more reliable when features are correlated.
32. How do you diagnose and address multicollinearity in a regression model?
✅ Click to read more
Diagnosis includes examining correlation matrices and calculating Variance Inflation Factor (VIF). A VIF above ~5–10 indicates problematic multicollinearity.
Solutions include removing one of the correlated features, combining features, applying PCA for orthogonal components, or using regularised models like Ridge regression.
33. What is the Gini index and how is it used to build a Decision Tree?
✅ Click to read more
The Gini index measures node impurity: Gini = 1 − Σ(p²), where p represents class probabilities. A pure node has Gini = 0.
At each split, the algorithm evaluates all possible splits and selects the one that maximises the reduction in Gini impurity (i.e., increases node purity). Entropy (information gain) is an alternative criterion.
34. Explain the concept of R-squared and its limitations as a regression metric
✅ Click to read more
R² measures the proportion of variance in the target variable explained by the model: R² = 1 − (SS_res / SS_tot).
Limitations: it always increases with more predictors (use Adjusted R² instead), does not indicate statistical significance of features, and a high R² does not guarantee good generalisation on unseen data.
35. What is the Apriori algorithm, and where is it applied?
✅ Click to read more
Apriori identifies frequent itemsets in transactional data using the anti-monotonicity property: if an itemset is infrequent, all its supersets are also infrequent. This enables efficient pruning.
It generates association rules evaluated by support, confidence, and lift. A common application is market basket analysis (e.g., identifying products frequently bought together).
36. How do you handle highly skewed data distributions before modelling?
✅ Click to read more
Common approaches include: - (1) Log transformation — effective for right-skewed, positive data; - (2) Square-root or cube-root transformations — milder compression; - (3) Box-Cox transformation — a parameterised family including log as a special case; - (4) Yeo-Johnson — extends Box-Cox to zero and negative values; - (5) Quantile transformation — maps data to a uniform or normal distribution.
Always fit transformations on training data only and apply the same parameters to validation/test sets.
37. What is Mean Average Precision (MAP) and in what context is it used?
✅ Click to read more
MAP is a metric for ranked retrieval systems (e.g., search engines, recommenders). For each query, Average Precision (AP) sums precision at each rank where a relevant item is retrieved. MAP is the mean of AP across all queries.
It rewards systems that rank relevant results higher and penalises those that push them lower in the list.
38. Explain the Euclidean and Cosine distance metrics — when does each apply?
✅ Click to read more
Euclidean distance measures straight-line distance between two points and is appropriate when magnitude matters (e.g., K-means on continuous data).
Cosine similarity measures the angle between vectors, ignoring magnitude, making it ideal for high-dimensional sparse data like text (e.g., TF-IDF vectors in NLP).
39. How do you handle non-linearly separable data in classification?
✅ Click to read more
Approaches include: (1) Kernel SVM — maps data to higher-dimensional space using RBF, polynomial, or sigmoid kernels; (2) Feature engineering — add polynomial or interaction terms; (3) Non-linear models — Decision Trees, Random Forests, and neural networks learn complex boundaries directly; (4) Dimensionality reduction methods (t-SNE, UMAP) for visualising non-linear structure.
40. What is the Chi-square test, and how is it used in feature selection?
✅ Click to read more
The Chi-square test checks whether two categorical variables are independent by comparing observed vs expected frequencies.
In feature selection, a high Chi-square statistic (low p-value) between a feature and the target indicates strong dependence, meaning the feature is informative and should be retained.
41. What is ANOVA, and when would a data scientist apply it?
✅ Click to read more
ANOVA (Analysis of Variance) tests whether the means of three or more groups differ significantly. It compares between-group variance to within-group variance using the F-statistic.
It is used to compare model performance, evaluate whether a feature differs across target classes (feature selection), and validate A/B tests with more than two variants.
42. How do you handle time-series data in a machine learning pipeline?
✅ Click to read more
Key steps: (1) Use chronological train-test splits — never use future data to predict the past; (2) Engineer time features such as lags, rolling statistics, and seasonality indicators; (3) Choose appropriate models — ARIMA (stationary), Prophet (trend + seasonality), LSTM (long dependencies); (4) Evaluate using time-series cross-validation (walk-forward validation) instead of random k-fold.
43. What is the K-Nearest Neighbours (KNN) algorithm, and what are its computational challenges?
✅ Click to read more
KNN predicts a new point using the majority label (classification) or average value (regression) of its K nearest neighbours based on a distance metric (e.g., Euclidean, cosine).
Challenges: (1) Prediction time is O(n·d), scaling poorly with dataset size and dimensionality; (2) Curse of dimensionality makes distances less meaningful; (3) Sensitive to irrelevant features and scale, requiring preprocessing.
44. Explain Log Loss and why it is preferred over accuracy for probabilistic classifiers.
✅ Click to read more
Log Loss (cross-entropy) measures the negative log-probability assigned to the true label: −(y·log(p) + (1−y)·log(1−p)).
It penalises confident wrong predictions much more than uncertain ones, rewarding well-calibrated probabilities rather than just correct classifications. This makes it more informative than accuracy for probabilistic models.
45. What techniques address high-dimensional data (the ‘curse of dimensionality’)?
✅ Click to read more
High dimensionality leads to sparse data and less meaningful distances. Solutions include:
- Feature selection — remove irrelevant/redundant features (Lasso, mutual information, RFE);
- Dimensionality reduction — PCA (linear), t-SNE/UMAP (non-linear);
- Regularization — penalize model complexity;
- Domain knowledge — engineer meaningful low-dimensional features.
46. What is the F1 score, and when should you optimise for it over precision or recall alone?
✅ Click to read more
F1 score is the harmonic mean of precision and recall:
F1 = 2 × (Precision × Recall) / (Precision + Recall).
It balances both metrics and is useful when false positives and false negatives are equally important, especially in imbalanced datasets. If costs are asymmetric, use Fβ= (1+ β2) × (Precision × Recall) / (β2 X Precision + Recall). where β > 1 emphasizes recall and β < 1 emphasizes precision.
47. Explain hyperparameter tuning — what methods exist and how do you avoid overfitting to the validation set?
✅ Click to read more
Hyperparameters control model structure (e.g., learning rate, tree depth, regularization strength). Common tuning methods include: (1) Grid Search — exhaustive but expensive; (2) Random Search — more efficient and often comparable results; (3) Bayesian Optimization (e.g., Optuna, Hyperopt) — builds a probabilistic model to guide search; (4) Early stopping on a validation set. To avoid overfitting to the validation set, use nested cross-validation or keep a completely untouched final test set that is only evaluated once after tuning.
48. What is data leakage, and why is it one of the most dangerous pitfalls in machine learning?
✅ Click to read more
Data leakage occurs when information from outside the training dataset is used to build the model — for example, scaling using test-set statistics, including features derived from the target after the event, or improperly splitting time-series data.
Leaked models often appear to perform exceptionally well during training but fail in real-world deployment because they rely on information that would not be available in practice. Prevention involves fitting preprocessing steps strictly on training data and properly separating training, validation, and test sets.
49. What is the purpose of the logistic function in logistic regression?
✅ Click to read more
The logistic (sigmoid) function σ(z) = 1 / (1 + e^-z) maps any real-valued linear combination of features into the interval (0, 1), allowing outputs to be interpreted as probabilities.
The decision boundary is typically set where σ(z) = 0.5 (i.e., z = 0). The model is trained using binary cross-entropy (log loss), which is convex and enables efficient gradient-based optimisation.
50. How do you approach feature engineering, and what impact does it typically have on model performance?
✅ Click to read more
Feature engineering transforms raw data into more informative representations. Key techniques include: (1) interaction and polynomial features; (2) aggregations (mean, standard deviation across related records); (3) date/time decomposition (hour, day-of-week, is_weekend); (4) target encoding for high-cardinality categorical variables; (5) domain-specific ratios (e.g., debt-to-income). Strong feature engineering often has a larger impact on performance than model choice, especially for structured/tabular data.
51. What is the difference between supervised and unsupervised learning?
✅ Click to read more
Supervised learning trains a model on labelled data where the target variable is known, allowing the model to learn a mapping from inputs to outputs (e.g., classification, regression).
Unsupervised learning works with unlabeled data and identifies underlying patterns or structure (e.g., clustering, dimensionality reduction).
52. Explain the concept of gradient descent and its role in optimizing machine learning models
✅ Click to read more
Gradient descent is an optimisation algorithm that minimises a loss function by iteratively updating model parameters in the direction of the steepest decrease (negative gradient).
It is fundamental to training many models, including linear models and neural networks, enabling them to converge toward optimal parameters.
53. What is a Convolutional Neural Network (CNN), and how is it applied in image recognition tasks?
✅ Click to read more
A CNN is a deep learning architecture designed for processing grid-like data such as images. It uses convolutional layers to learn spatial hierarchies of features (edges → textures → objects).
CNNs automatically extract features from images and achieve high performance in tasks such as image classification, detection, and segmentation.
54. How would you handle overfitting in a machine-learning model?
✅ Click to read more
Overfitting occurs when a model memorises training data but performs poorly on unseen data.
Mitigation strategies include regularisation (L1/L2), early stopping, cross-validation, reducing model complexity, adding more data, and applying data augmentation.
55. Explain the concept of transfer learning and its advantages in machine learning
✅ Click to read more
Transfer learning reuses a model trained on a large dataset and adapts it to a related task.
It reduces training time, requires less data, and improves performance, especially when labelled data is limited.
56. How would you evaluate the performance of a machine learning model?
✅ Click to read more
Evaluation depends on the task. Classification metrics include accuracy, precision, recall, F1-score, and ROC-AUC. Regression metrics include MSE, RMSE, and MAE.
Cross-validation provides robust estimates, and curves such as ROC or precision-recall help analyse performance across thresholds.
57. How would you handle imbalanced datasets in machine learning?
✅ Click to read more
Imbalanced datasets have unequal class distributions. Techniques include oversampling the minority class (e.g., SMOTE), undersampling the majority class, using class weights, and adjusting decision thresholds.
Models should be evaluated with metrics such as F1, PR-AUC, or MCC rather than accuracy.
58. Explain the Central Limit Theorem and its significance in statistics
✅ Click to read more
The Central Limit Theorem states that the sampling distribution of the sample mean approaches a normal distribution as sample size increases, regardless of the original distribution (under finite variance).
It enables statistical inference, supports confidence intervals, and justifies hypothesis testing methods.
59. What is hypothesis testing, and how would you approach it for a dataset?
✅ Click to read more
Hypothesis testing is a statistical framework for making decisions about a population using sample data.
Steps:
- define null (H₀) and alternative (H₁) hypotheses
- choose a test statistic
- set significance level (α)
- compute p-value
- reject or fail to reject H₀ based on the result.
60. Explain the concept of correlation and its interpretation in statistics
✅ Click to read more
Correlation measures the strength and direction of the linear relationship between two variables. It ranges from −1 to +1, where −1 indicates perfect negative correlation, +1 indicates perfect positive correlation, and 0 indicates no linear relationship.
The correlation coefficient helps quantify how strongly variables move together, but it does not imply causation.
61. What are confidence intervals, and how do they relate to hypothesis testing?
✅ Click to read more
A confidence interval provides a range of plausible values for a population parameter based on sample data. For example, a 95% confidence interval means that, over repeated samples, 95% of such intervals would contain the true parameter.
They relate to hypothesis testing because if a hypothesised value (e.g., μ₀) lies outside the interval, the null hypothesis would be rejected at the corresponding significance level.
62. What is the difference between Type I and Type II errors in hypothesis testing?
✅ Click to read more
A Type I error occurs when a true null hypothesis is incorrectly rejected (false positive), while a Type II error occurs when a false null hypothesis is not rejected (false negative).
Type I error is controlled by the significance level (α), whereas Type II error is related to statistical power (1 − β), which reflects the ability to detect real effects.
63. How would you perform hypothesis testing for comparing two population means?
✅ Click to read more
To compare two population means, use a t-test: an independent t-test for two separate groups, or a paired t-test for dependent samples.
Steps: - (1) define null (means are equal) and alternative hypotheses; - (2) choose the appropriate test; - (3) compute the test statistic; - (4) calculate p-value; - (5) decide whether to reject the null hypothesis based on α.
Assumptions include normality (or large samples) and, for independent t-tests, equal variance (or use Welch’s t-test).
64. Explain the concept of p-value and its interpretation in hypothesis testing
✅ Click to read more
The p-value is the probability of observing data as extreme as (or more extreme than) the observed result, assuming the null hypothesis is true.
A small p-value indicates strong evidence against the null hypothesis. If the p-value is less than the significance level (α), the null hypothesis is rejected.
65. What is ANOVA (Analysis of Variance), and when is it used in statistical analysis?
✅ Click to read more
ANOVA is used to compare the means of three or more groups to determine whether at least one group differs significantly.
It works by partitioning total variance into between-group and within-group components and evaluating their ratio using the F-statistic. ANOVA is commonly used in experiments, A/B/n testing, and feature analysis.
66. Imagine you are tasked with improving the search functionality of a search engine like Google. How would you approach this challenge?
✅ Click to read more
Improving search involves understanding user intent, analysing queries and behaviour, and improving relevance ranking.
Techniques include natural language processing (NLP), query expansion, semantic search, and learning-to-rank models. Continuous A/B testing, feedback loops, and monitoring metrics such as click-through rate (CTR) and dwell time are critical for iterative improvement.
67. How would you measure the impact and success of a new feature release in a mobile app?
✅ Click to read more
Measure success using metrics aligned with the feature’s goal, such as adoption rate, engagement (time spent, usage frequency), retention, and conversion.
Use A/B testing to compare users exposed to the feature versus a control group, and complement quantitative metrics with qualitative feedback (reviews, surveys).
68. Suppose you are tasked with improving the user onboarding process for a software platform. How would you approach this?
✅ Click to read more
Start by identifying user pain points through analytics and user research. Simplify workflows, provide guided tutorials, tooltips, and progressive onboarding experiences.
Measure success using activation rate, drop-off points, and time-to-value, and iterate based on user feedback and behavioural data.
69. How would you prioritize and manage multiple concurrent data science projects with competing deadlines?
✅ Click to read more
Prioritization should be based on business impact, urgency, dependencies, and resource availability. Break projects into milestones and use frameworks like Agile for iterative delivery.
Clear stakeholder communication, expectation setting, and regular progress tracking ensure deadlines are met effectively.70. Suppose you are asked to design a fraud detection system for an online payment platform. How would you approach this task?
✅ Click to read more
Design involves combining supervised learning (fraud classification) with anomaly detection for unseen patterns. Engineer features such as transaction amount, frequency, location, device, and user behaviour.
Handle class imbalance, use real-time scoring, and deploy monitoring systems. Continuously retrain models using new data and collaborate with domain experts to refine rules and improve detection accuracy.
Database Fundamentals
1. What is a primary key in a database?
✅ Click to read more
A primary key is a unique, non-null identifier for each record in a table. It ensures that every row can be uniquely identified and prevents duplicate entries. Primary keys are fundamental for maintaining data integrity and enabling efficient indexing and relationships between tables.
2. What is a foreign key in a database?
✅ Click to read more
A foreign key is a column that references the primary key of another table, establishing a relationship between tables.
It enforces referential integrity, ensuring that values in the foreign key column correspond to valid entries in the referenced table.
3. What is ETL?
✅ Click to read more
ETL stands for Extract, Transform, Load. It is a pipeline process where data is extracted from source systems, transformed into a usable format, and loaded into a data warehouse or database.
It is essential for integrating data from multiple sources and preparing it for analysis.
4. What is data analytics?
✅ Click to read more
Data analytics is the process of collecting, cleaning, transforming, and analysing data to extract insights, support decision-making, and drive business outcomes. It combines statistical, computational, and domain knowledge techniques.
5. What are the types of data analytics?
✅ Click to read more
The main types are:
Descriptive — summarises past data Predictive — forecasts future outcomes Prescriptive — recommends actions
Each type builds on the previous to provide increasing value.
✅ Click to read more
A pivot table is a data summarisation tool used to reorganise, filter, and aggregate data (e.g., sums, counts, averages).
It is widely used in spreadsheets and BI tools for quick exploratory analysis.
7. What is data normalization?
✅ Click to read more
Data normalization is the process of structuring or scaling data to reduce redundancy and improve consistency.
In databases, it removes duplication (normal forms). In machine learning, it scales features to a standard range.
8. Explain the concept of data warehousing
✅ Click to read more
A data warehouse is a centralised repository that stores integrated data from multiple sources, optimised for querying and analysis.
It enables reporting, dashboards, and historical analysis across an organisation.
Descriptive Stats & Viz
1. Explain the difference between qualitative and quantitative data
✅ Click to read more
Qualitative data is non-numeric (e.g., categories, text), while quantitative data is numeric (e.g., counts, measurements).
Quantitative data is typically easier to analyse statistically, while qualitative data provides richer context.
2. What is data cleaning?
✅ Click to read more
Data cleaning is the process of identifying and correcting errors such as missing values, inconsistencies, incorrect formats, and duplicates.
It is a critical preprocessing step to ensure data quality and reliable model performance.
3. What is an outlier?
✅ Click to read more
An outlier is a data point that significantly differs from the rest of the dataset.
Statistically, it is often defined as a value far from the mean (e.g., more than 3 standard deviations). Outliers may indicate errors or meaningful rare events.
4. What is a histogram?
✅ Click to read more
A histogram is a graphical representation of data distribution using bins to show the frequency of values. It helps identify patterns such as skewness, spread, and modality.
✅ Click to read more
A box plot visualises data distribution using quartiles, showing the median, interquartile range (IQR), and potential outliers.
It is useful for comparing distributions across groups.
✅ Click to read more
Correlation indicates a statistical relationship between two variables, while causation means one variable directly affects another.
Correlation alone does not imply causation due to confounding variables or coincidence.
✅ Click to read more
Missing data can be handled by: (1) Imputation (mean, median, mode); (2) Removing rows/columns; (3) Using models that handle missing values; or (4) Adding missing indicators.
The best approach depends on the context and the reason data is missing.
8. How do you deal with outliers in a dataset?
✅ Click to read more
Outliers can be handled by removing them, transforming the data (log, scaling), capping values (winsorization), or using robust models less sensitive to extreme values.
The first step is always to determine whether the outlier is an error or meaningful signal.