Data Science Interview Questions
Statistics, hypothesis testing, the bias-variance tradeoff, feature engineering, model evaluation, regularization, PCA, A/B testing and experiment design.
61 questions
Popular Searches
1. What is data science and how does it differ from data analytics and machine learning?
easyFoundations 6 mins2. What is the difference between supervised, unsupervised, and reinforcement learning?
mediumMachine Learning 7 mins3. What is the bias-variance tradeoff in data science?
mediumModel Evaluation 7 mins4. What is overfitting and how do you prevent it?
mediumModel Evaluation 7 mins5. What is the difference between correlation and causation in data analysis?
easyStatistics 6 mins6. What is feature engineering and why is it important?
mediumMachine Learning 7 mins7. What is the difference between a population and a sample in statistics?
easyStatistics Fundamentals 6 mins8. What is a p-value and how do you interpret it?
mediumHypothesis Testing 7 mins9. What is the difference between Type I and Type II errors?
mediumHypothesis Testing 7 mins10. What is the Central Limit Theorem and why does it matter for data science?
mediumStatistics 7 mins11. What is the difference between mean, median, and mode, and when do you use each?
easyStatistics 6 mins12. What is standard deviation and how does it differ from variance?
easyStatistics 6 mins13. What is a confidence interval and how do you interpret it?
mediumStatistical Inference 7 mins14. What is hypothesis testing and what are the null and alternative hypotheses?
mediumStatistical Inference 7 mins15. What is the difference between classification and regression problems?
easyMachine Learning 6 mins16. What are precision, recall, and F1-score, and when do you use each?
mediumModel Evaluation 7 mins17. What is a confusion matrix and what does it tell you?
easyModel Evaluation 6 mins18. What is the ROC curve and what does AUC represent?
mediumModel Evaluation 7 mins19. What is cross-validation and why is it used?
mediumModel Evaluation 7 mins20. What is the difference between L1 and L2 regularization?
mediumRegularization 7 mins21. What is dimensionality reduction and how does PCA work?
mediumUnsupervised Learning 8 mins22. What is the curse of dimensionality in data science?
mediumFeature Engineering 7 mins23. How do you handle missing data in a dataset?
mediumData Cleaning 7 mins24. How do you handle imbalanced datasets in classification?
mediumClassification 7 mins25. What is the difference between bagging and boosting in ensemble learning?
mediumEnsemble Learning 7 mins26. What is A/B testing and how do you design a valid experiment?
mediumExperimentation 7 mins27. What is the difference between a data warehouse and a data lake?
easyData Architecture 7 mins28. What is normalization vs standardization of features and when do you use each?
mediumFeature Scaling 7 mins29. What is exploratory data analysis (EDA) and what does it involve?
easyData Analysis 7 mins30. What is the difference between batch and online learning in data science?
mediumModel Training 7 mins31. What is Data Science?
easyFundamentals 5 mins32. Data Scientist vs Data Analyst: What's the Difference?
easyFundamentals 4 mins33. What is Exploratory Data Analysis (EDA)?
easyFundamentals 5 mins34. What is Feature Scaling?
easyFeature Engineering 5 mins35. What is a P-Value?
easyStatistics 6 mins36. What is Correlation?
easyStatistics 5 mins37. What is a Normal Distribution?
easyStatistics 5 mins38. What is Sampling in Statistics?
mediumStatistics 6 mins39. What is a Hypothesis Test?
mediumStatistics 7 mins40. What Is Regression in Data Science?
mediumStatistical Modeling 6 mins41. What Is ETL (Extract, Transform, Load)?
mediumData Engineering 6 mins42. What Is Data Wrangling?
mediumData Preparation 6 mins43. What Is a DataFrame?
mediumData Structures 5 mins44. What Is Imputation in Data Science?
hardData Preprocessing 7 mins45. What Is Standardization in Data Science?
hardFeature Scaling 7 mins46. What Is Dimensionality Reduction?
hardFeature Engineering 7 mins47. What Is a KPI (Key Performance Indicator)?
hardBusiness Analytics 6 mins48. How Does Data Leakage Creep Into a Cross-Validation Pipeline?
hardModel Validation 9 mins49. Why Do Random K-Fold Splits Mislead on Time-Series Data?
hardModel Validation 9 mins50. How Do You Decide the Sample Size an Experiment Needs?
mediumExperimentation 8 mins51. How Do You Handle the Multiple Comparisons Problem?
mediumStatistical Inference 8 mins52. How Do You Choose an Evaluation Metric for a Rare Event?
mediumModel Evaluation 8 mins53. What Is the Difference Between Calibration and Discrimination?
hardModel Evaluation 9 mins54. How Do You Reason About Confounding When You Cannot Run an Experiment?
hardCausal Inference 10 mins55. What Are the Most Common Ways an A/B Test Result Turns Out to Be Wrong?
hardExperimentation 10 mins56. How Does the Missing-Data Mechanism Change Your Imputation Choice?
mediumData Preparation 9 mins57. What Is Training/Serving Skew and How Does a Feature Store Reduce It?
hardMachine Learning Engineering 10 mins58. How Do You Detect and Respond to Drift in a Deployed Model?
hardMachine Learning Engineering 10 mins59. How Would You Explain a Model to a Non-Technical Stakeholder?
mediumCommunication 8 mins60. What Makes a Data Science Analysis Reproducible?
mediumWorkflow and Tooling 9 mins61. How Do You Tell a Stakeholder the Data Cannot Answer Their Question?
mediumCommunication 8 mins