← ALL_POSTS
TechnicalSeptember 28, 2026 · 12 min read

Data Science Interviews: Statistics, Machine Learning and Cases

Published by Ansh Modi · Vrenic

Preparing for a data science interview often feels like studying for three different degrees at once. The role title means entirely different things across the industry, leaving candidates guessing whether they will face a software engineering test, a pure statistics exam, or a high-level business strategy discussion. This ambiguity causes severe anxiety for candidates who are short on time and cannot afford to revise the entire machine learning syllabus. Understanding exactly what each round evaluates helps you direct your practice time toward the specific concepts and frameworks that actually determine a pass or fail.

The structure of a data science loop

Data science roles span a wide spectrum between product analytics and production algorithms. A product data scientist role leans heavily toward experiment design, setting up A/B tests and defining the business metrics that guide product decisions. Conversely, a machine learning engineer role focuses on training efficient algorithms, managing computational complexity and writing production-ready code. If the role involves building systems that serve predictions in real time, you may also face rounds similar to system design interviews, testing your ability to scale architecture.

You will usually face an initial technical screen followed by an on-site loop of three or four specialised rounds. The initial screen often involves SQL and basic Python logic to ensure you can manipulate data sets independently. Once in the loop, the statistical foundations round evaluates your grasp of probability, hypothesis testing and experiment design. The machine learning theory round tests your mathematical understanding of algorithms, checking that you know how models work under the hood rather than just how to import them from a library.

Finally, the modelling or product sense case round asks you to solve an open-ended business problem from scratch. Interviewers use these distinct rounds to build a complete profile of your technical depth and your commercial pragmatism. Passing the loop requires demonstrating that you can formulate rigorous mathematical solutions while maintaining a clear view of the end user's problem. Preparing effectively means knowing the specific mechanics each round tests and tailoring your answers accordingly.

Statistics, probability and experiment design

The statistical foundations round ensures you can make rigorous decisions using uncertain data. A major focus is always experiment design, specifically the nuances of A/B testing in a commercial environment. Interviewers will ask how you choose the correct test statistic, how you calculate a required sample size, and how long you should run an experiment to account for weekly seasonality. You must be able to state the null and alternative hypotheses clearly for any testing scenario they present.

You must also be ready to explain p-values and confidence intervals in plain language to a non-technical stakeholder. A common interview trap involves asking you to interpret a p-value as the probability that the null hypothesis is true, which is mathematically incorrect. You must clearly state that a p-value represents the probability of observing the data, or something more extreme, assuming the null hypothesis is completely true.

Experiment design questions often test your awareness of common pitfalls, such as network effects or the novelty effect. If you are testing a new messaging feature, users in the treatment group might message users in the control group, contaminating the experiment. A strong candidate will explicitly identify this interference and suggest solutions, such as randomising at the cluster level rather than the individual user level.

Conditional probability questions are another staple of this round, designed to test foundational mathematics and logical reasoning under pressure. Interviewers are looking for explicit steps and clearly stated assumptions, not just a final number. You must articulate your thought process out loud as you work through the formula, writing out Bayes' Theorem or the Law of Total Probability in plain text. For example, if asked for the probability of a user clicking an advert given they are on a mobile device, write out the formula before plugging in numbers, showing every step: P(Click | Mobile) = [P(Mobile | Click) * P(Click)] / P(Mobile).

Be meticulous with algebraic steps and assumptions. If an interviewer asks you to calculate a compound probability, clearly define your inputs as variables before substituting any values to avoid confusing the interviewer. Detail the assumptions you make during the calculation to show a rigorous mathematical approach. You must explicitly state whether you assume variables are independent, rather than treating their independence as an unstated fact about the world.

Worked example: designing an A/B test

To illustrate how the statistics round operates, consider a common experiment design question: evaluating a new checkout flow for an e-commerce platform. The prompt tests your ability to structure a rigorous experiment while avoiding statistical traps.

A weak candidate treats this as a trivial exercise in comparing two averages without establishing constraints:

"I would randomly assign half the users to the old checkout and half to the new checkout. I will let the test run for a few days to gather enough data. Once the test finishes, I would look at the conversion rate for both groups. If the new checkout has a higher conversion rate, we launch it to everyone, and if it is lower, we stick with the old one."

This answer fails because it lacks statistical rigour and ignores business context. The candidate does not define the significance level or power of the test, leaving the duration to guesswork. They also fail to account for seasonality, novelty effects, or how the randomisation will actually be implemented on the backend.

A strong candidate sets up the experiment with explicit mathematical constraints and commercial considerations:

"First, we must define the primary metric, which is the conversion rate, and a guardrail metric, such as average order value, to ensure we do not increase conversions at the expense of revenue. I will define a baseline conversion rate of p and set a minimum detectable effect of d. We will set the significance level at alpha and statistical power at 1 - beta.

Using these parameters, I would calculate the required sample size before starting the test. I will assume we get N daily visitors, which means the test needs to run for T full weeks to capture our required sample while accounting for day-of-week seasonality.

We will randomise at the user ID level to ensure a consistent experience across sessions. Once the test concludes, we will calculate the p-value. If it falls below our threshold, and our guardrail metric remains stable, we can confidently roll out the new checkout."

The strong answer succeeds because it formalises the experiment using standard statistical definitions. The candidate explicitly states their assumed baseline, lift, power and significance thresholds as variables, showing they understand the mathematical inputs required. They also address practical concerns by mentioning a guardrail metric and calculating a realistic duration that accounts for natural traffic cycles.

Machine learning concepts that always recur

The machine learning theory round separates candidates who understand how algorithms function from those who simply call library functions. You will rarely need to derive complex calculus on a whiteboard. Instead, interviewers test your intuition for the tradeoffs inherent in statistical modelling. The most common of these is the bias-variance tradeoff, which is central to model generalisation.

You must be able to explain the bias-variance tradeoff simply and practically. High bias means a model makes strong assumptions and underfits the training data, while high variance means a model captures noise and overfits. A strong answer explains how changing specific hyperparameters pushes the model along this spectrum. For instance, you should explain that increasing the maximum depth of a decision tree increases variance, while raising the regularisation penalty in logistic regression increases bias.

Interviewers also look for a deep understanding of evaluation metrics, testing whether you know exactly when to apply precision, recall, and the F-score. Stating that you would use accuracy is almost always a trap in an interview setting, because real-world datasets are rarely balanced. You need to identify whether the business problem requires optimising for false positives or false negatives. For a medical diagnosis model, missing a positive case is dangerous, meaning recall is paramount, whereas for a spam filter, a false positive frustrates users, making precision the priority.

Expect questions that test your ability to spot and prevent data leakage. Data leakage occurs when information from outside the training dataset is used to create the model, leading to artificially high performance during training that immediately collapses in production. You must demonstrate how to prevent this by strictly separating the training and test sets before applying any transformations. A practical example is imputing missing values: calculate the mean for imputation using only the training data, applying that exact same mean to the test set to avoid leaking future information.

Structuring a modelling case answer

The modelling case round presents a vague business problem and asks you to build a machine learning solution from the ground up. This is a test of structure as much as technical knowledge, as the interviewer wants to see how you translate a commercial objective into a mathematical framework. A disorganised answer that jumps straight into deep neural networks will fail, regardless of how advanced the mathematics might be.

Start by clarifying the business goal and defining the target variable. You must establish exactly what you are trying to predict and how the business will use that prediction to drive value. Ask clarifying questions to narrow the scope and explicitly state the definition of the target variable. Next, discuss data collection and feature engineering by proposing specific, realistic features based on the telemetry or transaction data the business would logically have available.

Instead of vaguely mentioning transaction history, specify that you would calculate the rolling average of basket size over the past three months. Explain how you would handle missing telemetry data, perhaps by creating a binary indicator feature that flags when a reading is absent. Only after defining the problem and detailing the features should you select a baseline model. Always start with the simplest algorithm that could possibly work, such as logistic regression for classification or linear regression for predicting continuous values.

Explain why this baseline makes sense and how it provides an interpretable benchmark. You should then discuss a more complex model, such as a gradient boosting machine, that you would test next if the baseline underperforms. Practising this structure requires repetition and feedback on your communication style. You can practise open-ended questions using Vrenic's Technical interview type on the Free plan, which provides a model answer and written feedback for every response.

Speaking out loud helps you build the habit of signposting your logical steps clearly. This ensures the interviewer can easily follow your framework from the initial business problem through to the final evaluation metric.

Worked example: predicting subscription churn

To see how this framework applies in practice, consider a classic modelling case: predicting subscription churn for a video streaming platform. The prompt is deliberately broad, giving you room to define the scope or trap yourself in unnecessary complexity.

A weak candidate treats this as a pure algorithmic challenge rather than a business problem:

"To predict churn, I would use a Random Forest or a gradient boosting model, because tree-based ensembles handle non-linear relationships well. I would feed it the user's demographic data, their viewing history, and their payment information, before splitting the data into a training set and a test set. Then, I would tune the hyperparameters using grid search to maximise the model's accuracy. Once the model is validated, we can deploy it and target the users who are predicted to leave."

This answer fails because it skips the definition phase entirely. The candidate does not define what churn actually means in this context, nor do they specify what the target variable looks like. They jump straight to complex models and choose accuracy as a metric, which is deeply flawed for an imbalanced dataset where most users do not churn in any given month.

A strong candidate approaches the same prompt by establishing a rigorous framework:

"First, we need to define the business objective and the target variable. I will assume the goal is to identify users at risk of cancelling so the marketing team can offer them a retention discount. Let us assume churn means a user cancels their paid subscription within the next thirty days. Our target variable is binary: one if they cancel in that window, zero otherwise.

For features, demographic data is useful, but behavioural data is highly predictive. I would engineer features like the number of days since the user last watched a video, the rolling average of watch time over the last two weeks, and whether their payment method is expiring soon.

I would start with a logistic regression baseline because it is interpretable, which helps the marketing team understand the risk factors. The dataset will be highly imbalanced, so accuracy is a poor metric. I would evaluate the model using the area under the precision-recall curve. We want to balance precision, to avoid wasting money on discounts for users who were staying anyway, with recall, ensuring we actually catch the users who are leaving."

The strong answer succeeds because it links every technical decision back to the commercial goal. The candidate explicitly defines the target variable before mentioning any algorithms. They suggest concrete, engineered features rather than vague data categories. Finally, they choose a sensible baseline and select an evaluation metric perfectly suited to an imbalanced classification problem.

Before your next data science round

Success in a data science loop comes down to structured communication. The interviewer needs to see that you can apply theory to messy, real-world problems without losing sight of the commercial objective.

  • Review core metric formulas and write them out by hand. You must be able to define precision, recall and the F-score from memory, without relying on external references.
  • State your assumptions explicitly during mathematical calculations. When solving probability questions, verbalise whether you are assuming independence or mutually exclusive events.
  • Practise case frameworks out loud. Take a broad business problem and walk through defining the target variable, brainstorming features, choosing a baseline model and setting an evaluation metric.
  • Tie every algorithm back to its business impact. Never suggest a complex model without first explaining the simple baseline that justifies its use, ensuring you always ground your mathematics in pragmatism.
  • Test your delivery under pressure. You can run unlimited interviews on Vrenic's Pro plan, answering by speaking to ensure your spoken structure is as rigorous as your mathematical logic.

Put Technical interview questions in front of the same rubric this article describes.

Practice a Technical interview →