20+ Questions to Test Your Skills on Logistic Regression
Logistic regression is a fundamental supervised learning algorithm for solving binary classification problems. While one of the easier machine learning techniques to understand, logistic regression has surprising depth and nuance. This post will challenge your knowledge with 20+ questions spanning core concepts, intuition, complexity analysis, and more. Whether you‘re preparing for data science interviews or want to solidify your understanding, this guide has you covered.
But first, let‘s start with the basics…
1. What is logistic regression?
Logistic regression is a classification algorithm that models the probability an input belongs to the default class (usually labeled 1). It‘s called "regression" because it fits a linear equation to the log-odds, or logit, of the data. But unlike linear regression which outputs continuous values, logistic regression transforms its output using the logistic sigmoid function to return a probability value.
2. Why is logistic regression so popular?
Logistic regression‘s popularity stems from a few factors:
- It‘s a simple and interpretable model
- Training is fast and efficient, even with large datasets
- Outputs well-calibrated probabilities useful for ranking examples
- Provides good results out-of-the-box without much tuning
- Works well when classes are linearly separable
As a result, logistic regression is often the first tool to reach for when approaching a new classification problem. More powerful models like neural networks and decision forests may yield better performance, but logistic regression remains a go-to baseline.
3. What are the types of logistic regression?
There are three main flavors of logistic regression:
-
Binary – The target variable has two possible classes, e.g. pass/fail, healthy/sick, click/no-click. This is the most common form.
-
Multinomial – The target variable has three or more unordered categories, e.g. food preference (vegetarian, vegan, meat-eater).
-
Ordinal – The target variable has three or more ordered categories, e.g. movie ratings, agreement scales (disagree, neutral, agree).
Each type requires a different loss function and final activation but otherwise follows a similar training procedure. We‘ll focus on binary logistic regression here, but many of the same concepts apply to the others.
4. How does logistic regression work under the hood?
At a high level, logistic regression:
1. Takes an input vector x
2. Computes a weighted sum z = wTx + b
3. Applies the sigmoid function to z to get the estimated probability y_hat = σ(z)
The sigmoid, defined as σ(z) = 1 / (1 + exp(-z)), squashes the output to be between 0 and 1.
During training, logistic regression learns the weights w and bias b that maximize the likelihood of the data (i.e. cross-entropy loss). This is typically done via gradient descent optimization.
At inference time, a decision boundary of 0.5 is commonly used. An input x is classified as class 1 if the estimated probability σ(wTx + b) > 0.5, and class 0 otherwise. This boundary can be shifted to trade-off precision and recall.
5. What are odds and what do they represent?
Odds represent the ratio of the probability an event happens to the probability it doesn‘t happen. If the probability of rain is 60%, then the odds of rain are 0.6 / (1 – 0.6) = 1.5, or 3:2. An odds > 1 means the event is more likely to occur than not.
The logit function, log(p / (1-p)), transforms probability into log-odds (a value between -inf and +inf). Logistic regression directly estimates the log-odds, which it then converts to probability via the sigmoid / logistic function.
6. Is the logistic regression decision boundary linear or nonlinear?
The decision boundary of logistic regression is linear. Recall logistic regression estimates the probability p(y=1|x) by fitting a linear equation wTx+b to the logit of the data. Applying the sigmoid function to this gives the estimated probability.
The decision boundary occurs where p(y=1|x) = p(y=0|x) = 0.5. Solving for this in terms of the linear equation wTx+b yields a straight line (or hyperplane in higher dimensions). Samples on one side of the line are predicted as class 1, samples on the other as class 0.
This linear decision boundary is one of the main limitations of logistic regression. If the classes are not linearly separable, logistic regression will struggle no matter how much data it‘s trained on. Methods like feature engineering, kernel tricks, or moving to a more complex model can help address cases where a nonlinear decision boundary is needed.
7. How does logistic regression handle outliers?
Logistic regression is sensitive to outliers, i.e. data points unusually far from the main cluster of their class. Outliers can exert high leverage and pull the decision boundary in their direction.
The sigmoid function offers some protection against outliers. As z = wTx + b gets very large or very small, the sigmoid asymptotically approaches 1 or 0 respectively. This limits how much influence any single point can have.
Nonetheless, outliers can still mislead logistic regression, especially if there are many of them. Possible solutions include removing outliers in a pre-processing step, using regularization to shrink weights of outlier-heavy features, or moving to a more robust model like support vector machines.
8. What‘s the difference between logits and probabilities?
Logits are the values logistic regression directly estimates, defined as the log-odds log(p / (1-p)). Probabilities are what we usually care about, the chance of a sample belonging to class 1.
Logistic regression fits a linear model to the logits. It then applies the sigmoid function to the logits to transform them into probabilities between 0 and 1. Many ML libraries hide this under the hood, but keeping the distinction in mind clears up a common point of confusion.
9. How are categorical variables handled?
Most machine learning models, including logistic regression, operate on numeric features. So categorical variables, e.g. color ∈ {red, green, blue}, need to be converted to numbers.
One common encoding is one-hot, which creates a binary column for each possible category value. For example:
red, green, blue
1, 0, 0
0, 1, 0
0, 0, 1
This encoding allows logistic regression to learn a separate weight for each category. The downside is it can greatly expand the dimensionality, especially for high cardinality features.
An alternative is label encoding, which maps each category to an integer, e.g. red=1, green=2, blue=3. This is more compact but implies an ordinal relationship between categories that may not exist.
The best encoding depends on the data. One-hot is a good default for nominal variables, while label encoding makes more sense for ordinal ones. Domain knowledge and experimentation can guide the choice.
10. What assumptions does logistic regression make about the data?
Logistic regression is a parametric model, meaning it assumes a functional form of the relationship between inputs and outputs. Specifically, logistic regression assumes:
-
Linearity – There‘s a linear relationship between the logit of the output and each input variable. However, it does not assume the data itself is linearly separable.
-
Independence – The training examples are independently sampled from the joint distribution of inputs and outputs. Data with correlated samples can bias weights and lead to overfitting.
-
No multi-collinearity – The input features are not highly correlated with each other. Multi-collinear variables cause numerical instability and make weights hard to interpret.
-
Large sample size – Logistic regression relies on maximum likelihood estimation. For the estimates to be reliable and close to the true values, there needs to be a sufficiently large number of samples. A rule of thumb is at least 10 samples per feature.
Violating these assumptions doesn‘t necessarily mean logistic regression will fail, but it can degrade performance. Testing for linearity between the logit and individual features, calculating correlation matrices, and collecting more samples are ways to validate and address issues.
11. Can logistic regression be used for multi-class problems?
Yes, though binary logistic regression cannot be directly applied to multi-class problems. Instead, the one-vs-rest (OVR) or multinomial strategies are used.
In OVR, a separate binary logistic regression model is trained for each class, learning to distinguish that class from all others. A sample‘s final class is assigned by the model with the highest probability. OVR scales linearly with the number of classes.
Multinomial logistic regression extends binary logistic regression to directly support multiple classes. It learns a set of weights for each class and normalizes the outputs so the class probabilities sum to one. The multinomial loss is then minimized with respect to all weights simultaneously. This is more expensive than OVR but can lead to better accuracy, especially when classes compete with each other.
12. What is the time and space complexity of logistic regression?
Let n be the number of samples, d the number of features, and i the number of training iterations.
Time complexity:
-
Training: O(n d i). For each iteration, logistic regression processes all samples and updates all feature weights. Techniques like stochastic gradient descent and mini-batching can lower the effective number of samples and speed up convergence.
-
Inference: O(d). To make a prediction, logistic regression computes the dot product of the learned weights with a sample‘s features. This happens in linear time with respect to the number of features.
Space complexity:
-
Training: O(n * d). Logistic regression needs to store the training set in memory to compute gradients and update parameters in each iteration.
-
Model: O(d). The learned logistic regression model consists of a weight for each feature and a bias term. This requires storing O(d) floating point values. The small model size is one of the appeals of logistic regression.
Overall, logistic regression is linear in both the number of samples and features. This makes it an efficient choice, especially for inference. However, as datasets grow larger, the O(n * d) training cost can become prohibitive and require moving to online or distributed learning.
13. Why is cross-entropy loss used instead of mean squared error?
Mean squared error (MSE) seems like a natural choice for a loss function. It directly measures how far off the model‘s predictions are from the true values. MSE is also simple to understand and optimize via gradient descent.
However, there‘s a problem using MSE for logistic regression. The sigmoid function is non-convex with respect to the model‘s raw predictions wTx + b. Put another way, the sigmoid has flat regions where large changes in the input lead to small changes in the output probability.
When MSE is combined with the sigmoid, the resulting loss function has local minima where gradient descent can get stuck. This makes finding the globally optimal weights difficult.
In contrast, the cross-entropy loss is convex with respect to the raw predictions. It yields a smooth optimization landscape with a unique minimum. Cross-entropy also has a probabilistic interpretation, measuring the dissimilarity between the true label distribution and the model‘s predicted distribution.
As a result, cross-entropy is almost always used as the loss function for logistic regression. Advanced optimizers like L-BFGS or ADAM can also help navigate non-convexity if MSE is needed for a specific problem.
14. How does logistic regression differ from linear regression for classification?
On the surface, logistic and linear regression appear quite similar. They both learn a linear function of the input features and can be used to make predictions. However, they differ in a few key ways:
Output – Linear regression outputs a continuous value, while logistic regression outputs a probability between 0 and 1. This gives logistic regression‘s predictions a clear interpretation and allows them to be easily thresholded into class labels.
Loss function – Linear regression typically minimizes mean squared error, while logistic regression minimizes cross-entropy loss. This reflects the different goals of the two models – linear regression aims to predict values as close to the true values as possible, while logistic regression aims to predict the correct class as often as possible.
Assumptions – Linear regression assumes the residuals are normally distributed and homoscedastic (having constant variance). Logistic regression assumes a binomial distribution of the output and does not make assumptions about the residuals.
Extrapolation – Linear regression can extrapolate to output values outside the range of the training data. Logistic regression probabilities are always bounded between 0 and 1, regardless of how extreme the inputs are.
These differences make logistic regression better suited for classification problems. Its probabilistic predictions, robustness to outliers, and lack of distributional assumptions on the input make it a go-to method for binary classification.
15. What are the strengths and weaknesses of logistic regression?
Strengths:
- Simple and easy to understand
- Fast to train and make predictions
- Outputs well-calibrated probabilities
- Weights are interpretable as feature importances
- Less prone to overfitting than more complex models
- Provides a good baseline for binary classification problems
Weaknesses:
- Assumes a linear relationship between inputs and the logit
- Requires careful feature engineering and scaling
- Can underfit if classes are not linearly separable
- Sensitive to outliers and multi-collinear features
- Performance plateaus as data size grows
- Not suited for data with many features or complex interactions
Despite its limitations, logistic regression remains popular due to its simplicity, interpretability, and strong benchmark results. It‘s a valuable tool to master for any data scientist or ML engineer.
16. What are some extensions or recent developments in logistic regression?
While logistic regression is an established algorithm, research continues to improve it. Some notable developments:
Elastic net regularization – Combines L1 and L2 penalties to balance feature selection and coefficient shrinkage. Particularly useful for high-dimensional data.
Kernel logistic regression – Applies the kernel trick to learn non-linear decision boundaries. Allows capturing feature interactions without manual feature engineering.
Bayesian logistic regression – Places a prior distribution over weights and finds their posterior distribution given data. Provides uncertainty estimates and guards against overfitting.
Online learning – Fits logistic regression to one sample at a time, enabling training on datasets too large to fit in memory.
Factorization machines – Generalize logistic regression to model pairwise feature interactions. Often used for sparse data like user-item matrices in recommender systems.
Neural additive models – Build on logistic regression by learning a non-linear function for each feature and adding the results. Preserves interpretability while capturing more complex patterns.
These advances aim to make logistic regression more flexible, scalable, and robust to violations of its assumptions. Understanding them can deepen your knowledge of logistic regression and expand its applications.
Wrapping Up
Logistic regression is a fundamental part of the data scientist and machine learning engineer toolkit. While easy to get started with, it takes effort to master the nuances and effectively apply it to real problems.
The questions covered here hit many of the key points – from mechanics to intuition to implementation considerations. Internalizing these concepts will serve you well in both theory and practice.
But logistic regression is just the tip of the iceberg. More complex models like support vector machines, decision trees, and neural networks offer even greater predictive power and flexibility. Building on a solid foundation of logistic regression will put you in good stead to learn and apply these more advanced techniques.
So crack open a dataset, prototype a logistic regression, and see what insights you can uncover. The path to mastery is lined with many a sigmoid curve. Happy learning!