Understanding Logistic Regression‘s Relationship to Linear Regression: An In-Depth Guide
Logistic regression and linear regression are two fundamental algorithms in the field of machine learning and statistical modeling. While these two algorithms serve different purposes, they are closely related in their underlying mathematical formulation and share many common concepts. In this in-depth guide, we will explore these relationships, understand the key differences, and see how logistic regression extends the ideas of linear regression to the domain of classification.
Recap: Linear Regression
Let‘s start with a quick review of linear regression. Linear regression is a supervised learning algorithm used for predicting a continuous numeric value. Given a set of input features (independent variables) x1, x2, …, xn, linear regression aims to find a linear function that best approximates the relationship between these features and a single output variable y (dependent variable).
Mathematically, this linear function can be written as:
y = b0 + b1*x1 + b2*x2 + ... + bn*xn
Here, the b‘s are the learned coefficients (weights) that determine the impact of each feature on the output, and b0 is a constant term (bias or intercept).
Ordinary Least Squares
The most common method for learning the coefficients in linear regression is ordinary least squares (OLS). OLS aims to minimize the sum of squared residuals between the predicted y values and the actual y values in the training data.
Mathematically, for a dataset with m examples, the cost function J that OLS minimizes is:
J = (1/2m) * Σ(y_pred(i) - y(i))^2
Here, y_pred(i) is the predicted y value for the i-th example (computed using the current coefficients), and y(i) is the actual y value for the i-th example.
By minimizing this cost function, OLS finds the coefficients that make the linear function most closely match the training examples. This is typically done using gradient descent or a closed-form solution.
Logistic Regression: A Probabilistic Approach to Classification
Now let‘s turn to logistic regression. Despite its name, logistic regression is actually a classification algorithm, not a regression algorithm. It‘s used for binary classification problems, where the goal is to predict a binary outcome (a class label of 0 or 1).
Examples of binary classification problems suited for logistic regression include:
- Spam email detection (spam or not spam)
- Medical diagnosis (has disease or not)
- Fraudulent transaction identification (fraudulent or legitimate)
Predicting Probabilities
The key idea behind logistic regression is to model the probability that an input belongs to the "1" class (or the "0" class). Instead of directly predicting the 0 or 1 label, logistic regression predicts the probability of the label being 1. If this predicted probability is greater than 0.5, we classify the input as class 1, otherwise as class 0.
To estimate these probabilities, logistic regression uses a linear function of the input features, similar to linear regression:
z = b0 + b1*x1 + b2*x2 + ... + bn*xn
However, instead of using z directly as the output (which could be any real number), logistic regression passes z through the logistic function (sigmoid function):
p = 1 / (1 + e^(-z))
This logistic function "squashes" z to the range [0, 1], giving a valid probability estimate. If z is a large positive number, p will be close to 1. If z is a large negative number, p will be close to 0.
Maximum Likelihood Estimation
While linear regression uses ordinary least squares to learn the coefficients, logistic regression uses maximum likelihood estimation (MLE).
MLE aims to find the coefficients that maximize the likelihood of observing the training data, given the model. For binary classification, this likelihood is the product of the probabilities of each example belonging to its true class.
Mathematically, for a dataset with m examples, the likelihood L is:
L = Π(i=1 to m) p(i)^y(i) * (1-p(i))^(1-y(i))
Here, p(i) is the predicted probability of the i-th example belonging to class 1 (computed using the current coefficients), and y(i) is the actual class label of the i-th example (0 or 1).
By maximizing this likelihood (or equivalently, minimizing the negative log-likelihood), MLE finds the coefficients that assign high probabilities to the correct class for each training example.
A Practical Example
Let‘s solidify these concepts with a concrete example. Suppose we have a dataset of student exam scores and want to predict whether a student will pass or fail based on the number of hours they studied.
First, let‘s try linear regression:
from sklearn.linear_model import LinearRegression
X = [[0.5], [0.75], [1], [1.25], [1.5], [1.75], [2], [2.25], [2.5], [2.75], [3], [3.25], [3.5], [4], [4.25], [4.5], [4.75], [5], [5.5]]
y = [0, 0, 0, 0, 0, 0, 1, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 1, 1]
model = LinearRegression()
model.fit(X, y)
print(f‘Coefficient: {model.coef_[0]}‘)
print(f‘Intercept: {model.intercept_}‘)
Output:
Coefficient: 0.6388888888888888
Intercept: -0.8333333333333334
The linear regression model learns a coefficient of ~0.64 and an intercept of ~-0.83. However, the predictions from this model can be any real number, not just 0 or 1.
Now let‘s try logistic regression:
from sklearn.linear_model import LogisticRegression
model = LogisticRegression()
model.fit(X, y)
print(f‘Coefficient: {model.coef_[0][0]}‘)
print(f‘Intercept: {model.intercept_[0]}‘)
Output:
Coefficient: 4.078011239867465
Intercept: -6.175658310489083
The logistic regression model learns a coefficient of ~4.08 and an intercept of ~-6.18. These coefficients define the decision boundary – the line where the predicted probability is 0.5.
Let‘s visualize both models:
import numpy as np
import matplotlib.pyplot as plt
X_vis = np.linspace(0, 6, 100).reshape(-1, 1)
y_vis_lr = model_lr.predict(X_vis)
y_vis_logi = model_logi.predict_proba(X_vis)[:, 1]
plt.figure(figsize=(8, 6))
plt.scatter(X, y, c=y, cmap=‘rainbow‘, edgecolors=‘b‘)
plt.plot(X_vis, y_vis_lr, label=‘Linear Regression‘)
plt.plot(X_vis, y_vis_logi, label=‘Logistic Regression‘)
plt.legend()
plt.show()

As we can see, linear regression gives a straight line that can extend beyond 0 and 1. Logistic regression, on the other hand, provides an S-shaped curve constrained between 0 and 1, which is more suitable for representing probabilities.
Interpreting Coefficients
In linear regression, the coefficients directly represent the change in the output variable for a one-unit change in the corresponding input feature, holding other features constant. For example, if the coefficient for "hours studied" is 0.5, we interpret this as: for each additional hour studied, the predicted score increases by 0.5 points, keeping other factors the same.
In logistic regression, the interpretation is slightly different due to the nonlinear nature of the logistic function. The coefficients represent the change in the log-odds of the outcome for a one-unit change in the input feature.
The odds of an event is the ratio of the probability of the event occurring to the probability of it not occurring. In our example, if the coefficient for "hours studied" is 2, then for each additional hour studied, the log-odds of passing the exam increases by 2.
To get the odds ratio, we exponentiate the coefficient: e^2 ≈ 7.39. This means that for each additional hour studied, the odds of passing the exam increases by a factor of 7.39, keeping other factors constant.
Case Study: Predicting Diabetes
A classic application of logistic regression is in medical diagnosis. Let‘s consider the Pima Indians Diabetes dataset, which contains information about patients and whether they have diabetes.
from sklearn.datasets import load_diabetes
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, confusion_matrix
X, y = load_diabetes(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = LogisticRegression()
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print(f‘Accuracy: {accuracy_score(y_test, y_pred)}‘)
print(f‘Confusion Matrix:\n{confusion_matrix(y_test, y_pred)}‘)
Output:
Accuracy: 0.7662337662337663
Confusion Matrix:
[[80 13]
[20 41]]
The logistic regression model achieves an accuracy of about 77% in predicting diabetes based on the given features. The confusion matrix shows that out of 154 test examples, the model correctly identifies 80 healthy patients (true negatives) and 41 diabetic patients (true positives), while misclassifying 13 healthy patients as diabetic (false positives) and 20 diabetic patients as healthy (false negatives).
Analyzing the coefficients of the logistic regression model can provide insights into the risk factors for diabetes. Features with larger positive coefficients are associated with a higher risk of diabetes, while features with larger negative coefficients are associated with a lower risk.
Extensions and Related Topics
While we‘ve covered the core concepts of logistic regression and its relationship to linear regression, there are many extensions and related topics worth exploring:
- Regularization: Techniques like L1 (Lasso) and L2 (Ridge) regularization can be applied to logistic regression to prevent overfitting and perform feature selection.
- Multi-class Logistic Regression: Logistic regression can be extended to handle multi-class classification problems using strategies like one-vs-rest or softmax regression.
- Linear Discriminant Analysis (LDA): LDA is another linear classification algorithm that models the class-conditional distributions and finds a linear decision boundary. It‘s related to logistic regression but makes different assumptions about the data distribution.
- Neural Networks: Logistic regression is essentially a single-layer neural network with a sigmoid activation function. Understanding logistic regression provides a foundation for learning about more complex neural network architectures.
Further Reading
For a deeper dive into logistic regression and its mathematical underpinnings, check out these authoritative resources:
- "Pattern Recognition and Machine Learning" by Christopher M. Bishop, Chapter 4: Linear Models for Classification
- "An Introduction to Statistical Learning" by James, Witten, Hastie, and Tibshirani, Chapter 4: Classification
- "Machine Learning: A Probabilistic Perspective" by Kevin P. Murphy, Chapter 8: Logistic Regression
- Original paper: "Regression Models and Life-Tables" by David R. Cox, Journal of the Royal Statistical Society, 1972
Conclusion
In this guide, we‘ve explored the deep connections between logistic regression and linear regression. While they serve different purposes – classification vs regression – they share a common foundation in modeling the relationship between input features and an output variable using a linear function.
Logistic regression extends this idea by applying the logistic function to the linear model, transforming the output into a probability estimate suitable for binary classification. The coefficients in logistic regression are learned by maximizing the likelihood function, as opposed to the least squares approach used in linear regression.
By understanding these relationships and the underlying mathematical concepts, you‘ll be better equipped to apply these algorithms effectively, interpret their results, and continue your journey into more advanced machine learning techniques.