Regression
Linear regression theory
Model
For linear regression, the hypothesis will be a linear equation with any amount of features. This can be univariate or multivariate linear regression:
The goal in linear regression is to minimize the squared error consisting of all the data points:
Loss functions
We have several choices of loss functions for linear regression:

Univariate linear regression
Evaluating the model
Coefficient of determination
We need some way to evaluate how well our model does and you can do that either with loss or something in statistics called the coefficient of determination, which only works in two-dimensional data (features X and output y).
- coefficient of determination: denoted via , and has bounds .
- use case: used as a measure of correlation strength between two variables
- correlation coefficient: denoted by , which is just the square root of the coefficient of determination, and has bounds .
-
use case: used as a measure of correlation strength
-
NOTE
The null model is just predicting the mean, which means that the null model has an value = 0
Null model
In linear regression, the Null Model is the simplest possible baseline. It assumes that the features () have no predictive power, so it simply predicts the mean () of the target variable () for every observation.
- Equation:
- Purpose: It serves as a benchmark. If your regression model isn't better than the null model, your features are not useful.
The null model will just predict that every point is equal to the mean, giving a straight line. Then from there, we can get the coefficient R^2, which is the coefficient of determination.
You can also think of it like this:
Logistic Regression
Problem and intuition
Model
Logistic regression is used for classification tasks, but like regression, it also outputs a single number.
The hypothesis uses the sigmoid function to make sure all probability outputs are between 0 and 1.
Logistic Regression predicts probabilities. To ensure the output is always between 0 and 1, it wraps a linear equation inside the Sigmoid (or Logistic) function:
Where .
What are Odds and Log Odds?
While we interpret the output as a probability (), the model itself is essentially a linear model for the Log Odds (also called the Logit).
-
Probability (): The chance of an event occurring (e.g., 0.8 or 80%).
-
Odds: The ratio of the probability of success to the probability of failure.
Example: If , then Odds = (meaning it is 4 times more likely to happen than not).
- Log Odds: The natural logarithm of the odds.
The Logistic Regression model assumes that this Log Odds value is a linear combination of your input features.
import numpy as np
# Let's see how Probabilities map to Log Odds
probabilities = np.array([0.1, 0.3, 0.5, 0.7, 0.9])
def calculate_log_odds(p):
odds = p / (1 - p)
log_odds = np.log(odds)
return odds, log_odds
print(f"{'Prob':<10} | {'Odds':<10} | {'Log Odds':<10}")
print("-" * 35)
for p in probabilities:
o, lo = calculate_log_odds(p)
print(f"{p:<10.2f} | {o:<10.2f} | {lo:<10.2f}")
Binary cross-entropy loss

Logistic regression with regularization
Logistic regression uses L2 regularization, the hyperparameter controls that behavior, acting as the penalty for regularization.
Multi-class logistic regression
There are two methods for multiclass logistic regression.
-
One vs rest: We do repeated loops of treating one class as the positive class, and lumping all other classes together as the negative class.
-
Multinomial: We do a softmax classification where all the classes have their own probability, and all together they sum to 1. This is the default for scikit-learn