The Wayback Machine - https://web.archive.org/web/20241128003406/https://www.geeksforgeeks.org/ml-logistic-regression-using-python/
Open In App

Logistic Regression using Python

Last Updated : 04 Dec, 2023
Summarize
Comments
Improve
Suggest changes
Like Article
Like
Save
Share
Report
News Follow

A basic machine learning approach that is frequently used for binary classification tasks is called logistic regression. Though its name suggests otherwise, it uses the sigmoid function to simulate the likelihood of an instance falling into a specific class, producing values between 0 and 1. Logistic regression, with its emphasis on interpretability, simplicity, and efficient computation, is widely applied in a variety of fields, such as marketing, finance, and healthcare, and it offers insightful forecasts and useful information for decision-making.

Logistic Regression

A statistical model for binary classification is called logistic regression. Using the sigmoid function, it forecasts the likelihood that an instance will belong to a particular class, guaranteeing results between 0 and 1. To minimize the log loss, the model computes a linear combination of input characteristics, transforms it using the sigmoid, and then optimizes its coefficients using methods like gradient descent. These coefficients establish the decision boundary that divides the classes. Because of its ease of use, interpretability, and versatility across multiple domains, Logistic Regression is widely used in machine learning for problems that involve binary outcomes. Overfitting can be avoided by implementing regularization.

How the Logistic Regression Algorithm Works

Logistic Regression models the likelihood that an instance will belong to a particular class. It uses a linear equation to combine the input information and the sigmoid function to restrict predictions between 0 and 1. Gradient descent and other techniques are used to optimize the model’s coefficients to minimize the log loss. These coefficients produce the resulting decision boundary, which divides instances into two classes. When it comes to binary classification, logistic regression is the best choice because it is easy to understand, straightforward, and useful in a variety of settings. Generalization can be improved by using regularization.

Key Concepts of Logistic Regression

Important key concepts in logistic regression include:

  • Sigmoid Function: The main function that ensures outputs are between 0 and 1 by converting a linear combination of input data into probabilities.
    The sigmoid function is denoted as \sigma(z)       , and is defined as:
    \sigma(z) = \frac{1}{1 + e^z}
    Where, z is linear combination of input features and coefficients.
  • Hypothesis Function: uses the sigmoid function and weights (coefficients) to combine input features to estimate the likelihood of falling into a particular class.
    In logistic regression, the hypothesis function is provided by:
    h_{\theta}(x) = \sigma(\theta^Tx)
    Where, h_{\theta}(x) is the predicted probability that y = 1, \theta       is the vector of coefficients, and x is the vector of input features.
  • Log Loss: The optimization cost function is a measure of the discrepancy between actual class labels and projected probability.
    The definition of the log loss for a single instance is:
    J(\theta) = -(y \log{h_{\theta}(x)} + (1 - y) \log {(1-h_{\theta}(x)))}
  • Decision Boundary: The surface or line used to divide instances into several classes according to the determined probability.
  • Probability Threshold: a number (usually 0.5) that is used to calculate the class assignment using the probabilities that are anticipated.
  • Odds Ratio: The likelihood that an event will occur as opposed to not, which sheds light on how characteristics and the target variable are related.

Prerequisite: Understanding Logistic Regression

Implementation of Logistic Regression using Python

Import Libraries

Python3

# Import necessary libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix, roc_curve, auc

                    

Read and Explore the data

Python3

# Load the diabetes dataset
diabetes = load_diabetes()
X, y = diabetes.data, diabetes.target
 
# Convert the target variable to binary (1 for diabetes, 0 for no diabetes)
y_binary = (y > np.median(y)).astype(int)

                    

This code loads the diabetes dataset using the load_diabetes function from scikit-learn, passing in feature data X and target values y. Then, it converts the binary representation of the continuous target variable y. A patient’s diabetes measure is classified as 1 (indicating diabetes) if it is higher than the median value, and as 0 (showing no diabetes).

Splitting The Dataset: Train and Test dataset

Splitting the dataset to train and test. 80% of data is used for training the model and 20% of it is used to test the performance of our model.  

Python3

# Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(
    X, y_binary, test_size=0.2, random_state=42)

                    

This code divides the diabetes dataset into training and testing sets using the train_test_split function from scikit-learn: The binary target variable is called y_binary, and the characteristics are contained in X. The data is divided into testing (X_test, y_test) and training (X_train, y_train) sets. Twenty percent of the data will be used for testing, according to the setting test_size=0.2. By employing a fixed seed for randomization throughout the split, random_state=42 guarantees reproducibility.

Feature Scaling

Python3

# Standardize features
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)

                    

This code uses StandardScaler from scikit-learn to achieve feature standardization:

The StandardScaler instance is created; this will be used to standardize the features. It uses the scaler’s fit_transform method to normalize the training data (X_train) and determine its mean and standard deviation. Then, itstandardizes the testing data (X_test) using the calculated mean and standard deviation from the training set. Model training and evaluation are made easier by standardization, which guarantees that the features have a mean of 0 and a standard deviation of 1.

Train The Model

Python3

# Train the Logistic Regression model
model = LogisticRegression()
model.fit(X_train, y_train)

                    

 Using scikit-learn’s LogisticRegression, this code trains a logistic regression model:

It establishes a logistic regression model instance.Then, itemploys the fit approach to train the model using the binary target values (y_train) and standardized training data (X_train). Following execution, the model object may now be used to forecast new data using the patterns it has learnt from the training set.

Evaluation Metrics

Metrics are used to check the model performance on predicted values and actual values. 

Python3

# Evaluate the model
y_pred = model.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy: {:.2f}%".format(accuracy * 100))

                    

Output:

Accuracy: 73.03%

This code predicts the target variable and computes its accuracy in order to assess the logistic regression model on the test set. The accuracy_score function is then used to compare the predicted values in the y_pred array with the actual target values (y_test).

Confusion Matrix and Classification Report

Python3

# evaluate the model
print("Confusion Matrix:\n", confusion_matrix(y_test, y_pred))
print("\nClassification Report:\n", classification_report(y_test, y_pred))

                    

Output:

Confusion Matrix:
 [[36 13]
 [11 29]]

Classification Report:
               precision    recall  f1-score   support

           0       0.77      0.73      0.75        49
           1       0.69      0.72      0.71        40

    accuracy                           0.73        89
   macro avg       0.73      0.73      0.73        89
weighted avg       0.73      0.73      0.73        89

Visualizing the performance of our model.

Python3

# Visualize the decision boundary with accuracy information
plt.figure(figsize=(8, 6))
sns.scatterplot(x=X_test[:, 2], y=X_test[:, 8], hue=y_test, palette={
                0: 'blue', 1: 'red'}, marker='o')
plt.xlabel("BMI")
plt.ylabel("Age")
plt.title("Logistic Regression Decision Boundary\nAccuracy: {:.2f}%".format(
    accuracy * 100))
plt.legend(title="Diabetes", loc="upper right")
plt.show()

                    

Output:

Screenshot-from-2023-12-04-16-09-19

Logistic Regression

To see a logistic regression model’s decision border, this code creates a scatter plot. An individual from the test set is represented by each point on the plot, which has age on the Y-axis and BMI on the X-axis. The points are color-coded according to the actual status of diabetes, making it easier to evaluate how well the model differentiates between those with and without the disease. An instant visual context for the model’s performance on the test data is provided by the plot’s title, which includes the accuracy information. The inscription located in the upper right corner denotes the colors that represent diabetes (1) and no diabetes (0).

Plotting ROC Curve

Python3

# Plot ROC Curve
y_prob = model.predict_proba(X_test)[:, 1]
fpr, tpr, thresholds = roc_curve(y_test, y_prob)
roc_auc = auc(fpr, tpr)
 
plt.figure(figsize=(8, 6))
plt.plot(fpr, tpr, color='darkorange', lw=2,
         label=f'ROC Curve (AUC = {roc_auc:.2f})')
plt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--', label='Random')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.title('Receiver Operating Characteristic (ROC) Curve\nAccuracy: {:.2f}%'.format(
    accuracy * 100))
plt.legend(loc="lower right")
plt.show()

                    

Output:

ROC Curve-Geeksforgeeks

Receiver Operating Characteristic (ROC) Curve


For the logistic regression model, this code creates and presents the Receiver Operating Characteristic (ROC) curve. The true positive rate (sensitivity) and false positive rate at different threshold values are determined using the probability estimates for positive outcomes (y_prob), which are obtained using the predict_proba method. Use of the roc_auc_score yields the area under the ROC curve (AUC). An illustration of the resulting curve is provided, and the legend shows the AUC value. The ROC curve for a random classifier is shown by the dotted line.

Frequently Asked Questions

Q1. What is Logistic Regression?

A statistical technique for binary classification issues is called logistic regression.It uses a logistic function to model the likelihood of a binary outcome occurring.

Q2. How is Logistic Regression different from Linear Regression?

The probability of a binary event is predicted by logistic regression, whereas a continuous outcome is predicted by linear regression. In order to limit the output between 0 and 1, logistic regression uses the logistic (sigmoid) function.

Q3. How to handle categorical variables in Logistic Regression?

Use one-hot encoding, for instance, to transform categorical information into numerical representation. Make sure the data has been properly preprocessed to prepare it for logistic regression.

Q4. Can Logistic Regression handle multiclass classification?

It is possible to use methods like One-vs-Rest or Softmax Regression to expand logistic regression for multiclass classification.

Q5. What is the role of the sigmoid function in Logistic Regression?

Any real integer can be mapped to the range [0, 1] using the sigmoid function. The linear equation’s output is converted into probabilities by it.



Similar Reads

ML | Linear Regression vs Logistic Regression
Linear Regression is a machine learning algorithm based on supervised regression algorithm. Regression models a target prediction value based on independent variables. It is mostly used for finding out the relationship between variables and forecasting. Different regression models differ based on – the kind of relationship between the dependent and
3 min read
Implementation of Logistic Regression from Scratch using Python
Introduction:Logistic Regression is a supervised learning algorithm that is used when the target variable is categorical. Hypothetical function h(x) of linear regression predicts unbounded values. But in the case of Logistic Regression, where the target variable is categorical we have to strict the range of predicted values. Consider a classificati
5 min read
Logistic Regression using PySpark Python
In this tutorial series, we are going to cover Logistic Regression using Pyspark. Logistic Regression is one of the basic ways to perform classification (don’t be confused by the word “regression”). Logistic Regression is a classification method. Some examples of classification are: Spam detectionDisease DiagnosisLoading Dataframe We will be using
3 min read
Logistic Regression using Statsmodels
Prerequisite: Understanding Logistic RegressionLogistic regression is the type of regression analysis used to find the probability of a certain event occurring. It is the best suited type of regression for cases where we have a categorical dependent variable which can take only discrete values. The dataset : In this article, we will predict whether
4 min read
Placement prediction using Logistic Regression
Prerequisites: Understanding Logistic Regression, Logistic Regression using Python In this article, we are going to discuss how to predict the placement status of a student based on various student attributes using Logistic regression algorithm. Placements hold great importance for students and educational institutions. It helps a student to build
4 min read
Text Classification using Logistic Regression
Text classification is the process of automatically assigning labels or categories to pieces of text. This has tons of applications, like sorting emails into spam or not-spam, figuring out if a product review is positive or negative, or even identifying the topic of a news article. In this article, we will see How logistic regression is used for te
5 min read
ML | Logistic Regression using Tensorflow
Prerequisites: Understanding Logistic Regression and TensorFlow. Brief Summary of Logistic Regression: Logistic Regression is Classification algorithm commonly used in Machine Learning. It allows categorizing data into discrete classes by learning the relationship from a given set of labeled data. It learns a linear relationship from the given data
6 min read
Heart Disease Prediction Using Logistic Regression in R
Machine learning can effectively identify patterns in data, providing valuable insights from this data. This article explores one of these machine learning techniques called Logistic regression and how it can analyze the key patient details and determine the probability of heart disease based on patient health data in R Programming Language. Logist
13 min read
ML | Heart Disease Prediction Using Logistic Regression
World Health Organization has estimated that four out of five cardiovascular disease (CVD) deaths are due to heart attacks. This whole research intends to pinpoint the ratio of patients who possess a good chance of being affected by CVD and also to predict the overall risk using Logistic Regression. What is Logistic Regression?Logistic Regression i
4 min read
Identifying handwritten digits using Logistic Regression in PyTorch
Logistic Regression is a very commonly used statistical method that allows us to predict a binary output from a set of independent variables. The various properties of logistic regression and its Python implementation have been covered in this article previously. Now, we shall find out how to implement this in PyTorch, a very popular deep learning
7 min read
ML | Kaggle Breast Cancer Wisconsin Diagnosis using Logistic Regression
Dataset:It is given by Kaggle from UCI Machine Learning Repository, in one of its challenge. It is a dataset of Breast Cancer patients with Malignant and Benign tumor. Logistic Regression is used to predict whether the given patient is having Malignant or Benign tumor based on the attributes in the given dataset. Code : Loading Libraries [GFGTABS]
5 min read
ML | Why Logistic Regression in Classification ?
Using Linear Regression, all predictions >= 0.5 can be considered as 1 and rest all < 0.5 can be considered as 0. But then the question arises why classification can't be performed using it? Problem - Suppose we are classifying a mail as spam or not spam and our output is y, it can be 0(spam) or 1(not spam). In case of Linear Regression, hθ(x
3 min read
ML | Logistic Regression v/s Decision Tree Classification
Logistic Regression and Decision Tree classification are two of the most popular and basic classification algorithms being used today. None of the algorithms is better than the other and one's superior performance is often credited to the nature of the data being worked upon. We can compare the two algorithms on different categories - CriteriaLogis
2 min read
Logistic Regression in R Programming
Logistic regression in R Programming is a classification algorithm used to find the probability of event success and event failure. Logistic regression is used when the dependent variable is binary(0/1, True/False, Yes/No) in nature. The logit function is used as a link function in a binomial distribution. A binary outcome variable's probability ca
8 min read
Differentiate between Support Vector Machine and Logistic Regression
Logistic Regression: It is a classification model which is used to predict the odds in favour of a particular event. The odds ratio represents the positive event which we want to predict, for example, how likely a sample has breast cancer/ how likely is it for an individual to become diabetic in future. It used the sigmoid function to convert an in
4 min read
Role of Log Odds in Logistic Regression
Prerequisite : Log Odds, Logistic Regression NOTE: It is advised to go through the prerequisite topics to have a clear understanding of this article. Log odds play an important role in logistic regression as it converts the LR model from probability based to a likelihood based model. Both probability and log odds have their own set of properties, h
4 min read
Logistic Regression on MNIST with PyTorch
Logistic Regression Logistic Regression is also known as Binary Classification is one of the most popular Machine Learning Algorithms. It comes under Supervised Learning Classification Algorithms. It is used to predict the probability of the target label. By binary classification, it means that the model predicts the label either 0 or 1. The target
4 min read
Logistic Regression Vs Random Forest Classifier
A statistical technique called logistic regression is used to solve problems involving binary classification, in which the objective is to predict a binary result (such as yes/no, true/false, or 0/1) based on one or more predictor variables (also known as independent variables, features, or predictors). Based on the values of the predictor variable
7 min read
Multinomial Logistic Regression with PyTorch
Logistic regression is a popular machine learning algorithm used for binary classification tasks. It models the probability of the output variable (also known as the dependent variable) given the input variables (also known as the independent variables). It is a linear algorithm that applies a logistic function to the output of a linear regression
11 min read
Cost function in Logistic Regression in Machine Learning
Logistic Regression is one of the simplest classification algorithms we learn while exploring machine learning algorithms. In this article, we will explore cross-entropy, a cost function used for logistic regression. What is Logistic Regression?Logistic Regression is a statistical method used for binary classification. Despite its name, it is emplo
10 min read
How to Run Binary Logistic Regression in SPSS?
Answer: To run binary logistic regression in SPSS, navigate to Analyze > Regression > Binary Logistic.Running binary logistic regression in SPSS involves several steps. Here's a detailed guide: Launch SPSS and Open Data:Open SPSS software and load your dataset containing the variables of interest, including the binary outcome variable and pre
2 min read
What Is the Difference Between SGD Classifier and the Logistic Regression?
Answer: The main difference is that SGDClassifier uses stochastic gradient descent optimization while logistic regression uses the logistic function to model binary classification.Explanation:Optimization Technique:Logistic Regression: In logistic regression, the optimization is typically performed using methods like gradient descent or Newton's me
2 min read
Logistic Regression vs K Nearest Neighbors in Machine Learning
Machine learning algorithms play a crucial role in training the data and decision-making processes. Logistic Regression and K Nearest Neighbors (KNN) are two popular algorithms in machine learning used for classification tasks. In this article, we'll delve into the concepts of Logistic Regression and KNN and understand their functions and their dif
4 min read
Weighted Logistic Regression for Imbalanced Dataset
In real-world datasets, it's common to encounter class imbalance, where one class significantly outnumbers the other(s). This class imbalance poses challenges for machine learning models, particularly for classification tasks, as models tend to be biased towards the majority class, leading to suboptimal performance. What are imbalanced datasets?Imb
6 min read
How to interpret odds ratios in logistic regression
Logistic regression is a statistical method used to model the relationship between a binary outcome and predictor variables. This article provides an overview of logistic regression, including its assumptions and how to interpret regression coefficients. Assumptions of logistic regressionBinary Outcome: Logistic regression assumes that the outcome
12 min read
Outlier Detection in Logistic Regression
Outliers, data points that deviate significantly from the rest, can significantly impact the performance of logistic regression models. In this article we will explore various techniques for detecting and handling outliers in Logistic regression. What are Outliers?An outlier is an observation that falls far outside the typical range of other data p
8 min read
Logistic Regression and the Feature Scaling Ensemble
Logistic Regression is a widely used classification algorithm in machine learning. However, to enhance its performance further specially when dealing with features of different scales, employing feature scaling ensemble techniques becomes imperative. In this guide, we will dive depth into logistic regression, its significance and how feature sealin
9 min read
How to Handle Missing Data in Logistic Regression?
Logistic regression is a robust statistical method employed to model the likelihood of binary results. Nevertheless, real-world datasets frequently have missing values, presenting obstacles while fitting logistic regression models. Dealing with missing data effectively is essential to prevent skewed estimates and maintain the model's accuracy. In t
9 min read
How to Optimize Logistic Regression Performance
Logistic Regression is a widely employed algorithm for binary classification tasks. However, the performance of Logistic Regression models can be significantly impacted by the choice of hyperparameters, which can lead to suboptimal results if not properly tuned. Therefore, it is crucial to explore the various hyperparameters that influence the perf
8 min read
Logistic Regression With Polynomial Features
Logistic regression with polynomial features is a technique used to model complex, non-linear relationships between input variables and the target variable. This approach involves transforming the original input features into higher-degree polynomial features, which can help capture intricate patterns in the data and improve the model's predictive
5 min read
three90RightbarBannerImg