The Wayback Machine - https://web.archive.org/web/20250306102731/https://www.geeksforgeeks.org/apriori-algorithm/
Open In App

Apriori Algorithm

Last Updated : 15 Jan, 2025
Summarize
Comments
Improve
Suggest changes
Like Article
Like
Share
Report
News Follow

Apriori Algorithm is a foundational method in data mining used for discovering frequent itemsets and generating association rules. Its significance lies in its ability to identify relationships between items in large datasets which is particularly valuable in market basket analysis.

For example, if a grocery store finds that customers who buy bread often also buy butter, it can use this information to optimize product placement or marketing strategies.

How the Apriori Algorithm Works?

The Apriori Algorithm operates through a systematic process that involves several key steps:

  1. Identifying Frequent Itemsets: The algorithm begins by scanning the dataset to identify individual items (1-item) and their frequencies. It then establishes a minimum support threshold, which determines whether an itemset is considered frequent.
  2. Creating Possible item group: Once frequent 1-itemgroup(single items) are identified, the algorithm generates candidate 2-itemgroup by combining frequent items. This process continues iteratively, forming larger itemsets (k-itemgroup) until no more frequent itemgroup can be found.
  3. Removing Infrequent Item groups: The algorithm employs a pruning technique based on the Apriori Property, which states that if an itemset is infrequent, all its supersets must also be infrequent. This significantly reduces the number of combinations that need to be evaluated.
  4. Generating Association Rules: After identifying frequent itemsets, the algorithm generates association rules that illustrate how items relate to one another, using metrics like support, confidence, and lift to evaluate the strength of these relationships.

Key Metrics of Apriori Algorithm

  • Support: This metric measures how frequently an item appears in the dataset relative to the total number of transactions. A higher support indicates a more significant presence of the itemset in the dataset. Support tells us how often a particular item or combination of items appears in all the transactions (“Bread is bought in 20% of all transactions.”)
  • Confidence: Confidence assesses the likelihood that an item Y is purchased when item X is purchased. It provides insight into the strength of the association between two items.
  • Confidence tells us how often items go together. (“If bread is bought, butter is bought 75% of the time.”)
  • Lift: Lift evaluates how much more likely two items are to be purchased together compared to being purchased independently. A lift greater than 1 suggests a strong positive association. Lift shows how strong the connection is between items. (“Bread and butter are much more likely to be bought together than by chance.”)

Lets understand the concept of apriori Algorithm with the help of an example. Consider the following dataset and we will find frequent itemsets and generate association rules for them:

Screenshot-2025-01-07-120244

Transactions of a Grocery Shop

Step 1 : Setting the parameters

  • Minimum Support Threshold: 50% (item must appear in at least 3/5 transactions). This threeshold is formulated from this formula:

[Tex]\text{Support}(A) = \frac{\text{Number of transactions containing itemset } A}{\text{Total number of transactions}} [/Tex]

  • Minimum Confidence Threshold: 70% ( You can change the value of parameters as per the usecase and problem statement ). This threeshold is formulated from this formula:

[Tex]\text{Confidence}(X \rightarrow Y) = \frac{\text{Support}(X \cup Y)}{\text{Support}(X)} [/Tex]

Step 2: Find Frequent 1-Itemsets

Lets count how many transactions include each item in the dataset (calculating the frequency of each item).

Screenshot-2025-01-07-120700

Frequent 1-Itemsets

All items have support% ≥ 50%, so they qualify as frequent 1-itemsets. if any item has support% < 50%, It will be ommited out from the frequent 1- itemsets.

Step 3: Generate Candidate 2-Itemsets

Combine the frequent 1-itemsets into pairs and calculate their support.

For this usecase, we will get 3 item pairs ( bread,butter) , (bread,ilk) and (butter,milk) and will calculate the support similiar to step 2

Screenshot-2025-01-07-121028

Candidate 2-Itemsets

Frequent 2-itemsets:

  • {Bread, Butter}, {Bread, Milk} both meet the 50% threshold but {butter,milk} doesnt meet the threeshold, so will be ommited out.

Step 4: Generate Candidate 3-Itemsets

Combine the frequent 2-itemsets into groups of 3 and calculate their support.

for the triplet, we have only got one case i.e {bread,butter,milk} and we will calculate the support.

Screenshot-2025-01-07-121350

Candidate 3-Itemsets

Since this does not meet the 50% threshold, there are no frequent 3-itemsets.

Step 5: Generate Association Rules

Now we generate rules from the frequent itemsets and calculate confidence.

Rule 1: If Bread → Butter (if customer buys bread, the customer will buy butter also)

  • Support of {Bread, Butter} = 3.
  • Support of {Bread} = 4.
  • Confidence = 3/4 = 75% (Passes threshold).

Rule 2: If Butter → Bread (if customer buys butter, the customer will buy bread also)

  • Support of {Bread, Butter} = 3.
  • Support of {Butter} = 3.
  • Confidence = 3/3 = 100% (Passes threshold).

Rule 3: If Bread → Milk (if customer buys bread, the customer will buy milk also)

  • Support of {Bread, Milk} = 3.
  • Support of {Bread} = 4.
  • Confidence = 3/4 = 75% (Passes threshold).

The Apriori Algorithm, as demonstrated in the bread-butter example, is widely used in modern startups like Zomato, Swiggy, and other food delivery platforms. These companies use it to perform market basket analysis, which helps them identify customer behavior patterns and optimize recommendations.

Applications of Apriori Algorithm

Below are some applications of Apriori algorithm used in today’s companies and startups

  1. E-commerce: Used to recommend products that are often bought together, like laptop + laptop bag, increasing sales.
  2. Food Delivery Services: Identifies popular combos, such as burger + fries, to offer combo deals to customers.
  3. Streaming Services: Recommends related movies or shows based on what users often watch together, like action + superhero movies.
  4. Financial Services: Analyzes spending habits to suggest personalized offers, such as credit card deals based on frequent purchases.
  5. Travel & Hospitality: Creates travel packages (e.g., flight + hotel) by finding commonly purchased services together.
  6. Health & Fitness: Suggests workout plans or supplements based on users’ past activities, like protein shakes + workouts.

For implementing apriori algorithm, please refer to Apriori algorithm in Python


Get IBM Certification and a 90% fee refund on completing 90% course in 90 days! Take the Three 90 Challenge today.

Master Machine Learning, Data Science & AI with this complete program and also get a 90% refund. What more motivation do you need? Start the challenge right away!


Next Article
Article Tags :
Practice Tags :

Similar Reads

Implementing Apriori algorithm in Python
Prerequisites: Apriori AlgorithmApriori Algorithm is a Machine Learning algorithm which is used to gain insight into the structured relationships between different items involved. The most prominent practical application of the algorithm is to recommend products based on the products already present in the user's cart. Walmart especially has made g
4 min read
ML | T-distributed Stochastic Neighbor Embedding (t-SNE) Algorithm
T-distributed Stochastic Neighbor Embedding (t-SNE) is a nonlinear dimensionality reduction technique and it is suited for visualizing high-dimensional data in a lower-dimensional space typically in 2D or 3D. It is a widely used dimensionality reduction technique and in this article we will learn about it.Dimensionality reduction is a process that
4 min read
Encoding Methods in Genetic Algorithm
Biological Background : Chromosome: All living organisms consist of cells. In each cell, there is the same set of Chromosomes. Chromosomes are strings of DNA and consist of genes, blocks of DNA. Each gene encodes a trait, for example, the color of the eye. Reproduction: During reproduction, combination (or crossover) occurs first. Genes from parent
3 min read
Crossover in Genetic Algorithm
Crossover is a genetic operator used to vary the programming of a chromosome or chromosomes from one generation to the next. Crossover is sexual reproduction. Two strings are picked from the mating pool at random to crossover in order to produce superior offspring. The method chosen depends on the Encoding Method. Crossover mask: The choice of whi
2 min read
ML | K-means++ Algorithm
Prerequisite: K-means Clustering – IntroductionDrawback of standard K-means algorithm:One disadvantage of the K-means algorithm is that it is sensitive to the initialization of the centroids or the mean points. So, if a centroid is initialized to be a "far-off" point, it might just end up with no points associated with it, and at the same time, mor
5 min read
Inductive Learning Algorithm
In this article, we will learn about Inductive Learning Algorithm which generally comes under the domain of Machine Learning. What is Inductive Learning Algorithm? Inductive Learning Algorithm (ILA) is an iterative and inductive machine learning algorithm that is used for generating a set of classification rules, which produces rules of the form “I
5 min read
Upper Confidence Bound Algorithm in Reinforcement Learning
In Reinforcement learning, the agent or decision-maker generates its training data by interacting with the world. The agent must learn the consequences of its actions through trial and error, rather than being explicitly told the correct action. Multi-Armed Bandit Problem In Reinforcement Learning, we use Multi-Armed Bandit Problem to formalize the
6 min read
Implementation of Perceptron Algorithm for AND Logic Gate with 2-bit Binary Input
In the field of Machine Learning, the Perceptron is a Supervised Learning Algorithm for binary classifiers. The Perceptron Model implements the following function: \[ \begin{array}{c} \hat{y}=\Theta\left(w_{1} x_{1}+w_{2} x_{2}+\ldots+w_{n} x_{n}+b\right) \\ =\Theta(\mathbf{w} \cdot \mathbf{x}+b) \\ \text { where } \Theta(v)=\left\{\begin{array}{cc
2 min read
Implementation of Perceptron Algorithm for OR Logic Gate with 2-bit Binary Input
In the field of Machine Learning, the Perceptron is a Supervised Learning Algorithm for binary classifiers. The Perceptron Model implements the following function: \[ \begin{array}{c} \hat{y}=\Theta\left(w_{1} x_{1}+w_{2} x_{2}+\ldots+w_{n} x_{n}+b\right) \\ =\Theta(\mathbf{w} \cdot \mathbf{x}+b) \\ \text { where } \Theta(v)=\left\{\begin{array}{cc
2 min read
Implementation of Perceptron Algorithm for NAND Logic Gate with 2-bit Binary Input
In the field of Machine Learning, the Perceptron is a Supervised Learning Algorithm for binary classifiers. The Perceptron Model implements the following function: \[ \begin{array}{c} \hat{y}=\Theta\left(w_{1} x_{1}+w_{2} x_{2}+\ldots+w_{n} x_{n}+b\right) \\ =\Theta(\mathbf{w} \cdot \mathbf{x}+b) \\ \text { where } \Theta(v)=\left\{\begin{array}{cc
3 min read