About
Activity
6K followers
Experience & Education
Licenses & Certifications
Volunteer Experience
-
High School Teacher
Girija Welfare Association - India
- 2 years 3 months
Children
• Taught Mathematics, Science, History and Geography to Primary and Higher Secondary School Children at Girija Orphanage managed by Girija Welfare Association
• Worked on children's personal development skills -
Technical Event Head
Computer Society of India
- 3 years 4 months
Science and Technology
• Organized technical events during national level technical festivals
• Worked in DBA team to store and manage large amount of events data
Courses
-
Communication and Visualization for Data Analytics
ALY6070 21347
-
Data Mining Applications
ALY6040 21289
-
Foundation of Informatics
ALY6400 70856
-
Intermediate Analytics
ALY6015 21280
-
Introduction to Analytics
ALY6000 71151
-
Probability and Statistics
ALY6010 70942
Projects
-
Image Classification, Object Detection, and Semantic Segmentation using Amazon Sagemaker and S3
-
See project• Trained and Deployed an Image Classifier, object detector and semantic segmentation by using the image classification algorithm, SSD Object Detection algorithm and semantic segmentation algorithm from Amazon Sagemaker.
• The model classified 37 breeds, localizes the faces and segments image pixels as foreground, background or transition image of dogs and cats from the famous IIIT-Oxford Pets Dataset with 79% accuracy, 96% and 89%, respectively -
Startup Evaluator - Capstone Project
-
See project• Analyzed the factors affecting the startup’s and crowdfunding campaign’s success; and predict success rate
• Web scraped the startups data and performed data cleaning, feature engineering techniques to transform data; implemented Decision Tree, Random Forest, and Gradient Boosting to predict success rate; Survival Analysis models to to predict survival rate over time
• A web-based decision support system backed with descriptive and predictive analytics which estimates the operational…• Analyzed the factors affecting the startup’s and crowdfunding campaign’s success; and predict success rate
• Web scraped the startups data and performed data cleaning, feature engineering techniques to transform data; implemented Decision Tree, Random Forest, and Gradient Boosting to predict success rate; Survival Analysis models to to predict survival rate over time
• A web-based decision support system backed with descriptive and predictive analytics which estimates the operational survival probability of startups based on funding patters; Also suggested how social media presence of startups can influence audience and raise targeted fund amount.
GitHub Link: https://github.com/Sagar401/Startup_Assessment_using_ML -
GE Aviation IP Risk Analysis and Modeling
-
See project• Developed IP Theft detection model and address the factors causing risks over time
• Performed unsupervised learning using K-Mean and built Decision Tree Classifier, Random Forest classifier, Gradient Boosting using python
• Analyzed factors like high-risk unit 11, indicator email 13, career band, and job functions which were clustered with High Risk Alerts; Using the Decision Tree classifier achieved an accuracy of 92.37% with a smaller number of false negatives and false positives…• Developed IP Theft detection model and address the factors causing risks over time
• Performed unsupervised learning using K-Mean and built Decision Tree Classifier, Random Forest classifier, Gradient Boosting using python
• Analyzed factors like high-risk unit 11, indicator email 13, career band, and job functions which were clustered with High Risk Alerts; Using the Decision Tree classifier achieved an accuracy of 92.37% with a smaller number of false negatives and false positives thus saving the cost and efforts of the GE team to manually classify the alerts.
-
Predicting Molecular Properties (Kaggle Live Competition)
-
See project• Predicted the magnetic interactions between atoms in a molecule called Scalar Coupling Constant.
• Implemented linear regression with gradient boosting to optimize a model.
• Analyzed the molecule distance of different scalar coupling constant and used correlated variables such as dipole moments, magnetic shielding tensors, potential energy, etc. to predict scalar coupling constant. Using XGBoost received 92% accuracy.
-
Big Data Project-Energy Consumption Trends in the Netherlands
-
• Analyzed the spread of smart meters, a number of connections and energy consumptions patterns over the years in the cities of the Netherlands.
• Worked on 2.5 million rows data, created Data Lake Storage Gen2 on Microsoft Azure Portal, cleaned the data using PySpark and SQL query processing on databricks and visualized the cleaned data by connecting it to Power BI.
• Observed that the number of connections and smart meters is growing linearly; however, the energy consumption in cities…• Analyzed the spread of smart meters, a number of connections and energy consumptions patterns over the years in the cities of the Netherlands.
• Worked on 2.5 million rows data, created Data Lake Storage Gen2 on Microsoft Azure Portal, cleaned the data using PySpark and SQL query processing on databricks and visualized the cleaned data by connecting it to Power BI.
• Observed that the number of connections and smart meters is growing linearly; however, the energy consumption in cities seemed to decrease.
-
Heart Disease UCI Analysis and Prediction, Northeastern University, Boston
-
See project• Examined the features correlating with heart disease and predicted whether a person would have heart disease or not.
• Implemented Exploratory Data Analysis (EDA) on the various attributes and used the Logistic Regression in R with train function in the caret package.
• Observed that patients with chest pain type 2, diabetes, fixed defect thalassemia are more likely to have heart disease. Patient with a higher number of blood vessels that are visible by fluoroscopy and exercise-induced…• Examined the features correlating with heart disease and predicted whether a person would have heart disease or not.
• Implemented Exploratory Data Analysis (EDA) on the various attributes and used the Logistic Regression in R with train function in the caret package.
• Observed that patients with chest pain type 2, diabetes, fixed defect thalassemia are more likely to have heart disease. Patient with a higher number of blood vessels that are visible by fluoroscopy and exercise-induced angina have lower chances of having heart Disease — predicted on test data with ROC 88%, Sens 74%, Spec 85% and Accuracy 89.47%.
-
Time Series Data Analysis using LSTM Neural Network
-
• Cleaned and visualized the household electric power consumption dataset using NumPy, Pandas, Seaborn, and Matplotlib.
• Applied LSTM neural network to predict global active power.
• Visualized the electricity consumption quarter and month wise. Observed that the global intensity and global active power are correlated. Used MAE loss function to get the optimal point and minimize the error.
-
Olympic Historical Data Analysis
-
See project• Cleaned, visualized and analyzed the change in several athletes, events, nations, female participation, factors related to winning medals using python.
• Applied K-means clustering algorithm to form clusters of Medal Counts with GDP, Population, Height and Weight distribution.
• Analyzed the Olympics development, trends of participation, winning attributes and performance of women over the years from cluster analysis observed that the countries with the highest GDP have won the highest…• Cleaned, visualized and analyzed the change in several athletes, events, nations, female participation, factors related to winning medals using python.
• Applied K-means clustering algorithm to form clusters of Medal Counts with GDP, Population, Height and Weight distribution.
• Analyzed the Olympics development, trends of participation, winning attributes and performance of women over the years from cluster analysis observed that the countries with the highest GDP have won the highest number of medals. Medal winners and non-winners follow the same distribution, and all players have almost equal chance to win a medal.
-
Data Analysis and Statistics Modeling using R programming and Microsoft Excel
-
See project• Worked on projects of Probability & Statistics on the data of a manufacturing company, US occupational data, and US vehicles per household data using R programming and Microsoft Excel.
• Used the concepts of Descriptive Statistics, Probability Theory & Distribution, One-Sample, and Two-Sample Confidence Interval Testing, Hypothesis Testing, Chi-squared Test of Goodness of Fit and Independence.
-
Search Engine Optimization Audit tool: (Undergraduate Final Year Project)
-
Made Search Engine Optimization Audit Tool which helped new web developers know their weaknesses and ways to improve them. Worked on Frond end & Backend development, including the use of JavaScript, HTML, CSS & MySQL
Recommendations received
4 people have recommended Arvind
Join now to viewOther similar profiles
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content