United Kingdom
6K followers 500+ connections

Join to view profile

About

I'm an AI researcher at Project Prometheus, working to fundamentally change how we design…

Articles by Andrei

  • Life lessons from my sophomore year doing a Computer Science degree

    This article is a collection of the thoughts and events which have shaped who I am today starting from the beginning of…

    1 Comment
  • The world needs heroes

    Problem: It’s tough for the young generation to choose a path in life and especially a career since there is complete…

  • The lessons from my CS degree

    It’s collection of thoughts, lessons and events that happened in my first year as a Computer Science (CS) student at…

    1 Comment
  • How to prepare for competitive programming ?

    This is how I won 3 out of 4 Gold medals in the Computing Olympiad. I started learning C++ from scratch during my first…

    4 Comments

Activity

6K followers

See all activities

Experience & Education

  • Project Prometheus

View Andrei’s full experience

See their title, tenure and more.

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

Volunteer Experience

  • President

    Interact: Rotary Sponsored Club

    - 2 years

    Civil Rights and Social Action

    • Founded a non-profit community service organisation and grew it to more than 30 members
    • Demonstrated the ability to take initiative by convincing the Town Hall Council to plant 500 trees
    • Acquired organisational skills by managing competitions with 150+ participants and raised 1000€ for the paediatric department of the hospital

Publications

  • TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based Models

    Neural Information Processing Systems (NeurIPS), 2024

    Data collection is often difficult in critical fields such as medicine, physics, and chemistry. As a result, classification methods usually perform poorly with these small datasets, leading to weak predictive performance. Increasing the training set with additional synthetic data, similar to data augmentation in images, is commonly believed to improve downstream classification performance. However, current tabular generative methods that learn either the joint distribution p(x,y) or the…

    Data collection is often difficult in critical fields such as medicine, physics, and chemistry. As a result, classification methods usually perform poorly with these small datasets, leading to weak predictive performance. Increasing the training set with additional synthetic data, similar to data augmentation in images, is commonly believed to improve downstream classification performance. However, current tabular generative methods that learn either the joint distribution p(x,y) or the class-conditional distribution p(x∣y) often overfit on small datasets, resulting in poor-quality synthetic data, usually worsening classification performance compared to using real data alone. To solve these challenges, we introduce TabEBM, a novel class-conditional generative method using Energy-Based Models (EBMs). Unlike existing methods that use a shared model to approximate all class-conditional densities, our key innovation is to create distinct EBM generative models for each class, each modelling its class-specific data distribution individually. This approach creates robust energy landscapes, even in ambiguous class distributions. Our experiments show that TabEBM generates synthetic data with higher quality and better statistical fidelity than existing methods. When used for data augmentation, our synthetic data consistently improves the classification performance across diverse datasets of various sizes, especially small ones.

    See publication
  • GCondNet: A Novel Method for Improving Neural Networks on Small High-Dimensional Tabular Data

    Transactions on Machine Learning Research (TMLR), 2024

    Neural networks often struggle with high-dimensional but small sample-size tabular datasets. One reason is that current weight initialisation methods assume independence between weights, which can be problematic when there are insufficient samples to estimate the model's parameters accurately. In such small data scenarios, leveraging additional structures can improve the model's performance and training stability. To address this, we propose GCondNet, a general approach to enhance neural…

    Neural networks often struggle with high-dimensional but small sample-size tabular datasets. One reason is that current weight initialisation methods assume independence between weights, which can be problematic when there are insufficient samples to estimate the model's parameters accurately. In such small data scenarios, leveraging additional structures can improve the model's performance and training stability. To address this, we propose GCondNet, a general approach to enhance neural networks by leveraging implicit structures present in tabular data. We create a graph between samples for each data dimension, and utilise Graph Neural Networks (GNNs) to extract this implicit structure, and for conditioning the parameters of the first layer of an underlying predictor network. By creating many small graphs, GCondNet exploits the data's high-dimensionality, and thus improves the performance of an underlying predictor network. We demonstrate GCondNet's effectiveness on 12 real-world datasets, where it outperforms 14 standard and state-of-the-art methods. The results show that GCondNet is a versatile framework for injecting graph-regularisation into various types of neural networks, including MLPs and tabular Transformers.

    See publication
  • ProtoGate: Prototype-based Neural Networks with Global-to-local Feature Selection for Tabular Biomedical Data

    International Conference on Machine Learning (ICML), 2024

    Tabular biomedical data poses challenges in machine learning because it is often high-dimensional and typically low-sample-size (HDLSS). Previous research has attempted to address these challenges via local feature selection, but existing approaches often fail to achieve optimal performance due to their limitation in identifying globally important features and their susceptibility to the co-adaptation problem. In this paper, we propose ProtoGate, a prototype-based neural model for feature…

    Tabular biomedical data poses challenges in machine learning because it is often high-dimensional and typically low-sample-size (HDLSS). Previous research has attempted to address these challenges via local feature selection, but existing approaches often fail to achieve optimal performance due to their limitation in identifying globally important features and their susceptibility to the co-adaptation problem. In this paper, we propose ProtoGate, a prototype-based neural model for feature selection on HDLSS data. ProtoGate first selects instance-wise features via adaptively balancing global and local feature selection. Furthermore, ProtoGate employs a non-parametric prototype-based prediction mechanism to tackle the co-adaptation problem, ensuring the feature selection results and predictions are consistent with underlying data clusters. We conduct comprehensive experiments to evaluate the performance and interpretability of ProtoGate on synthetic and real-world datasets. The results show that ProtoGate generally outperforms state-of-the-art methods in prediction accuracy by a clear margin while providing high-fidelity feature selection and explainable predictions.

    See publication
  • TabMDA: Tabular Manifold Data Augmentation for Any Classifier using Transformers with In-context Subsetting

    In-context Learning Workshop at the International Conference on Machine Learning (ICML), 2024

    Tabular data is prevalent in many critical domains, yet it is often challenging to acquire in large quantities. This scarcity usually results in poor performance of machine learning models on such data. Data augmentation, a common strategy for performance improvement in vision and language tasks, typically underperforms for tabular data due to the lack of explicit symmetries in the input space. To overcome this challenge, we introduce TabMDA, a novel method for manifold data augmentation on…

    Tabular data is prevalent in many critical domains, yet it is often challenging to acquire in large quantities. This scarcity usually results in poor performance of machine learning models on such data. Data augmentation, a common strategy for performance improvement in vision and language tasks, typically underperforms for tabular data due to the lack of explicit symmetries in the input space. To overcome this challenge, we introduce TabMDA, a novel method for manifold data augmentation on tabular data. This method utilises a pre-trained in-context model, such as TabPFN, to map the data into an embedding space. TabMDA performs label-invariant transformations by encoding the data multiple times with varied contexts. This process explores the learned embedding space of the underlying in-context models, thereby enlarging the training dataset. TabMDA is a training-free method, making it applicable to any classifier. We evaluate TabMDA on five standard classifiers and observe significant performance improvements across various tabular datasets. Our results demonstrate that TabMDA provides an effective way to leverage information from pre-trained in-context models to enhance the performance of downstream classifiers.

    See publication
  • Weight Predictor Network with Feature Selection for Small Sample Tabular Biomedical Data

    AAAI Conference on Artificial Intelligence, 2023

    Tabular biomedical data is often high-dimensional but with a very small number of samples. Although recent work showed that well-regularised simple neural networks could outperform more sophisticated architectures on tabular data, they are still prone to overfitting on tiny datasets with many potentially irrelevant features. To combat these issues, we propose Weight Predictor Network with Feature Selection (WPFS) for learning neural networks from high-dimensional and small sample data by…

    Tabular biomedical data is often high-dimensional but with a very small number of samples. Although recent work showed that well-regularised simple neural networks could outperform more sophisticated architectures on tabular data, they are still prone to overfitting on tiny datasets with many potentially irrelevant features. To combat these issues, we propose Weight Predictor Network with Feature Selection (WPFS) for learning neural networks from high-dimensional and small sample data by reducing the number of learnable parameters and simultaneously performing feature selection. In addition to the classification network, WPFS uses two small auxiliary networks that together output the weights of the first layer of the classification model. We evaluate on nine real-world biomedical datasets and demonstrate that WPFS outperforms other standard as well as more recent methods typically applied to tabular data. Furthermore, we investigate the proposed feature selection mechanism and show that it improves performance while providing useful insights into the learning task.

    See publication

Projects

  • Decode Your 20s (YouTube series)

    -

    • Decode Your 20s is a video-documentary series with over 40 episodes on Youtube that presents my views about the world and how they change as I navigate my 20s.
    • I focus on understanding human psychology and how the world works, covering topics such as self-discipline, empathy and mental models.

    https://youtube.com/playlist?list=PLEE21reVyOXWr-9waKm07QStz1waBUIb5

  • Introduction to Algorithms and Data structures in C++

    -

    • I created a course teaching basic algorithms and data structures in C++ to over 100,000 students from 160+ countries

    See project

Honors & Awards

  • Best Romanian PhD student in the UK

    The Romanian Embassy in the UK

    Awarded by the Romanian Ambassador in the UK to only one PhD student.

  • Top 3 Romanian Undergraduate Students in Europe

    League of Romanian Students Abroad

  • Bloomberg CodeCon Global Finalist

    -

  • Best Romanian Undergraduate Student in the UK

    The Romanian Embassy in the UK

    Awarded by the Romanian Ambassador in the UK to only one undergraduate student.

  • UCL Engineering Award for Excellence

    -

    The Honour is awarded by the Dean to the Best First-year student in the whole Engineering Faculty.

  • World Finalist in Google Hashcode

    -

    • The World Finals of the largest team-based algorithmic competition organised by Google
    • I led the youngest team in the Finals
    • Had to optimise the placement of WI-FI routers in a building in order to get the most coverage and reduces the costs and our solution was just 1.5% slower than the winning one

  • 1st team from London in NWERC 2016

    -

    Ranked 1st team from London in the largest algorithmic competition Northwestern Europe Regional Contest part of ACM.

  • Gold medal - Computing Olympiad C/C++

    Ministry of Education and Research

    Ranked 4th in Romania.

  • 1st - Info-Oltenia C/C++ programming contest

    Ministry of Education and Research

    Ranked 1st in both individual and team tasks.

  • Gold medal - Computing Olympiad C/C++

    Ministry of Education and Research

    Ranked 2nd in Romania.

  • 1st - National Computing Olympiad C/C++ Online

    Ministry of Education and Research

  • Gold medal - National Computing Olympiad C/C++

    Ministry of Education and Research

    Ranked 4th in Romania.

Languages

  • Engleză

    Native or bilingual proficiency

  • Romanian

    Native or bilingual proficiency

  • French

    Limited working proficiency

Recommendations received

View Andrei’s full profile

  • See who you know in common
  • Get introduced
  • Contact Andrei directly
Join to view full profile

Other similar profiles

Explore collaborative articles

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Explore More