The next generation of drug discovery will be built on integrated, AI-native systems that connect automation, experimentation, computation, and machine learning into a continuous cycle of learning where every experiment makes the next one smarter. A recent article in Scientific American by Patrick Sisson explores how this new research infrastructure is beginning to take shape and highlights Recursion's pioneering work in this space. At Recursion, we run up to 2.2 million experiments each week and leverage more than 50 petabytes of proprietary biological, chemical, and patient data as part of an end-to-end learning engine for drug discovery and development. But scale alone isn't the differentiator. The real opportunity lies in transforming multimodal data into biological understanding and ultimately into new medicines. Take our neuroscience collaboration with Roche and Genentech. For decades, neuroscience drug discovery has been constrained by repeatedly investigating the same well-studied targets. To move beyond those limitations and explore entirely new biology, our teams developed advanced cell manufacturing capabilities to produce more than 100 billion human iPSC-derived microglia, the brain's resident immune cells, which are notoriously difficult to generate and study at scale. The result is a first-of-its-kind whole-genome Microglia Map comprising 46 million cellular images across 17,000 genes. This systems-level view of biology allows our AI models to move beyond traditional approaches, uncover novel biological insights, and identify therapeutic opportunities that may have otherwise remained hidden. What excites me most is what comes next. These maps – and the AI models trained on them – are the foundation. The real opportunity is translating them into novel, first-in-class therapeutic programs. That's the frontier we're pioneering: turning systems-level biological understanding into medicines for patients. There's still important work ahead, but we're making meaningful progress, and I'm excited about what's possible as we continue to push the boundaries of AI-native drug discovery. Stay tuned. #AI #DrugDiscovery #TechBio #Biotechnology #MachineLearning
AI in Molecular Prediction
Explore top LinkedIn content from expert professionals.
-
-
Big breakthrough: A few months my lab at MIT introduced SPARKS, our autonomous scientific discovery model. Since then we have demonstrated applicability to broad problem spaces across domains from proteins, bio-inspired materials to inorganic materials. SPARKS learns by doing, thinks by critiquing itself & creates knowledge through recursive interaction; not just with data, but with the physical & logical consequences of its own ideas. It closes the entire scientific loop - hypothesis generation, data retrieval, coding, simulation, critique, refinement, & detailed manuscript drafting - without prompts, manual tuning, or human oversight. SPARKS is fundamentally different from frontier models. While models like o3-pro and o3 deep research can produce summaries, they stop short of full discovery. SPARKS conducts the entire scientific process autonomously, generating & validating falsifiable hypotheses, interpreting results & refining its approach until a reproducible, fully validated evidence-based discovery emerges. This is the first time we've seen AI discover new science. SPARKS is orders of magnitude more capable than frontier models & even when comparing just the writing, SPARKS still outperforms: in our benchmark evaluation, it scored 1.6× higher than o3-pro and over 2.5× higher than o3 deep research - not because it writes more, but because it writes with purpose, grounded in original, validated compositional reasoning from start to finish. We benchmarked SPARKS on several case studies, where it uncovered two previously unknown protein design rules: 1⃣ Length-dependent mechanical crossover β-sheet-rich peptides outperform α-helices—but only once chains exceed ~80 amino acids. Below that, helices dominate. No prior systematic study had exposed this crossover, leaving protein designers without a quantitative rule for sizing sheet-rich materials. This discovery resolves a long-standing ambiguity in molecular design and provides a principle to guide the structural tuning of biomaterials and protein-based nanodevices based on mechanical strength. 2⃣ A stability “frustration zone” At intermediate lengths (~50- 70 residues) with balanced α/β content, peptide stability becomes highly variable. Sparks mapped this volatile region and explained its cause: competing folding nuclei and exposed edge strands that destabilize structure. This insight pinpoints a failure regime in protein design where instability arises not from randomness, but from well-defined physical constraints, giving designers new levers to avoid brittle configurations or engineer around them. This gives engineers and biologists a roadmap for avoiding stability traps in de novo design - especially when exploring hybrid motifs. Stay tuned for more updates & examples, papers and more details.
-
For 50 years, a key protein behind heart disease, among the leading cause of death worldwide remained a scientific mystery. It was too large and complex for traditional methods; its structure was invisible to us. Now, researchers have combined cryo-electron microscopy with DeepMind's AlphaFold to reveal the atomic structure of that protein: apoB100, the very scaffold of "bad cholesterol." This marks a deeper shift in how we approach science. When we can see biology at this level of detail, healthcare moves from managing symptoms to engineering interventions at the molecular root. AI starts to function as a new kind of microscope, one that reveals the invisible machinery of life and allows entirely new questions to be asked. This is the kind of progress that matters. AI as an instrument for understanding, precision, and prevention. It’s a glimpse into a future where compute and science converge to tackle humanity’s hardest health challenges at their source. Read the full story: https://lnkd.in/gbum2dKu #AIInHealthCare #AIForGood
-
I’m often asked where I see AI make a tangible, real impact in the world today. To that, I answer with #AlphaFold, the revolutionary AI model from Google DeepMind, that is able to predict the structure of a protein simply from its amino acid sequence. 5 years ago, AlphaFold solved the 50-year grand challenge of protein folding, followed by the equally meaningful decision to make 200 million protein structures freely available to the scientific community. Since then, Demis Hassabis and John Jumper have been recognized with a Nobel Prize for their work on AlphaFold, and we see over 3.3 million users of it globally, with more than a third of users right here in Asia-Pacific. Here is just a snapshot of those applications: 🔬 Dr. Su Datt Lam at the National University of Malaysia (UKM) is learning more about Melioidosis to better fight the silent killer. 🧬 Researchers Lim Jackwee lim and Yinxia Chao at Singapore’s A*STAR - Agency for Science, Technology and Research and National Neuroscience Institute (NNI) are visualizing proteins linked to Parkinson’s. 🔍 Professor Ji-Joon Song’s team at the Korea Advanced Institute of Science and Technology lead to cancer and other diseases. 🪢Dr. Danny Hsu at Academia Sinica, Taiwan is advancing our understanding of exceptionally complex protein “knots”. ♨️ Dr. Syun-ichi Urayama’s team is uncovering new evolutionary insights from microbes in Japan’s hot springs! Listen to one of their stories below, and read more about all of them here: https://lnkd.in/d7wyACpK #GoogleDeepMind #AIforGood
-
Five years ago, AlphaFold solved the protein structure prediction problem at CASP14, cracking a 50-year grand challenge in biology. It has been an absolute honour and privilege to have been part of this journey alongside Demis and John. Over 3 million researchers across 190 countries have since used AlphaFold to predict the structure of more than 200 million proteins. The impact spans from revealing apoB100's structure, advancing heart disease research, to supporting endangered honeybee conservation in Europe. Protein structure prediction was the root node problem in structural biology. By solving it, we opened up entirely new avenues for discovery. What AlphaFold demonstrated is that AI can accelerate scientific progress when applied to the right foundational challenges. We've since expanded this approach across biology. AlphaMissense and AlphaGenome are helping researchers understand genetic mutations and disease. AlphaProteo is designing new protein binders for targets in cancer and diabetes. We're applying similar thinking to challenges in fusion energy, materials discovery and climate science. Today, we're sharing The Thinking Game, following our team through the journey that made AlphaFold possible. To understand more about AlphaFold's impact, see the blog here: https://lnkd.in/eiPSAeKc #AlphaFold #AIforScience
-
Biotech company Insilico Medicine is using AI to rethink drug discovery. Now, with a fresh $110 million in funding pushing its valuation past $1 billion, the startup is considering a Hong Kong IPO. Developing a new drug is notoriously slow and expensive—it can take over a decade and billions of dollars before a single treatment reaches patients. Insilico wants to change that with AI, making the process faster, cheaper, and more precise. ► At the core of Insilico’s approach is Pharma.AI, which analyzes vast biological datasets and predicts which molecules are most likely to work, reducing the need for excessive trial and error. ► Insilico is already delivering results—its leading drug candidate, Rentosertib, for a serious lung disease, reached early clinical trials in just 2.5 years, a process that normally takes up to six. ► The company is making big deals by licensing its AI-generated drugs to pharmaceutical giants like Sanofi and Fosun Pharma, securing $3.5 billion in contract value. The startup’s pipeline now includes 30 drug candidates, with 10 receiving clearance from the US Food and Drug Administration (FDA) to proceed with human trials. Beyond discovery, Insilico is using AI to speed up lab work. The company is testing humanoid robots in its China lab to automate repetitive tasks, collect data, and reduce human error, all in an attempt to speed up research. If Insilico succeeds, it could reshape the entire drug discovery process. By combining AI, real drug candidates, and even lab robots, the company is tackling the whole process, not just one piece of it. Faster, cheaper, and maybe even better—this could be a glimpse of how future drugs come to life. #artificialintelligence #innovation
-
AI and Protein Localization! 🧬🔬 The journey from understanding protein structures to predicting their precise locations within cells has taken a monumental leap forward. Introducing ProtGPS, a cutting-edge machine-learning model developed by researchers at the Whitehead Institute and Massachusetts Institute of Technology's CSAIL, led by Professor Richard Young and his team. Why is this a game-changer? 🤔 🔹 Predictive power: ProtGPS accurately forecasts where proteins will localize in cells, crucial for understanding both their functions and the mechanisms of diseases. 🔹 Disease insight: By examining over 200,000 proteins with disease-associated mutations, ProtGPS uncovers profound links between mis-localization and disease, paving the way for novel therapeutic strategies. 🔹 Generative potential: Beyond predictions, ProtGPS creates new proteins, designing sequences to target specific cellular locales. This innovation could revolutionize drug design by enhancing precision and minimizing side effects. 🔹 Experimental validation: Unlike many AI models, ProtGPS's predictions have been validated in real cell experiments, bridging the gap between computational design and biological application. The potential applications? Endless.....From developing targeted therapies to uncovering fundamental cellular mechanisms, the implications of this research are vast. ProtGPS isn't just a tool; it’s the start of a new era in biological exploration and therapeutic innovation. #AI #MachineLearning #Biotechnology #Proteomics #ResearchInnovation #Therapeutics #MIT #WhiteheadInstitute
-
Exciting progress in AI x Biology you should know about: EvolutionaryScale's ESM3, a new language model, simulates 500 million years of protein evolution. As someone working in AI x Biology, I dove into this right away. (1) This model, trained on an extensive dataset, can generate diverse protein sequences, structures, and functions. ESM3 has already demonstrated its capability by creating esmGFP, a green fluorescent protein analogous to evolving over half a billion years. This achievement underscores ESM3’s potential to revolutionize programmable biology and protein design. (2) ESM3 integrates multimodal reasoning, allowing precise control over protein creation. This opens doors to significant advancements in medicine, biological research, and sustainable energy solutions. The model’s ability to reason across different modalities sets a new benchmark for AI in scientific research. (3) Moreover, EvolutionaryScale’s commitment to open science ensures that ESM3’s models and data are accessible, fostering collaboration and responsible AI development. This transparency is vital for accelerating scientific discoveries and practical applications. (A) I find the quality of ESM3’s work impressive, showcasing a sophisticated understanding of protein biochemistry. Its capacity to generate high-quality, functional proteins far removed from existing variants illustrates its transformative potential. Future research could explore ESM3’s application in developing specific therapeutic proteins or industrial biocatalysts, paving the way for innovative solutions across various fields. (B) However, again with powerful generative AI models, a key area for improvement is optimizing the model’s efficiency to balance complexity and performance, making it more accessible for broader scientific and industrial use. I believe ESM3 stands as a testament to the power of AI in advancing biological research and technology. Its implications are far-reaching, promising a future of accelerated scientific breakthroughs and innovative applications. Paper: https://lnkd.in/gvv8FKPA Blog Post: https://lnkd.in/gQRQKExG #GenAI #Biology #ArtificialIntelligence
-
OpenAI launches GPT-Rosalind to bring specialised AI reasoning into drug discovery: 🔘OpenAI has launched GPT-Rosalind, its first purpose-built AI model for life sciences, designed specifically for biology, drug discovery, and translational medicine rather than adapting a general model to scientific use 🔘The model is positioned as a reasoning engine for science, combining chemistry, genomics, and protein biology with the ability to navigate data, tools, and literature in a single workflow rather than treating each step in isolation 🔘In practical terms, it targets the hardest parts of early R&D such as understanding protein function, identifying drug targets, and predicting interactions, areas where failure rates are high and timelines can stretch to a decade or more 🔘Unlike typical AI tools, GPT-Rosalind is designed to actively support scientific workflows by generating hypotheses, retrieving evidence, and even suggesting experimental or chemical optimizations, effectively acting as a co-pilot for researchers 🔘Access is restricted to pharma, biotech, and research institutions, reflecting both the sensitivity of biological research and the need for expert validation in high-stakes domains like drug development 💬This signals a shift from general-purpose AI toward domain-specific reasoning systems in pharma R&D, where the competitive edge will come less from having AI and more from embedding it deeply into end-to-end scientific workflows #digitalhealth #ai #pharma
-
Delighted to share new Arc Institute work from our group on AI-accelerated lab-in-the-loop, in Science Magazine today. One of the most remarkable things about biology is that it's digital. DNA, RNA, proteins: these are all sequences, and their function is directly encoded in their sequence of letters. But a protein of length N has 20^N possible variants and the vast majority are non-functional. Evolution spent billions of years finding the functional needles in this haystack through random exploration and natural selection. For modern biomedicine, we need to solve this in days to weeks. The process of scientific research is fundamentally a search problem, and we basically do guess and check. We've trained predictive models of biology, like our Evo series of DNA language models, to learn the evolutionary constraints on biological sequences. Such models learn a fitness landscape of what evolution has explored. But the fitness landscape is not the same as the function you actually care about: whether an enzyme catalyzes faster, whether an antibody binds tighter, whether a CRISPR tool edits better. The core question is how do you connect the knowledge of these models to the functional search that has to happen in the physical lab? MULTI-evolve is one of our first answers. It's a full-stack, AI-lab-in-the-loop framework that "jumps" directly to hyperactive multi-mutant proteins via ML-guided evolution. We combine an ensemble of protein language models pretrained on all proteins across evolution to discover beneficial mutations, then systematically measure pairwise combinations to learn the epistatic landscape (e.g. which mutations are synergistic vs. antagonistic), and extrapolate to predict powerful higher-order combinations of 5-7+ mutations. We also built MULTI-assembly, a molecular biology method to physically construct these complex multi-mutants cheaply and quickly, regardless of protein length (previously a major bottleneck). We applied this to three very different proteins: APEX (a proximity labeling enzyme), CRISPR-Cas13d (for RNA trans-splicing), and a therapeutic anti-CD122 antibody, achieving up to 256-fold improvement. For CD122, we simultaneously optimized both binding affinity and antibody expression, navigating real trade-offs between competing developability objectives. We built Arc Institute to be a full-stack biology and AI research organization. We've been training frontier AI models for biology, like Evo, Evo 2, State, Stack, etc. But the whole point of having frontier AI capabilities and experimental biologists under a single physical roof is to close the loop between computation and the wet lab. MULTI-evolve is one of Arc's first examples of AI lab-in-the-loop, and there will be much more to come. This work was a wonderful collaboration with Silvana Konermann and Brian Hie, and led by the remarkably driven and persistent Vincent Tran with Matthew Nemeth, Liam Bartie, Sita C., Alison Fanton, and Chad Moon.