Sign in to view Matthew’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Matthew’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
McLean, Virginia, United States
Sign in to view Matthew’s full profile
Matthew can introduce you to 10+ people at Nebius
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
26K followers
500+ connections
Sign in to view Matthew’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Matthew
Matthew can introduce you to 10+ people at Nebius
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Matthew
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Matthew’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Websites
- Personal Website
-
http://www.matthewzeiler.com
- Company Website
-
http://www.review-mate.com
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Articles by Matthew
-
The hashtag is dead, and machine learning killed it
The hashtag is dead, and machine learning killed it
In this series, retail experts at Shoptalk discuss the most pressing issues facing their industries today. Write your…
2,255
130 Comments
Activity
26K followers
-
Matthew Zeiler shared thisExcited to see the Nebius Research Grant Program continue to support AI research. We hope this program helps more researchers explore ambitious ideas, share their work with the community, and accelerate progress for everyone. Nebius doesn't just want to provide compute, Nebius Research is looking forward to collaborating with grant recipients on novel research projects and helping turn promising ideas into impactful results. If you're working on cutting-edge AI research, I'd encourage you to apply. Looking forward to working with you!Matthew Zeiler shared thisHow do you join the Nebius Research Grant Program? Ambitious projects shouldn't be limited by infrastructure. To support innovation and discovery, we provide researchers working in AI and AI-driven science with the compute resources they need. 👉 Our main goal is to boost open science and open source – learn which projects we prioritize and what research areas we actively support in our guide for researchers: https://lnkd.in/gxE5Bpe8 👉 Apply for a grant: https://lnkd.in/eFCbmEVu
-
Matthew Zeiler posted thisThis past week has been a whirlwind as we started at Nebius. The pace and scale at which we are already diving in alongside other exceptional talent makes each day energizing. We're excited to join the team and push AI innovations further at this rocket ship of a company! Huge shout‑out to the Clarifai team for over a decade of breakthroughs in the AI space. Your talent, grit and tenacity is unmatched and I’m grateful to know you all. Although it was an meaningful and fulfilling time building Clarifai, I truly can’t wait to see what we do at our new home. I am thrilled to go back to my roots and lead the research efforts for Nebius at a much larger scale. If you’re a researcher interested in doing novel cutting edge work at the next Hyperscaler, send me a DM. To Danila Shtan, Roman Chernin, Ophir Nave, Arkady Volozh and the rest of Nebius team thanks for the warm welcome, let’s go!!! 🚀
-
Matthew Zeiler shared thisThrilled to officially announce that Clarifai has entered into an agreement with Nebius to license our AI inference and compute orchestration IP, along with the patent portfolio behind it. Our core engineering and research team will be joining Nebius to keep building. For over a decade, we’ve obsessed over making AI inference faster and more cost-efficient. Now we get to take that work to a much bigger stage. Nebius is building toward becoming the world's next hyperscale-grade AI cloud, and our technology becomes a core part of their Token Factory — the inference platform powering their full-stack AI cloud. Personally, I'm joining as SVP Research at Nebius, leading a new unit focused on multimodal agentic reasoning, world models, token efficiency, and long-term memory. To every teammate, customer, and partner who was part of the Clarifai journey — thank you. Can't wait to ride this rocket ship and build the ultimate foundation for the next decade of AI inference. 🚀 Full announcement here: https://lnkd.in/gksaJGPs
-
Matthew Zeiler shared thisExcited to share that NVIDIA Nemotron 3 Nano Omni is now available on Clarifai with Zero Day support. This is a strong multimodal model from NVIDIA for developers building agentic systems that need to work across documents, images, video, audio, and text without stitching together separate stacks for each modality. On Clarifai Reasoning Engine, it runs at 400 tokens per second, giving developers the throughput needed for real production workflows. You can try it in the Clarifai Playground or access it through an OpenAI-compatible API. Try it out here: https://lnkd.in/gCiAbKKK
-
Matthew Zeiler shared thisClarifai is heading to HumanX! Last week at GTC, we announced 414 TPS on Kimi K2.5 - first provider to reach this performance. Now leading on Qwen3.5-397B at 290 TPS, benchmarked by Artificial Analysis. Clarifai Reasoning Engine is purpose-built for production reasoning workloads. Stop by booth #405, April 6-9. I'll be there - let's connect.
-
Matthew Zeiler shared thisClarifai just hit 290 tokens per second on Qwen3.5-397B. Fastest on Artificial Analysis. This follows our 414 TPS on Kimi K2.5 from earlier this week. Clarifai Reasoning Engine is purpose-built for reasoning workloads.
-
Matthew Zeiler shared thisClarifai just showed up in Jensen's GTC keynote. 🚀 We hit 418 tokens per second on Kimi K2.5. First provider to break 400 TPS on a trillion-parameter reasoning model. If you're at GTC building with reasoning models, let's talk: https://lnkd.in/e3P_jirx
-
Matthew Zeiler shared this🚀 Clarifai featured on Ramp’s Top Software Vendors for February Ramp's list is based on real spend and new customer adoption across thousands of companies. Clarifai’s inclusion here reflects how teams are choosing platforms they rely on to run AI in production. We’re seeing strong adoption from AI-native and digital-native teams building beyond a single request or a single model. Long-running workflows, agentic systems, and high-performance inference under real cost constraints. That’s where we’ve been focused. Over the past year, we’ve invested in the foundations required to support this shift: compute orchestration that runs models at scale across any infrastructure, pipelines for asynchronous and long-running AI workflows, and the Clarifai Reasoning Engine built for agentic and reasoning workloads. We’ve also aligned pricing with reality with simple, predictable pay-as-you-go billing. This reflects teams choosing Clarifai not just to try things, but to operate AI systems in production. Proud of the team and the builders choosing Clarifai. 💪 Ramp blog: https://lnkd.in/eXh7i7nF
-
Matthew Zeiler shared thisScaling agentic systems is no longer just about better models. It is about coordinating compute, data, and execution across local and cloud environments so agents can actually run end to end. I’ll be speaking at NeurIPS today about how we approach this with Compute Orchestration and why this architecture matters for teams building real agent workloads. If you're at NeurIPS, stop by: 📅 Dec 4, 2025 ⏰ 4:00 PM 📍 San Diego Convention Center, Exhibit Hall A,B
-
Matthew Zeiler reacted on thisSeven years ago, I dove into the world of AI. My time at Clarifai was the true definition of a startup journey. It demanded grit, and it was rarely easy. What stands out, by far, is the people — being surrounded by a team that shared a common vision and never stopped. Now, that journey is taking an exciting new direction. I am thrilled to share that I will be joining Nebius along with Matt Zeiler and our technical team. We are taking what we built and applying it to an incredible new mission: Building the 4th Hyperscaler. To everyone who believed in us and helped us along the way—thank you. A special shout out to Matthew Zeiler - for always setting a high bar and finding the right balance. Joshua Tepper Michael Gormish Brock Dusome Arman C Kizilkale Chris Mulder Douglas Shapiro Sajai K. Matthew Metzger Rob Spectre and so many more. ❤️Matthew Zeiler reacted on thisToday, Nebius announced that it is licensing Clarifai’s full-stack orchestration platform for enterprise AI, along with its inference and compute orchestration technology, and welcoming core engineering and research talent from Clarifai. The transaction strengthens Nebius Token Factory as a full-stack inference platform following Nebius’s recent acquisition of Eigen AI. While Eigen AI optimizes at the model level, Clarifai’s technology optimizes the system, creating the end-to-end infrastructure required to run complex AI models reliably in production. Clarifai founder and CEO Matthew Zeiler will join Nebius as SVP, Research, where he will lead a team focused on frontier AI innovation across areas including multimodal agentic reasoning, world models, token efficiency, and long-term memory. Roman Chernin, co-founder and Chief Business Officer of Nebius, said: “We are building a complete inference ecosystem, because delivering efficient execution at scale is a system optimization game: model optimization, system design, and compute orchestration all have to work together. The integration of Clarifai’s advanced system-building capabilities and proven team will further strengthen Nebius Token Factory, offering customers the infrastructure they need to run models reliably and cost-effectively in production.” Read the full announcement: https://lnkd.in/enjsAaK9
-
Matthew Zeiler liked thisMatthew Zeiler liked thisNVIDIA Nemotron 3 Nano Omni is now available on Clarifai with Zero Day support. This new 30B A3B multimodal reasoning model from NVIDIA is built for agent workflows that need to work across documents, images, video, audio, and text within a single reasoning loop. Why it stands out: → Built for multimodal sub-agents: Nano Omni supports text, image, video, and audio inputs with text output, making it well suited for computer use, document intelligence, and audio-video reasoning. → Efficient by design: With a hybrid Mixture-of-Experts architecture, Transformer-Mamba design, 3D convolution layers, and Efficient Video Sampling, Nano Omni is designed to deliver strong multimodal reasoning with efficient video and temporal understanding. → Agent-friendly deployment: Nano Omni can run on a single H100, H200, or B200, making it practical for production sub-agent workloads. On Clarifai Reasoning Engine, NVIDIA Nemotron 3 Nano Omni runs at 400 tokens per second, giving developers the throughput needed for production multimodal agent workflows. Try the model: https://lnkd.in/gvfH3tQt Read the blog: https://lnkd.in/gUSw_-uk
-
Matthew Zeiler liked thisMatthew Zeiler liked thisClarifai is now leading on Qwen3.5-397B performance. 🚀 290 tokens per second, fastest on Artificial Analysis benchmarks. Qwen3.5 is a 397-billion-parameter reasoning model built for complex agentic workflows. Combined with our 414 TPS on Kimi K2.5 from earlier this week, we're showing consistent performance across reasoning models. We're at GTC this week. If you're working on reasoning models or agentic workloads, meet the team here: https://lnkd.in/e-3VMqJA
Experience & Education
-
Nebius
*** ********
-
********
******* *** ***
-
******
******** *********** ******
-
*** **** **********
****** ** ********** ****** ******** ******* undefined
-
-
********** ** *******
******** ** ******* ******* *********** *******
-
View Matthew’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Honors & Awards
-
NSERC Canadian Graduate Scholarship (Doctorate)
National Science and Engineering Research Council of Canada
-
MacCracken Fellowship
New York University
Full tuition and stipend support for PhD duration.
-
NSERC Canadian Graduate Scholarship (Masters)
National Science and Engineering Research Council of Canada
-
Entrepreneurship Stream
University of Toronto Engineering Science
Completed the Entrepreneurship Stream during my Engineering Science program.
-
5T3 Engineering Award
University of Toronto Alumni Class of 5T3
-
Chown Centennial Scholarship
University of Manitoba
Entrance scholarship awarded for academic excellence.
-
Governor General's Award
Canada
Awarded for having the highest academic average in final year of high school.
-
McDonald's BEST Scholarship
McDonald's
Awarded for basketball skills and academic excellence by McDonald's.
-
North Eastman Health Association Scholarship
North Eastman Health Association
Awarded for academic excellence.
-
University of Manitoba Entrance Scholarship
University of Manitoba
Entrance scholarship to University of Manitoba based on academic excellence.
View Matthew’s full profile
-
See who you know in common
-
Get introduced
-
Contact Matthew directly
Explore more posts
-
SciPulse
32 followers
𝐓𝐡𝐞 𝐁𝐫𝐞𝐚𝐤𝐭𝐡𝐫𝐨𝐮𝐠𝐡 𝐢𝐧 𝐒𝐞𝐥𝐟-𝐃𝐢𝐬𝐭𝐢𝐥𝐥𝐚𝐭𝐢𝐨𝐧 Current RL methods for LLMs often rely on a simple binary signal: Success or Failure. But in complex domains like math and coding, a "No" doesn't tell the model why it failed. This creates a massive credit-assignment bottleneck. At SciPulse, we just finished a deep dive into a revolutionary approach: Self-Distillation Policy Optimization (SDPO). The Core Shift: Instead of a simple scalar reward, SDPO leverages Rich Textual Feedback (like compiler errors or judge evaluations). It essentially teaches the model to become its own tutor, distilling its ability to retrospectively identify mistakes back into its policy. Why this matters for the industry: - 3x Fewer Attempts: Achieving the same discovery probability as best-of-k sampling with significantly less compute. - Sample Efficiency: Outperforms standard RLVR baselines across scientific reasoning and competitive programming (LiveCodeBench v6). - Beyond Code: The methodology suggests a future where AI can learn from any "rich" environment without needing an external teacher or explicit reward model. The transition from binary "right/wrong" signals to dense, feedback-driven learning is a fundamental step toward more autonomous, reasoning-capable AI. 🎥 𝐏𝐨𝐝𝐜𝐚𝐬𝐭: https://lnkd.in/eU9gXnKK 🎥 𝐕𝐢𝐝𝐞𝐨: https://lnkd.in/e8eR4FSP 📄 𝐎𝐫𝐢𝐠𝐢𝐧𝐚𝐥 𝐏𝐚𝐩𝐞𝐫: https://lnkd.in/eubMwsfp #AIResearch #MachineLearning #ReinforcementLearning #LLMs #ComputerScience #SelfDistillation #SciPulse #TechInnovation
1 Comment -
Ovadya Menadeva
Deep Algorithm • 10K followers
The most important AI conference Key Highlights Worth Knowing: NeurIPS 2025 Best Paper Awards: NeurIPS just announced the 7 winning papers for 2025 and the themes say a lot about where AI research is heading. This year’s awards spotlight breakthroughs across LLM diversity, attention mechanisms, RL scaling, diffusion theory, reasoning limits, online learning, and neural scaling laws. 🏆 Best Papers : • Artificial Hivemind — A massive 26K-prompt benchmark revealing how LLMs collapse into homogenous outputs, raising long-term concerns about creativity and value-plurality. • Gated Attention for LLMs — A simple “sigmoid gate” after softmax attention consistently boosts performance, stability, and long-context handling. Already adopted in Qwen3-Next. • 1000-Layer Self-Supervised RL — Shatters the belief that RL can’t scale deep. Networks up to 1024 layers learn complex behaviors without rewards or demos. • Why Diffusion Models Don’t Memorize , Shows diffusion models generalize because of implicit dynamical regularization with predictable generalization/memorization timescales. Runner-Up Papers: • Does RL Really Create New Reasoning? Surprising result: RLVR improves sampling efficiency but does not generate new reasoning abilities beyond the base model. • Optimal Mistake Bounds (30-year open problem solved) Shows transductive online learning enjoys a quadratic advantage thanks to unlabeled data. • Superposition & Neural Scaling Laws Evidence that LLMs operate in a strong superposition regime, explaining why loss scales predictably with model size. --- The 2025 winners highlight a shift toward: * Understanding why scaling works * Addressing AI safety risks (diversity collapse) * Building more efficient, deeper, and interpretable architectures * Challenging assumptions in RL-based reasoning NeurIPS continues to prove that the future of AI isn’t just bigger models — it’s smarter theory, cleaner mechanisms, and better alignment with human diversity. --- #NeurIPS2025 #NeurIPS #AIResearch #MachineLearning #DeepLearning #LLMs #ReinforcementLearning #DiffusionModels #NeuralScaling #AIAlignment #AIEthics #AttentionMechanisms #Benchmarking #Datasets #ArtificialIntelligence #GenerativeAI #RLVR #Superposition #ScalingLaws #TechConference #AI
2
-
Yuvraj Singh Bhadoria
Bank of America • 8K followers
SpecActor, Explained — Why Rollout, Not Math, Is the Real Bottleneck in LLM Post-training A paper from ByteDance Seed and SJTU that's the sharpest systems argument I've read this year: you don't lose RL post-training time to compute — you lose it to waiting. The core insight: rollout eats 75–80% of post-training time, and on real traces ~50% of that GPU time is idle — every worker stalled on the slowest one still solving its task. Generation is sequential and memory-bound, so you can't buy your way out: 2× the GPUs nets 1.2–1.3×. The fix: take speculative decoding — the inference trick you already know — and retrofit it to training. Fast draft path, parallel verification, lossless by construction. But naive speculation fails at training batch sizes: at per-worker batch 128, verification hits the compute limit and the gain goes to zero or negative. SpecActor fixes that with two moves. Move 1 — Decoupled speculation. Put drafter and verifier on separate GPUs. The drafter runs ahead without waiting for verification, so the compute-hungry verifier gets real GPU time instead of idling behind a tiny draft model. Aggression is capped by a draft window (w): at most 2w−1 tokens wasted on a rejected guess. Move 2 — Fastest-of-N speculation. Stop committing to one draft method per request. Build a draft ladder offline (which drafter wins at which acceptance rate), pick the best guess by the batch's average acceptance rate — statistically stable even when individual requests vary — then, as workers finish, launch additional drafters (0.5B, 1.5B, n-gram…) at the long-tailed stragglers on the freed GPUs: 1. Pass the request through every live drafter in parallel 2. Accept the first draft method that emits an accepted EOS 3. Remove the request everywhere — the batch finishes when the fastest finishes Benefits: - 2.0–2.4× mean rollout speedup, up to 2.7× - 1.1–2.6× faster than vanilla speculative rollout - 1.4–2.3× faster end-to-end post-training - Lossless — exact token matching, zero accuracy change - Algorithm-agnostic: works on GRPO, DAPO, and PPO; dense and MoE - Drop-in replacement of the inference component in veRL Key takeaways for anyone scaling post-training: 1. Buy hardware last — sequential waits don't yield to silicon; 2× GPUs bought 1.2× 2. Decouple the dependency you took as law: draft-then-verify coupling was starving the verifier of GPU time 3. Don't pick one drafter — when the batch shrinks, bet on the tail with parallel drafters 4. Measure batch-completion time, not token throughput — those optimize different systems Source: Fast LLM Post-training via Decoupled and Fastest-of-N Speculation, arXiv:2511.16193
17
3 Comments -
Armin Parchami
Adobe • 31K followers
Evaluating agentic LLMs (especially coding agents) is much harder than most benchmarks suggest. Spending the past year working on various agentic benchmarks, a few uncomfortable truths stand out: • 𝗡𝗼𝗻-𝗱𝗲𝘁𝗲𝗿𝗺𝗶𝗻𝗶𝘀𝗺 𝗰𝗼𝗺𝗽𝗼𝘂𝗻𝗱𝘀 𝘄𝗶𝘁𝗵 𝘀𝗲𝗾𝘂𝗲𝗻𝗰𝗲 𝗹𝗲𝗻𝗴𝘁𝗵 Agents are stochastic, and the longer the reasoning / tool-use chain, the wider the outcome distribution. One run succeeds, the next fails. Metrics like pass@k hide reliability; stricter notions (pass^k, repeated success) collapse quickly. • 𝗘𝘃𝗲𝗻 𝘀𝗶𝗻𝗴𝗹𝗲-𝘁𝘂𝗿𝗻 𝗲𝘃𝗮𝗹𝘂𝗮𝘁𝗶𝗼𝗻 𝗶𝘀 𝗮𝗺𝗯𝗶𝗴𝘂𝗼𝘂𝘀 What does “correct” mean? Exact match breaks on trivial equivalences (e.g., sin(π) vs sin(2π)), rubric scoring doesn’t scale, and LLM-as-judge introduces bias, variance, and sometimes self-preference—creating evaluation loops with unknown error. • 𝗛𝘆𝗽𝗲𝗿𝗽𝗮𝗿𝗮𝗺𝗲𝘁𝗲𝗿𝘀 𝗾𝘂𝗶𝗲𝘁𝗹𝘆 𝗱𝗲𝗳𝗶𝗻𝗲 𝘁𝗵𝗲 𝗿𝗲𝘀𝘂𝗹𝘁 Temperature, max tokens, reasoning depth, tool budgets, retries, and k in pass@k all materially change outcomes. Two evals of the same model with different settings can tell opposite stories about capability. • 𝗧𝗵𝗲 𝗱𝗮𝘁𝗮 𝗽𝗿𝗼𝗯𝗹𝗲𝗺 𝗶𝘀 𝗿𝗲𝗮𝗹 Dataset diversity, edge cases, and distributional match to real user prompts matter and contamination is increasingly obvious (e.g., SWE-bench variants). Many “improvements” vanish on newer or harder task splits. • 𝗠𝘂𝗹𝘁𝗶-𝘁𝘂𝗿𝗻 𝗲𝘃𝗮𝗹 𝗮𝗱𝗱𝘀 𝗮𝗻𝗼𝘁𝗵𝗲𝗿 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 𝘀𝘂𝗿𝗳𝗮𝗰𝗲 Simulated users can induce errors, drift from the original task, or mask agent weaknesses. Tool failures, missing data, and policy violations get conflated with model capability unless carefully disentangled (τ-Bench makes this painfully clear). On the other hand single-turn is not evaluating agnents in a true agentic setting. • 𝗘𝘃𝗮𝗹 𝗵𝗮𝘀 𝗯𝗲𝗰𝗼𝗺𝗲 𝗺𝗼𝗿𝗲 𝗼𝗳 𝗮 𝗱𝗲𝘀𝗶𝗴𝗻 𝗰𝗵𝗼𝗶𝗰𝗲, 𝗻𝗼𝘁 𝗮 𝗳𝗮𝗰𝘁 Every vendor chooses how to sample, judge, aggregate, and report. That doesn’t make results useless, but it does mean comparisons are fragile and often non-apples-to-apples. 𝗧𝗟;𝗗𝗥: agentic eval isn’t “just accuracy.” It’s stochastic systems, subjective judging, sensitive knobs, imperfect data, and long-horizon behavior all interacting at the same time in a highly complex system. Progress is real, but measuring it rigorously is still an open problem.
123
5 Comments -
🛸Steven Joseph
DamageBDD • 3K followers
🧠 Geometry as Acceleration vs Algebra as Substrate There’s been some understandable confusion in recent discussions between geometric compute architectures (e.g. Hologram-style systems) and what ECAI is actually doing. They operate at very different layers of the stack. Geometric compute architectures use structured spaces, canonicalization, resonance classes, and lookup tables to accelerate classical computation. The goal is performance: reduce routing complexity, achieve constant-time lookup, improve cache locality, and compress computation into efficient deterministic pipelines. Geometry is used as an optimization mechanism. This is legitimate engineering. It improves throughput, latency, and scalability of existing computational models. ECAI operates in a different category entirely. Rather than using geometry to accelerate computation, ECAI uses algebraic structure — specifically elliptic curve group dynamics — as the computational substrate itself. State transitions are not table-driven or heuristic. They are governed by strict mathematical invariants that admit formal proof. In this model: Computation is constrained by algebra, not optimized by geometry. Correctness emerges from mathematical structure, not from architectural tuning. Verification is intrinsic rather than layered on top. Trust shifts from engineering assumptions to provable properties. Geometric accelerators answer the question: > “How do we compute faster?” Substrate-level algebra answers the question: > “What computations are even allowed to exist?” That distinction matters. Acceleration improves performance inside the existing paradigm. Substrate change alters the paradigm itself — how correctness, safety, composability, and trust are enforced. Both approaches are valuable, but they solve fundamentally different problems and should not be conflated. When people see geometry intersect computation, it’s easy to assume these systems are equivalent. They’re not. One optimizes execution. The other constrains reality. That gap is where most of the misunderstanding currently lives. #ECAI #SystemsEngineering #Cryptography #Verification #DeterministicAI #Computation #Architecture #Bitcoin
3
-
Yariv Adan
ellipsis • 14K followers
Your startup is a program.md file and a metric! 😱 Andrej Karpathy shared a weekend project called “AutoResearch.” It’s a practical example of how autonomous AI agents can manage iterative work through “Agentic Loops.” Sharing my view: 🤔 What exactly did Karpathy launch? #AutoResearch is a minimal, self-contained repository designed to autonomously train a small language model. Instead of a human researcher manually adjusting parameters and running experiments, an AI agent manages the entire iteration cycle. The system relies on two components: 👩🏽💼 program.md: A plain text file where a human sets the research strategy, guiding the AI on what to experiment with and how to behave. 🤖 train.py: The Python code the AI agent is permitted to modify, including model architecture and hyperparameters. The engine is a strict 5-minute autonomous loop. The agent reads the human strategy, modifies the code, and kicks off a 5-minute training run. It then evaluates the outcome against a single objective metric . If the score improves, the agent commits the code. If not, it reverts and tries a different approach. The system can run indefinitely without human intervention. 🪄 🔭 The Broader View: The Agentic Loop AutoResearch highlights the practical application of an agentic loop, sometimes called a “Ralph Wiggum loop.” This is a persistent cycle where an AI generates output, tests it against a benchmark, and iterates to improve it. 💰 What makes a process a good candidate for an Agentic Loop? Not every task is ready for this level of automation. The initial successes will be found where five conditions are met: 🏉 Objective Scoring: A measurable score so the loop can distinguish better from worse without human judgment. ⚡️Fast, Cheap Iterations: Quick cycles where a bad attempt wastes minutes, not months. ⚪️ Bounded Environment: A clearly defined work and action space for the agent. 🧯Low Cost of Failure: Consequences of a bad iteration are minimal. ↩️ Traceability: The agent can leave a clear trail of its work and decisions. Implications for the Future of Work This approach can theoretically be applied to any business process that meets these criteria. A few areas where we might see it first: 📈 Marketing: Agents generating ad variations, testing them against live audiences, and iterating toward a target metric. 📥 Sales: Loops pointed at cold outreach, autonomously testing subject lines to maximize positive reply rates. 💻 Engineering: #Agents running continuous loops to generate code and apps based on acceptance and success criteria. But it's easy to come with many more examples - it's a highly scalable generic approach. Once again, we see how #AI capabilities commoditize vast areas - your startup is a program.md file and a metric! 🤓 The next interesting frontier is expanding the single agent loop to a swarm of agents, with different skills and capabilities - looping autonomously together towards a brighter future 🤩
23
3 Comments -
Russ Salakhutdinov
Sooth Labs • 10K followers
Check out the new Machine Learning Department at CMU blogpost: How to Explore to Scale RL Training of LLMs on Hard Problems? https://lnkd.in/egihtUvu Current on-policy RL methods fail to learn from hard problems as they rarely generate a single correct rollout, producing no reward signal and no learning. Including easy problems can also be harmful, as models tend to overfit to them and fail to improve on harder tasks. Distilling human-written solutions is not only costly, but also provides difficult targets for fine-tuning. This blogpost discusses various approaches and introduces a framework that uses existing human or model solutions as privileged guidance to unlock learning on hard problems. The key idea is simple: Prepend a minimal solution prefix to difficult prompts, enabling on-policy RL to obtain reward and learn behaviors that generalize back to the original, unconditioned tasks. This expands the set of solvable problems and results in significant gains on challenging reasoning benchmarks. With YUXIAO QU, Amrith Setlur, Virginia Smith, and Aviral Kumar. Paper/Code is coming soon.
172
1 Comment -
Nishantha Ruwan
IWROBOTX Software Inc. • 2K followers
The authors identify a phenomenon they call latent over‑thinking in large language models (LLMs): after one standard forward pass, further latent iteration on all tokens sometimes hurts performance, because “easy” tokens that were already correct get needlessly refined and become wrong. To address this, they propose the Think‑at‑Hard (TaH) method: a lightweight decider module predicts which output tokens are likely to be wrong after the first pass, and only those tokens are subject to additional latent thinking (i.e., further iterations). During those iterations, they use Low‑Rank Adaptation (LoRA) modules to shift the objective away from general next‑token prediction toward focused refinement of hard tokens, and they introduce a duo‑causal attention mechanism to allow information to flow across both token sequence and iteration depth, while preserving parallelism. Empirically, TaH demonstrates consistent improvements across five challenging reasoning benchmarks: when compared to a baseline that applies a second iteration to all tokens, TaH achieves accuracy gains of about 8.1‑11.3% while skipping iteration on ~94% of tokens (thus saving compute). Against a strong single‑iteration model (Qwen3) finetuned on the same data, TaH still delivers a 4.0‑5.0% advantage. Allowing <3% extra parameters (for the decider + LoRA), gains increase further to 8.5‑12.6% and 5.3‑5.4% respectively. Thus, TaH offers a more efficient and targeted way to improve reasoning in LLMs without needing to apply expensive iteration uniformly across all tokens. https://lnkd.in/g8fhxtg7
-
Thomas Wolf
Hugging Face • 194K followers
Despite all the big funding rounds and flashy demos in US robotics, K-Scale’s inability to raise more money should worry us. We're at risk of replaying the LLM story all over again in robotics: - Chinese companies are going open-source and collaborating across the value chain (from EV suppliers to downstream integrators) - most western teams are going full-stack proprietary, closed-source, all in-house Guess which robots the next wave of research labs and startups will actually be able to build on when they want to invent new algorithms or tackle unseen real-world use cases?
273
33 Comments -
Ben Dickson
TechTalks Media • 14K followers
"Attention Matching" compacts the KV cache of LLMs by 50x with very minimal impact on the accuracy at at a fraction of the time of other compaction methods. It took me a while to figure out how it works, but it makes so much sense. To compact the context of the LLM: - You create a set of "reference queries" for the context. These are proxies for the kind of tasks the LLM is expected to do - You choose a subset of keys in the KV cache (e.g., highest attention values) to preserve - To choose the values (along with an accompanying bias term), you use the reference queries and the keys to compare and adjust your compacted KV cache with the full attention cache. This formulation makes it possible to compact the KV cache using simple and super-fast algebraic optmization tricks (e.g., OLS) as opposed to slow gradient-based learning methods. Result: 50x compaction, very high accuracy, very fast compaction (seconds vs hours for baseline methods that preserve accuracy). There are some caveats, including how you compose the reference queries and the key values, which enable you to adjust the speed-compaction-accuracy tradeoffs. https://lnkd.in/efPDhUfB
4
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content