LlamaIndex’s cover photo
LlamaIndex

LlamaIndex

Technology, Information and Internet

San Francisco, California 288,904 followers

Turn any document into agent-ready context.

About us

LlamaParse is the most accurate agentic OCR platform for production AI — purpose-built for the documents agents actually encounter in the real world. Unlike general-purpose models that guess at structure, LlamaParse is engineered for complex layouts, dense tables, handwritten annotations, and scanned pages. Every page is automatically routed to the optimal model, so accuracy and cost are optimized without manual configuration. Trusted by teams at Lovable, 8am, Tabs, KPMG, and others running document-intensive workflows across legal, finance, healthcare, and more.

Website
https://www.llamaindex.ai/
Industry
Technology, Information and Internet
Company size
11-50 employees
Headquarters
San Francisco, California
Type
Public Company

Locations

Employees at LlamaIndex

Updates

  • LlamaIndex reposted this

    Today we’re introducing Extract v2.5 - a series of frontier agents tuned for document extraction. The agents (cost-effective, agentic, agentic plus) are tuned for value accuracy and grounding. Our extraction agents outperform Opus 5.5 and GPT-6 Sol while being 30%-4x cheaper. We’ve made massive improvements on complex extraction over * long lists (86.1% -> 95.5% on our agentic tier) * records spanning pages (85.5% -> 96.5%) * scanned forms (90.9% -> 95.7%) We’ve also launched the following features: ✅ Advanced citations: we locate bounding boxes for all supporting values for an inferred field, even if there's not an exact match. ✅ Structural Reasoning: we tailor document extraction algorithms depending on the type, layout, and information Our extraction agents are SOTA in price-performance on ExtractBench, across a wide range of cost points. We are the best tool for document extraction across documents of any complexity. Blog: https://lnkd.in/gPfUnNna All of these are available on LlamaParse: https://lnkd.in/g9Wpqn7w . Come check it out!

  • LlamaIndex reposted this

    Today we’re introducing Extract v2.5 - a series of frontier agents tuned for document extraction. The agents (cost-effective, agentic, agentic plus) are tuned for value accuracy and grounding. Our extraction agents outperform Opus 5.5 and GPT-6 Sol while being 30%-4x cheaper. We’ve made massive improvements on complex extraction over * long lists (86.1% -> 95.5% on our agentic tier) * records spanning pages (85.5% -> 96.5%) * scanned forms (90.9% -> 95.7%) We’ve also launched the following features: ✅ Advanced citations: we locate bounding boxes for all supporting values for an inferred field, even if there's not an exact match. ✅ Structural Reasoning: we tailor document extraction algorithms depending on the type, layout, and information Our extraction agents are SOTA in price-performance on ExtractBench, across a wide range of cost points. We are the best tool for document extraction across documents of any complexity. Blog: https://lnkd.in/gPfUnNna All of these are available on LlamaParse: https://lnkd.in/g9Wpqn7w . Come check it out!

  • How many people show up on a Tuesday night to talk about document processing for AI agents? 🤔 In New York, enough to fill the room and start a waitlist. In San Francisco, almost 600. We were in both cities on the same night to talk about one problem: agents don't read documents the way people do. Look at an invoice and you can tell the monthly rate from the total in a second. An agent often gets the words without the layout, so it has to guess. In New York, we partnered with KERNEL on Braintrust's Agent Builders Night. In San Francisco, we partnered with WorkOS on Daytona's AI Builders. Thanks to them for having us, and to Abrar Mahi and Yong Park from our team for speaking!

    • No alternative text description for this image
    • No alternative text description for this image
  • LlamaIndex reposted this

    Yesterday I hosted a fun dinner conversation with Vincent Sunn Chen from Snorkel AI on evals and RL environments. The "data and RL env" companies (like Snorkel) have seen massive growth in the past few years. There's been an explosion of interest in evals. At the same time, models are ripping through benchmarks with each new release. We talked about the evals everyone is defining, what evals are still left unsolved, what’s left up to frontier models vs. intelligence that you own, and more: - A big challenge for building RL environments is “fairness” - when the model fails on a given environment, can you attribute it to the input, harness, or reward model? - Building proper rewards is hard. Some tasks are not easily quantifiable. You also want to discourage reward hacking. At the same time, you don’t want to be too prescriptive with intermediate rewards. - Long horizon evals are still extremely hard, some business processes can take up to weeks or months before the final outcome - Most regulated industries still need human in the loop to guarantee ~100% accuracy, “80%” accuracy is not good enough - As models get more intelligent, there will be a barbell of boutique data vendors (e.g. any SMB) any scaled up data providers. - Models still exhibit “jagged intelligence” where they still fail on a long tail of edge cases. - There might always be opportunities to gather unique data for a given task and posttrain models for lower cost and higher accuracy. This marks #003 in our founder dinner series. What topic should we discuss next? Let us know your thoughts below! Preston Yadegar Faraz Siddiqi Akshaya Jagadeesh Ali Ebrahim Shahul Elavakkattil Shereef Robert Stewart Jithin James Abhijeet Shenoi Vincent Sunn Chen

    • No alternative text description for this image
  • Document parsing requires a lot of on-the-fly decision making. Can Jev make those decisions for you? How closely can OSS alternatives land? We explored the early approaches to Jev and Jev-like models on several documents tasks like orientation detection, language detection, routing, and more. https://lnkd.in/gHuWievF And of course, LlamaParse already handles these same tasks for you 🦙

  • Parsing forms still trips up the latest frontier VLMs. A form isn't text on a page. It's a set of fields, grouped into sections, each tied to a specific box. That's why forms need purpose-built parsing, not a bigger general model: ✅️ Detect every field, not just the obvious ones ✅️ Keep the hierarchy of sections and fields ✅️ Tie every value to the exact box it came from ✅️ Understand handwriting and checkmarks Our latest blogpost breaks down where VLMs fail on real W-2s, 1040s, W-9s and scanned W-4s, and a custom cookbook for LlamaParse to handle them at a fraction of the cost 👇 https://lnkd.in/gzBPmtBa

    • No alternative text description for this image

Similar pages

Browse jobs