𝗡𝗮𝗶𝘃𝗲 𝗥𝗔𝗚 𝘄𝗼𝗿𝗸𝘀 𝗶𝗻 𝗮 𝗱𝗲𝗺𝗼. 𝗜𝘁 𝗳𝗮𝗶𝗹𝘀 𝘁𝗵𝗲 𝗺𝗼𝗺𝗲𝗻𝘁 𝗿𝗲𝗮𝗹 𝘂𝘀𝗲𝗿𝘀 𝘀𝗵𝗼𝘄 𝘂𝗽. Embed → retrieve → generate looks clean in a notebook. Real requirements break it: → Questions whose answer is spread across many documents → Industry terms that embeddings get wrong → Bad chunks the pipeline never catches → Answers that live in how things connect, not in any single chunk → PDFs full of tables and images a text-only index cannot read These 5 architectures are how serious teams stay ahead in the agentic AI era: 𝟬𝟭 𝗛𝘆𝗯𝗿𝗶𝗱 𝗥𝗔𝗚 → Dense vectors find meaning. BM25 finds exact words. → Reciprocal Rank Fusion combines both ranked lists. → A safe baseline for almost every team. 𝟬𝟮 𝗚𝗿𝗮𝗽𝗵𝗥𝗔𝗚 → Pull entities and their relationships into a knowledge graph. → Retrieve subgraphs and community summaries, not chunks. → Best when the answer lives in how things connect. 𝟬𝟯 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗥𝗔𝗚 → A planner agent picks the right tool: vector, web, or SQL. → A reasoner agent keeps trying until the answer is solid. → Retrieval becomes a plan, not a single step. 𝟬𝟰 𝗖𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝘃𝗲 𝗥𝗔𝗚 (𝗖𝗥𝗔𝗚) → Grade every retrieval before you trust it. → Correct → answer. Unclear → rewrite the query. Wrong → search the web. → This is what production RAG actually looks like. 𝟬𝟱 𝗠𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹 𝗥𝗔𝗚 → One embedding model (CLIP, ColPali) for text, images, and tables. → One vector index. One multimodal LLM. → No more separate pipelines for PDFs with charts. I built a runnable example for each of the five patterns. GitHub link in the first comment. The best teams in 2026 do not pick one. They combine them — hybrid retrieval inside an agentic loop, with a corrective grader, over a multimodal index. Naive RAG is a starting point, not a finish line. That is why most enterprise GenAI projects stall at the demo. Which of these five becomes the default RAG stack in the next 18 months — and which stays a specialized tool?
Multimodal AI Developments
Explore top LinkedIn content from expert professionals.
-
-
I've been exploring ways to make Vision-Language Models (VLMs) more efficient during inference, and I came up with a simple but effective technique: dynamic token selection. Here’s the idea: instead of using the same number of tokens for every part of an image, we dynamically assign tokens based on region complexity. Why waste compute on blank spaces or redundant details when you can focus on what’s important? For example, in this image of Manhattan, some regions are compressed significantly while others retain more detail depending on their complexity: 👇 How does it work? 1️⃣ Extract visual features as usual. 2️⃣ Measure cosine similarity between features. 3️⃣ Cluster similar features using DBSCAN. 4️⃣ Replace clusters with their means as tokens. Check out the code in the comments! The result? At 25% compression, there’s almost no performance loss across tasks. Even at 50% compression, some tasks remain viable, although others (like OCR) begin to struggle. I'll put a full evaluation plot in the comments. This is just something I tried out while thinking about test-time compute for VLMs. It’s model-agnostic, doesn’t require retraining, and can potentially work for other modalities like video and audio. I might write more about it in the future if we decide to do a SmolVLM paper. For now, I’ve linked the full evaluation results and code in the comments. Some related work worth exploring: VisionZip, FastV, and Feather. Let me know what you think!
-
Can we use AI agents for stock market prediction? 😮 Recently, LLM-based agents have demonstrated remarkable advancements in handling multi-modal data, enabling them to execute complex, multi-step decision-making tasks. This research introduces a multi-modal multi-agent system designed specifically for financial trading tasks. The framework employs a team of specialized LLM-based agents, each adept at processing and interpreting various forms of financial data, such as textual news reports, candlestick charts, and trading signal charts. The framework comprises four primary components: the Summarize Module, the Technical Analyst Module, the Prediction Module, and the Reflection Module. The Summarize Module condenses large volumes of textual news data into concise summaries that highlight factual information influencing stock trading decisions. The Technical Analyst Agent leverages the visual reasoning capabilities of LLMs to analyze candlestick charts with technical indicators, providing interpretations for next-day trading strategies. The Reflection Module consists of two parts: one assesses the short-term and medium-term performance of previous trades, while the other plots past trading signals, generates charts, and offers insights into the effectiveness of trades. The Prediction Agent integrates information from these components to forecast trading actions, determine position size as a percentage of the portfolio, and provide a detailed explanation of the decision. Based on the Prediction Agent’s output, the Reward Agent executes trades and calculates performance metrics. These metrics are then used by the Reflection and Prediction Agents in the subsequent iterations. The detailed flow of our framework is illustrated in the Figure. Know more about the framework in this practical research paper: https://lnkd.in/gxvEUGAA Here is my simple video explaining how AI agents work: https://lnkd.in/d_V9DqbH This is my practical hands-on guide on building multi-agent AI system: https://lnkd.in/gdaA5s3Z
-
Excited to share insights from Walmart 's groundbreaking semantic search system that revolutionizes e-commerce product discovery! The team at Walmart Global Technology(the team that I am a part of 😬) has developed a hybrid retrieval system that combines traditional inverted index search with neural embedding-based search to tackle the challenging problem of tail queries in e-commerce. Key Technical Highlights: • The system uses a two-tower BERT architecture where one tower processes queries and another processes product information, generating dense vector representations for semantic matching. • Product information is enriched by combining titles with key attributes like category, brand, color, and gender using special prefix tokens to help the model distinguish different attribute types. • The neural model leverages DistilBERT with 6 layers and projects the 768-dimensional embeddings down to 256 dimensions using a linear layer, achieving optimal performance while reducing storage and computation costs. • To improve model training, they implemented innovative negative sampling techniques combining product category matching and token overlap filtering to identify challenging negative examples. Production Implementation Details: • The system uses a managed ANN (Approximate Nearest Neighbor) service to enable fast retrieval, achieving 99% recall@20 with just 13ms latency. • Query embeddings are cached with preset TTL (Time-To-Live) to reduce latency and costs in production. • The model is exported to ONNX format and served in Java, with custom optimizations like fixed input shapes and GPU acceleration using NVIDIA T4 processors. Results: The system showed significant improvements in both offline metrics and live experiments, with: - +2.84% improvement in NDCG@10 for human evaluation - +0.54% lift in Add-to-Cart rates in live A/B testing This is a fantastic example of how modern NLP techniques can be successfully deployed at scale to solve real-world e-commerce challenges!
-
AI agents and physical AI are shifting industrial automation from equipment supply to autonomous, self-optimizing systems. The most mature vendors are moving from pilots to production, with robots navigating complex environments and digital twins optimizing the value chain. This CB Insights brief gives a good view of where the top 20 industrial automation companies stand on AI maturity. Three key trends. 1. Leaders like Siemens Industry and ABB are linking AI systems across design, logistics, manufacturing, and maintenance creating compounding benefits. 2. Optimization dominates near-term priorities, while digital twins are emerging as the backbone for connecting hardware and software. 3. Partnerships with tech companies like Microsoft, Google, and Nvidia are essential, but they create new dependencies that must be managed. Siemens at the top of the ranking, combining copilots, edge platforms, and digital twins. Its work with Microsoft and Nvidia expands capabilities but increases reliance on external tech. Honeywell takes a more focused approach, embedding AI into devices and workflows. Its Qualcomm partnership highlights product-level integration over broad system building. ABB advances through its OmniCore platform and acquisitions such as Sevensense and SensorFact, blending robotics, software, and energy management. Schneider Electric pushes AI in energy management, using digital twins and partnerships with Nvidia, Microsoft, and Itron to extend from factory optimization into grid intelligence. The path forward in industrial AI is moving beyond pilots or isolated tools. It will depend on how well vendors embed AI into their platforms, link technologies across domains, and balance the benefits of external partners with the need for strategic independence. Those that will get it right will turn AI from experimentation into durable advantage. Just as critical is how their customers adopt these technologies. Industrial firms must shift from isolated use cases to embedding AI in design, production, energy, and logistics. Success requires not only advanced tools, but also the data, skills, and processes to make AI scale in complex operations.
-
Multi-Head Attention (MHA) is the engine of LLMs. But over the years, we have added several tweaks to make it more efficient for long-context settings, especially when using KV caching during inference. I implemented the most common variants from scratch: 1) Grouped-Query Attention (GQA): Instead of having a unique key and value for each query head, multiple queries share the same key and value. As long as the sharing ratio is not too extreme, this has minimal impact on model quality. It is the most widely used variant today and found in almost every modern LLM including Llama 2-4, GPT-OSS, Gemma 3, Qwen 3, GLM 4.6, and many others. 2) Multi-Head Latent Attention (MLA): This variant introduces a compressed latent representation for the keys and values that are stored in the KV cache. During inference, these latent keys and values are up-projected back to the full dimension. The extra projection adds a small computational cost, but the memory savings make it worth it. This approach is currently used by DeepSeek V3 and Kimi K2. 3) Sliding-Window Attention (SWA): SWA restricts each token’s attention span to a fixed local window, which reduces memory needs by shrinking the KV cache in long-context regimes. It is usually applied selectively, for example every other layer, or in Gemma 3's case, five SWA layers for each full-attention layer. While less common today, it remains an important optimization, notably in Gemma 3. All three variants, GQA, MLA, and SWA, can also be combined freely within the same model. Here's a link to check them out: 1️⃣ GQA: https://lnkd.in/grDPXUUi 2️⃣ MLA: https://lnkd.in/gm4FzE32 3️⃣ SWA: https://lnkd.in/g7x-fdgn
-
Meta just dropped SAM 3D, but more interestingly, they basically cracked the 3D data bottleneck that's been holding the field back for years. Manually creating or scanning 3D ground truth for the messy real world is effectively impossible at scale. But what if you just have humans rank model outputs? Route the weird edge cases to actual 3D artists to model, loop it back in. Suddenly you can annotate like a million images. It's basically RLHF for 3D reconstruction. Synthetic data is pretraining, real world ranking is alignment. They borrowed the whole damn playbook and it actually works. Two models - one for objects/scenes, one for humans. They're already shipping it in FB Marketplace so you can see if that lamp or chair looks good in your room before buying. Also they're releasing everything - models, code, their human body rig under commercial license. And they built an eval set of actual messy real-world images to help bridge the sim-to-real gap. The data engine thing is the most interesting though. 3D has been bottlenecked by ground truth forever. If verification scales easier than creation, suddenly the whole game changes.
-
🚀 NVIDIA just dropped 707 GB of real-world robot training data 🤖. And it’s now available on Hugging Face. Cosmos3-DROID isn’t just another video dataset. 📦 707 GB of robot data 🤖 71,907 episodes 🎥 22.4M+ frames 🌎 564 scenes across 52 buildings 🦾 86 manipulation tasks 🏭 18 labs across 13 institutions 📹 3 synchronized camera views + depth ⚙️ Robot states, torques, velocities & actions ✅ Success and failure trajectories Even better: NVIDIA converted the raw DROID data to LeRobotDataset v3.0, making it much easier to plug into modern robot-learning pipelines. This is the kind of dataset that helps move robotics from: “AI understands what the robot should do.” to: “AI learns how to actually do it.” 🤯 And the original DROID research reported roughly 20% average improvement in policy performance, robustness and generalization when co-training with DROID. The race for Physical AI just got a lot more interesting. 👀 🔗 Dataset: https://lnkd.in/gdg3xrrk 🔗 DROID Project: https://lnkd.in/g2RVGcQv 🔗 Research Paper: https://lnkd.in/giJ6KPFW 🔗 DROID GitHub: https://lnkd.in/gGrcArDM #NVIDIA #Robotics #PhysicalAI #AI #RobotLearning #MachineLearning
-
I've been saying for over a year that multimodal large language models will become the ultimate interface between physicians and a range of AI-based solutions. Here is the proof! In this study, the authors developed and evaluated an autonomous clinical AI agent leveraging GPT-4 with multimodal precision oncology tools to support personalized clinical decision-making. They used multiple sources such as histopathology slides, radiological images and search tools like OncoKB, PubMed and Google. "Evaluated on 20 realistic multimodal patient cases, the AI agent autonomously used appropriate tools with 87.5% accuracy, reached correct clinical conclusions in 91.0% of cases and accurately cited relevant oncology guidelines 75.5% of the time. Compared to GPT-4 alone, the integrated AI agent drastically improved decision-making accuracy from 30.3% to 87.2%." Source: https://lnkd.in/dwjGvxcH
-
Excited to announce the public release of the LLaVA-Rad training data! 🎉 Medicine is inherently multi-modal – accurate patient understanding requires AI that can process diverse biomedical data. While multi-modal LLMs hold great potential for medical applications, publicly available datasets remain scarce. To bridge this gap, we processed 239,025 additional X-ray image-text pairs from MIMIC-CXR using GPT-4, expanding the dataset to 400,042 pairs – more than doubling its original size. This LLaVA-Rad dataset (https://lnkd.in/gpPUX2Yw) features: ✅ More accurate structuring of radiology reports into key sections such as reason for exam, findings, and impression ✅ Cleaning of broken words and repetitive phrases for clarity ✅ Removing temporal mentions to focus on information relevant to the current image This dataset is a foundation of our LLaVA-Rad model (https://lnkd.in/gRAJufy3), a compact yet powerful 7B vision-language model. Despite its small size, LLaVA-Rad outperforms frontier models like OpenAI’s GPT-4V and Google’s Med-PaLM M (84B). Its smaller size makes it easier for researchers to experiment and develop faster. The data is released in the popular LLaVA format, making it easy to train your model with the LLaVA-Med codebase (https://lnkd.in/gBPnrGwM). While LLMs and LRMs like o1 and R1 mark significant progress for text-only problems, there’s still much to improve in multi-modal AI, especially in biomedicine. We hope LLaVA-Rad resources empower researchers and drive real-world breakthroughs. Looking forward to seeing how the community builds on this! 🔥LLaVA-Rad paper: https://lnkd.in/gRAJufy3 📚LLaVA-Rad Data: https://lnkd.in/gpPUX2Yw 💡LLaVA-Med Code: https://lnkd.in/gBPnrGwM 🔍CheXprompt (Expert-level report evaluation using GPT-4) : https://lnkd.in/gdTjgiUS Juan Manuel Zambrano Chaves Mars Huang @Yanbo Xu Hanwen Xu Sheng Zhang Fei Wang Yujia Xie Mahmoud Khademi Ziyi Yang Hany Awadalla Julia Gong Houdong Hu Jianwei Yang Chunyuan Li Jianfeng Gao Yu Gu Cliff Wong Mu Wei Tristan Naumann Muhao Chen Matthew Lungren MD MPH Akshay Chaudhari Serena Yeung Curtis Langlotz Sheng Wang Hoifung Poon Microsoft Research