Voice is the next frontier for AI Agents, but most builders struggle to navigate this rapidly evolving ecosystem. After seeing the challenges firsthand, I've created a comprehensive guide to building voice agents in 2024. Three key developments are accelerating this revolution: -> Speech-native models - OpenAI's 60% price cut on their Realtime API last week and Google's Gemini 2.0 Realtime release mark a shift from clunky cascading architectures to fluid, natural interactions -> Reduced complexity - small teams are now building specialized voice agents reaching substantial ARR - from restaurant order-taking to sales qualification -> Mature infrastructure - new developer platforms handle the hard parts (latency, error handling, conversation management), letting builders focus on unique experiences For the first time, we have god-like AI systems that truly converse like humans. For builders, this moment is huge. Unlike web or mobile development, voice AI is still being defined—offering fertile ground for those who understand both the technical stack and real-world use cases. With voice agents that can be interrupted and can handle emotional context, we’re leaving behind the era of rule-based, rigid experiences and ushering in a future where AI feels truly conversational. This toolkit breaks down: -> Foundation layers (speech-to-text, text-to-speech) -> Voice AI middleware (speech-to-speech models, agent frameworks) -> End-to-end platforms -> Evaluation tools and best practices Plus, a detailed framework for choosing between full-stack platforms vs. custom builds based on your latency, cost, and control requirements. Post with the full list of packages and tools as well as my framework for choosing your voice agent architecture https://lnkd.in/g9ebbfX3 Also available as a NotebookLM-powered podcast episode. Go build. P.S. I plan to publish concrete guides so follow here and subscribe to my newsletter.
Designing For Multimodal Interactions
Explore top LinkedIn content from expert professionals.
-
-
To build enterprise-scale, production-ready AI agents, we need more than just a large language model (LLM). We need a full ecosystem. That’s exactly what this AI Agent System Blueprint lays out. 🔹 1. Input/Output – Flexible User Interaction Agents today must go beyond text. They take multimodal inputs—documents, images, audio, even video—so users can interact naturally and contextually. 🔹 2. Orchestration – The Nervous System Frameworks like LangGraph, Guardrails, Google ADK sit at the orchestration layer. They handle: Context management Streaming & tracing Deployment and evaluation Guardrails for safety & compliance Without orchestration, agents remain fragile demos. With it, they become scalable and reliable. 🔹 3. Data and Tools – Context is Power Agents get smarter when connected to enterprise data: Vector & semantic DBs Internal knowledge bases APIs from Stripe, Slack, Brave, and beyond This ensures every decision is grounded in context, not hallucination. 🔹 4. Reasoning – Brains of the System Multiple model types collaborate here: LLMs (Gemini Flash, GPT-4o, DeepSeek R1) SLMs (Gemma, PiXtral 12B) for lightweight use cases LRMs (OpenAI o3, DeepSeek) for specialized reasoning Agents analyze prompts, break them down, and decide which tools or APIs to call. 🔹 5. Agent Interoperability – Teams of Agents No single agent does it all. Using protocols like MCP, multiple agents—Sales Agent, Docs Agent, Support Agent—communicate and collaborate seamlessly. This is where multi-agent ecosystems shine. Why This Blueprint Matters When you combine these layers, you get AI agents that: ✅ Adapt to any input ✅ Make reliable decisions with enterprise context ✅ Collaborate like real teams ✅ Scale safely with guardrails and orchestration This is how we move from fragile prototypes → production-ready agent ecosystems. The big question: Which layer do you see as the hardest bottleneck for enterprises—Orchestration, Reasoning, or Data & Tools?
-
The Future of Immersion is Headset-Free? 😇 We often talk about the Metaverse being accessible via VR/AR headsets, but what happens when the most powerful immersive experience is shared, device-free, and right in front of you? The Shanghai Natural History Museum's "China's Dinosaur World" exhibition offers a powerful answer. They're using large-scale Projection Mapping and physical space to immerse 118 dinosaur specimens. Visitors literally walk through a prehistoric world, without a single tether or headset on their face. This isn't just a cool effect; it's a profound demonstration of how to scale presence and communal engagement. The Key Insight? True immersion isn't always about personal isolation. It's about collective experience. We need to stop framing immersive tech solely through the lens of hardware. The real innovation lies in the experience design—leveraging technologies like Projection Mapping and Spatial AR to make digital content accessible and communal for a massive audience. It democratizes the experience, making the 'Metaverse' a space for everyone, not just early adopters. Think about corporate training, product showcases, or massive-scale entertainment. The museum's approach proves that "shared reality" is perhaps the most impactful reality of all. What's the most compelling headset-free immersive experience you've encountered? 💡 ¿El Futuro de la Inmersión es Sin Auriculares? A menudo hablamos del Metaverso accesible a través de dispositivos VR/AR, pero ¿qué pasa cuando la experiencia inmersiva más potente es compartida, sin necesidad de dispositivos y está justo frente a ti? La exposición "China's Dinosaur World" en el Museo de Historia Natural de Shanghái ofrece una respuesta contundente. Están utilizando Projection Mapping a gran escala y el espacio físico para dar vida a 118 especímenes de dinosaurios. Los visitantes caminan literalmente a través de un mundo prehistórico, sin ataduras ni cascos. Esto no es solo un efecto visual genial; es una demostración profunda de cómo escalar la presencia y el compromiso comunitario. #AugmentedReality #VirtualReality #SpatialComputing #ExperientialDesign #Museums #EmergingTechnology
-
Shopping centres must become experiential arenas! The term ‘experiential arenas’ comes from Diana Teixeira Pinto and aligns with my view of how to design worlds not spaces. So how do we transform spaces into worlds? Here are some of my top design principles for executing successful Experiential Arenas: Build Worlds, Not Spaces Design destinations that transport people into new realities, not just corridors of commerce. Colour as Energy Bold, surprising palettes and patterns that lift mood and inject personality into every corner. Wellness in Motion Seating that heals, greenery that breathes, zones that invite pause and reset through biophilic design. Shopping should restore, not exhaust. Fill the Forgotten Atriums, rooftops, stairwells, and voids become playgrounds for art, light, and imagination. Sensory Immersion Use sound, scent, light, and texture as storytelling layers to spark memory and emotion. Everywhere’s a canvas Turn escalators, walkways, and food courts into theatres for entertainment, surprise, and play. Participation Over Passivity Invite people to co-create through interactive art, digital play, gamified shopping, and communal rituals. Play is Serious Business Design joy into the architecture: swings as benches, slides as shortcuts, playful touchpoints everywhere. Local Stories, Global Scale Embed local culture, artists, and narratives, then amplify them into experiences with global resonance. Micro-Magic Surprise through small details like bins that talk, ceilings that glow, restrooms that delight. Fluid & Ever-Changing Keep spaces alive with rotating installations, seasonal scenography, and pop-up moments of wonder. Sustainable Spectacle Awe doesn’t need waste: design modular, reusable, and eco-conscious experiences that wow responsibly. Community as Stage Curate experiences where people become part of the show — from live performance to collaborative design. Memory is the Metric Success isn’t footfall, it’s stories: people leave with moments worth retelling, not just receipts. Elena Knezović #retail #architecture #interior #design
-
'Immersive' is one of those words that means everything and nothing. Ask ten people in AV what it means and you'll get ten different answers - and all of them will be partially right. That ambiguity is a real barrier, both for integrators trying to scope projects and for end users trying to articulate what they actually want. We've put together a short guide that maps the territory properly. Six practical definitions, real-world examples for each, and a section on how to actually approach the design of these spaces - from audience and objectives through to content and future-proofing. It's written for AV integrators, technology managers, and end users who want a clear framework before they start specifying or buying. No jargon, no vendor pitch. If you're exploring immersive for the first time - or trying to explain it to a client - this is a good place to start. Please download the PDF using the link in the Comments section below. #ImmersiveAV #ProAV #AVintegration #ImmersiveDisplays #ProjectionMapping #VisualDisplays #DisplayTechnology #Twinmotion #ArchViz #CAVEsystems #Stereo3D
-
🔊 Have you ever stayed on a customer‑service call simply because the person on the other end sounded trustworthy? 🎧 Researchers from Beijing University of Technology , the The University of Texas at Austin and the University of Memphis recently tested how different AI voices affect persuasion. Their findings were: • 𝗙𝗹𝗶𝗿𝘁𝘆 𝗱𝗼𝗲𝘀𝗻’𝘁 𝘄𝗼𝗿𝗸. A playful “coquetry” voice actually decreased persuasion, especially for male chatbots. • 𝗦𝘁𝗲𝗿𝗻 𝗶𝗻𝘃𝗶𝘁𝗲𝘀 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀. Stern voices were just as effective as gentle ones and, in male voices, even increased customer questions. • 𝗔𝗴𝗲 𝗶𝘀𝗻’𝘁 𝘁𝗵𝗲 𝗶𝘀𝘀𝘂𝗲. 𝗲𝗻𝗴𝗮𝗴𝗲𝗺𝗲𝗻𝘁 𝗶𝘀. There was no significant difference between “young” and “old” voices. What mattered was that older‑sounding voices kept people talking longer. • 𝗪𝗼𝗿𝗱𝘀 𝗺𝗮𝘁𝘁𝗲𝗿. Using affirmative sentences - particularly in female voices - prompted more customer inquiries, whereas rhetorical questions were less effective. For leaders in banking and finance, this isn’t just academic. Voice is the new front door of your brand. A gentle but confident tone can build trust with high‑net‑worth clients. An affirmative female voice can reassure anxious SME owners. Conversely, a playful chatbot might unintentionally undermine credibility. 𝗦𝗼𝗺𝗲 𝗾𝘂𝗶𝗰𝗸 𝗮𝗰𝘁𝗶𝗼𝗻𝘀 𝘁𝗼 𝗰𝗼𝗻𝘀𝗶𝗱𝗲𝗿: 1. Audit your AI voice scripts. Are you using affirmative statements that invite dialogue? 2. Experiment with different voice personas. Avoid flirty tones and observe how clients react. 3. Treat voice as part of your CX strategy. Integrate data from calls, chatbots and apps so you can personalize the experience for each customer, because customer empathy is your competitive moat. We’ve moved from building “voices” metaphorically to designing them intentionally. The tone of your AI isn’t just a detail, it’s part of the customer experience. Link to research in comments below. #AI #Voice
-
The best design will soon be invisible. Interfaces used to ask: what do you need? Agentic AI flips it to: what have I already handled? The surface shrinks. We’ll gesture less and grant more permission. Location, calendar, biometrics, preference history: these signals replace tap-and-type. The UI only shows up when confidence drops and the agent needs clarity. The foreground becomes explanation: “Here’s what I did, veto if wrong.” The background is silent execution. Multimodal stops being a demo trick. Voice for speed. Text for precision. Glanceable cards for audit. Users glide across modes instead of switching apps. Design shifts from fetching tasks to negotiating autonomy. Micro-copy matters more than motion. Reversible actions matter more than dark-mode flair. If an agent moves money or publishes words, it owes the user a trail they can scan in seconds. Solving for who makes invisible work feel trustworthy is the edge. Build the layer that hides the work and surfaces the proof. Boundless Ventures
-
Enterprises today are drowning in multimodal data - text, images, audio, video, time-series, and more. Large multimodal LLMs promise to make sense of this, but in practice, embeddings alone often collapse nuance and context. You get fluency without grounding, answers without reasoning, “black boxes” where transparency matters most. That’s why the new IEEE paper “Building Multimodal Knowledge Graphs: Automation for Enterprise Integration” by Ritvik G, Joey Yip, Revathy Venkataramanan, and Dr. Amit Sheth really resonates with me. Instead of forcing LLMs to carry the entire cognitive burden, their framework shows how automated Multi Modal Knowledge Graphs (MMKGs) can bring structure, semantics, and provenance into the picture. What excites me most is the way the authors combine two forces that usually live apart. On one side, bottom-up context extraction - pulling meaning directly from raw multimodal data like text, images, and audio. On the other, top-down schema refinement - bringing in structure, rules, and enterprise-specific ontologies. Together, this creates a feedback loop between emergence and design: the graph learns from the data but also stays grounded in organizational needs. And this isn’t just theoretical elegance. In their Nourich case study, the framework shows how a food image, ingredient list, and dietary guidelines can be linked into a multimodal knowledge graph that actually reasons about whether a recipe is suitable for a diabetic vegetarian diet - and then suggests structured modifications. That’s enterprise relevance in action. To me, this signals a bigger shift: LLMs alone won’t carry enterprise AI into the future. The future is neurosymbolic, multimodal, and automated. Enterprises that invest in these hybrid architectures will unlock explainability, scale, and trust in ways current “all-LLM” strategies simply cannot. Link to the paper -> https://lnkd.in/gv93znbQ #KnowledgeGraphs #MultimodalAI #NeurosymbolicAI #EnterpriseAI #KnowledgeGraphLifecycle #MMKG #AIResearch #Automation #EnterpriseIntegration
-
Designing for the Senses: Why multi-sensory spaces create deeper human connection Ever wondered why some spaces stay with you long after you leave them? It’s not just what you see , it’s what you feel, hear, and sense. Multi-sensory design is reshaping how we approach interiors and architecture. Instead of focusing only on aesthetics, it engages multiple senses at once: • The texture of materials under your hand • The soundscape that shapes focus or calm • The scent that evokes memory and emotion • The lighting that influences mood and energy When senses interact, the experience becomes more memorable, meaningful, and human-centered. The impact of sensory design is powerful: • Stronger emotional connection to a space • Heightened awareness and perception • Greater engagement and social interaction • Improved wellbeing through thoughtful stimuli As designers, our challenge is to move beyond the visual to create immersive environments that truly resonate with people. 👉 The question: Are we designing just for the eye, or for the whole human experience? #SensoryDesign #InteriorDesign #Architecture #MultiSensoryExperience #HumanCentricDesign #DesignInnovation #WellbeingThroughDesign #ExperientialDesign
-
Exciting Research Alert: Multimodal Semantic Retrieval Revolutionizing E-commerce Product Search! Just came across a fascinating paper from Amazon researchers that tackles a crucial challenge in e-commerce search - integrating both text and image data for better product discovery. >> Key Innovations The researchers developed two groundbreaking architectures: - A 4-tower multimodal model combining BERT and CLIP for processing both text and images - A streamlined 3-tower model that achieves comparable performance with reduced complexity >> Technical Deep Dive The system leverages dual-encoder architecture with some impressive components: - Bi-encoder BERT model for processing text queries and product descriptions - Visual transformers from CLIP for image processing - Advanced fusion techniques including concatenation and MLP-based approaches - Cosine similarity scoring for efficient large-scale retrieval >> Real-world Impact The results are remarkable: - Up to 78.6% recall@100 for product retrieval - Over 50% exact match precision - Significant reduction in irrelevant results to just 11.9% >> Industry Applications This research has major implications for: - E-commerce search optimization - Visual product discovery - Large-scale retrieval systems - Cross-modal product recommendations What's particularly impressive is how the system handles millions of products while maintaining computational efficiency through smart architectural choices. This work represents a significant step forward in making online shopping more intuitive and accurate. The researchers from Amazon have demonstrated that combining visual and textual information can dramatically improve search relevance while maintaining scalability.