Scatter Lab's Zeta app lets people become the main character in a story and talk to AI characters like they're real. And they now boast 2.1 million users, with 2.5-hour average daily sessions. That level of engagement requires massive LLM serving infrastructure, thousands of simultaneous interactions with low latency, every second. When the big cloud providers couldn't get them the GPUs they needed to scale, they built a multi-region orchestration layer on Runpod. Which is why they now handle 1,000+ inference requests per second at nearly half the infrastructure cost. Read their full story here: https://lnkd.in/eSeSq7Xj
About us
Runpod is the AI Developer Cloud for teams building, training, and scaling AI applications. Developers use Runpod to access GPUs, run Pods, deploy Serverless inference endpoints, and move from prototype to production without managing infrastructure from scratch. Runpod gives AI builders the primitives they need to ship faster: GPU Cloud, Serverless, persistent storage, templates, and tools built for real production workloads.
- Website
-
https://www.runpod.io
External link for Runpod
- Industry
- Software Development
- Company size
- 51-200 employees
- Headquarters
- San Francisco, CA
- Type
- Privately Held
- Founded
- 2022
- Specialties
- Machine Learning, Artificial Intelligence, Deep Learning, AI Infrastructure, GPU Cloud, Serverless AI, and GPU Computing
Employees at Runpod
Locations
-
Primary
Get directions
San Francisco, CA 94107, US
Updates
-
This is the kind of story we love hearing. Not just the technical decision, but the engineering maturity behind it. We're happy to be part of your stack Nodalview!
What does a scale-up do? It scales. Up. Fast. Six years ago, Nodalview served 2,000 active users. Today: 20,000, and ~1.5M photos processed every month. The engineering team behind that? 12 people. Scaling like this is a Ship of Theseus exercise. You replace the ship plank by plank while sailing faster and carrying more. It only works when you're surrounded by people who care as much about how you grow as how fast. About a year ago, a few of our engineers came to me with a problem we couldn't ignore. Up to 70% of our images arrive in a single one-hour window every afternoon. Spinning GPU servers up on demand meant cold starts that couldn't clear the rush. Keeping them warm meant paying for idle machines most of the day. They took it end to end. They benchmarked providers, challenged our network and software architecture, and landed on Runpod Serverless. It scales with the afternoon spike, costs nothing when idle, and keeps the stability, performance and observability we need in production. Then they went one step further and negotiated a commercial agreement that today looks more like a partnership than a volume discount. That's what I'm proudest of. Not just the architecture, but the maturity: engineers who spot the problem, own the solution, and handle the business side too. That maturity was built over years, and it shows. Philippe Gilles, Deniz Engin, Margaux Mansanarez, Sébastien Carbain : this is your work. Thank you. And thank you Brooke Spath (Gracey) for telling our story. Full case study in the comments 👇
-
-
Congrats to the team that created Peel, which won the 1st-place grand prize at HackMIT out of hundreds of projects! They built it on Runpod. Here's what the team said about their stack: "We used Runpod to host our backend on a CPU pod with a Network volume. The network volumes helped us avoid repeating setup. It also made it very simple to switch between GPU and CPU pods. We also used the Runpod Secrets manager instead of a .env file on the Pod itself." Can't wait to see what you build next! Leo Pozhenko, Alan Tai, Viktor Minchev, Nikhil Ramlukan Take a look at the project here: https://lnkd.in/dgMnz9mK
-
-
🚀 Excited to announce that we've partnered with Comfy to launch Comfy API deployments!
Comfy API is live Until now, shipping a ComfyUI workflow to production meant rebuilding its environment somewhere else. Renting GPUs, reinstalling every custom node and model, untangling Python dependencies, and writing your own scaling logic. Comfy API handles that for you: → Builder reads your workflow JSON, finds the models and custom nodes it needs, and helps resolve dependency conflicts → Each Build pins the ComfyUI version, nodes, models, and Python deps together → Releases are immutable, so the environment you test is the one you deploy → Endpoints autoscale, down to zero when idle or with warm workers when latency matters → GPU time is billed by the second on RTX PRO 6000, H100, H200, or B200 Your graph stays yours, the engine stays open source, and your Builds stay portable. Click the link below to deploy your first workflow today
-
What can 1.2 million developers tell you about where AI is actually going? Charlotte Daniels dug into the data and pulled out seven things we didn't expect. You can find them here: https://lnkd.in/esAHxBkK
-
-
Hardware supply tightened in 2026. Developers answered in software. For our second State of AI Compute report, we analyzed anonymized activity across 1.2 million Runpod users from April 2025 through August 2026. Developers are tuning the entire stack to the workload: model size, precision, serving pattern, and GPU memory. Here's what the data shows: → Agents are 8.4% of Runpod users, but resources they create generate 24% of revenue. That’s roughly 2.9× the platform average revenue per user. → 80% of Pods that reference a frontier model also host local open weights or serving stacks. → Quantization rises with model size: 11.6% of Pods below 8B parameters use a quantized build, compared with 40.7% above 70B. → 85.2% of video Pods touch post-production. Just 11.2% generate video without working on existing footage. → Eight of Runpod’s ten fastest-growing industries are not AI-specific. We also checked last year’s GPU supply forecasts against what happened. H100 SXM supply roughly doubled, as predicted. H200 grew 1.6x. B200 tripled rather than nearly quadrupled. New: B300 grew 2.4x. Download the full report: https://lnkd.in/eQBH_y39
-
Faceless.video hit $1M+ ARR without raising a dollar or hiring an infrastructure team. Jacob Seeger built it as a solo founder, moved his video generation workload onto Runpod Serverless, and cut generation costs by more than half. Since their launch, 2.5 million creators have signed up. Want to know more about their story? You can find it in the link below: https://lnkd.in/e4xxktCS
-
-
We brought LEGO to WeAreDevelopers last Thursday, and it turns out everyone's a builder. Thanks to everyone who joined us and to our co-hosts, Apify & Flox. We'll be back for more!
-
-
Just spun up your first Pod and wondering how to get your files onto it? We just wrote a guide that covers everything from scratch. SSH keys, SCP transfers, your FileZilla setup. You can find it here: https://lnkd.in/eU7bphep
-
Transcribing a thousand hours of audio on the wrong GPU costs $507. Using the right one could bring that down to just $7. We benchmarked Whisper large-v3-turbo across 23 GPUs, and the results aren't what most people expect. Datacenter cards like the H200 and B300 finished last, both slower and more expensive than the more efficient cards for this. Read the entire breakdown here: https://lnkd.in/e45PCTaW
-