Next week, Meryem is at World Summit AI in Amsterdam talking about what it takes to move AI from experimentation into reliable production. She’ll be joining the panel “Scaling AI beyond models: Engineering reliable, production-ready systems”, discussing the infrastructure behind production AI, from inference and orchestration to observability, evaluation and failure modes. 📅 Wednesday 7 October, 11:30am
Doubleword
Technology, Information and Internet
Let there be tokens. Open-weight inference for token-hungry workloads.
About us
Open-weight inference built for high-volume workloads. Doubleword helps teams run long-horizon agents, batch pipelines and large-scale inference with better caching, higher throughput and dramatically lower token costs. For one customer, that meant cutting inference costs by 4x on the same model and workload. Let there be tokens.
- Website
-
https://doubleword.ai/landing
External link for Doubleword
- Industry
- Technology, Information and Internet
- Company size
- 11-50 employees
- Headquarters
- London
- Type
- Privately Held
- Founded
- 2021
- Specialties
- LLMs, LLMOps, GenerativeAI, Deployment, and Inference
Employees at Doubleword
Locations
-
Primary
Get directions
London, GB
-
Get directions
New York, US
Updates
-
More efficient inference does not just lower cost, it can improve outcomes too. On SWE-bench Pro, a single DeepSeek-V4-Pro agent scored 51.2%. Running 64 independent attempts on Doubleword raised that to 70.7%. Because each completed attempt cost just $0.049 of node time, all 64 attempts per problem came to $3.13 overall, roughly the spend of a single attempt on the SGLang baseline. That is what cheaper inference makes possible: more agent loops, increased accuracy and better results. Read more: https://lnkd.in/eUjJc-MS
-
-
32.6M tokens in 10 minutes 57 seconds for just $0.58 on Doubleword.
By running the workloads in parallel, peak token processing throughput reached: 128000 tokens per second 🚀 Scaling at production grade traffic and keeping AI workloads cost-effective has many levers: We just processed 32.6 million LLM tokens in 10 minutes 57 seconds. Total cost: $0.58 or $18/billion token (for comparison Jev would be $42/billion and much slower on total throughput). The workload itself was straightforward: 30,000 classifications, each mapped into one of 80 categories (the underlying inference was hosted via an extremley scalable platform Doubleword). What interested me was not the classification task, It was the economics. That matters because a lot of AI business cases are still evaluated with the wrong mental model: one request, one model call, one cost. At scale, the real question is different: How much useful work can the system complete per minute, and at what total cost? Once you optimise for throughput, batching and reuse rather than individual requests, the cost curve starts to look very different. And that changes which use cases are commercially viable. Things like: - classifying large document estates - enriching millions of records - content and catalogue processing - compliance and quality checks - large-scale evaluation - back-office workflows that previously needed people simply because the volume was too high The model wasn’t the interesting part. The interesting part was turning inference into an operational system with very different unit economics.
-
-
Doubleword reposted this
By running the workloads in parallel, peak token processing throughput reached: 128000 tokens per second 🚀 Scaling at production grade traffic and keeping AI workloads cost-effective has many levers: We just processed 32.6 million LLM tokens in 10 minutes 57 seconds. Total cost: $0.58 or $18/billion token (for comparison Jev would be $42/billion and much slower on total throughput). The workload itself was straightforward: 30,000 classifications, each mapped into one of 80 categories (the underlying inference was hosted via an extremley scalable platform Doubleword). What interested me was not the classification task, It was the economics. That matters because a lot of AI business cases are still evaluated with the wrong mental model: one request, one model call, one cost. At scale, the real question is different: How much useful work can the system complete per minute, and at what total cost? Once you optimise for throughput, batching and reuse rather than individual requests, the cost curve starts to look very different. And that changes which use cases are commercially viable. Things like: - classifying large document estates - enriching millions of records - content and catalogue processing - compliance and quality checks - large-scale evaluation - back-office workflows that previously needed people simply because the volume was too high The model wasn’t the interesting part. The interesting part was turning inference into an operational system with very different unit economics.
-
-
64 DeepSeek agents solved 70.7% of SWE-bench Pro in under a day. We ran 46,784 attempts across all 731 problems on a single 8× B300 node. On the Doubleword stack, every agent completed the full benchmark in 20h 23m, delivering ~30× the throughput of throughput-oriented SGLang. That extra throughput mattered: a single agent solved 51.2% on average, while accepting a solution from any of the 64 pushed the score to 70.7%. Rushil breaks down the results and what’s driving that ~30× throughput gain in the latest Doubleword Inference Lab blog. Link in the comments 👇
-
-
Doubleword reposted this
A week after release, DeepSeek V4.1 Flash is already one of the most popular models on Doubleword!! The latest cost per task figures from Artificial Analysis help explain why... 🐋 Try it here: https://lnkd.in/evemyvED
-
-
Doubleword reposted this
📣 Speaker Announcement 07: Agent Engineering 🇬🇧 London! 🎙️ We're excited to announce Meryem Arik, Co-founder & CEO of Doubleword as a speaker at Agent Engineering London Founders Edition 2026 to cover the Inference Engineering track: ⚡ Meryem co-founded the company in 2021, originally as TitanML, to make open-weight models practical to run in production. Doubleword serves them at frontier quality for a fraction of closed-source cost, either as a hosted API with a batch pipeline or as an inference stack a company deploys inside its own private environment. Doubleword is shipping something amazing in the inference stack from London 🇬🇧.. We can't think of a better company than Doubleword to join the Inference Engineering track. 📚 She read theoretical physics and philosophy at Oxford, is a Forbes 30 Under 30 honouree, and speaks regularly at QCon and TEDx. 🎙️Talk: "Inference Engineering for Long-Running Agents" She joins us on the Inference ⚡ Engineering track. 🙏 Delighted to welcome her to the stage at Everyman Canary Wharf on 16 October 2026. 👉 Meet Meryem and follow the speaker lineup: https://lnkd.in/eBe-iKBj More speaker announcements coming soon. Stay tuned! #AgentEngineering #InferenceEngineering #Doubleword #LLMInference #AIEngineering #AIAgents #London
-
-
DeepSeek V4.1-Flash now on Doubleword
We've just added DeepSeek V4.1-Flash to Doubleword - available now across high-throughput async, 24-hour batch and real-time (with input caching enabled)! To try it in real time, you can start on our testing endpoints. Try it here 👉 https://lnkd.in/eRRgnyUu For larger production workloads, we can also spin up dedicated infrastructure around your requirements, with guaranteed throughput and no shared rate limits - DMs open :)
-
-
The demand for intelligence is unbounded. Unfortunately, the costs of tokens are also unbounded. Models are getting better, agents are getting more ambitious and the number of tokens they consume is exploding. We don’t think developers should have to ration intelligence because tokens are too expensive. That’s what we spend our time on at Doubleword: making leading open models dramatically more efficient through better speculation, caching, scheduling and inference systems research. All this means we’re able to offer dramatically lower token costs for your highest volume by workloads. Better token economics don’t just mean lower bills, they change what developers can afford to build. Let there be tokens. Get started at: https://app.doubleword.ai
-
✨ Let there be tokens! ✨
We’re launching Doubleword publicly today. We made a very unhinged video to tell you something important: You’re probably overpaying for your tokens. The era we’re entering is incredible. Models are getting better, agents are getting more ambitious, and the demand for intelligence is effectively unbounded. Unfortunately, the number of tokens required to power it is exploding too. At Doubleword, we’ve built our entire stack around one goal: maximum intelligence per dollar. Better caching, speculation, scheduling and high-throughput infrastructure means our customers are saving up to 4x per task, while getting more throughput and fewer rate limits. We’ve actually been live for a while. Today is just the first time we’re talking about it publicly. Try Doubleword: https://app.doubleword.ai And if you’re running a very high-volume workload, send me a message with your requirements. We’d love to see what we can do. 🙏 ✨ Let there be tokens! ✨ 🙏