I solved 6 open Erdős problems in 5 days, using @OpenAI GPT-5.6 Sol.
I have a math background, but the Codex workflow I used does not require deep mathematical knowledge.
Here’s exactly how I approached it, including my prompts 🧵
This feels like the future of research: humans choose meaningful questions, contribute ideas and frameworks, and decide where to spend compute; AI explores at scale.
The more AI helps us discover, the more important it becomes for humans to understand those discoveries and ask
yes, nonsofic groups exist: this statement is one of many new beautiful results proved by Astra, our next major model.
We're releasing 10 such Astra proofs, complete with lean certificates and CoT walkthroughs for each of them. The results are wide-ranging, from von Neumann
🚨New Paper with @Weiye_Xi , @ciamac and @Qiaoqiao2001
Here’s two different prediction markets priced at ~6.5%:
1. Will the US confirm the existence of Aliens in 2026?
2. *that* Spurs @ Knicks Game 4, ~4th quarter.
They suggest that both these events are the same probability,
As agents move into real deployment, static benchmarks stop being enough.
What matters is whether an agent remains robust when the other side is learning how to exploit it for profit.
In our paper, we study this through profit-driven red teaming.
Here’s a simple example of how we compute the breakdown:
2 liquidated longs are taken over by vault at 95 while fair price is 100
It then unwinds:
• Sell 1 via ADL at 97 (fair price is 100)
• Sell 1 via market at 96 (fair price is 99)
Realized PnL = 97 + 96 − 2 × 95 = 3