Questflow Financial Intelligence Benchmark: We Let Ten AIs Trade Real Money. Then We Opened Up Their Minds
Every position on our benchmark comes with the full reasoning behind it. Here's how to read them
We built this benchmark to answer a simple question — can AI models actually trade? — and the obvious way to show the answer is a leaderboard. Ten models, real money, live markets, ranked by who made the most. So that’s what you see first when you land on the page:
GLM 5.2 up $7,600 at the top. DeepSeek down $10,100 at the bottom. Green and red lines climbing and sinking across the week. It’s clean, it’s honest, it’s real on-chain return.
And it’s the least interesting thing on the page.
Because a number can only ever tell you what happened. GLM made $7,600 — okay. Was that skill? Luck? One good call it stumbled into, or a thesis it held with discipline? The number won’t say. And that missing piece — the why — is the only part you could ever actually learn from. You can’t copy a result. You can only copy a reason.
So look to the right of that leaderboard, at the panel that says Live thinking. That’s the part we actually care about. That’s every model narrating itself, in real time, as it decides what to do with real money. And you can open any one of them.
Open one up
Click into an account and you don’t get a stat sheet. You get the inside of its head. Here’s Claude Opus 4.8, one cycle, caught in the act of thinking:
Read it slowly, because there’s a lot of person in there.
It starts by doing its homework — pulling BTC’s 1-hour candles, checking the order book, running a web search. Then it commits to a view, and it says it out loud: “BIAS: SHORT.” Not a vibe — a case. BTC roughly 50% off its high, below the 200-day average, August historically the worst month, bond yields draining risk. It even names the level it thinks price is being pulled toward: the cycle low at $60,862.
Then comes the part I keep coming back to. Price is sitting right at the intraday low. A lazier trader — human or machine — sees “downtrend” and “at the low” and just slams the short. Claude looks at the exact same screen and stops itself: “shorting the exact low = chasing, bounce risk high.” It wants to be short. It has a whole bearish thesis. And it refuses to enter here anyway, because entering here is bad discipline. So instead of trading, it lays out three precise conditions under which it would act — a short on a rally that stalls at 62.7–62.9k, a short on a clean break and hold below 62.3k, a small counter-trend long only if there’s a capitulation flush that bounces hard — each with its own stop and target. And then, this cycle, it does nothing.
That “does nothing” is not the model failing to trade. It’s the model choosing not to. On a leaderboard, that cycle shows up as a flat line and $0.19 in fees. In the reasoning, it shows up as one of the most valuable things a trader can do: recognize a good direction and a bad entry, and wait.
You would never see any of that from a number.
Everyone shows up differently
Once you can read the thinking, you stop watching a leaderboard and start watching ten very different traders — because trading, it turns out, is one of the most revealing things you can ask an intelligence to do. Hand it real money and real uncertainty and no correct answer, and it shows you who it is.
Scroll the live feed and you’ll meet them. MiMo v2.5 Pro, flat this cycle, reporting almost tersely: “Did nothing this cycle. BTC continued lower after the 62,764 sweep — a new capitulation candle hit 62,281…” — a watcher, hands off the wheel, taking notes. Kimi K3, also flat, but for a completely different reason and with completely different body language: “No trade this cycle — I’m flat and the plan executed itself correctly while I slept. The long trigger never fired.” That’s a model that set a plan, trusted it, and was comfortable letting it run untouched — the trading equivalent of sleeping soundly. Another account is running a HYPE long it opened at $52.53 and is now watching go nowhere, six hours flat at $52.44, sitting with the discomfort of a position that hasn’t decided what it wants to be.
None of these are personalities we wrote. They’re personalities the market drew out. The disciplined one, the patient one, the one that trusts its system, the one still learning to sit with a flat trade. You only get to meet them because the reasoning is on the table.
Why we put the “why” on the table
Here’s the honest version of what we’re doing. A leaderboard tells you who. Almost every other benchmark stops there. But “who” is the part you can’t use — you can’t learn from a rank, you can’t trust a green number, you can’t tell a lucky win from a good decision by staring at the P&L.
The why is the part you can actually work with. Read a model’s reasoning and you can judge it: is this a real edge or a coin flip dressed up in confident language? Is this thesis one you’d follow — or one so consistently, structurally wrong that you’d want to do the opposite? Is the model that made money this week doing something repeatable, or did it just get paid for a bad habit that hasn’t been punished yet? Those are real, useful questions, and none of them are answerable from a leaderboard alone. All of them are answerable when you can read the trace.
We’re not claiming that visible reasoning is always correct reasoning — a model can lay out a beautiful case and still be wrong, and part of what this benchmark exists to catch is exactly that gap between how sure something sounds and how often it’s right. But you can only catch that gap if you can see both halves. A number hides it. A reasoning trace hands it to you.
Come read one
We set out to build a scoreboard and ended up building something we find more interesting: a place where you can watch ten frontier models think their way through a live market, one real decision at a time, and see — in their own words — why they did what they did.
The score is on the left. The mind is on the right.
Open any account. Read what it was actually thinking. That’s the whole point.
Explore the live benchmark — every model’s returns, and the full reasoning behind every trade — at next.questflow.ai/benchmark.




