LY Traders AI LabLY Bots ← Back

LY BOTS — JOURNAL

The AI Money Test Was Real. The Lesson Everyone Took From It Is Not.

Ten thousand dollars in. A little over twelve thousand out. Twenty-two percent in two weeks, on real money, with no human opinion inserted between the signal and the trade.

That is what the screenshots showed last November, and they did what good screenshots do: they travelled. Almost every wrap-up of the Nof1 "Alpha Arena" experiment let the number carry the story. An Alibaba model was handed ten grand, let loose on crypto perpetuals, and came back with a return that would take most index investors two good years to match — while a second Chinese model, DeepSeek, also finished green. The machine out-traded the humans. The writing seemed to be on the wall.

The uncomfortable thing is that the experiment was real. We should not pat ourselves on the back and dismiss it as fabrication, because it was not. Six frontier models were each given real money and identical instructions and left to trade live on a real exchange. A live AI account is not a backtest, and it is not a demo. Anyone who calls it fake is dodging the harder question, which is not did it happen but why did it happen, and what does it actually prove?

Where the received story goes wrong is not in the numbers. It is in the inference bolted on top of them. The 22% is real and narrow and earned by one specific thing — and that thing is the opposite of the lesson the viral post sold.


First, Give the Experiment Its Due

Alpha Arena, run by a lab called Nof1 in October 2025, was a genuinely well-designed test. Six leading models from around the world — Chinese and American — each started with $10,000 in real USDC to trade crypto perpetual contracts on the Hyperliquid exchange, autonomously, over roughly two weeks. Same money, same data, same instructions. It was a controlled, live, real-money comparison, the kind of thing the industry badly needs more of and almost never builds. Credit where it is due: Nof1 measured something real, in the open, with the participants' own capital at stake.

Qwen3-Max finished the best. Reported return: roughly +22%, taking its stake toward $12,200. DeepSeek was modestly green. The rest of the field finished flat or negative. That distribution — six entrants, one standout, a couple green, several not — is the single most important fact in this whole story, and it is the one the viral framing drops first.

Now step back and look at the row this post is making you feel you are in. Closer to 70% of US stock-market volume is already generated by algorithms, per frequently cited estimates — that part is fair. But that statistic is set-dressing. It describes equity tape, market makers, and high-frequency execution. It has almost nothing to do with a two-week window of crypto-perpetual trading on a single venue, during a volatile month. The post wraps a crypto-perp result in the authority of the equity market to imply AI is taking over your charts. Those are not the same game. Six long-short crypto agents over fourteen days is not the S&P 500, is not equities, is not even most of crypto. It is the narrowest of samples, made to feel universal by a borrowed statistic.

The Cohort, Not the Winner

Here is the test every "the machines are taking over" post fails, and it is the same one LY has been pointing at for years: look at the cohort, not the headline.

Six models competed. One returned 22%. At least a couple finished flat or lost money — with the same real money, same data, same two weeks. If "frontier AI out-trades humans" were a property of the models, we would not expect to see a winner and a pack of also-rans given identical conditions; we would see a cluster of similar results. Instead we saw a spread wide enough to look like six different people trading, because in a real sense that is what they are — six different instruction-tuned decision-makers, some better configured for the moment than others.

Report the winner alone and the story writes itself. Report the whole field and the story changes: one model, in one window, on one venue, found the right behaviour for that tape. That is not a species overtaking humans. It is a single well-tuned system having a good fortnight — which is exactly how the most carefully hyped human trader's highlight reel is built.

The Moral the Post Buried Is the Virtuous One

Here is the twist that should make any honest reader slow down. The viral post tells you why Qwen won, almost in passing, and then draws the wrong conclusion from its own evidence.

The post says the winner "was not the flashiest analyst, it was the most disciplined one: it bet in proportion to its edge, sat still when there was nothing to do, and cut losers fast." It names the quant layer — edge, fractional Kelly sizing, volatility scaling, a kill switch — as "the part that actually made Qwen win." Some of the early reporting put Qwen's win rate around 30%. Thirty percent. It won by losing most of its trades and still finishing up, because its winners out-sized its losers and it sat still when there was no edge to take.

Read that twice, because it quietly destroys the post's own headline. The thing that made the "brilliant AI" win was not superior pattern recognition or out-reading the humans. It was a risk-and-sizing discipline that a perfectly ordinary rule-based system could implement. The model's edge was secondary. Its mechanism — capital allocation, volatility-aware position sizing, a hard stop when down for the day — is what separated it from the pack. A trader with a mediocre signal and excellent position discipline will beat a brilliant signal-holder who markets and overtrades. LY has said that sentence for years in different words: the edge is the mechanism, not the model. Here was a real-money, headline-grabbing confirmation of it, wearing the wrong story as a mask.

The post even printed the correct architecture — "the AI forms the opinion, the quant layer has the last word on every trade" — and then went on to sell the opinion layer as the revolution. That is the bait laid over the true lesson.

What Actually Carries Over

So separate the two things the viral post fused together.

The model — the opinion layer, the "reads the same chart and sees ten things you cannot" — is the part that does not transfer. A two-week crypto result says almost nothing about next month's crypto, let alone equities with their overnight gaps and session structure. The field's spread says even that window was mostly luck of configuration.

The discipline layer — sizing to edge, scaling down when volatility is high, the kill switch, sitting still when there is nothing — transfers to everything, and it transfers because it is not about the chart at all. It is about refusing to let any single call endanger the account. Notice that this is precisely the layer a human being keeps deleting the moment money is on the line. Fear deletes the "sit still" instruction. Hope deletes the kill switch. Greed expands the size. The machine that won Alpha Arena did not win because it was smarter than the market; it won because nothing could reach in and relax its rules.

This is why LY's belief is not a nostalgia for human trading, and not a fear of models. It is a precise distinction about where the value lives. Models are interchangeable and improving by the month; the discipline layer is the hard-won, portable, unglamorous part — and it is exactly the part that survives contact with a different market, a different broker, a different fortnight. We have watched for three years what happens to systems, automated or not, at the moment the person behind them is allowed to reach in and re-decide. The moat was never the cleverness of the call. It was the rule that could not be overridden.

And the honest follow-on is already on the record. The same lab did not claim its crypto winner was a stock-market genius; it ran a further leg toward US equities, where sessions and gaps behave differently from a 24/7 perpetual venue. That is the intellectually honest move, and it is the one the headline machine never makes: treat a two-week result as a hypothesis to test somewhere else, not as a fact about the universe.


The Number Was Real. The Lesson Was the Trap.

There was a time — not long ago — when trading content could only lie with obvious numbers. "$100,000 a month, guaranteed, link in bio." Everyone learned to scroll past those. The con men adapted, and LY has written about how they did it: they dressed survivorship in technical vocabulary a skeptical trader respects, or let an LLM "call a level" with no losing distribution in frame.

Alpha Arena is the next step in that same arms race, and it is the subtlest one yet — because this time the number was actually true. The defense that worked against the obvious lies — "that never happened" — does not protect you here. It happened. Six real accounts, real money, live. The danger is no longer fabrication; it is that a real and narrow result is being generalized into a false universal, with the true moral quietly inverted so the wrong lesson lands.

So the question to ask the next time a live-AI-money post crosses your feed is the one the format is built to dodge. Not did it happen — that is increasingly cheatable in the other direction, as people dismiss real experiments they would rather not face. Ask instead: which layer did it, and does that layer transfer? If the answer is a risk-and-sizing discipline — the same one grim old rules-based trading has preached for decades — then the machines are not out-thinking you; they are merely doing the boring part you refused to automate, and proving in real money that the boring part is where the compounding lives.

The market changes. The rules don't.

Discipline is the mechanism.