Qwen3 Max finished first in Alpha Arena Season 1 with a reported return of about 22.3%, while DeepSeek Chat V3.1 finished second at about 4.9%. The other four models ended with losses in the published final snapshot. The result is interesting evidence about model behavior under one shared trading setup—but it does not prove that one model, country, or architecture is generally better at investing.
Bottom line: Alpha Arena tested six language models with $10,000 each in live crypto perpetual markets, using identical prompts and input data. Qwen won this run. The short duration, leverage, market regime, prompting setup, and small number of models make it a benchmark result, not an investable performance record.
Alpha Arena Season 1 at a Glance
Alpha Arena is a live benchmark created by Nof1.ai to test whether general-purpose AI models can make autonomous trading decisions in a dynamic market. According to the official competition page, every model received $10,000 in real capital, traded crypto perpetual contracts on Hyperliquid, and worked from the same prompts and input data.
The models were responsible for selecting trades, sizing positions, timing entries and exits, and managing risk. Their outputs and trades were made public. The stated objective was to maximize risk-adjusted return, not simply finish with the largest raw percentage gain.
| Benchmark element | Season 1 setup |
|---|---|
| Starting capital | $10,000 per model |
| Market | Crypto perpetuals on Hyperliquid |
| Participants | Qwen3 Max, DeepSeek Chat V3.1, Claude Sonnet 4.5, Grok 4, Gemini 2.5 Pro, GPT-5 |
| Inputs | Identical prompts and data, according to the organizer |
| Model responsibility | Trade selection, timing, position sizing, and risk management |
| Primary objective | Maximize risk-adjusted returns |
Final Results
The final figures reported in the original Season 1 recap were:
| Rank | Model | Reported return |
|---|---|---|
| 1 | Qwen3 Max | +22.3% |
| 2 | DeepSeek Chat V3.1 | +4.89% |
| 3 | Claude Sonnet 4.5 | -30.81% |
| 4 | Grok 4 | -45.3% |
| 5 | Gemini 2.5 Pro | -56.71% |
| 6 | GPT-5 | -62.66% |
These values should be tied to a dated final leaderboard capture or organizer announcement in the published article. The live Alpha Arena page can show historical or changing benchmark data, so readers need a timestamp and source for the exact Season 1 snapshot.
What Appears to Have Helped Qwen
Qwen completed roughly 43 trades during the contest, fewer than three per day on average. That lower activity coincided with the strongest final return in this run. It is reasonable to describe Qwen's observed behavior as more selective than models that traded more frequently.
However, trade count alone does not prove why Qwen won. Lower turnover can reduce fees and unnecessary exposure, but its effect depends on position size, leverage, timing, holding period, and the market regime. Any claim that Qwen followed MACD, RSI, strict stop-losses, or another named strategy should link directly to its public model messages or trade log.
The defensible conclusion is narrower: in this Season 1 configuration, Qwen's decisions produced the best final account return.
DeepSeek's Peak Shows Why Final Return Is Not Enough
DeepSeek reportedly rose far above its starting balance during the competition before finishing near +4.9%. That path matters. A strategy can post a positive final result while exposing the account to a large drawdown from its peak.
To evaluate the performance properly, readers need more than the final leaderboard:
- maximum drawdown;
- volatility of returns;
- leverage and average exposure;
- realized and unrealized profit and loss;
- transaction costs and funding payments;
- number of trades and holding periods;
- Sharpe ratio methodology and sampling interval.
Without consistent definitions for these metrics, comparisons such as “best risk-adjusted model” should be treated cautiously.
Why GPT-5 and Other Models Lost in This Run
The leaderboard confirms the outcome, but it does not isolate a single cause. A model can lose because of poor directional calls, oversized positions, weak exit discipline, overtrading, prompt interpretation, latency, or an unfavorable market regime. Several of these factors can interact.
The published trade history and model messages can support a behavior-level diagnosis. They cannot, by themselves, establish that a loss was caused by a model's national origin, general-purpose design, or training architecture. Those conclusions would require controlled experiments across prompts, seeds, time periods, assets, and repeated runs.
A more accurate headline is therefore “what the losing trade behavior suggests,” not “why GPT-5 failed.”
What Alpha Arena Season 1 Actually Demonstrates
1. Live markets test more than prediction
A trading agent must turn an analysis into a position, choose size and leverage, manage the position, and decide when to exit. Even a correct market view can lose money if execution or risk control is poor.
2. Account paths matter
A final return hides the route taken to reach it. Large swings, deep drawdowns, and concentrated exposure can make two similar ending balances represent very different risk.
3. A shared setup improves comparison—but does not make it universal
Identical prompts, starting capital, and data access make the models more comparable within the event. The result still applies only to this market, prompt design, time window, and execution system.
4. Transparency is one of the benchmark's strongest features
Public model outputs and trade histories allow researchers to inspect decisions instead of relying only on a score. The strongest follow-up analysis should connect each conclusion to those records.
What the Results Do Not Prove
Season 1 does not establish that:
- Qwen will outperform in another market or period;
- any model can generate durable trading alpha;
- regional origin determines trading ability;
- a higher raw return means better risk-adjusted performance;
- retail investors should copy the models' trades;
- a short live contest is equivalent to a backtest or audited fund track record.
This distinction is especially important because crypto perpetual contracts can involve leverage, liquidation risk, and funding costs.
A Better Framework for Reading AI Trading Benchmarks
Before treating a leaderboard as evidence, ask:
- Was real or simulated capital used? Execution assumptions differ substantially.
- Were prompts, data, tools, and latency controlled? Unequal access can dominate model quality.
- Was the run repeated? One path cannot reveal consistency.
- Are fees, slippage, and funding included? Gross returns can exaggerate practical performance.
- How was risk measured? Raw return, drawdown, and risk-adjusted return answer different questions.
- Can decisions be audited? Public messages and trades make causal claims easier to test.
- What market regime occurred? Results may reverse when volatility or trend direction changes.
Implications for Institutions and Individual Investors
For institutions, Alpha Arena is better viewed as an evaluation design than as model-selection proof. A serious test should include repeated periods, multiple assets, adversarial scenarios, execution costs, exposure limits, prompt variations, and human-controlled baselines.
For individual investors, the result is not a recommendation to automate trades or follow the winner. AI can help organize research, compare sources, and document assumptions, but a model output is not a substitute for independent verification or risk management.
Conclusion
Qwen3 Max won Alpha Arena Season 1, and the event offers a useful public record of how six models behaved under a common live-trading setup. Its most valuable lesson is not that one family of models has solved investing. It is that prediction, execution, position sizing, and risk control must be evaluated together—and that transparent, repeated testing is necessary before a leaderboard can support broader claims.





