Zilla Arena record / H1, sealed six months
AUDUSDAustralian Dollar / US Dollar
A win with an asterisk the record itself supplies: the profit is real, but the concentration check shows it leaning on its best trades.
Grok 4 (xAI) · rsi_rev_choppy_v4
graded positive
The design, in plain words
The same model and nearly the same design as the euro cell: Grok, RSI 14 mean reversion at default levels, with only a cooldown between entries. The training window here measured as chop as well, and the rule traded it the same way.
The honest wrinkle is in the concentration check. Remove the five best trades and the expectancy that remains is barely above zero, so the aussie result depends on its top trades in a way the euro result does not. The runner up on this market, a near identical rule from DeepSeek entered on limit orders, made slightly less money but was green in six months of seven and keeps a clearly positive expectancy after the same check. Both are shown below; a reader is free to prefer the runner up, and the page makes that comparison possible on purpose.
The exact rule, as the engine ran it
RSI 14 30/70 · stop 1.5xATR(14) · 2R · market
Every parameter the engine received
- strategy
- "rsi_reversal"
- atr_period
- 14
- order_type
- "market"
- targets
- []
- valid_bars
- 12
- cooldown_bars
- 10
- rsi_period
- 14
- rsi_oversold
- 30
- rsi_overbought
- 70
- sl_atr_mult
- 1.5
- rr
- 2
The model’s own stated reasoning, verbatim
“Choppy regime favors reversal. Cooldown 10 to moderate frequency. Avoid fitting to seen results.”
How it got there: the accepted configurations, in order
- 01rsi_rev_choppy-1.000R on 4 trades in training
- 02rsi_rev_choppy_v2-0.413R on 20 trades in training
- 03rsi_rev_choppy_v3-0.451R on 16 trades in training
- 04rsi_rev_choppy_v4-0.583R on 14 trades in training
- 05rsi_rev_session-1.000R on 4 trades in training
- 06rsi_rev_choppy_v4-0.583R on 14 trades in training
- 07rsi_rev_choppy_v4-0.583R on 14 trades in training
The measurement
Month by month
| Month | Trades | Net P&L |
|---|---|---|
| Feb 2026 | 2 | +$97 |
| Mar 2026 | 13 | -$125 |
| Apr 2026 | 12 | +$215 |
| May 2026 | 10 | +$759 |
| Jun 2026 | 9 | -$421 |
| Jul 2026 | 12 | +$511 |
| Aug 2026 | 3 | -$24 |
The rest of the scorecard
- Win rate
- 41.0% (25W 36L)
- Expectancy per unit of risk
- +0.169R
- Profit factor
- 1.26
- Longest losing streak
- 8 trades
- Average risk per trade, delivered
- 1.1% (peak 1.9%)
- Expectancy without the 5 best trades
- +0.007R
- Share of gross wins in the 5 best
- 21.3%
- Final balance
- $11,012
Every entrant, not just the winner
All entrants. Three designs lost money, one of them on a single trade, which the arena rules rank below any honest sample.
| Rank | Model | Design | Closed | Expectancy | Net P&L | Max DD | Verdict |
|---|---|---|---|---|---|---|---|
| 1 | Grok 4 (xAI) | rsi_rev_choppy_v4 RSI 14 30/70 · stop 1.5xATR(14) · 2R · market Its stated reasoning“Choppy regime favors reversal. Cooldown 10 to moderate frequency. Avoid fitting to seen results.” | 61 small sample | +0.169R | +$1,012 | 8.1% | graded positive |
| 2 | DeepSeek Reasoner | rsi-limit-meanrevert-final RSI 14 30/70 · stop 1xATR(14) · 1R · limit · no entry 15m before/15m after high news Its stated reasoning“Final config stays with the standard RSI 14 crossing 30/70 and a limit entry 0.25R in favour, which fits the choppy/mean-reverting training regime. Trade management is deliberately unchanged from the configuration that had positive expectancy: 1.0 ATR stop and 1.0R target so the bounce only needs one ATR of room. A short high-impact news buffer is kept to avoid gap risk around scheduled events, while the limit order is given slightly longer to fill and cooldown is short enough to allow distinct RSI crossings. The RSI period and levels are not being curve-fit to the training window.” | 62 small sample | +0.143R | +$911 | 4.3% | graded positive |
| 3 | GPT-4.1 (OpenAI) | AUDUSD H1 Fast EMA Cross, Wide Session EMA 5/15 · stop 1xATR(10) · 1.2R · market · session 06-18 UTC Its stated reasoning“This configuration produced positive expectancy (+0.295R) and a high win rate (61.5%) in the training window, with a simple, robust approach that fits the choppy, mean-reverting regime. The fast EMA cross (5/15) captures frequent swings, and the tight stop (1.0 ATR) with a low RR (1.2) is appropriate for short, mean-reverting moves. The session filter (06-18 UTC) covers the most active hours, and minimal filtering maximizes participation and avoids overfitting. The sample size is still small, so the result is noisy, but the logic is robust and regime-appropriate. No further changes are made to avoid fitting to the training window.” | 105 | +0.019R fragile: fails the best five check (-0.039R without them) | +$481 | 11.4% | graded break even |
| 4 | Claude Opus 4.5 (Anthropic) | AUDUSD_EMA_Cross_Broader EMA 12/26 · stop 2xATR(14) · 1.5R · market · session 00-20 UTC Its stated reasoning“Keeping this configuration unchanged. The results show +0.139R expectancy with 13 trades, 46.15% win rate, and 1.47R average win. While the sample is small (13 trades), the parameters are grounded in sound reasoning rather than curve-fitting: 12/26 EMA is a classic momentum relationship used across markets, 2.0 ATR stop accommodates the documented volatility expansion (+42.77% second half vs first), 1.5 RR is conservative and achievable given the market's limited run lengths (max 10 bars up). The broad session (0-20 UTC) captures active hours while avoiding the dead zone. No trend filter because directional efficiency of 0.0998 indicates a choppy market where both directions should be traded. The temptation here is to tweak parameters to improve the training result, but that would be fitting to this specific window. The current configuration is based on market structure observations (volatility, session activity, regime) rather than optimizing to the scorecard. A modest positive expectancy with reasonable parameters is more likely to persist out-of-sample than a spectacular in-sample result from over-optimization.” | 81 small sample | -0.018R | -$88 | 7.8% | graded break even |
| 5 | MiniMax M3 | EMA Cross Minimal EMA 20/50 · stop 1.5xATR(14) · 2R · market · no entry 30m before/30m after high news Its stated reasoning“Same submission as my previous attempt. Across five tries the scorecards have spanned -0.35R to -0.68R with sample sizes of 5 to 12 — all within each other's noise, none large enough to be ranked above another by the brief's own rule. The most generalisable version of any rule is the one closest to the engine defaults; every parameter I tweaked to chase this window's scorecard made the result worse or no different. Keeping the published ema_cross defaults (20/50, 1.5 ATR, 2.0R) plus a 24-bar cooldown (so a fast whipsaw doesn't read as separate crosses) and the news 'avoid' filter on high-impact AUD/USD events (the one filter driven by public scheduled-release data, not by this window's price action). No trend filter, no session filter, no volatility filter, no targets ladder — every additional gate risks being fit to this window rather than robust to the next. This is what I'm scored on, out-of-sample.” | 46 small sample | -0.160R | -$1,043 | 13.5% | graded negative |
| 6 | Qwen3-Max (Alibaba) | AUDUSD RSI Reversal with Filters RSI 14 30/70 · stop 1.8xATR(14) · 2.2R · market · session 12-20 UTC · EMA200 with-trend · ATR 0.1-inf% · no entry 60m before/30m after high news Its stated reasoning“The training window shows a choppy, mean-reverting regime (directional efficiency 0.0998), which favors counter-trend entries like RSI reversals over trend-following EMA crosses. AUDUSD also exhibits higher volatility and range during UTC 12-20 (mean range peaks at 0.00183 in hour 14), aligning with London/NY overlap—ideal for reversal setups with sufficient movement. Using 'with' trend filter on EMA(200) ensures longs only occur above the dominant level (73.7% of bars closed above it), avoiding shorts into strength. Volatility floor at 0.1% ATR excludes dead zones where reversals lack follow-through. News avoidance around high-impact events reduces adverse spikes; the training window had 24 such events. Stop at 1.8x ATR and RR 2.2 balance risk/reward given the house record’s avg win 1.58R and loss -0.99R. Cooldown prevents overtrading in sideways markets.” | 1 noise sample | -1.000R | -$105 | 1.1% | graded negative |
| Gemini 2.5 Pro (Google) | Not called: no API key was configured for this provider. Listed so the size of the field is not overstated. | ||||||
| Kimi K2.5 (Moonshot) | Not called: no API key was configured for this provider. Listed so the size of the field is not overstated. | ||||||
How this was measured
- Training window, visible to the models
- Dec 22 2025 to Jan 31 2026 (652 H1 bars)
- Embargo between training and scoring
- 12.5 days, equal to the engine's settle margin, so no training trade can still be open when scoring starts
- Sealed scoring window
- Feb 13 2026 to Aug 14 2026 (4368 H1 bars), opened exactly once per locked design
- Search before the seal broke
- 150 candidate configurations scored on training windows across all 8 markets; 30 on this market; every candidate is stored in the record with its training score
- Capital and intended risk
- $10,000 per design, 1% of account intended per trade
- Costs charged
- spread, once per position, taken from the ask recorded on the fill bar, with a measured fallback
- Costs not charged
- commission, swap and slippage beyond the next bar open; thin margins over break even do not survive them, and the verdicts say so
- Record
- AUDUSD H1, ref 579b4414, committed under backend/arena_matrix_out
Everything above is a measurement of one past window on our own candles, with spread charged and commission and swap not charged. It is history, not a forecast. None of it annualises, and none of it says anything about the next window.