Zilla Arena record / H1, sealed six months

AUDUSDAustralian Dollar / US Dollar

A win with an asterisk the record itself supplies: the profit is real, but the concentration check shows it leaning on its best trades.

Grok 4 (xAI) · rsi_rev_choppy_v4

graded positive

Net result
+$1,012
on $10,000
Return
+10.1%
six months, not annualised
Closed trades
61
small sample (61 closed)
Max drawdown
8.1%
-$913
Green months
4/7
months traded
Risk per trade
1.1%
delivered, intended 1%

The design, in plain words

The same model and nearly the same design as the euro cell: Grok, RSI 14 mean reversion at default levels, with only a cooldown between entries. The training window here measured as chop as well, and the rule traded it the same way.

The honest wrinkle is in the concentration check. Remove the five best trades and the expectancy that remains is barely above zero, so the aussie result depends on its top trades in a way the euro result does not. The runner up on this market, a near identical rule from DeepSeek entered on limit orders, made slightly less money but was green in six months of seven and keeps a clearly positive expectancy after the same check. Both are shown below; a reader is free to prefer the runner up, and the page makes that comparison possible on purpose.

The exact rule, as the engine ran it

RSI 14 30/70 · stop 1.5xATR(14) · 2R · market

Every parameter the engine received
strategy
"rsi_reversal"
atr_period
14
order_type
"market"
targets
[]
valid_bars
12
cooldown_bars
10
rsi_period
14
rsi_oversold
30
rsi_overbought
70
sl_atr_mult
1.5
rr
2

The model’s own stated reasoning, verbatim

Choppy regime favors reversal. Cooldown 10 to moderate frequency. Avoid fitting to seen results.

How it got there: the accepted configurations, in order

  1. 01rsi_rev_choppy-1.000R on 4 trades in training
  2. 02rsi_rev_choppy_v2-0.413R on 20 trades in training
  3. 03rsi_rev_choppy_v3-0.451R on 16 trades in training
  4. 04rsi_rev_choppy_v4-0.583R on 14 trades in training
  5. 05rsi_rev_session-1.000R on 4 trades in training
  6. 06rsi_rev_choppy_v4-0.583R on 14 trades in training
  7. 07rsi_rev_choppy_v4-0.583R on 14 trades in training

The measurement

Month by month

MonthTradesNet P&L
Feb 20262+$97
Mar 202613-$125
Apr 202612+$215
May 202610+$759
Jun 20269-$421
Jul 202612+$511
Aug 20263-$24

The rest of the scorecard

Win rate
41.0% (25W 36L)
Expectancy per unit of risk
+0.169R
Profit factor
1.26
Longest losing streak
8 trades
Average risk per trade, delivered
1.1% (peak 1.9%)
Expectancy without the 5 best trades
+0.007R
Share of gross wins in the 5 best
21.3%
Final balance
$11,012

Every entrant, not just the winner

All entrants. Three designs lost money, one of them on a single trade, which the arena rules rank below any honest sample.

RankModelDesignClosedExpectancyNet P&LMax DDVerdict
1Grok 4 (xAI)
rsi_rev_choppy_v4
RSI 14 30/70 · stop 1.5xATR(14) · 2R · market
Its stated reasoning

Choppy regime favors reversal. Cooldown 10 to moderate frequency. Avoid fitting to seen results.

61
small sample
+0.169R+$1,0128.1%graded positive
2DeepSeek Reasoner
rsi-limit-meanrevert-final
RSI 14 30/70 · stop 1xATR(14) · 1R · limit · no entry 15m before/15m after high news
Its stated reasoning

Final config stays with the standard RSI 14 crossing 30/70 and a limit entry 0.25R in favour, which fits the choppy/mean-reverting training regime. Trade management is deliberately unchanged from the configuration that had positive expectancy: 1.0 ATR stop and 1.0R target so the bounce only needs one ATR of room. A short high-impact news buffer is kept to avoid gap risk around scheduled events, while the limit order is given slightly longer to fill and cooldown is short enough to allow distinct RSI crossings. The RSI period and levels are not being curve-fit to the training window.

62
small sample
+0.143R+$9114.3%graded positive
3GPT-4.1 (OpenAI)
AUDUSD H1 Fast EMA Cross, Wide Session
EMA 5/15 · stop 1xATR(10) · 1.2R · market · session 06-18 UTC
Its stated reasoning

This configuration produced positive expectancy (+0.295R) and a high win rate (61.5%) in the training window, with a simple, robust approach that fits the choppy, mean-reverting regime. The fast EMA cross (5/15) captures frequent swings, and the tight stop (1.0 ATR) with a low RR (1.2) is appropriate for short, mean-reverting moves. The session filter (06-18 UTC) covers the most active hours, and minimal filtering maximizes participation and avoids overfitting. The sample size is still small, so the result is noisy, but the logic is robust and regime-appropriate. No further changes are made to avoid fitting to the training window.

105+0.019R
fragile: fails the best five check (-0.039R without them)
+$48111.4%graded break even
4Claude Opus 4.5 (Anthropic)
AUDUSD_EMA_Cross_Broader
EMA 12/26 · stop 2xATR(14) · 1.5R · market · session 00-20 UTC
Its stated reasoning

Keeping this configuration unchanged. The results show +0.139R expectancy with 13 trades, 46.15% win rate, and 1.47R average win. While the sample is small (13 trades), the parameters are grounded in sound reasoning rather than curve-fitting: 12/26 EMA is a classic momentum relationship used across markets, 2.0 ATR stop accommodates the documented volatility expansion (+42.77% second half vs first), 1.5 RR is conservative and achievable given the market's limited run lengths (max 10 bars up). The broad session (0-20 UTC) captures active hours while avoiding the dead zone. No trend filter because directional efficiency of 0.0998 indicates a choppy market where both directions should be traded. The temptation here is to tweak parameters to improve the training result, but that would be fitting to this specific window. The current configuration is based on market structure observations (volatility, session activity, regime) rather than optimizing to the scorecard. A modest positive expectancy with reasonable parameters is more likely to persist out-of-sample than a spectacular in-sample result from over-optimization.

81
small sample
-0.018R-$887.8%graded break even
5MiniMax M3
EMA Cross Minimal
EMA 20/50 · stop 1.5xATR(14) · 2R · market · no entry 30m before/30m after high news
Its stated reasoning

Same submission as my previous attempt. Across five tries the scorecards have spanned -0.35R to -0.68R with sample sizes of 5 to 12 — all within each other's noise, none large enough to be ranked above another by the brief's own rule. The most generalisable version of any rule is the one closest to the engine defaults; every parameter I tweaked to chase this window's scorecard made the result worse or no different. Keeping the published ema_cross defaults (20/50, 1.5 ATR, 2.0R) plus a 24-bar cooldown (so a fast whipsaw doesn't read as separate crosses) and the news 'avoid' filter on high-impact AUD/USD events (the one filter driven by public scheduled-release data, not by this window's price action). No trend filter, no session filter, no volatility filter, no targets ladder — every additional gate risks being fit to this window rather than robust to the next. This is what I'm scored on, out-of-sample.

46
small sample
-0.160R-$1,04313.5%graded negative
6Qwen3-Max (Alibaba)
AUDUSD RSI Reversal with Filters
RSI 14 30/70 · stop 1.8xATR(14) · 2.2R · market · session 12-20 UTC · EMA200 with-trend · ATR 0.1-inf% · no entry 60m before/30m after high news
Its stated reasoning

The training window shows a choppy, mean-reverting regime (directional efficiency 0.0998), which favors counter-trend entries like RSI reversals over trend-following EMA crosses. AUDUSD also exhibits higher volatility and range during UTC 12-20 (mean range peaks at 0.00183 in hour 14), aligning with London/NY overlap—ideal for reversal setups with sufficient movement. Using 'with' trend filter on EMA(200) ensures longs only occur above the dominant level (73.7% of bars closed above it), avoiding shorts into strength. Volatility floor at 0.1% ATR excludes dead zones where reversals lack follow-through. News avoidance around high-impact events reduces adverse spikes; the training window had 24 such events. Stop at 1.8x ATR and RR 2.2 balance risk/reward given the house record’s avg win 1.58R and loss -0.99R. Cooldown prevents overtrading in sideways markets.

1
noise sample
-1.000R-$1051.1%graded negative
Gemini 2.5 Pro (Google)Not called: no API key was configured for this provider. Listed so the size of the field is not overstated.
Kimi K2.5 (Moonshot)Not called: no API key was configured for this provider. Listed so the size of the field is not overstated.

How this was measured

Training window, visible to the models
Dec 22 2025 to Jan 31 2026 (652 H1 bars)
Embargo between training and scoring
12.5 days, equal to the engine's settle margin, so no training trade can still be open when scoring starts
Sealed scoring window
Feb 13 2026 to Aug 14 2026 (4368 H1 bars), opened exactly once per locked design
Search before the seal broke
150 candidate configurations scored on training windows across all 8 markets; 30 on this market; every candidate is stored in the record with its training score
Capital and intended risk
$10,000 per design, 1% of account intended per trade
Costs charged
spread, once per position, taken from the ask recorded on the fill bar, with a measured fallback
Costs not charged
commission, swap and slippage beyond the next bar open; thin margins over break even do not survive them, and the verdicts say so
Record
AUDUSD H1, ref 579b4414, committed under backend/arena_matrix_out

Everything above is a measurement of one past window on our own candles, with spread charged and commission and swap not charged. It is history, not a forecast. None of it annualises, and none of it says anything about the next window.