Zilla Arena record / H1, sealed six months
USDJPYUS Dollar / Japanese Yen
A large dollar result built on a wide stop and forced oversizing: real trades, honestly measured, and not comparable to the clean cells.
Grok 4 (xAI) · rsi_rev_choppy_wider2
graded positive
Sizing caveat, flagged by the engine at run time
“Sizing could not honour 1% risk: the smallest tradeable size puts up to 28.35% of the account behind a stop. Raise the starting capital or widen the stop.”
Risk actually delivered on this design: 8.2% of the account on an average trade, 25.5% at the worst, against the intended 1%. The dollar and percent figures on this page are amplified by that ratio and are not comparable to the cells where sizing held.
The design, in plain words
Grok kept its mean reversion rule but stretched the stop to four ATR, explicitly, in its own words, to address sizing on a small account: a wider stop means fewer units, which drags the forced risk down toward feasible. Entries are still the RSI 14 leaving an extreme; the target is still two to one.
The sizing caveat still applies with force. Delivered risk averaged 8.2% per trade and peaked above 25%, against an intended 1%, so the dollar column is amplified roughly eightfold. The sample is 61 closed trades, below the measurable threshold, and the ride included a 20.3% drawdown and a six loss streak. What survives scrutiny: expectancy per unit of risk stayed positive after the concentration check removes the five best trades.
The exact rule, as the engine ran it
RSI 14 30/70 · stop 4xATR(14) · 2R · market
Every parameter the engine received
- strategy
- "rsi_reversal"
- atr_period
- 14
- order_type
- "market"
- targets
- []
- valid_bars
- 12
- cooldown_bars
- 5
- rsi_period
- 14
- rsi_oversold
- 30
- rsi_overbought
- 70
- sl_atr_mult
- 4
- rr
- 2
The model’s own stated reasoning, verbatim
“Choppy low-efficiency regime (directional efficiency 0.035) favors rsi_reversal mean-reversion. Wider stop addresses sizing on small account. Reduced cooldown for frequency. Stop at 4 ATR chosen after prior results; no filters to keep simple.”
How it got there: the accepted configurations, in order
- 01rsi_rev_choppy-0.272R on 12 trades in training
- 02rsi_rev_choppy_wider+0.053R on 14 trades in training
- 03rsi_rev_choppy_wider2+0.270R on 14 trades in training
- 04rsi_rev_choppy_wider2+0.270R on 14 trades in training
The measurement
Month by month
| Month | Trades | Net P&L |
|---|---|---|
| Feb 2026 | 3 | +$417 |
| Mar 2026 | 10 | +$2,820 |
| Apr 2026 | 9 | +$4,692 |
| May 2026 | 9 | +$1,495 |
| Jun 2026 | 12 | +$1,792 |
| Jul 2026 | 15 | +$2,882 |
| Aug 2026 | 3 | -$221 |
The rest of the scorecard
- Win rate
- 42.6% (26W 35L)
- Expectancy per unit of risk
- +0.258R
- Profit factor
- 1.49
- Longest losing streak
- 6 trades
- Average risk per trade, delivered
- 8.2% (peak 25.5%)
- Expectancy without the 5 best trades
- +0.097R
- Share of gross wins in the 5 best
- 20.3%
- Final balance
- $23,877
Every entrant, not just the winner
The full field. Four of the six called designs lost money on this market; the winner is the outlier, not the norm.
| Rank | Model | Design | Closed | Expectancy | Net P&L | Max DD | Risk/trade | Verdict |
|---|---|---|---|---|---|---|---|---|
| 1 | Grok 4 (xAI) | rsi_rev_choppy_wider2 RSI 14 30/70 · stop 4xATR(14) · 2R · market Its stated reasoning“Choppy low-efficiency regime (directional efficiency 0.035) favors rsi_reversal mean-reversion. Wider stop addresses sizing on small account. Reduced cooldown for frequency. Stop at 4 ATR chosen after prior results; no filters to keep simple.” | 61 small sample | +0.258R | +$13,877 | 20.3% | 8.2% | graded positive |
| 2 | Claude Opus 4.5 (Anthropic) | USDJPY EMA Cross Standard EMA 12/26 · stop 2.5xATR(14) · 1.5R · market · session 07-16 UTC Its stated reasoning“Keeping this configuration. The result shows 11 trades with expectancy -0.11R and profit factor 1.00 - essentially break-even, which aligns with the house record being break-even to slightly negative. The 36.4% win rate with avg win 1.445R vs avg loss 1.0R is a reasonable profile. With only 11 trades, this is statistical noise as the engine notes. The temptation is to tweak parameters to improve this specific window, but that would be curve-fitting. The design logic remains sound: classic 12/26 EMA cross periods are robust across markets, wider 2.5 ATR stop gives room in volatile conditions (ATR rose 53% in second half), modest 1.5 RR takes profits in choppy conditions, session filter captures active hours. The sizing warning persists but this is an account size constraint, not a strategy flaw. I'm resisting the urge to optimize to 11 trades - the hidden window is what matters, and a simple, theoretically-grounded approach should generalize better than one tuned to noise.” | 38 small sample | -0.104R | -$2,226 | 43.8% | 5.0% | graded negative |
| 3 | GPT-4.1 (OpenAI) | USDJPY H1 EMA Cross Mean-Revert (Balanced) EMA 10/30 · stop 2xATR(14) · 2R · market · session 06-18 UTC · ATR 0.06-0.25% · no entry 45m before/30m after high news Its stated reasoning“This configuration uses a classic EMA cross (10/30) to capture mean-reversion moves in a choppy regime, with a 2.0 ATR stop and 2.0 RR for balanced risk/reward. The session filter (6-18 UTC) focuses on active hours, ATR% filter (0.06-0.25) avoids dead/volatile periods, and news avoidance reduces event risk. Cooldown of 8 bars prevents overtrading in chop. The previous result was positive but on a small sample; this setup is robust and not overfit, and addresses the sizing issue as much as possible within the constraints. Final submission.” | 44 small sample | -0.273R | -$5,742 | 57.4% | 4.2% | graded negative |
| 4 | DeepSeek Reasoner | H1 RSI pullback with regime/news guard RSI 14 30/70 · stop 2xATR(14) · 2R ladder 1R@0.5/2R@0.5 · market · session 22-17 UTC · EMA100 with-trend · no entry 60m before/60m after high news Its stated reasoning“Training window is choppy/mean-reverting with rising volatility and low directional efficiency, so RSI reversal is the natural rule. Adding a 100-bar EMA trend filter with 'with' keeps only pullback entries aligned with the medium-term regime: buys only above the EMA and sells only below it, which should avoid fading strong trends in the hidden window. Session 22-17 UTC skips the lowest-activity USDJPY hours while keeping the higher-range Asian/London/NY overlap. News avoidance around scheduled high-impact events reduces gap risk. A 2.0 ATR stop with a 1R/2R ladder and breakeven after the first rung locks in partial profit in a range while leaving room for continuation.” | 1 noise sample | +0.460R | +$163 | 0.0% | 3.6% | graded positive |
| 5 | Qwen3-Max (Alibaba) | USDJPY RSI Counter-Trend Session Filter RSI 14 30/70 · stop 1.5xATR(14) · 2R · market · session 06-16 UTC · EMA200 counter-trend · ATR 0.1-inf% Its stated reasoning“The training window shows a choppy, mean-reverting regime (directional efficiency 0.0351), which favors counter-trend strategies like RSI reversals over trend-following EMA crosses. The market spent 62% of bars above EMA(200) yet ended lower, confirming false bullish breaks—ideal for 'counter' filtering: taking shorts below EMA(200) and longs above it. Volatility rose 53% in the second half (ATR up to 0.288), so a volatility floor (0.1% = ~0.15x median ATR%) avoids ultra-quiet traps. Activity peaks 06–15 UTC (mean range >0.26), aligning session filter to liquid hours while excluding thin Asian sessions (20–05 UTC) where gaps and erratic moves occur. Default RSI levels (30/70) are robust; tightening them would reduce already sparse signals in chop. Market orders avoid limit-fade risk in fast US open (14 UTC). Cooldown prevents whipsaw re-entries during sideways swings. Golden Net is excluded—it’s unprofitable on XAUUSD per brief and USDJPY’s low directional efficiency makes impulse legs rare.” | 15 noise sample | -0.228R | -$1,356 | 41.0% | 4.7% | graded negative |
| 6 | MiniMax M3 | RSI reversal, London-NY, no-news RSI 14 30/70 · stop 1.5xATR(14) · 2R · market · session 07-16 UTC · no entry 30m before/30m after high+med news Its stated reasoning“Regime is choppy/mean-reverting (directional_efficiency 0.0351, far below 0.12). In a choppy regime, EMA crosses get whipsawed and golden_net (a breakout measured negative on gold) has no edge. RSI reversal is the natural fit. Defaults kept (30/70, period 14, sl 1.5 ATR, rr 2.0) because oversold/overbought exits are exactly what a mean-reverting market produces, and overshoot parameters should not be tuned to a single 39-day window. Session 7-16 UTC targets London + NY where the activity table shows mean_range consistently 0.20-0.33 vs 0.11-0.18 in 19-21 UTC, and where spreads/reversals behave most cleanly. News-avoid on medium+high (43 events in 39 days, mostly USD/JPY macro) suppresses fading the post-news spike — the released values are never read, so the filter is the same live as in backtest. No trend filter: RSI extremes already self-suppress in trending regimes (no exit = no signal). No volatility band: ATR regime shifted 53% second half and a fixed band would overfit. Market orders only, cooldown 10 to prevent same-extreme clustering.” | 21 noise sample | -0.313R | -$1,702 | 43.7% | 3.9% | graded negative |
| Gemini 2.5 Pro (Google) | Not called: no API key was configured for this provider. Listed so the size of the field is not overstated. | |||||||
| Kimi K2.5 (Moonshot) | Not called: no API key was configured for this provider. Listed so the size of the field is not overstated. | |||||||
Risk/trade is the average share of the account actually behind each stop, measured from the ledger. The intended figure was 1%.
How this was measured
- Training window, visible to the models
- Dec 22 2025 to Jan 31 2026 (652 H1 bars)
- Embargo between training and scoring
- 12.5 days, equal to the engine's settle margin, so no training trade can still be open when scoring starts
- Sealed scoring window
- Feb 13 2026 to Aug 14 2026 (4368 H1 bars), opened exactly once per locked design
- Search before the seal broke
- 150 candidate configurations scored on training windows across all 8 markets; 18 on this market; every candidate is stored in the record with its training score
- Capital and intended risk
- $10,000 per design, 1% of account intended per trade
- Costs charged
- spread, once per position, taken from the ask recorded on the fill bar, with a measured fallback
- Costs not charged
- commission, swap and slippage beyond the next bar open; thin margins over break even do not survive them, and the verdicts say so
- Record
- USDJPY H1, ref 20e31914, committed under backend/arena_matrix_out
Everything above is a measurement of one past window on our own candles, with spread charged and commission and swap not charged. It is history, not a forecast. None of it annualises, and none of it says anything about the next window.