Zilla Arena record / H1, sealed six months

USDCADUS Dollar / Canadian Dollar

One design made money and five lost. This is what a market without an easy edge looks like, and the one win is a small sample.

Claude Opus 4.5 (Anthropic) · USDCAD_EMACross_v1

graded positive

Net result
+$1,194
on $10,000
Return
+11.9%
six months, not annualised
Closed trades
32
small sample (32 closed)
Max drawdown
3.9%
-$390
Green months
4/7
months traded
Risk per trade
1.0%
delivered, intended 1%

The design, in plain words

Claude took the opposite road from the mean reversion crowd: a momentum rule that buys when a 12 period average crosses over a 26 period average and sells the reverse cross, but only during the hours around the North American session where this pair actually moves. A two ATR stop gives the cross room to breathe, and a modest one and a half to one target takes profit before the chop takes it back.

It closed 32 trades in six months, which is below the hundred the platform treats as a measurable sample, so the headline number deserves suspicion on size alone. What the record does support: the design was green in four of seven months, its worst run was three losses, and its result survives the concentration check. The five losing designs below are the context that makes the one that won credible.

The exact rule, as the engine ran it

EMA 12/26 · stop 2xATR(14) · 1.5R · market · session 12-18 UTC

Every parameter the engine received
strategy
"ema_cross"
atr_period
14
order_type
"market"
targets
[]
valid_bars
12
cooldown_bars
12
fast
12
slow
26
sl_atr_mult
2
rr
1.5
session_start_hour
12
session_end_hour
18

The model’s own stated reasoning, verbatim

Keeping this configuration unchanged. The result shows +0.466R expectancy with profit factor 2.10, but critically only 5 trades - this is pure noise statistically. However, the parameters themselves are grounded in market structure rather than fitted to outcomes: 12/26 EMAs are standard responsive periods, 2.0 ATR stop accommodates the choppy regime (directional_efficiency 0.0343), 1.5 RR takes profits before mean-reversion kicks in, and 12-18 UTC session captures documented peak volatility (0.00174-0.00247 mean range). Any change now would be fitting to 5 data points, which is worse than maintaining a principled design. The house record shows -0.0246R expectancy - even break-even would beat it. This design is simple, mechanically sound, and based on observable market characteristics rather than in-sample optimization. I am marking final because further iteration with n=5 would be curve-fitting, not improvement.

How it got there: the accepted configurations, in order

  1. 01USDCAD_MeanRevert_v1-1.000R on 1 trades in training
  2. 02USDCAD_MeanRevert_v2-0.225R on 11 trades in training
  3. 03USDCAD_EMACross_v1+0.466R on 5 trades in training
  4. 04USDCAD_EMACross_v1+0.466R on 5 trades in training

The measurement

Month by month

MonthTradesNet P&L
Feb 20263-$305
Mar 20265-$9
Apr 20266+$130
May 20264-$172
Jun 20265+$232
Jul 20266+$870
Aug 20263+$448

The rest of the scorecard

Win rate
56.3% (18W 14L)
Expectancy per unit of risk
+0.356R
Profit factor
1.83
Longest losing streak
3 trades
Average risk per trade, delivered
1.0% (peak 1.1%)
Expectancy without the 5 best trades
+0.147R
Share of gross wins in the 5 best
29.2%
Final balance
$11,194

Every entrant, not just the winner

The full field: five of the six called designs lost money. The losses are printed at the same weight as the win.

RankModelDesignClosedExpectancyNet P&LMax DDVerdict
1Claude Opus 4.5 (Anthropic)
USDCAD_EMACross_v1
EMA 12/26 · stop 2xATR(14) · 1.5R · market · session 12-18 UTC
Its stated reasoning

Keeping this configuration unchanged. The result shows +0.466R expectancy with profit factor 2.10, but critically only 5 trades - this is pure noise statistically. However, the parameters themselves are grounded in market structure rather than fitted to outcomes: 12/26 EMAs are standard responsive periods, 2.0 ATR stop accommodates the choppy regime (directional_efficiency 0.0343), 1.5 RR takes profits before mean-reversion kicks in, and 12-18 UTC session captures documented peak volatility (0.00174-0.00247 mean range). Any change now would be fitting to 5 data points, which is worse than maintaining a principled design. The house record shows -0.0246R expectancy - even break-even would beat it. This design is simple, mechanically sound, and based on observable market characteristics rather than in-sample optimization. I am marking final because further iteration with n=5 would be curve-fitting, not improvement.

32
small sample
+0.356R+$1,1943.9%graded positive
2DeepSeek Reasoner
RSI Counter-Trend Limit
RSI 14 35/65 · stop 1.5xATR(14) · 2R ladder 1R@0.5/2R@0.5 · limit · session 08-18 UTC · EMA100 counter-trend · no entry 30m before/30m after high news
Its stated reasoning

The unfiltered RSI test lost because 30/70 extremes were rare and often marked the start of a move rather than a reversal. This revision uses 35/65 for more crossings, a counter-trend EMA(100) filter so longs are taken below the medium-term average and shorts above it, a limit entry with finite validity so fills require a small retracement instead of chasing the cross, and session/news guards to avoid the least liquid and most event-driven bars.

42
small sample
-0.088R-$3837.6%graded negative
3GPT-4.1 (OpenAI)
USDCAD H1 RSI Fast Mean-Revert FINAL
RSI 8 36/64 · stop 1xATR(10) · 1.2R · market · session 07-17 UTC · ATR 0.03-0.18%
Its stated reasoning

Final revision: The choppy, mean-reverting regime on USDCAD H1 is best addressed with a fast RSI reversal. I lengthened the RSI period to 8 and widened the reversal bands to 36/64 to further reduce noise and false signals, while still capturing frequent reversals. ATR stop is set to 1.0 and RR to 1.2, balancing room for trades to develop with a modest reward, aiming for a higher win rate in a market that rarely trends. Session and ATR percent filters remain to focus on active hours and avoid volatility extremes. Cooldown is kept at 4 bars for more frequent entries. No news filter to maximize sample size and generalizability. This configuration is designed to be robust for a choppy, mean-reverting USDCAD H1 environment and to avoid overfitting to the training window.

143-0.177R-$1,75021.6%graded negative
4Qwen3-Max (Alibaba)
USDCAD RSI Counter-Trend Reversal
RSI 14 30/70 · stop 1.8xATR(14) · 2.2R · market · session 12-20 UTC · EMA200 counter-trend · ATR 0.07-inf% · no entry 45m before/30m after high+med news
Its stated reasoning

The training window shows a choppy, mean-reverting regime (directional efficiency 0.0343), favoring counter-trend entries over trend-following. RSI reversal aligns with this by capturing pullbacks from extremes. The session filter (12-20 UTC) targets the highest volatility hours (mean range peaks at 14-15 UTC), improving signal reliability. A 'counter' trend filter against EMA(200) exploits the slight bearish bias (price closed below start) while avoiding false signals in sideways zones. Volatility floor (ATR% >= 0.07) ensures trades occur only when movement justifies risk; no ceiling avoids missing valid setups as volatility rose 30% in the second half. News avoidance around medium+ events reduces noise during scheduled releases (44 such events in training). Stop (1.8x ATR) and RR (2.2) balance realistic H1 swings (mean ATR 0.00119) with positive expectancy, while cooldown (15 bars) prevents overtrading in ranging conditions.

16
noise sample
-0.033R-$605.0%graded break even
5MiniMax M3
Choppy Regime RSI Reversal
RSI 14 30/70 · stop 1.5xATR(14) · 2R ladder 1R@0.5/2R@0.5 · market · session 11-17 UTC · ATR 0.05-0.18% · no entry 30m before/30m after high news
20
noise sample
-0.161R-$3425.8%graded negative
6Grok 4 (xAI)
rsi_rev_session
RSI 14 30/70 · stop 1.5xATR(14) · 2R · market · session 11-16 UTC
Its stated reasoning

Choppy regime produced too many low-quality signals. Added session filter to higher-range hours (11-16 UTC) and raised cooldown to cut frequency while keeping core reversal logic.

18
noise sample
-0.205R-$3946.3%graded negative
Gemini 2.5 Pro (Google)Not called: no API key was configured for this provider. Listed so the size of the field is not overstated.
Kimi K2.5 (Moonshot)Not called: no API key was configured for this provider. Listed so the size of the field is not overstated.

How this was measured

Training window, visible to the models
Dec 22 2025 to Jan 31 2026 (652 H1 bars)
Embargo between training and scoring
12.5 days, equal to the engine's settle margin, so no training trade can still be open when scoring starts
Sealed scoring window
Feb 13 2026 to Aug 14 2026 (4368 H1 bars), opened exactly once per locked design
Search before the seal broke
150 candidate configurations scored on training windows across all 8 markets; 19 on this market; every candidate is stored in the record with its training score
Capital and intended risk
$10,000 per design, 1% of account intended per trade
Costs charged
spread, once per position, taken from the ask recorded on the fill bar, with a measured fallback
Costs not charged
commission, swap and slippage beyond the next bar open; thin margins over break even do not survive them, and the verdicts say so
Record
USDCAD H1, ref c5b07138, committed under backend/arena_matrix_out

Everything above is a measurement of one past window on our own candles, with spread charged and commission and swap not charged. It is history, not a forecast. None of it annualises, and none of it says anything about the next window.