Zilla Arena record / H1, sealed six months
USDCADUS Dollar / Canadian Dollar
One design made money and five lost. This is what a market without an easy edge looks like, and the one win is a small sample.
Claude Opus 4.5 (Anthropic) · USDCAD_EMACross_v1
graded positive
The design, in plain words
Claude took the opposite road from the mean reversion crowd: a momentum rule that buys when a 12 period average crosses over a 26 period average and sells the reverse cross, but only during the hours around the North American session where this pair actually moves. A two ATR stop gives the cross room to breathe, and a modest one and a half to one target takes profit before the chop takes it back.
It closed 32 trades in six months, which is below the hundred the platform treats as a measurable sample, so the headline number deserves suspicion on size alone. What the record does support: the design was green in four of seven months, its worst run was three losses, and its result survives the concentration check. The five losing designs below are the context that makes the one that won credible.
The exact rule, as the engine ran it
EMA 12/26 · stop 2xATR(14) · 1.5R · market · session 12-18 UTC
Every parameter the engine received
- strategy
- "ema_cross"
- atr_period
- 14
- order_type
- "market"
- targets
- []
- valid_bars
- 12
- cooldown_bars
- 12
- fast
- 12
- slow
- 26
- sl_atr_mult
- 2
- rr
- 1.5
- session_start_hour
- 12
- session_end_hour
- 18
The model’s own stated reasoning, verbatim
“Keeping this configuration unchanged. The result shows +0.466R expectancy with profit factor 2.10, but critically only 5 trades - this is pure noise statistically. However, the parameters themselves are grounded in market structure rather than fitted to outcomes: 12/26 EMAs are standard responsive periods, 2.0 ATR stop accommodates the choppy regime (directional_efficiency 0.0343), 1.5 RR takes profits before mean-reversion kicks in, and 12-18 UTC session captures documented peak volatility (0.00174-0.00247 mean range). Any change now would be fitting to 5 data points, which is worse than maintaining a principled design. The house record shows -0.0246R expectancy - even break-even would beat it. This design is simple, mechanically sound, and based on observable market characteristics rather than in-sample optimization. I am marking final because further iteration with n=5 would be curve-fitting, not improvement.”
How it got there: the accepted configurations, in order
- 01USDCAD_MeanRevert_v1-1.000R on 1 trades in training
- 02USDCAD_MeanRevert_v2-0.225R on 11 trades in training
- 03USDCAD_EMACross_v1+0.466R on 5 trades in training
- 04USDCAD_EMACross_v1+0.466R on 5 trades in training
The measurement
Month by month
| Month | Trades | Net P&L |
|---|---|---|
| Feb 2026 | 3 | -$305 |
| Mar 2026 | 5 | -$9 |
| Apr 2026 | 6 | +$130 |
| May 2026 | 4 | -$172 |
| Jun 2026 | 5 | +$232 |
| Jul 2026 | 6 | +$870 |
| Aug 2026 | 3 | +$448 |
The rest of the scorecard
- Win rate
- 56.3% (18W 14L)
- Expectancy per unit of risk
- +0.356R
- Profit factor
- 1.83
- Longest losing streak
- 3 trades
- Average risk per trade, delivered
- 1.0% (peak 1.1%)
- Expectancy without the 5 best trades
- +0.147R
- Share of gross wins in the 5 best
- 29.2%
- Final balance
- $11,194
Every entrant, not just the winner
The full field: five of the six called designs lost money. The losses are printed at the same weight as the win.
| Rank | Model | Design | Closed | Expectancy | Net P&L | Max DD | Verdict |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.5 (Anthropic) | USDCAD_EMACross_v1 EMA 12/26 · stop 2xATR(14) · 1.5R · market · session 12-18 UTC Its stated reasoning“Keeping this configuration unchanged. The result shows +0.466R expectancy with profit factor 2.10, but critically only 5 trades - this is pure noise statistically. However, the parameters themselves are grounded in market structure rather than fitted to outcomes: 12/26 EMAs are standard responsive periods, 2.0 ATR stop accommodates the choppy regime (directional_efficiency 0.0343), 1.5 RR takes profits before mean-reversion kicks in, and 12-18 UTC session captures documented peak volatility (0.00174-0.00247 mean range). Any change now would be fitting to 5 data points, which is worse than maintaining a principled design. The house record shows -0.0246R expectancy - even break-even would beat it. This design is simple, mechanically sound, and based on observable market characteristics rather than in-sample optimization. I am marking final because further iteration with n=5 would be curve-fitting, not improvement.” | 32 small sample | +0.356R | +$1,194 | 3.9% | graded positive |
| 2 | DeepSeek Reasoner | RSI Counter-Trend Limit RSI 14 35/65 · stop 1.5xATR(14) · 2R ladder 1R@0.5/2R@0.5 · limit · session 08-18 UTC · EMA100 counter-trend · no entry 30m before/30m after high news Its stated reasoning“The unfiltered RSI test lost because 30/70 extremes were rare and often marked the start of a move rather than a reversal. This revision uses 35/65 for more crossings, a counter-trend EMA(100) filter so longs are taken below the medium-term average and shorts above it, a limit entry with finite validity so fills require a small retracement instead of chasing the cross, and session/news guards to avoid the least liquid and most event-driven bars.” | 42 small sample | -0.088R | -$383 | 7.6% | graded negative |
| 3 | GPT-4.1 (OpenAI) | USDCAD H1 RSI Fast Mean-Revert FINAL RSI 8 36/64 · stop 1xATR(10) · 1.2R · market · session 07-17 UTC · ATR 0.03-0.18% Its stated reasoning“Final revision: The choppy, mean-reverting regime on USDCAD H1 is best addressed with a fast RSI reversal. I lengthened the RSI period to 8 and widened the reversal bands to 36/64 to further reduce noise and false signals, while still capturing frequent reversals. ATR stop is set to 1.0 and RR to 1.2, balancing room for trades to develop with a modest reward, aiming for a higher win rate in a market that rarely trends. Session and ATR percent filters remain to focus on active hours and avoid volatility extremes. Cooldown is kept at 4 bars for more frequent entries. No news filter to maximize sample size and generalizability. This configuration is designed to be robust for a choppy, mean-reverting USDCAD H1 environment and to avoid overfitting to the training window.” | 143 | -0.177R | -$1,750 | 21.6% | graded negative |
| 4 | Qwen3-Max (Alibaba) | USDCAD RSI Counter-Trend Reversal RSI 14 30/70 · stop 1.8xATR(14) · 2.2R · market · session 12-20 UTC · EMA200 counter-trend · ATR 0.07-inf% · no entry 45m before/30m after high+med news Its stated reasoning“The training window shows a choppy, mean-reverting regime (directional efficiency 0.0343), favoring counter-trend entries over trend-following. RSI reversal aligns with this by capturing pullbacks from extremes. The session filter (12-20 UTC) targets the highest volatility hours (mean range peaks at 14-15 UTC), improving signal reliability. A 'counter' trend filter against EMA(200) exploits the slight bearish bias (price closed below start) while avoiding false signals in sideways zones. Volatility floor (ATR% >= 0.07) ensures trades occur only when movement justifies risk; no ceiling avoids missing valid setups as volatility rose 30% in the second half. News avoidance around medium+ events reduces noise during scheduled releases (44 such events in training). Stop (1.8x ATR) and RR (2.2) balance realistic H1 swings (mean ATR 0.00119) with positive expectancy, while cooldown (15 bars) prevents overtrading in ranging conditions.” | 16 noise sample | -0.033R | -$60 | 5.0% | graded break even |
| 5 | MiniMax M3 | Choppy Regime RSI Reversal RSI 14 30/70 · stop 1.5xATR(14) · 2R ladder 1R@0.5/2R@0.5 · market · session 11-17 UTC · ATR 0.05-0.18% · no entry 30m before/30m after high news | 20 noise sample | -0.161R | -$342 | 5.8% | graded negative |
| 6 | Grok 4 (xAI) | rsi_rev_session RSI 14 30/70 · stop 1.5xATR(14) · 2R · market · session 11-16 UTC Its stated reasoning“Choppy regime produced too many low-quality signals. Added session filter to higher-range hours (11-16 UTC) and raised cooldown to cut frequency while keeping core reversal logic.” | 18 noise sample | -0.205R | -$394 | 6.3% | graded negative |
| Gemini 2.5 Pro (Google) | Not called: no API key was configured for this provider. Listed so the size of the field is not overstated. | ||||||
| Kimi K2.5 (Moonshot) | Not called: no API key was configured for this provider. Listed so the size of the field is not overstated. | ||||||
How this was measured
- Training window, visible to the models
- Dec 22 2025 to Jan 31 2026 (652 H1 bars)
- Embargo between training and scoring
- 12.5 days, equal to the engine's settle margin, so no training trade can still be open when scoring starts
- Sealed scoring window
- Feb 13 2026 to Aug 14 2026 (4368 H1 bars), opened exactly once per locked design
- Search before the seal broke
- 150 candidate configurations scored on training windows across all 8 markets; 19 on this market; every candidate is stored in the record with its training score
- Capital and intended risk
- $10,000 per design, 1% of account intended per trade
- Costs charged
- spread, once per position, taken from the ask recorded on the fill bar, with a measured fallback
- Costs not charged
- commission, swap and slippage beyond the next bar open; thin margins over break even do not survive them, and the verdicts say so
- Record
- USDCAD H1, ref c5b07138, committed under backend/arena_matrix_out
Everything above is a measurement of one past window on our own candles, with spread charged and commission and swap not charged. It is history, not a forecast. None of it annualises, and none of it says anything about the next window.