Joshua Dunnink · Trading Systems

← Captain Karat-Smash

Captain Karat-Smash — Research Note

What the machine does, what was changed on the way to the shipped version, and what each change did — with the tests that decided it.

Version 1.11 · October 2026

All figures in this note are Strategy Tester runs on real broker ticks at 100 % History Quality, compiled defaults, 10,000 starting deposit unless stated otherwise. They are backtests and do not represent trading on a live account. Every number names the build, the broker feed and the window it came from; a number that cannot be traced that way does not appear here.


0. The result in one table

Version 1.11, compiled defaults. Backtest, MetaTrader 5 Strategy Tester, real ticks at 100 % history quality, 10,000 starting deposit, 2 January 2024 to 26 September 2026, three live broker feeds (Vantage's real-tick history starts in July 2025, so its row covers July 2025 on):

feed net % of deposit profit factor trades relative equity drawdown
Fusion - the published figure, the weaker of the two feeds that cover the whole window 15,456.67 +154.6 % 1.88 2,119 8.72 %
BlackBull 16,866.36 +168.7 % 2.16 1,933 6.32 %
Vantage (from July 2025) 8,155.62 +81.6 % 2.55 727 5.19 %

Captain Karat-Smash 1.11 at its default settings on three live brokers

Pepperstone, used in earlier measurements, is left out of the current figures: its real-tick archive on our test terminal has gaps in 2026, and a figure built on filled-in ticks is not published.

Earlier measurement

Captain Karat-Smash runs out of the box on its Medium risk level. One 14-month window run identically on four live broker feeds, compiled defaults, 10,000 starting deposit (backtest, MetaTrader 5 Strategy Tester, real ticks, 2025-07-01 to 2026-09-07):

feed net % of deposit profit factor trades equity drawdown
Fusion (worst qualifying live feed) 4,935.14 +49.4 % 1.73 896 7.88 %
Vantage 5,184.50 +51.8 % 2.14 705 5.65 %
BlackBull 5,234.33 +52.3 % 1.99 786 5.99 %
Pepperstone 6,907.96 +69.1 % 2.11 924 6.56 %
OANDA (demo, spread-survival check only) 7,618.23 +76.2 % 1.76 902 9.71 %

The published claim is the worst of the four live feeds, Fusion. Read as: the same seven-strategy machine takes closely related trades on every broker and still earns different money on each one, which is the honest range rather than a defect. How the machine got from a raw idea to that table is the rest of this note.


1. How it was tested

Before the mechanisms, the rules the research followed, because the numbers are only worth what the method is worth.


2. The base mechanism: one object, seven strategies

Captain Karat-Smash is a level-break system. Before a burst, gold sits under a recent high or above a recent low it has not managed to close through — an unbroken swing. Seven strategies run side by side on one chart, each watching that same kind of object on its own timeframe (two on daily bars, five on four-hour bars), each with its own entry offset inside the level and its own bracket:

  1. Level. A swing is a bar whose high or low is more extreme than the bars on both sides of it, and which no closed bar has since traded through.
  2. Entry. A stop order rests a small, strategy-specific percentage of price inside the level, so price has to carry through the level to trigger the order — a break, not a touch.
  3. Bracket. Every fill receives a stop loss and a take profit set as a percentage of the entry price. Because the geometry is proportional to price rather than a fixed dollar distance, it reads the same whether gold trades near 1,800 or near 4,300 — confirmed by measuring the same percentages holding constant, to two decimal places, across sixteen calendar quarters and four different price eras.
  4. Exit. Once a trade is far enough ahead its stop locks, then trails behind each closed one-minute bar and is never loosened. Two of the seven strategies also pull their target closer when a trade goes far enough against them, so they can leave on a bounce instead of riding to the full stop.

What the level is worth. A control that replaces the real level with a random price at the same distance and the same placement cadence keeps only a minority of the net that the real level earns, on two independent random seeds. The level is the source of the edge; the bracket geometry alone is not enough.

Nine variants, seven shipped. The underlying research identified nine timeframe-and-strictness variants of this one object. Two were set aside on causal, not curve-fit, grounds: one earned a few cents on average across nearly two thousand trades — practically inside the noise of trading costs — and lost money in the quiet and middling-volatility thirds of the sample; the other traded only a handful of times a year and lost money in two of the three volatility thirds. The seven that shipped were each profitable in at least two of the three volatility thirds, several in every one.


3. Building the exit: bracket, then lock, then trail

The bracket alone is not the finished machine. Three exit arms were compared on identical entries:

A separate control removed the trail entirely and let every position ride to its bracket stop or target: this also raised net but at a large multiple of the drawdown, which is why the close-anchored trail is the shipped risk control rather than an option.

Per-tick versus per-bar trailing. A per-tick version of the same trail moved most exits earlier by roughly two minutes on average, but the net effect across the trade population was negligible. The one-minute-close trail is what ships — cheaper to compute and no worse for it.

3.1 Version 1.07-1.09: widening and conditioning the trail

Three further changes were made to the trail after the roster and lock-then-trail exit were fixed, each proved deal-for-deal against the version before it on a real-tick window before release:


4. Sizing: the risk facade

Position size is set by one input, the balance behind each 0.01 lot; the account scales the same shape of exposure up or down with that single number. The calibration was measured, not assumed: the boundary between risk levels was set from the drawdown actually produced on the feed with the longest history and the only full flat year in the record, then re-measured on the same feed at the shipped default to confirm the prediction held before it was used. The exact figures for each level, and what each one has produced historically, are in the manual (chapter 3) rather than repeated here, because they are safety information for choosing a setting, not a growth claim.


5. Prop-firm accounts

An account-type switch sets the position size and turns off two of the seven strategies for prop-firm modes. Those two strategies were identified, by measuring the worst single day across the record, as the pair that puts the most loss into the worst days; removing them for prop accounts was decided on that measurement, not by searching for the combination that produced the most favourable backtest. A day-replay simulation that reapplies a firm's daily-loss and maximum-loss rules to the record's own trading days was used to check the prop sizing against real limits before it shipped; the pass-rate table and the rule-guard mechanics are in the manual, chapters 3a and 5a.


6. Out-of-sample: the eight-year test

The seven strategies' constants were measured on 2022-2026 data from three brokers. A fourth broker's archive reaches back to 2018 and was deliberately not used while the constants were being fixed, so it functions as a genuine out-of-sample test rather than more fitting data (backtest, MetaTrader 5 Strategy Tester, real ticks, OANDA GOLD.pro, 2018-06-13 to 2026-09-05, Medium risk level):


7. What was tested and rejected

So the reader knows what is not in the box, and why:


8. Why results differ by broker

The same machine, run in the same Strategy Tester backtest on the same window (2025-07-01 to 2026-09-07) on four different live brokers (Fusion, Vantage, BlackBull, Pepperstone; section 0), produces different net and different drawdown. The entries are closely related across feeds; what differs is the cost of getting filled and stopped — spread and execution. A percentage-of-price bracket does not remove that difference, it only keeps it proportional to the price level rather than letting it drift as gold's price changes over the years. This is why every figure in this note and in the listing names its broker and its window, and why the published claim is always the worst of the qualifying live feeds rather than the most favourable one.


9. Limits


Appendix — version history and how each change was proved

Every change below was proved deal for deal in a Strategy Tester backtest, real ticks (Pepperstone and the other feeds named in the sections above), 2025 to 2026, before it shipped.

version change proof
1.00 product build from the seven-strategy research engine fixed-lot parity: 1,293 fills, entry match 1.0000, net matched to the cent
1.01-1.03 internal constant and research-only controls exposed as switchable knobs, all off by default defaults reproduce the prior version deal for deal on a real-tick window before each release
1.04 a scheduling fix so expired orders are not repeatedly evaluated while the market is closed reproduces the version before it deal for deal on a real-tick window
1.05 prop-firm rule group added (daily loss, maximum loss, profit target, Friday close), all off unless a number is entered defaults reproduce the version before it deal for deal; each rule fires at its intended trigger in a functional test
1.06 account-type switch for prop-firm sizing and the two-strategy exclusion (section 5) personal mode reproduces the version before it deal for deal; each prop mode reproduces the strategy-off configuration it is built from, to the cent
1.07 wider trail on four strategies (section 3.1) shipped defaults reproduce the tested wide-trail configuration deal for deal on an independent real-tick window
1.08 trend-conditioned trail width (section 3.1) shipped defaults reproduce the tested trend-conditioned configuration deal for deal on the same independent window
1.09 London-session profit lock and wider with-trend trail (section 3.1) shipped defaults reproduce the tested configuration deal for deal on the same independent window

Every figure in this note names its feed and window; the compiled binary that produced each row is recorded by its hash in the product manifest.