THE CREW - Research Note
What the machine does, what was changed on the way to the shipped version, and what each change did - with the tests that decided it.
Version 2.84.
All figures in this note are Strategy Tester runs on real broker ticks at 100 percent history quality. They are backtests, not results from live trading. Every number names the build, the broker feed and the window it came from; a number that cannot be traced that way does not appear here. Ideas that were tried and set aside are described in words, without numbers, because the number that mattered was a historical build's, not this one's.
0. The result in one table
The shipped build runs on its default configuration - Balanced 6 (six of the eleven specialists), Conservative risk tier (a 10 percent floating-exposure budget). Backtest, MetaTrader 5 Strategy Tester, real ticks, 100 percent history quality, XAUUSD, 2025-07-01 to 2026-09-07, 10,000 starting balance, measured on build 2.83 (the shipped 2.84 build was verified behaviour-identical to 2.83 by a trade-for-trade parity test on the same feed and window - identical deal count, profit matched to the cent):
| feed | net | % of deposit | profit factor | equity drawdown |
|---|---|---|---|---|
| Pepperstone | +7,569.28 | +75.7% | 4.42 | 6.5% |
| BlackBull | +8,083.94 | +80.8% | 4.12 | 6.17% |
| Fusion | +7,965.23 | +79.7% | 4.87 | 6.11% |
| Vantage | +8,011.91 | +80.1% | 4.43 | 6.13% |
| OANDA (demo, spread-survival check only) | +1,827.46 | +18.3% | 4.31 | 2.01% |
Same configuration, four different live cost structures, four different amounts of money. That gap is the subject of section 7. How the machine got from a raw idea to this table is the rest of this note.
1. How it was tested
Before the mechanism, the rules the testing followed - because the numbers are only worth what the method is worth.
- Real ticks only. Every figure in this note comes from a Strategy Tester run on a broker's own tick archive at 100 percent history quality.
- One binary, byte for byte. The build is compiled once and the identical bytes are copied to every terminal used for testing; a guard refuses to run a leg whose on-disk binary does not match the recorded build, so a stale binary can never be mistaken for a change that did something.
- Behaviour-neutral changes are proved, not asserted. Whenever the compiled surface changed without an intended change to trading logic, a trade-for-trade parity check confirmed the deal count and the money matched the previous build to the cent, on the same feed and window, before the change shipped.
- Four brokers are one market, read through four cost structures. Pepperstone, Vantage, Fusion and BlackBull see the same gold market. Agreement across them shows a mechanism survives the broker; it does not multiply the evidence by four, and disagreement across them is itself a finding (section 7).
- A demo feed is a spread-survival check, not a result to lead with. OANDA appears throughout this note and the manual for that reason only.
2. The base mechanism: eleven specialists, one warden
The machine does not send one method into every hour of the trading day. It runs eleven, each built for one hour window and one direction, and puts one warden above all of them who is not herself a trader.
Each specialist waits for its own hour and takes at most one position a day. When a position moves against it, the specialist adds another position of the same size - never larger - a measured distance further out, which pulls the group's average entry price toward the market. The whole group closes together once price reaches that shared average target. There is no per-trade stop loss; the only brake is the warden, who can flatten the entire book if the account's floating loss passes a configured share of its peak.
Three of the eleven specialists can want the same hour and the same direction. A fixed priority order, not a race, decides which one is allowed to act when that happens - a consistency device, not a risk control (its effect on the worst measured drawdown was found to be negligible; a narrower roster or a lower risk tier is what actually reduces risk).
The number of specialists enabled is the primary risk control, and the risk tier - the share of the account budgeted for floating exposure - is the second. Both are exposed as inputs; nothing about a specialist's own hour, gate or ladder geometry is, because those interlock and a partial edit produces a configuration nobody tested.
3. Building the risk architecture
The ladder mechanism alone is not the whole story of what ships. Three further layers were added on top of it, in this order, each one tested before it shipped.
A per-specialist pre-seed check. Before a specialist opens a new group, a feature drawn from that specialist's own recent trading history is read, and the group is refused - or sized down - if the reading sits in the dangerous tail of that specialist's own distribution. This shipped switched off at first, which was later found to be a mistake: two of the eleven specialists carry almost all of the book's worst-case floating loss between them, and with the check off, both of those specialists ran fully unguarded in every roster that included them. Turning the check on by default cut both of their worst cases sharply, without touching any other specialist's behaviour.
An absolute threshold, not a relative one. The check was first built to rank today's reading against that specialist's own recent history; it was rebuilt to compare against a fixed distance in market terms instead, and then widened slightly once testing located a plateau in front of a cliff - the gate now sits inside that plateau rather than at its edge, and a per-direction cap across the whole book was added alongside it as a second, independent brake.
The warm-up problem. A check that calibrates against its own history does not work on day one of a fresh account, because the history it needs has not been recorded yet. The same configuration, run from two different start dates on the same feed, produced very different measured floating losses for exactly this reason - the shorter history understated how deep the book could go, because its own safety check had not yet learned where the tail was. The fix pre-loads each specialist's check from measured reference values at start-up, so a fresh account is calibrated from its first decision rather than starting blind. This is the harmful direction to get wrong: a backtest run over many years accumulates history as it goes and looks safer than a fresh deployment actually is, not the other way round. Anyone re-measuring this engine's numbers on a short window should treat the result as the more honest one for a new account, not the long backtest.
A further correction followed from the same change: once the pre-seed check began refusing some entries, the reference figures the position-sizing arithmetic divided by were measured again, because the old reference values were stale by a wide margin - sizing had been equalising against a number that no longer described the specialist it was attached to.
4. What was tested and rejected
So that the reader knows what is not in the box, and why - described in words, because the specific numbers belong to superseded builds and do not describe the shipped one.
- A hard cap on how many layers a specialist's ladder may add. This was tested as an obvious risk control and made things worse, not better: capped at a shallower depth, a group could no longer pull its average price toward the market, so it stayed open far longer and its floating loss grew larger, not smaller. Adding layers is what lets a group close; capping them is what strands it open.
- Closing a group once its floating loss passes a fixed amount. This converts a position that was on track to recover into a realised loss. Measured on one specialist, a group that would have finished ahead finished behind instead. No input for this exists in the shipped build.
- Scaling the ladder's spacing with a volatility measure. The idea was that a wider spacing in a volatile period should behave more consistently across regimes. Measured, it did the opposite: it cost a large share of the net result and made the outcome more variable across regimes, not less, because widening the spacing in a volatile period starves the ladder of the layers it needs exactly when it needs them most. The lesson kept from this experiment: for a recovery grid, high volatility calls for more layers, not wider ones.
5. Compliance and hygiene passes
Two further passes shipped without touching any trading decision, confirmed by the same trade-for-trade parity method described in section 1.
The adjustable surface was collapsed. What had been dozens of independently adjustable levers - gate constants, ladder geometry, calendar filters, research toggles - was reduced to the twelve inputs a user sees today. Everything else is compiled as a fixed constant at the value it was validated on. The full adjustable surface still exists in one source file and can be rebuilt for research use, but the shipped product and any research finding describe the same engine, because they come from the same file.
An affordability check on every order, not only the first of a group. A market-readiness validator surfaced two mechanical edge cases: on a very small account, later layers of a recovery were being sent without checking whether the account could actually afford them, and the engine's own one-minute clock was occasionally trying to send or close an order while the symbol's market was shut. Both are now checked before every order the engine sends: available margin and the broker's own per-symbol volume limit for the first, the symbol's own trading session for the second. An order the broker would decline is simply not sent, rather than sent and refused.
6. Sizing: turning a risk share into a lot size
The engine is linear in traded volume: flat lot sizes inside a group, no per-trade stop loss, no escalation while a recovery is running. That means doubling the lot doubles both the profit and the floating exposure, and nothing else about the trade set changes. A risk setting can therefore be stated as one honest number - the share of the balance a user is willing to see floating - and the lot is derived by dividing that budget by the specialist roster's own measured worst-case floating exposure at the smallest tradable lot. The full mechanism, including what happens on an account too small to reach the requested budget, is in the manual, section 5b.
7. Why results differ by broker
The section 0 table runs the same configuration on four live cost structures and gets four different amounts of money for it, because a fixed-distance ladder rests its orders at prices that a wider or narrower spread reaches, or does not reach, at slightly different moments. Backtest, MetaTrader 5 Strategy Tester, real ticks, measured across several broker tick archives spanning 2022 to 2026: the worst-case floating exposure at a fixed 0.01 lot ranges roughly from the low hundreds to the mid-300s of account currency per 0.01 lot depending on the feed. The shipped sizing model uses the less forgiving end of that range rather than the friendlier one, because a user cannot know in advance which kind of feed their own broker will turn out to be, and planning against the friendlier number would be planning to be lucky.
8. Limits
- The evidence behind this note runs through a period in which gold rose substantially. A sustained decline is the condition this design has the least evidence about, because a no-stop recovery grid has not faced a full one in the tested history.
- There is no per-trade stop loss anywhere in the engine. The account-level breaker is the only brake, and it acts by realising the loss it was holding, not by preventing one.
- One specialist - the one working the hours just before the Western session opens - accounts for a large share of both the trade count and the net result. When the broker clock setting is wrong, that same specialist is also the one most affected. This is a property of the design, independent of broker or feed.
- Internally, this build's trade set is close to, but not byte-identical with, an earlier reference version of the same engine: one specialist is not fully isolated from the others by a rule that has been named precisely. The measured effect is small, and it is disclosed rather than rounded away.
- These are backtest figures on historical data. There is no live or forward track record yet.
Appendix - build order
| stage | change | how it was confirmed |
|---|---|---|
| base engine | eleven session specialists, shared warden, equal-lot ladder, scoped priority arbiter | the mechanism this note describes throughout |
| risk lever, on by default | per-specialist pre-seed check against that specialist's own dangerous tail | two specialists' worst-case floating loss cut sharply; every other specialist unaffected |
| absolute gate + book cap | fixed-distance threshold, widened into a measured plateau; per-direction exposure cap added | cliff located and avoided by testing across the threshold's own range |
| warm-start | pre-load each specialist's check from measured reference values at start-up | fresh-account and long-backtest measurements converge instead of diverging |
| reference re-measurement | position-sizing divisor recalculated after the gate changed which groups seed at all | stale divisor identified and corrected |
| facade collapse | dozens of research levers reduced to twelve shipped inputs; the rest fixed as compiled constants | trade-for-trade parity to the prior build, same feed and window |
| affordability and session checks | every order checked against margin, volume limit and the symbol's own trading session before it is sent | validator edge cases closed; canonical result unchanged to the cent |
| behaviour-neutral rename and geometry declaration | product renamed; per-specialist ladder geometry exposed through the same compiled-constant mechanism as everything else | trade-for-trade parity to the prior build, same feed and window |
Every table in this note names its feed and window; the compiled binary that produced each row is recorded by its hash in the product manifest.