Verification: the full 23.5-year record

This page holds the whole record for BBF_Cycle 4.0: how the test was run, what happened in each of the 23.5 years, how often it loses on a day, a week and a month, what changes when the account is small, where the diversification stops working, and the original MetaTrader 5 files behind every number. Read it critically. That is what it is here for.

How the test was run

Period
January 2003 – June 2026 (23.5 years)
Instruments
Four engines on USDJPY, EURJPY, AUDJPY and GBPJPY (GBPJPY at half size), each with its own recovery leg — eight synchronised symbols in total
Starting balance
¥100,000 (≈ $650) for the tables below; the capital ladder further down runs the same test at ¥200,000 up to ¥1,000,000
Costs
Spread fixed per pair at 1.5–3.0 pips; swap applied on the unfavourable side
Leverage
500:1. Margin was never the binding constraint (lowest margin level in the test: 10,685%)
Data quality
3% (MetaTrader's own history-quality figure for the eight-symbol synchronised run)

Why the costs are set against us

A backtest can be made to look like anything by choosing friendly assumptions. We chose unfriendly ones. Each pair's spread is held fixed at a level at or above what a real account usually sees, and the swap is applied on the side that costs money rather than the side that pays. If the result survives that, the gap between test and reality is more likely to work in your favour than against you. It is still a simulation, and a simulation is not a promise.

The test starts in 2003 because that is where the synchronised data for all eight symbols begins.

One number that is easy to misread. The risk setting below is a percentage of account equity, and 4.0 splits it across four engines. It is not the same ruler as the single-pair Generation 1 system described further down this page: 0.25% there and 0.25% here are different amounts of risk. Whenever the two generations appear together, treat the percentages as belonging to their own generation.

The whole period

Mode (¥100,000 start)ResultAnnual returnMax drawdownProfit factor
Fixed size (0.01 lot)+¥1,908,34812.34%1.22
Compounding 0.135% (default)+¥58,268,11631.1%12.34%1.26
Compounding 0.15%+¥101,687,92534.3%12.34%1.26
Compounding 0.175% (top of range)+¥248,574,51339.5%12.34%1.26

Read the compounded numbers with care. Compounding assumes the trade size keeps growing with the account, without limit, for 23.5 years. Real accounts hit broker lot limits and liquidity limits, so the later part of any compounded run is not achievable as printed. Notice also that the drawdown column does not move: on a ¥100,000 account the risk setting changes the return but not the worst loss — the reason is in how much capital changes the risk, and it is not a good thing.

The chart below is the same system at the same default setting, started at ¥1,000,000 (≈ $6,500) — the conservative reference run: 31.8% a year with a maximum drawdown of 9.07%.

¥1,000,000¥656M
BBF_Cycle 4.0 · account equity · 0.135% compounding · ¥1,000,000 (≈ $6,500) start · January 2003 – June 2026 · vertical axis is logarithmic.

The vertical axis is logarithmic. On a linear axis the last few years would flatten everything before them into a straight line at zero. The same lot-ceiling caveat as above applies to this curve.

Year by year

Each row below is a separate test: the account is started at ¥100,000 on 1 January of that year and stopped on 31 December, at a fixed trade size. Fixed size — not compounding — is used here so that a good year and a bad year are measured with the same ruler. The drawdown column is the worst fall inside that year.

YearNet (JPY)On ¥100,000Worst drawdown in yearPFTrades
2003+74,824+74.8%10.90%1.201,772
2004+91,713+91.7%18.87%1.212,071
2005+77,626+77.6%13.20%1.291,340
2006+61,017+61.0%15.61%1.211,401
2007+98,236+98.2%10.17%1.222,045
2008+34,929+34.9%19.59%1.043,523
2009+116,601+116.6%11.25%1.182,993
2010+30,472+30.5%21.88%1.062,311
2011+106,929+106.9%9.24%1.311,696
2012+74,891+74.9%8.71%1.41998
2013+109,007+109.0%12.66%1.291,844
2014+60,513+60.5%10.95%1.231,269
2015+79,533+79.5%16.76%1.201,868
2016+97,557+97.6%17.37%1.222,095
2017+62,929+62.9%14.65%1.231,428
2018+48,895+48.9%12.85%1.181,385
2019+39,838+39.8%16.68%1.21993
2020+70,142+70.1%23.18%1.231,467
2021+83,437+83.4%5.42%1.54851
2022+155,449+155.4%9.69%1.362,081
2023+108,138+108.1%12.78%1.321,647
2024+107,302+107.3%12.53%1.291,769
2025+43,834+43.8%23.50%1.102,050
2026 (H1)+69,191+69.2%8.72%1.59616

The percentage column is the four engines' combined result on a ¥100,000 account at a fixed trade size — a simple, uncompounded reference, not a compounded return. The thin years are highlighted: 2010 (+30.5%), 2008 (+34.9%) and 2025 (+43.8%). 2025 also contains the deepest fall inside any single year, 23.50%. A year can finish well up and still have been unpleasant to sit through for months.

The yearly figures add up to +¥1,903,003, while the single continuous 23.5-year test returned +¥1,908,348 — a difference of ¥5,345, or 0.28%. This is not rounding and not a bug. The yearly figures come from 24 independent tests, and a position still open when a test ends is force-closed by the tester (it appears as “end of test” in the raw reports). That resets the position state, so the following year's sequence of trades diverges slightly from the continuous run: 41,513 trades in total across the yearly tests against 41,325 in the continuous one. Both numbers are published exactly as measured.

If a record with no losing year makes you suspicious

It should. A system with no losing year in 24 years is exactly what a curve-fitted result looks like, and the normal way to produce one is to keep adjusting parameters until the bad years disappear. So here is the evidence that this is not what happened.

Generation 1 — the single-pair engine that is now one quarter of this system — did have a losing year, and we published it for months. It is still published, further down this page. 2010 is the year in question, and the two generations ran the same rules over the same data:

2010 (¥100,000, fixed size)Result
Generation 1 — single pair−¥4,369 (−4.4%)
Generation 2 — four engines+¥30,472 (+30.5%)

The engine that lost money in 2010 still loses money in 2010 inside 4.0. Nothing was tuned to remove it. What changed is that three other engines were trading their own pairs that year, and their combined result absorbed it. That is the whole idea: the parts lose, the portfolio absorbs. It is also why we can show you the losing part instead of hiding it.

This is not a promise that no year will ever finish negative. It says that in 23.5 years of history, at this configuration, the portfolio's losses landed inside years rather than across them. A future year can be different, and the mechanism that absorbed 2010 has a limit, described in where the diversification stops.

What losing looks like

“No losing year” hides how the money actually moves. Below is the same default run (0.135% compounding, ¥100,000) measured on shorter clocks. A period counts as a loss when the realised result for it — closed trades, swap and costs — is negative.

ClockPeriodsLosing periodsAverage lossWorstLongest losing run
Day5,8132,426 (41.7%)−0.41%−2.92%11 days
Week1,226436 (35.6%)−0.87%−7.40%5 weeks
Month28264 (22.7%)−1.53%−6.01%3 months

So: roughly two days in five lose money, one month in five loses money, and there have been stretches of three consecutive losing months. The longest the account went without setting a new high was 246 days (measured on the fixed-size run) — eight months of watching a number that does not improve. Percentages are measured against the balance at the start of each period; the figures exclude open positions, whose swings are covered by the drawdown columns elsewhere on this page.

If you intend to judge this system after four weeks, this table is the reason not to. Four weeks is inside the noise.

How much capital changes the risk

The same test, same default setting, started with different amounts:

Starting balanceAnnual returnMax drawdownProfit factor
¥100,000 (≈ $650)31.1%12.34%1.26
¥200,000 (≈ $1,300)29.1%8.69%1.26
¥300,000 (≈ $1,950)29.6%8.85%1.26
¥500,000 (≈ $3,250)30.9%9.08%1.26
¥1,000,000 (≈ $6,500)31.8%9.07%1.26

The smallest account is the risky one, which is the opposite of what most people assume. The minimum trade size in MetaTrader is fixed, and four engines each trading that minimum is already more exposure than 0.135% of ¥100,000. The setting cannot make the position smaller than the platform allows, so on a ¥100,000 account the drawdown floor of 12.34% is set by the platform, not by your choice — which is also why every risk setting in the first table produced the same drawdown.

From about ¥200,000 the floor stops binding and the drawdown settles into the 8.7–9.1% band, where the setting does what it says. Our reading of our own data: ¥100,000 is the minimum only if a 12% drawdown is acceptable to you; ¥200,000 (≈ $1,300) upward is where the risk control actually works.

Where the diversification stops

Four engines on four pairs sounds like four independent bets. It is not, and the honest version is this: all four pairs are Japanese yen crosses. They share a common factor, and a sharp move in the yen hits all four at once. At the peak of the 2022 yen weakness the combined yen-short exposure reached about 249% of account equity — roughly 2.5 times effective leverage. Margin was never close to binding (the lowest margin level in the whole test was 10,685%), but a violent yen reversal is the scenario in which this portfolio's four parts stop cancelling each other out and start moving together.

Diversification here smooths the differences in behaviour between the pairs. It does not remove the one thing they have in common. We publish this because it is the most likely way for a future year to look worse than any year in the table above.

Why the risk setting is capped

The setting that decides trade size can be raised beyond the recommended range. We tested what happens when it is, past the point where a reasonable person would stop, and we publish the whole frontier rather than the flattering part of it. The right-hand column is the same test with the automatic drawdown throttle disabled — the mechanism that shrinks trade size while the account is falling.

Risk setting (¥100,000)Annual returnMax drawdownSame run, throttle disabled
0.135–0.175% (recommended range)31.1–39.5%12.34%12.34–13.48%
0.25%54.6% (ref.)17.06%22.07%
0.5%73.3% (ref.)26.80%36.29%
1.0%78.6% (ref.)34.37%56.09%

“(ref.)” marks reference values: at these settings the compounded account grows so large inside the test that it runs into the lot ceiling, so the return figure is an artefact of that ceiling rather than something an account could reproduce. The drawdown column is the part worth reading — 12% → 17% → 27% → 34% — and the throttle column shows what those settings would do without the safety mechanism. In the shipped software the throttle is built in and cannot be switched off; the disabled runs exist so that the mechanism can be judged rather than believed.

Download the raw reports

These are the original MetaTrader 5 Strategy Tester files, unedited except for one thing: the section listing the software's internal input parameters has been removed, because those values are the product. Every result, statistic and chart in them is untouched.

The whole-period reports contain every one of the 41,000-plus trades, which makes each file roughly 78 MB — too large for a browser to open. Those are served as zip files of about 4 MB; unzip and open the report inside. Nothing is removed to make them smaller. The yearly files are a few megabytes and open directly, so on a metered connection start with one of those.

One test per year (fixed size, ¥100,000)

The continuous 23.5-year runs (zip · ≈ 4 MB each)

Throttle disabled — the runs behind the last column

The last file is a control: the research build used for the throttle-disabled tests, run at the product's own settings. It returns the identical net result and trade count as the shipped build, which is what makes the comparison in the previous section meaningful rather than decorative.

Generation 1 — the archive

4.0 replaced a single-pair system that we published in exactly the same way. Its record stays up: it is where the losing year is, and it is the earlier half of our own track record.

Generation 1 (single pair, ¥100,000, 23.5 years)Result
Year by year, fixed size23 winning years, 1 losing year (2010: −4.4%)
Whole period, fixed size+¥473,185
Whole period, 0.25% compounding14.5% a year · max drawdown 9.05%

Generation 1's percentages are on its own ruler — a single engine, not four — so they are not comparable with the 4.0 settings above. Its yearly totals and its continuous run differ by ¥2,253 (0.48%) for the same “end of test” reason described earlier.

Generation 1 raw reports (the whole-period files are zipped for the same reason):

The live record

Generation 2 (4.0) — live and public from trade one

4.0 runs on a real account, and that account is published as a free signal on MQL5: equity curve, drawdown and every closed trade, including the days it loses. The signal carries no trades from any earlier version — the record you see there is 4.0's alone, from its first trade onwards.

View the 4.0 live signal on MQL5 →

It went live on 28 July 2026, so that record is currently measured in days. Until it has had time to accumulate, treat 4.0 as what it is: a system with a 23.5-year test record and a live history that has only just started.

Generation 1 — verified differently

Generation 1 ran its forward verification before this signal existed: a demo account matched against the backtest trade by trade, then a live account. It still runs today as an internal test system, and its full test reports remain downloadable above — but its trades were never published as a public signal, and we would rather say so plainly than let you assume a record that is not there.

Comparing a backtest with a live account

The honest way to judge software like this is not the size of the backtest number — it is the size of the gap between the backtest and the live account running the same rules. Different brokers, different spreads and different execution all widen that gap.

We built a free tool that measures it: you feed it a report and a live history, and it tells you how closely the two curves track each other. It is currently published in Japanese on the main site (backtest–forward deviation checker) and the interface is being translated; the calculation is language-independent if you want to try it now.