PROPSURVIVAL

A simulated pass probability has two errors.
Only one of them is for sale.

Six trades, arranged 720 ways, add to $1,400 every time — and in 144 of those arrangements the account is closed before it collects. Counting how often that happens is the whole job of a simulation, and it counts very precisely. But its precision belongs to the counting, not to the trades it was handed, and only one of those two errors falls when you add paths.

What this establishes
  • 01The rule reads the first moment a path touches a line. It never reads the ending balance.
  • 02An account can be closed without ever having been behind: two wins, three losses, and the floor arrives at exactly the $50,000 it opened with.
  • 03That 20% belongs to the example, not the mechanism: matched sets give 0% · 0% · 20% · 80% · 80%, ordered by the allowance divided by the worst single loss.
  • 04Simulation error falls as one over the square root of the path count. Parameter error does not fall at all.
  • 05With a 200-trade history the two cross at about 15 paths. It reverses at 13,273 for a 200,000-trade history, which is why this is a finding and not a tautology.
01

What does a pass probability actually measure?

Definitional

A pass probability measures how often a path touches one line before it touches another. It is not a summary of where accounts finished, and the rule that ends an account never consults an ending balance.

Write the rule down and the shape of the problem is forced. An account starts at E₀. Its equity E moves trade by trade. M is the highest equity the account has ever reached, and the elimination floor sits a fixed allowance A beneath it:

The rule, in full
floorF = M − A
eliminated whenE ≤ F
equivalentlyM − E ≥ A
passed whenE ≥ E₀ + T

The second line is the whole difficulty. M − E is the distance from the account's best moment to wherever it stands now, and it is a property of the path, not of the trades. Two accounts can hold identical trades, identical totals and identical statistics, and differ in M − E at every step.

A pass probability is the frequency with which a path reaches one boundary before the other. It is a first-passage question, and first passage is decided by order.

This is why the question cannot be answered by arithmetic on your results. Win rate, expectancy, profit factor and total return are all properties of the multiset of trades — they are unchanged by shuffling. The floor is not.

Symbols
E₀starting equity
Eequity now
Mhighest equity so far — it never falls
Athe allowance: how far below M the floor sits
Tthe profit target
nnumber of simulated paths
A stated convention
Everything here tests for elimination once per trade. A real account is tested continuously, so a per-trade test understates how often the floor is reached. Every number below is therefore optimistic in the same direction.
02

Why can't win rate and expectancy give you the answer?

Definitional · exhaustive

Because win rate and expectancy are properties of the trades, and elimination is a property of the order. This is not an argument from a model; it can be settled by counting.

Take six trades, and test only the floor for now — +$1,500, +$1,200, −$900, −$1,000, −$800, +$1,400 — on a $50,000 account with a $2,000 allowance. Three wins, three losses. They sum to $1,400, so every ordering has a 50% win rate, an expectancy of $233.33 per trade, and the same $1,400 result. There are 720 orderings. Walk all of them.

144 of the 720 orderings are eliminated — 20% — and the 576 that survive collect the full $1,400. Nothing a trader would quote about these trades distinguishes the two groups.

The mechanism is visible in the losing case. Two wins arrive first, so M ratchets to $52,700 and drags the floor up to $50,700 behind it. The floor never comes back down. Three losses then walk the equity into it on trade 5 — at $50,000.

Read that number again. The account is eliminated at exactly the balance it opened with. It has not lost a cent overall, and the $1,400 winner still to come never arrives, because there is no longer an account to place it in.

Check it yourself
Six items have 720 arrangements — small enough for a spreadsheet. Run a cumulative sum, keep a running maximum, and flag any row where the running maximum minus the equity reaches $2,000.
If your count differs from 144, this article is wrong and the rest of it does not survive.
The scope of that count
This enumeration tests the floor alone — no profit target — because the question is whether order by itself decides survival. Apply the $3,000 target used later in this article and the count changes: 144 orderings reach the target first, 108 are eliminated (15%), and 468 finish in between.
Both are printed because the scope is what makes them differ, and a reader who assumed the other one would have every right to think this page was wrong.
When elimination happens
t336 orderings
t436 orderings
t536 orderings
t636 orderings
Never on trades 1 or 2: the two largest losses together are $1,900, and the floor is $2,000 away.
Same six trades. Does the order change the outcome?
48,00050,00052,00012345TRADEORDERING A — WINS FIRSTOUT ON TRADE 5 AT $50,000FLOOR STOOD AT $50,700123456TRADEORDERING B — INTERLEAVEDSURVIVES AT $51,400ROOM NEVER USED $1,000
Every ordering of these six trades has the same 50% win rate, the same $233.33 expectancy and the same $1,400 total. Ordering A is eliminated on trade 5 at exactly $50,000 — the balance it started with, having lost nothing overall. Ordering B finishes $1,400 up with $1,000 of room it never used. Of all 720 orderings, 144 are eliminated.
Enumeration · exhaustive
Orderings720
Eliminated144
Share20%
Earliest eliminationtrade 3
Allowance$2,000

Definitional · no simulation

03

How much of that 20% is the mechanism, and how much is the example?

Definitional · exhaustive

Almost none of it: the 20% is a property of the six trades clause 02 happened to choose, not of the rule. Five trade sets with an identical win rate, an identical total and an identical expectancy put the eliminated share anywhere from 0% to 80%.

All five are matched on every statistic a trader would quote: three wins, three losses, $1,400 total, $233.33 expectancy. Only the shape differs. The middle column is the allowance divided by the set's worst single loss — how many of its own bad days the account can absorb before the floor arrives.

Matched on win rate, total and expectancy
Trade setA ÷ worst lossEliminated
+800  +700  -300  -400  -500  +1100 4.00 0%
+400  +350  -200  -250  -150  +1250 8.00 0%
+1500  +1200  -900  -1000  -800  +1400 2.00 20%
+2500  +1000  -1400  -1500  -600  +1400 1.33 80%
+1500  +1900  -1500  -1600  -1300  +2400 1.25 80%

Same win rate, same expectancy, same total — and the share of orderings eliminated runs from 0% to 80%. The statistics do not merely fail to give you the number. They fail to bound it.

Now hold the trades fixed and move only the allowance. The first column is what changes; the middle column is what is doing the work:

One trade set, five allowances
AllowanceA ÷ worst lossEliminated
$1,5001.5080%
$1,8001.8060%
$2,0002.0020%
$2,5002.5020%
$3,0003.000%

Across all ten cases enumerated above, the share never rises as that middle column rises: at 1.25 it is 80%, at 2.00 it is 20%, and by 8.00 it is 0%. That is a description of these ten cases and not a law. But it is the quantity the summary statistics cannot see, and it is the only number on this page a reader can compute from their own account: the allowance, divided by the worst single loss they have taken.

Order is the one property of a trade log that no statistic carries, so the only way to price it is to count it. That is what a simulator is for. Everything below is about how much of its count can be believed.

What survives the correction
The enumeration in clause 02 is exact and its conclusion stands: order decides elimination, and the statistics cannot see it. What does not survive is the magnitude. Any single worked example fixes a share, and the share it fixes is a property of that example.
So the quantity to carry out of this clause is not a percentage. It is the middle column.
04

How much does the simulation's own randomness move the number?

Definitional

At 25,000 paths the 95% interval on the pass probability is 1.15 points wide — and the simulation's randomness is the only part of this problem that money can solve.

Each simulated path either passes or does not, so the count of passes is a binomial draw and the estimate carries the standard error of a proportion:

Simulation error
standard errorse = √( p(1−p) / n )
95% interval width2 × 1.96 × se
thereforequadruple n, halve the error

Run the illustrative model below at 25,000 paths and it returns 68.976% — 17,244 passes, 7,756 eliminations, 0 accounts that ran out of trades. The 95% interval is 1.15 points wide.

Interval width by path count · points
PathsWidth
1,0005.73
5,0002.56
25,0001.15
100,0000.57
1,000,0000.18

This error has a price list. It falls without limit, it never stops falling, and you buy it with compute you already own.

Which is exactly why it is the error every tool displays — and exactly why displaying it alone is misleading.

The illustrative model
E₀$50,000
A$2,000
T$3,000
win$500, 55% of trades
loss−$400
E[·]$95 per trade
cap120 trades
seed4417
Invented inputs, chosen to be legible rather than realistic. Every number they produce is illustrative and reproducible from the generator shipped with this article — never a measurement of anyone's account.
Note on the horizon
0 of 25,000 paths hit the 120-trade cap, so the deadline is doing no work here. A binding deadline could only lower the pass rate.
05

How much does your trade history move the number?

Illustrative · seed 4417

Far more than the path count does: on a 200-trade history the identical model spans 40.088% to 87.472%, and no amount of compute narrows that by a hundredth of a point.

The model above was handed a 55% win rate as though it were known. It is not known; it is estimated, from however many trades you have logged. A win rate of 55% observed over 200 trades carries a 95% Wilson interval of 48.08% to 61.74%. Running the identical model at each end of that interval, and at every other history length, gives:

Win-rate interval, propagated · % passing
Trades loggedWin ratePass rateWidthPaths to parity
40 39.83–69.29 11.90–96.59 84.69 5
80 44.12–65.42 24.45–93.15 68.70 7
200 48.08–61.74 40.09–87.47 47.38 15
500 50.62–59.31 51.03–82.08 31.05 35
2,000 52.81–57.17 60.29–76.07 15.78 133
10,000 54.02–55.97 65.06–72.16 7.11 652
200,000 54.78–55.22 67.95–69.52 1.57 13,273

The width is the result, not the midpoint. At 200 logged trades the band is 47.384pp wide, and running ten times as many paths moves it by nothing at all.

The last column is where the two errors change rank. Simulation error falls as paths are added and the parameter band does not, so somewhere the falling quantity crosses the flat one. With a 40-trade history that crossing arrives at five paths, because the band the simulation has to get under is 84.688pp wide. Five, not 25,000.

It is not symmetric, and it is not centred on the number you were given. Around 68.976% the arms are −28.89 and +18.50 points. Writing it as "±24" would be wrong in both directions at once, which is why both arms are printed wherever this band appears.

At short histories the interval contains a losing system. At 40 logged trades the win rate could honestly be 39.83%, which on these payoffs is an expectancy of −$41.54 per trade, and the pass rate runs from 11.899% to 96.587%. A band that wide is not a precise answer with error attached. It is the absence of an answer.

What this band prices
Only the win rate. The payoff sizes and the dependence between trades are held at point estimates, and both carry uncertainty of their own — so the published band understates parameter error.
In the other direction, running the model at the two ends of an interval is a worst-case reading rather than a posterior: it asks how far the answer could move, not how likely each move is. Both distortions are stated; neither is netted against the other.
What the crossing is not
The crossing marks where the two errors swap ranks, not where either becomes small. At 40 logged trades they are equal at five paths because the parameter band is enormous, not because the simulation has become good.
Common random numbers
Path i draws the identical random stream at both ends of the interval, so the difference between the two runs is a parameter effect rather than the noise of two unrelated simulations. Without this, the reversal in clause 06 cannot be measured at all.
A trade count is not an amount of information
The table's left column counts trades, and the Wilson interval treats every one of them as a draw from the same coin. A history spanning two market regimes, or one taken before a change in position sizing, holds fewer independent observations than its length suggests.
The band is therefore an upper bound on how much a history of that length can tell you, not a description of any particular history of that length.
Which of the two errors can you buy your way out of?
0204060pp101001,00010k100k1000kPATHS SIMULATEDPARAMETER BAND — 200,000-TRADE HISTORY · 1.574ppPARAMETER BAND — 200-TRADE HISTORY · 47.384pp15 PATHS13,273 PATHSSIMULATION ERROR
Simulation error is the black curve: it falls as one over the square root of the path count and keeps falling for as long as you are willing to pay for compute. The parameter band is a horizontal line — adding paths does not move it by a hundredth of a point. With a 200-trade history the two cross at about 15 paths; every path after that buys precision in the error that is not deciding the answer. The dashed line is the same quantity for a 200,000-trade history: still flat, but low enough that simulation error is the larger of the two until roughly 13,273 paths. That reversal is why this is a finding and not an arithmetic tautology.
95% interval width · pp
Simulation · 1,000 paths5.73
Simulation · 25,000 paths1.15
Simulation · 1,000,000 paths0.18
Parameter · 200 trades47.38
Parameter · 200,000 trades1.57

Illustrative · seed 4417 · 400,000 paths per endpoint · common random numbers

06

Which error can you buy your way out of — and when does that reverse?

The finding

Simulation error is the one for sale, and on a 200-trade history it stops being the larger of the two after about 15 paths — a crossing that only reaches the path counts anyone actually runs, 13,273, at a history of 200,000 trades.

Simulation error is for sale: spend compute and it falls without limit. Parameter error is not for sale at any price. The only thing that moves it is more trades.

The threshold is where this becomes actionable. With a 200-trade history, simulation error drops below the parameter band after roughly 15 paths. Past that point every additional path is refining a quantity that is not deciding the answer. The last column of the table in clause 05 gives that crossing for every history length, and it is the history that sets it rather than the simulator: the shorter the history, the earlier the crossing, because the band the simulation has to get under is wider.

Which leaves two things worth acting on, and the path count is neither: the length of the history, and the decision to read the output as an ordering rather than as a level — which is clause 08.

And it reverses. The claim above would be an arithmetic tautology — an estimate from 200 observations is noisier than one from 25,000 draws — if there were no regime in which it failed. There is one. With a 200,000-trade history the parameter band narrows to 1.574pp, and simulation error is the larger of the two until about 13,273 paths.

That reversal is measured, not asserted. Repeating the 200,000-trade case under six independent seeds gives 1.574, 1.557, 1.594, 1.596, 1.567, 1.605 points — a spread of 0.048pp on a reading of 1.574pp.

One thing this article will not tell you is that one error is some number of times larger than the other. A ratio between a quantity that falls as one over the square root of n and a quantity that does not fall at all is a number chosen by whoever picks n: at 1,000 paths it is 8×, at 25,000 it is 41×, at a million it is 261×. Same data, three headlines. A finding you can tune is not a finding.

Precision is not accuracy
Weigh yourself ten times and average the readings. The number steadies, and every decimal place gained is real. Not one of those weighings can tell you the scale itself reads heavy. Adding paths is adding weighings.
Why the seeds are printed
A reversal measured with an instrument noisier than the reading is not a measurement. The spread across six seeds here is 0.048pp on a reading of 1.574pp — about 3% of the quantity being claimed.
Every seed is printed rather than their mean, so the reader can see the spread instead of taking the stability on the publisher's word.
What would falsify this
Re-run the shipped generator. If the parameter band at 200 trades is not close to 47.384pp, or the crossings are not near 15 and 13,273 paths, the finding is wrong.
07

Does reshuffling your real trades preserve what matters?

Declared assumption

Not by default: reshuffling discards the ordering, which is precisely what clause 02 showed the floor charges for.

Resampling a trade log in random order is attractive because it appears to assume nothing: no distribution is fitted, the trades are your own. But drawing them in random order is itself an assumption. It assumes the trades are exchangeable — that the order carries no information.

Shuffling your trades assumes the order does not matter. A trailing floor is a price on order. The one assumption the shuffle makes is the one the rule was written to collect on.

The direction of the error follows from one assumption, stated rather than buried: that losses in real trading arrive in runs. Runs are the mechanism that walks equity down to a floor, so shuffling breaks them apart, the resampled paths reach the floor less often than the real process would, and the pass probability comes back too high.

This is a direction, not a magnitude, and it has a condition under which it reverses: if a history's losses are anti-clustered — a loss making the next loss less likely — then shuffling concentrates them instead, and the bias runs the other way. How common such histories are is not measured here, which is why the reversal is named rather than dismissed.

Block resampling, which draws short consecutive runs rather than single trades, preserves dependence over the length of the block and is the standard repair. It does not fix what it cannot see: a regime that has not occurred in your sample, a change in your own behaviour after a drawdown, or the fact that the block length itself is estimated from the same short history.

The survivorship problem underneath
A trade log exists because the account it came from was not closed. Fitting inputs to surviving accounts inflates the estimated edge before any resampling begins, and no resampling scheme repairs it.
08

What can the number honestly be used for?

Definitional + illustrative

As an ordering, not as a level: a pass probability compares two configurations reliably and describes one account badly.

Comparison survives the errors in clause 05 because both configurations inherit the same mistaken inputs. If a smaller position size raises the modelled pass rate, that ordering holds across the whole parameter band, and the direction is trustworthy even though neither level is.

The level does not survive. Look at where accounts actually stop and there is a region nothing occupies — not rarely, but by construction.

Why the gap must exist
an account that has not passedM < E₀ + T
elimination requiresE ≤ M − A
so every non-passing account stops belowE₀ + T − A = $51,000

The gap between the highest non-passing outcome and the lowest passing one is exactly the allowance, for any distribution of trade outcomes, whenever the deadline does not bind. The mean of the run — $51,931 — falls inside it.

So the average outcome is an outcome that 25,000 accounts managed to miss 25,000 times out of 25,000. Checked against the identity: the highest non-passing account stopped at $50,900, against a bound of $51,000. With continuous payoffs over 40,000 paths it reaches $50,984.79 — closer to the bound, still under it, still nothing in the gap.

Which settles the grammar of the result. The model produces a frequency across an ensemble of accounts under assumed inputs, so the sentence it supports is "of accounts with these inputs, 68.976% pass". It does not support "you have a 69% chance" — that is a personal probability, and it would need a calibration study nobody has published.

Where the gap closes
Add a deadline that actually binds and accounts start stopping mid-range, on the day they run out of time rather than at a boundary. The gap fills in, and the mean becomes a fairer summary. In this run 0 paths timed out, so it does not.
The usable form
Compare two settings under one set of inputs and trust the ordering. Quote a level and you have quoted the midpoint of a band you did not show.
Where do the paths actually end?
48,00050,00052,000EQUITY WHEN THE ACCOUNT STOPPEDNO PATH CAN END HERE$2,000 = THE ALLOWANCEMEAN $51,9310 OF 25,000 PATHSBREACHED · 31.024%PASSED · 68.976%
Nothing can finish in the wine bracket, and no simulation was needed to know it: an account that has not passed never had a high-water mark above $53,000, and it cannot be eliminated until it is $2,000 below that mark — so it must stop below $51,000. The gap is the allowance, exactly, for any distribution of trade outcomes. The mean of this run, $51,931, lands inside it — in a $250 bin holding 0 of the 25,000 paths.
Run · seed 4417
Paths25,000
Passed17,244
Eliminated7,756
Ran out of trades0
Mean$51,931
Median$53,100
Highest non-passing end$50,900

Bracket definitional · bars illustrative, seed 4417

09

Where does this argument stop?

Limits

At the model itself: every claim above is a claim about a toy, and the limits below say where the toy stops resembling an account.

  • The model is a toy and says so. Two fixed payoffs, independent trades, one position size, elimination tested once per trade. Real trading has none of those properties. It is built to be legible and reproducible, not realistic, and every number it produces is illustrative.
  • The rule modelled is generic. A floor at a fixed distance below a running maximum, plus a target. No firm is named, no firm's rules are quoted, and nothing here should be taken as a description of any particular programme. Real rules add daily limits, lock points, consistency requirements and deadlines, each of which changes the arithmetic.
  • The parameter band prices one input. Win rate only, as clause 05 states. Payoff sizes and dependence are held fixed, so the true parameter uncertainty is wider than what is drawn.
  • Clause 07 is a direction, not a measurement. The upward bias from i.i.d. resampling is argued from mechanism and stated with its reversal condition. No magnitude is claimed, because none was measured.
  • Execution is free in this model. Both payoffs are gross: no commission is deducted, and no fill is ever worse than the price named. Charging either would shrink the two payoffs, which lowers every path and shortens its distance to the floor — so the eliminated share above is a bound on the real one rather than an estimate of it. Only the direction is claimed here; no magnitude is, because none was measured.
  • The enumeration is exact; everything else is a run. Clauses 02, 03 and the gap identity in 08 are counting and algebra, and hold regardless of what any simulation says. The rest depends on the shipped generator and a stated seed.

And the refusal this argument leads to: if your trade history is short, no simulator can answer the question you came with. Not this one, not a better one, not one with more paths. Find the row in clause 05 that matches the number of trades you have logged; that band is a property of the sample, and the only thing that narrows it is more trading.

That conclusion deserves a disclosure. It is published by a company that sells a simulator, and it is a conclusion that flatters any tool willing to print parameter bands. The arithmetic is checkable and the incentive is not, so the honest response is to re-run the generator rather than to trust the publisher — and corrections to the arithmetic will be applied to this article rather than appended to it.

Not financial advice
Nothing here is a recommendation to trade, to size a position, or to attempt an evaluation. It is a description of how a class of calculation behaves.
10

What else does this article answer?

Verbatim in structured data

What does a simulated pass probability actually measure?

A simulated pass probability measures how often accounts survived across the sequences the simulator generated, under the inputs it was handed. That makes it a statement about an ensemble and about those inputs — not a forecast, and not a property of the strategy on its own. Two separate errors sit inside such a number: the sampling error of the count, which falls as paths are added, and the error in the inputs themselves, which does not.

Answers
Each answer below is quoted verbatim in this page’s structured data. The markup asserts nothing the page does not say.

Can a Monte Carlo simulation predict whether I will pass a prop firm evaluation?

No. It reports how often accounts pass under a set of assumed inputs, which is a statement about an ensemble and not about one account. The output is only as good as the inputs, and the inputs are estimated from a finite trade history that carries its own uncertainty.

How many simulation runs are enough?

Simulation error falls as one over the square root of the path count, so 25,000 paths give a 95% interval about 1.15 points wide in the illustrative model below, seed 4417. But past the point where that width drops below the uncertainty in your inputs, additional paths buy precision in the error that is not deciding the answer. With a 200-trade history that point arrives at roughly 15 paths.

Why does my backtest pass but the simulation says I fail?

A backtest is one ordering of your trades. Under a trailing floor the order is priced: in the worked example in this article, 144 of the 720 orderings of the same six trades are eliminated while the rest survive, with the win rate, the expectancy and the total identical in every one.

Does reshuffling my trades destroy real information in them?

Yes, if the trades are serially dependent. Resampling in random order assumes the order does not matter, and a trailing floor is a charge on order. Destroying loss runs removes the mechanism that reaches the floor, which biases the pass probability upward. Block resampling preserves short runs; it does not fix regime change.

How many trades do I need before the number is meaningful?

Enough that the interval stops containing a losing system. At a 40-trade history the 95% interval on the win rate is wide enough to contain one — an expectancy of −$41.54 per trade at its lower bound — and the pass probability across that interval runs from 11.899% to 96.587% in the illustrative model below, seed 4417, 400,000 paths per endpoint. No number of simulated paths narrows that.

Provenance
Published2026-07-27
Every number emitted byfigures.data.js
Generatormulberry32 · seed 4417 · per-path streams
Base run25,000 paths
Endpoint runs400,000 paths · common random numbers
Enumeration720 orderings · exhaustive
Gates shipped with itgates.js · gates.render.js
Firm rules usednone

Clauses 01–03 and the gap identity in 08 are definitional. Every other number is illustrative, reproducible from the generator shipped beside this article, and describes an invented model — never an account, a firm, or a measured outcome.