← Back to blog

The End-to-End Run: How One Hypothesis Passes Every Gate of the Cycle — and Dies at the Last

Published 2026-08-01 · By Drew Shelem

The End-to-End Run: How One Hypothesis Passes Every Gate of the Cycle — and Dies at the Last

The finale. We gather the cycle into one checklist — and prove one thing rigorously: a strategy can pass the significance test and still be a losing bet on the regime

This is the last article in the cycle. Each previous one gave one tool: how to look for an edge with a reason, how to count degrees of freedom, how to budget for execution, how to catch dependence on the regime, how to calibrate to resolution, how to demand a sufficient sample, how to bet by the lower bound. Now we’ll gather all of it into one list of checks — and prove one fact that holds the whole cycle together. The fact is this: passing the statistical-significance test is necessary but not sufficient. A strategy with a confident, significant edge in the backtest can lose money live, if its “edge” is really a bet on the state of the market. And this can be shown honestly, without a single fabricated number.


Passed the test — and lost anyway

A finale article has a temptation: to invent an impressive run. Something like “the raw backtest gave plus six cents, and I subtract check after check, and the plus melts before your eyes.” I won’t do that. Let me explain why — it’s a lesson in itself.

Every number in such a “waterfall” would be invented by me. Whatever I put into the model is what I’d then “find.” It’s a closed loop: I hide the answer and pull it back out myself. Such a proof can’t be checked — and the whole cycle is about checking only what’s measurable. A made-up waterfall shows my assumptions, not a real mechanism. So it has no place in an article under the banner “written to be checked.”

That’s why the finale is built honestly, in two parts. The first part is a list of checks. Each of them we already took apart in a separate article, on its own data; here I simply line them up into a queue that any hypothesis must pass through. The second part is one thing that can be proven rigorously. It’s the reason this article exists: even a passed significance test doesn’t save you from a bet on the regime. I’ll show it on a simple construction with round numbers — and say honestly that I’m choosing the numbers. Because what’s proven isn’t “momentum has such-and-such an edge,” but that significance alone isn’t enough.


The list of checks: nine gates, each from its own article

Here’s the queue of questions any hypothesis goes through before you risk money. Each item is a short question and an explanation of what it catches. We’ve already covered the mechanics of each in detail; here’s the essence, so you can walk the list without rereading anything.

1. Is there a reason? Ask yourself: why should this edge exist at all? If you can name a reason in advance — say, the crowd overpays for the obvious favorite, or the contract price lags the spot — that’s a good sign. But if the edge just “showed up” when you tried a hundred combinations of indicators, it’s most likely chance, not an edge. A reason is thought up before you dig into the data, not after.

2. How many settings did you try? Count how many thresholds, windows, and filters you fitted to the data. And compare that not to the number of days, but to the number of trades. The rule is simple: there must be many times fewer settings than trades, or you’re fitting noise. And count honestly: five indicators on one bitcoin are almost one setting, not five, because they look at the same price.

3. The best of how many variants? If you tried eight versions of the strategy and kept the prettiest, part of its result is just the luck of the luckiest of eight. Even at a zero edge, the best of many random tries will show a noticeable plus. So from a pretty result you have to mentally subtract this selection markup — or you’ll take luck for skill.

4. At what price do you enter in the backtest? Here’s a common trap. Suppose you’ve already accounted for the commission, and you take the entry price at the middle of the spread (the “mid”). But you’ll never fill at the mid in real life. To buy immediately, you cross the spread and pay about half of it. And if you post a limit at the mid and wait, you get filled more often on the unfavorable side, and you miss some of the good trades. So an edge computed at the mid is overstated by the real cost of entry. Subtract that cost — your own, measured — not abstract “costs.”

5. Where do you get the outcome of a trade? The contract resolves at the oracle’s price (an aggregate of several exchanges at a preset moment), not at the Binance spot you’re watching. If your backtester decides the winner itself, by the spot, it gets it wrong near the strike. Take the outcome from Polymarket’s actual settlement: that settlement is the oracle’s verdict, and you have it for every completed contract.

6. Are there enough trades to believe the plus? A small plus over fifty trades is most likely noise. Compute the confidence interval of the edge: it’s the range the true value lies in. If the lower end of the range is above zero, the plus is probably real. If the range covers zero, you haven’t yet told the edge from luck — you need more trades.

7. Is the edge everywhere — or only in one regime? Split your trades by the state of the market: calm periods separately, stormy periods separately. Compute the edge in each. If it exists only in one state, while in the other you lose, you don’t have a strategy — you have a bet that the needed state will last. This is the most insidious check, and the whole second half of this article is about it.

8. What size do you bet? Live, the edge is almost always smaller than in the backtest (because you selected lucky strategies). So bet not by the edge you measured but by the lower end of the confidence range. Then the gap between backtest and life won’t ruin you.

9. How will you know the edge has died? Prepare two things in advance: the boundary of a normal drawdown (how deep your edge can legitimately sink) and a breakdown of the edge into parts (win rate at your price, quality of execution). Without this you won’t tell an ordinary drawdown from the real death of the edge, and you’ll either abandon the working thing or hold the dead one.

Skipping any gate means not catching a whole class of errors. And the most dangerous gates aren’t the ones that catch a bad backtest. They’re the ones that catch a good one. Those are what I’ll show.


The one thing proven rigorously: significance is not yet a strategy

Here’s the question almost everyone stumbles on. My edge is statistically significant — a big t, the confidence range above zero. Isn’t that enough to launch?

No. And this can be proven cleanly. I’ll invent, openly, a hypothesis that is really a bet on the regime. And I’ll see what the significance test does with it. The numbers below are my construction, I choose them on purpose. What I’m proving isn’t that some strategy has this edge, but which checks see this setup and which miss it.

Suppose the hypothesis brings plus 5 cents per trade when the market is stormy, and minus 5 cents when it’s calm. Those two numbers I set myself. And suppose in my backtest stormy periods were 60% and calm ones 40%. What will an ordinary significance test show if we look at all the trades together?

Significance test (all trades together) Value
Average edge +1.04 cents per trade
t-statistic 3.53
95% confidence range [+0.46, +1.62]

Look at what came out. The average edge is plus one cent. The range is entirely above zero, t is more than three. By every rule of statistics the edge is significant. A naive analyst here passed the check and would launch a bot — and formally he’d be right.

Now apply the seventh gate — split the same trades by regime:

The same edge, but by regime Value
When the market is stormy +4.96 cents per trade
When the market is calm −5.09 cents per trade

The picture flips. All the plus sits in the stormy regime. In the calm one — a confident minus. That “significant” average plus of one cent was just a blend of a big plus and a big minus, and it came out positive only because there were more stormy periods in the sample. (The numbers repeat what I set — that’s the point: the significance test didn’t see this layering, while the regime split showed it at once.)

And now the main thing — what happens live, if the regime shifts. Suppose stormy periods become not 60% but 40%:

Live edge at 40% stormy regime: −1.00 cent per trade.

The very same hypothesis that passed the significance test with t=3.5 and a range above zero loses money live. Nothing broke. The share of regimes simply shifted, and with it the sign of the edge flipped. A naive analyst wouldn’t have seen this: his significance was real, he accounted for costs. He just didn’t split the trades by regime.

Here’s what’s proven, and what isn’t. Proven rigorously: a significant plus in the backtest can coexist with a loss live — if it’s a bet on the regime. So the significance test is necessary, but alone it’s not enough; a regime check must stand beside it. This is a logical fact, shown on a live counterexample. And here’s what I did not prove: that momentum has an edge of exactly +5 cents, or that stormy periods are exactly 60%. That’s my construction, an instrument of demonstration. The numbers here are like a blueprint explaining a principle, not a measurement from a real market.


The verdict

For such a hypothesis the honest outcome is one: kill it as an all-weather strategy. Its significant edge turned out to be an average across regimes, and live, when the regime shifted, it went negative.

Can anything be saved? Only in a narrow, heavily-caveated form. It’s a conditional bet, traded only in the stormy regime, at a very small size, and only under two conditions: you can determine the regime in real time, and you’re ready for it to change suddenly. This is no longer a “strategy” but a narrow wager on the state of the market. For most, it’s more honest to just say: kill it and move on.

And here’s the most important thing to understand about what you just did. You found out that this hypothesis can’t be traded — by running an analysis, not by losing a deposit. A naive trader would learn exactly the same thing, only expensively: he’d trade the stormy period in the black, believe in himself, raise the size — and give it all back when the regime shifted. This is the whole cycle. The tools don’t find you an edge. They kill the false ones cheaply. Every gate is a way to part with an illusion at the cost of analysis, not at the cost of a trading account.


What would refute this — and the honest limits

Let me start with what the conclusion would collapse on. If a significant average plus could not coexist with a loss live — then the significance test would be enough, and the regime split would be superfluous. But the construction shows this coexistence directly: a significant plus of +1.04 and a live minus of −1.00 from one and the same hypothesis. The counterexample works — so significance alone isn’t enough. The conclusion held.

Now the limits — without them I’d repeat the very error I’m fixing here.

First. The numbers of the regime construction — plus-or-minus 5 cents, shares of 60 and 40 — were set by me, not measured. They show the mechanism: significance is blind to the regime. They say nothing about the properties of a real strategy. What’s actually computed in the run is only the t and the confidence range; everything else is a transparent input that I named out loud.

Second. I deliberately did not build a waterfall of invented subtractions. How much execution or the oracle takes off a specific strategy is a measurement on your data, not a number the author has the right to assign. So the gates are given as questions with a link to their article, not as pretty but empty arithmetic.

Third. Survivors do exist. The same list of checks on a real edge ends differently: the plus holds even after honest execution, and in the regime split too (or you knowingly trade one regime, aware of it). The method is a filter, not a death sentence. It kills a lot, because there’s a lot of false around; but what survives is real.


What to do tomorrow

Run every hypothesis through all nine gates in order, not selectively. Passed them all — including the regime split and sizing by the lower bound — launch it, but with live monitoring and with trust. Died at any gate — you saved your deposit, and that’s a win, not a defeat.

And one rule above all the others: don’t confuse significance with a working strategy. A significant edge is only one gate of nine. What tells a real strategy from a significant bet on the regime is only the regime split — and it must always stand in your list.

You can run a hypothesis through these gates on real data, without risking money, in paper mode — free, before any subscription.


FAQ

My edge is statistically significant — a big t, the range above zero. Is that not enough? Not enough. Significance says the plus isn’t random in this sample. But if the sample was mostly from one regime, the plus can just be a bet on that regime. In our example a significant plus of one cent turns into a minus of one cent when the regime shifts. Significance is needed, but a regime split must stand beside it.

Why didn’t you show a waterfall with numbers for each gate? Because those numbers would be invented by me. How much execution or the oracle takes off a specific strategy is measured on data, not assigned by the author. A made-up waterfall would look convincing and prove nothing — exactly what the whole cycle is against. So the gates are given as questions, and one fact is shown rigorously — the one that can be shown honestly.

My commission is already in the backtest, and the price is the mid. What about the costs gate? You don’t need to subtract separate “costs” — you’ve accounted for them. But you won’t fill at the mid. The real entry is worse: crossing the spread, you pay about half of it; posting a limit, you lose on adverse selection and missed trades. Measure your real entry price against the mid and subtract that difference. That’s the execution gate for your case.

If the method kills everything, what’s the point? It kills not everything but the false — and does it cheaply. The point is where exactly you part with the illusion: in analysis over a day, or on the deposit over months. The rare hypothesis that passes every gate you launch with trust — precisely because you honestly tried to kill it and couldn’t.

Why is the regime the decisive gate? Because it catches what passes all the others, significance included. A strategy can be significant, cheap to execute, correctly labeled by the oracle — and still turn out to be a bet on the regime. Skipping this gate costs the most, so in the list it’s mandatory.


Disclaimer

This material is educational and is not financial advice. Past results do not predict future results. Trading on prediction markets carries real risk, up to and including the total loss of capital. Access to Polymarket is restricted or prohibited in some jurisdictions — verify legality where you live before trading.