Verdicts
What we tested, and what we threw out
Almost everyone publishes what worked. Here is the other half of the job: the rules we measured, and most often dropped. Protocol, numbers and date every time.
- 19studies published
- 17with no rule kept
- 4prove us wrong
Why we publish this
The most common complaint among retail traders is that backtests lie. It is a fair one: you can make a price series say almost anything if you look for long enough. The only honest answer we have found is to also show the research that led nowhere, with the protocol that lets anyone contradict us.
4 of these verdicts count against us: they describe rules we had put into production and that were doing harm. They stay online. Without them, the other 15 would be worth nothing.
One editorial rule: rejected research is published in full. Research that ended in a rule running in production is published as a verdict, without its parameters.
- 01Rejected31 July 2026
When several stocks qualify on the same day, which one should you take?
Protocol
Portfolio of 5 shared slots, 2017–2026, production exits, accounting in R (a constant risk per trade, so position size is neutralised).
Result
The arm with no information at all, an order drawn at random, returns +0.229 R per trade over 522 trades. Seven ranking criteria were put up against it: score, high momentum, low momentum, proximity to the 52-week high, leading sector, low volatility, stop width. None beats it significantly. Two do clearly worse. And our own ranking score comes in slightly BELOW the random draw. Same conclusion at 10 slots, where high momentum and low momentum both beat chance: noise, then, not a signal.
Link to this verdictWhat it changes
No candidate ranking has ever been deployed. The reason is structural: the portfolio is occupied 87% of the time, so a ranking only has the remaining 13% of slots to express itself on. That is a mechanical ceiling, independent of how good the criterion is.
- 02Framing adopted12 August 2026
How many positions should you hold at the same time?
Protocol
3, 5, 8, 12 then 20 slots, real universe of 2,338 stocks, production stop and exits, random ranking so that only capacity is measured.
Result
The ratio of cumulative return to maximum drawdown peaks at 8 slots (3.84), with 5 slots close behind (3.77), then degrades clearly: 3.01 at 12 slots and 2.76 at 20. Going from 5 to 8 slots adds 225 R of cumulative return, or +58%, without degrading the return-to-drawdown ratio. It is the only free gain in the whole grid. Beyond 8, you make the bet bigger without making it better. Occupancy stays at 90% at every level, and between 380,000 and 1.6 million candidates are turned away for lack of a slot.
Link to this verdictWhat it changes
What limits this kind of portfolio is neither the entry nor the exit: it is the number of slots. Stated reservation: the bench ignores the correlation between simultaneous positions. 20 slots in the same sector are not 20 independent bets. 8 is a cautious ceiling, not a target.
- 03Rejected13 August 2026
A red candle touching a moving average, then two green ones, with RSI, volume and MACD: does that let you separate candidates?
Protocol
The pattern turned into a ranking score, not a filter. Union of six signal families, shared slots, production stop, verdict delivered out of sample on closes after 2024, and one inverted control per component.
Result
The “moving-average touch” pattern returns −0.079 R over the training period and falls below neutral ranking in test, with only 2 winning years out of 9. But the decisive result is elsewhere: EVERY component, RSI, volume, MACD and their combination, has its inverted control beating neutral ranking in test as well. The best arm overall is volume reversed (+0.615 R in test), which is worth nothing in training.
Link to this verdictWhat it changes
When a rule and its exact opposite both “win” over the same period, the bench is not measuring an edge: it is measuring the variance of the draw. Third confirmation that nothing beats neutral ranking when slots are shared. Closed, not to be reopened without new data.
- 04Rejected8 August 2026
Buying when price comes down to its moving average and turns back up: does it work?
Protocol
Bare bounces off the 20, 50 and 200-session moving averages, full universe, shared-slot portfolio, production exits.
Result
+0.006 R on the 20-session moving average, −0.063 R on the 50, −0.036 R on the 200. All three sit below the families already in place. An earlier test on a reduced sample of 60 stocks gave the opposite and flattered bounces outright: the full universe reversed the verdict.
Link to this verdictWhat it changes
It is the most widely taught setup in swing trading, and it does not survive measurement. The explanation fits in one sentence: the trader's eye only sees the bounces that worked. The ones that failed belong to companies that fell, sometimes disappeared, and are no longer on the chart being looked at.
- 05Refuted8 August 2026
Is a breakout on heavy volume more likely to work?
Protocol
89,815 breakouts analysed, sorted into volume quintiles, measuring the share that then produces a move of at least 6 ATR.
Result
36%, 36%, 37%, 37%, 36%. Volume separates nothing. Two other indicators, ATR and distance from the 50-session moving average, looked strongly predictive, from 5% to 56%: that was a unit artefact, a threshold expressed as a percentage being cleared mechanically by the jumpiest stocks. Only the 14-period RSI and the ADX survive, with +7 to +9 points, and in the opposite direction to intuition, a HIGH RSI giving the better outcome. The 2-period RSI, the MACD and the length of the base predict nothing.
Link to this verdictWhat it changes
A volume filter kept in one of our families was measured again on its own: its inverted control matches it, which is to say that requiring low volume does just as well as requiring high volume. It is filed as decorative and documented as the first candidate for removal.
- 06Premise reversed31 July 2026
The flag after a 30% run-up: is it really on small caps that it pays?
Protocol
200 reference stocks plus a dedicated set of 236 small caps, production exits, the same setup on both sides.
Result
The opposite of what is taught. On mega caps: 279 trades, +0.127 R. On small caps: 413 trades, −0.039 R. The +90% variant on small caps shows +0.496 R, but over 20 trades. Nothing conclusive, and this is exactly the ground where the absence of companies that have disappeared distorts the result the most. Volume contraction, the criterion the literature insists on, adds nothing: it halves the sample with expectancy unchanged.
Link to this verdictWhat it changes
Nothing deployed. The case is kept as a reminder: a setup can be described correctly and assigned to the wrong ground.
- 07Trap avoided31 July 2026
Should you favour the setups whose stop is closest?
Protocol
First as a ranking, then, and this is where it is decided, as a filter, the only form that can actually be deployed.
Result
As a ranking, sorting by tightest stop costs 490 R at 5 slots and 693 R at 10. That is by far the largest effect of the whole campaign, and the temptation was to turn it into a rule. Tested as a filter: every floor loses, consistently, and the inverted control loses as well. So the effect was not “setups with a tight stop are bad” but the over-rotation that this sorting causes: 1,652 trades instead of 522, and an average holding time of 6 days instead of 20.
Link to this verdictWhat it changes
Method rule adopted since: a spectacular effect seen in a ranking does not carry over to a filter. You always have to retest in the form you would actually deploy.
- 08Premise false31 July 2026
Exiting when price closes below its short moving average, rather than on a trailing stop: does it give winners more room?
Protocol
Five entry families, entries and stops strictly identical, only the trailing method changes. 200 stocks.
Result
The premise is false: an exit on a short moving average is not more permissive, it is tighter. The signature is clear and identical across all five families: the win rate goes up, the average gain goes down, and the share of trades above 3 R falls everywhere. A price drops back under its short average all the time; a trailing stop placed several ATR below the high is far more tolerant. The trailing stop wins or ties on all five families.
Link to this verdictWhat it changes
A reading trap avoided along the way: the moving-average variants showed a spectacular maximum gain, up to 50 R against 6 R. On checking, it was a single trade, in March 2021. We now look at the share of trades above 3 R, never at the maximum.
- 09We were wrong8 August 2026
Did our own loss ceiling, expressed as a percentage, really protect anything?
Protocol
Six families replayed over nine years, with and then without the ceiling that was in production at the time.
Result
Under our ceiling, three families out of six LOST money over nine years. With no ceiling at all, five families out of six do better: +269 R against +107 R, and with less maximum drawdown, not more. A barrier we presented as a protection was in fact cutting the trades that paid. A per-family variant, tighter on some and wider on others, was tested straight after: it came last, and its exactly inverted version beat it.
Link to this verdictWhat it changes
The ceiling was replaced by a bound expressed in volatility rather than as a percentage of price. The replacement did not last the day: it is told in full further down, along with what made it fail. Second admission, more awkward than the first: the test bench was bounding its own stops, so the improvement announced from the backtest could not carry over to production as it stood. The percentage barrier is therefore still the one running, even though this very verdict shows that it cuts trades that paid. The gap is documented and the work remains open.
- 10Rejected14 August 2026
Crossing above the 200-session moving average, the most widely followed reversal signal: does it pay?
Protocol
Real universe of 2,338 stocks, 8 shared slots, production exits. Ten simple filters tested, then the combinations of the best four. Finally a cross-validation on a second universe of mid caps.
Result
The raw event does not pay: −0.011 R, profit factor 0.98, 4 positive years out of 9. The best formula found, two confirmation closes plus sustained volume, gets back to +0.097 R, but stays below the families already in place. The cross-validation is final: that formula collapses on the second universe, and the volume filter CHANGES SIGN from one universe to the other. The symmetric test gives the same result the other way round, the best formula of the second universe falling back on the first.
Link to this verdictWhat it changes
Definitively rejected on two universes. And a method rule: when no filter behaves the same way on two samples, the ranking of the filters is sampling variance, not a result.
- 11Rejected14 August 2026
Sticking to large, low-volatility names, the approach known as “safe and steady”: does it improve a bounce signal?
Protocol
The bounce signal is strictly identical in every arm, only the eligibility of the stocks changes. Real universe, 5 slots, production exits.
Result
Full universe: +0.057 R over 583 trades. Mega caps only: +0.114 R, the best arm, with a smaller maximum drawdown. Mega caps AND low volatility: back down to +0.066 R, with a maximum drawdown TWICE as bad. Low volatility alone, all sizes: +0.012 R. The edge has gone.
Link to this verdictWhat it changes
It is SIZE that pays, not caution. The low-volatility filter removes exactly the stocks capable of bouncing hard, which is to say the ones the signal exists for. Lead not deployed: a single configuration was tested, and it needs revalidating before any use.
- 12Rejected14 August 2026
Do the companies that have raised their dividend for twenty-five years beat the market?
Protocol
634 US large caps, 2018–2026, equal-weighted portfolios of 20 stocks, rebalanced every quarter, coupons credited on the ex-dividend date. Four metrics × three seniority filters × three overlays, so 36 arms, plus the ablations.
Result
The benchmark index returns 14.4% a year with dividends reinvested; the equal-weighted universe, 15.6%. The true “aristocrats”, filtered on twenty-five years of consecutive increases, return 9 to 13.5% a year, below both references, their only gain being a slightly smaller maximum drawdown. The one arm that clearly outperforms is raw yield with no filter, at 26.6% a year: but its average yield of 15.2% points to a basket of companies in trouble, with a 56% maximum drawdown. The most instructive ablation: removing explicit reinvestment changes almost nothing, the “reinvestment advantage” amounting to staying invested.
Link to this verdictWhat it changes
Rejected as a product direction. The trend overlays tested on top improve nothing and degrade almost everywhere. Stated limits: the 2018–2026 window is a rising market, the universe contains survivors only, and tax is not modelled.
- 13Effect of opposite sign13 August 2026
Should you avoid entering just before an earnings release?
Protocol
1,105 US large caps and 57,989 release dates, portfolio of 5 shared slots, production exits. Coverage of roughly 70% of trades: the collection of dates was cut short by the provider's quota, and validation on mid caps is still to be done.
Result
The most universal piece of advice in the business, never hold through earnings, collapses as soon as you separate by signal family. A pullback bought less than three sessions before a release returns +0.333 R; at four or five sessions, +1.116 R with 83% winning trades. It is the best pocket in the whole family. The same proximity costs the breakout (−0.104 R) and wrecks a volume signal (−0.443 R). The direction of the effect depends on what the signal is buying: the families that buy strength ahead of a breakout suffer from the price gap at the open, the ones that buy the pullback benefit from the anticipation move that precedes the release.
Link to this verdictWhat it changes
The result that decides, though, is not that one. Turned into a filter and measured at portfolio level, almost any blocking degrades, including on the families whose diagnosis shows a negative pocket. The reason is mechanical: when the filter frees a slot, the replacement trade is worse than the one just refused. Only the breakout benefits from a widened exclusion window. Nothing was changed in production, where the existing window stays a deliberate compromise: justified for one family, costly for another, neutral elsewhere. A window differentiated by family is the lead, and it is waiting for validation on a second universe.
- 14Deployed, then withdrawn8 August 2026
Since our percentage loss ceiling was cutting the trades that paid, does a bound expressed in volatility do better?
Protocol
Immediate follow-up to the previous verdict. Six families over nine years, a grid of bounds expressed in volatility, measured first trade by trade, then on a shared-slot portfolio. The bound was deployed the same day. The bench was then replayed out of sample on closes after 2022, with random controls.
Result
Three things gave way, in this order. First the bench itself: its stop construction was already bounding its own stops between 1 and 3.5 ATR, so that the “wide bound” arm was measuring something close to “no bound at all”: 3,040 trades against 3,022 with no ceiling whatsoever. The production stop construction, for its part, has no upper bound: on the breakout, the median stop there is 4.7 ATR against 3.4 on the bench. The improvement announced from the backtest, a “wait for a pullback” warning that was supposed to fall from 57% to 9% on the breakout, therefore could not carry over: measured in production, it went from 77% to 72%. Then the family-by-family reading: the optimal value jumps from one family to the next and the curves are not monotonic, one family showing +37 R at a setting flanked by −46 R and −48 R at the neighbouring settings. The overall gain of +286 R rested on that single cell. Finally significance: no arm clears two standard deviations, and on the family that had prompted the whole exercise, the best arm of the campaign is the INVERTED CONTROL, keeping only the widest stops, at +0.164 R and a profit factor of 1.38. The out-of-sample rerun finished the subject off: the worst learned setting matches or beats the best, the gap between best and worst changes sign on four families out of six, and the hybrid version loses both against the percentage ceiling and against its own random controls. Over that period, the percentage ceiling returns 85 R where the volatility bound returns 2.
Link to this verdictWhat it changes
The percentage ceiling was restored the same day, and the per-family bound is filed as definitively abandoned. That makes three successive admissions on the same rule: the original barrier was cutting what paid, its replacement was measured on a bench that was already bounding its own stops, and that replacement died out of sample. The method rule that comes out of it is the most expensive of all those published here: before testing a constraint, you have to check that the bench is not already applying it itself. Otherwise you compare a rule with itself and call it a result. And the benefit of the doubt goes to the rule in place, not to the one you are hoping for.
- 15Refuted by its own controls8 August 2026
Should you only buy stocks whose price structure is bullish, higher highs and higher lows?
Protocol
The sequence of highs and lows, read on confirmed pivots and with no look-ahead, turned into an eligibility condition. Union of the signal families, 5 shared slots, production stop construction, verdict delivered out of sample on closes after 2022. And two control arms that require exactly the opposite of the rule under test.
Result
With no filter at all, the portfolio returns −1 R over the test period. Accepting only bullish structures lifts it to +17 R, that is 0.7 standard deviations, which is to say nothing. Excluding only bearish structures gives +8 R. And the two inverted controls beat the rule: everything EXCEPT bullish structures returns +33 R, and accepting ONLY bearish structures, the thesis turned upside down point for point, returns +46 R, the best arm of the campaign. Two side observations show that the bench was not measuring what we believed: the “undetermined structure” value comes out on no signal in the universe, so the tolerance clause written for it never had any effect and two arms thought to be distinct were in fact the same; and the family that buys pullbacks produces 23% of its signals in bearish structure.
Link to this verdictWhat it changes
No structure filter was deployed. The structure reading is still shown as information in the analysis, never as a veto: that is the form in which it costs nothing. One curiosity was left open, and it is still not deployed: “bearish structure only” is the best arm over BOTH periods, +44 R in training and +46 R in test, three quarters of it carried by the pullback family. That is consistent with another result from the same day, where requiring a price above a moving average degraded a mean-reversion signal: those signals need the deep pullback. But it is a CONTROL arm that wins, over 240 trades, at 1.5 standard deviations and with no monotonic ordering. A hypothesis that was not stated before it was seen to win is not a result: it has to be investigated separately, on the window that did not produce it.
- 16Rejected4 September 2026
An up day whose volume exceeds every recent down day: is it the early buy signal it is said to be?
Protocol
Real universe of 2,338 stocks, closing years 2018 to 2026, 5 shared slots, production exits, structural stop, neutral ranking, expectancy in capped R. Three definitions of the pattern, from the crudest to the most canonical, and a fourth arm that applies its volume signature alone to our breakout family. From 381 to 495 trades per arm. The campaign dates from August 2026 and was replayed on 4 September with the production exits as they run today, without a single conclusion moving.
Result
In its raw form, an up day whose volume exceeds the largest down-day volume of the previous ten sessions on a trending stock, the pattern returns −0.081 R per trade over 476 trades, with a profit factor of 0.86 and 2 winning years out of 9. Its maximum drawdown, 44 R, is twice that of our two reference families. The tidied-up pattern turns positive again: adding contact with a short moving average gives +0.072 R, and adding an extension bound on top of that gives +0.088 R over 495 trades, with 7 winning years out of 9. That tidying still leaves it below the two families already in place, +0.146 R for the pullback and +0.132 R for the breakout. The arm that settles it is the fourth. Grafting the volume signature onto our breakout removes almost no trades, 381 instead of 391, and drops expectancy from +0.132 R to +0.080 R, with 4 winning years out of 9 instead of 7. The trade count barely moves but its composition changes: a refused slot is immediately taken by another candidate, and the replacement is worse.
Link to this verdictWhat it changes
Nothing was deployed. The result that matters is not the rejection, it is where the improvement comes from. What takes the pattern from negative to positive is not its volume signature, which is precisely what the method claims as its discovery: it is the two conditions added to make it presentable, a pullback to a short moving average and an extension bound, and those two conditions are already the ones our production families use. The pattern only survives by becoming what we already do, less well. The volume signature, for its part, can be measured on its own: grafted onto an entry that works, it degrades it. That is the only thing this pattern brings of its own, and it costs.
- 17Rejected in cross-validation9 August 2026
Should you only buy on the daily chart if the weekly is already bullish?
Protocol
Real universe of 2,338 stocks, closing years 2018 to 2026, 5 shared slots, production exits, structural stop, neutral ranking, expectancy in capped R. Two levels of filter: the first requires a weekly close above its ten-week moving average, the second adds the alignment of two weekly averages. Closed weeks only, with no look-ahead. Both levels applied to each of our signal families, then the one surviving lead replayed on two other capitalisation bands.
Result
On the families already in place, the filter costs. The clearest case is the one that buys volume surges: expectancy falls from +0.165 R to +0.055 R per trade once the filter is applied. Pullbacks and breakouts degrade as well. The second, more demanding level kills everything it touches, without exception. One combination stood out: applied to the family that buys moving-average crossovers, the filter took expectancy from +0.081 R to +0.148 R over 483 trades, with a profit factor of 1.29 against 1.15, a maximum drawdown brought down from 49 R to 32 R and 6 winning years out of 9. The number held for half a day. Replayed on mid caps, the same filter gives +0.002 R against +0.084 R without it, and beats its reference in only 2 years out of 9. On small caps, +0.002 R against +0.030 R. Nothing was left. A second lead, on the mean-reversion family, was already too thin to investigate: +0.194 R against +0.156 R, but 5 winning years out of 9, which is a coin toss.
Link to this verdictWhat it changes
No weekly filter was deployed, in either of its two forms. What this campaign really produced is a method rule, and it was expensive to learn: a gain measured on a single universe is not a result, it is a hypothesis. Here the gap looked good, twice the expectancy and a third less maximum drawdown, over nearly 500 trades. It disappeared entirely as soon as the capitalisation band was changed, which means it was not measuring the idea under test but a peculiarity of the segment it had been read on. Since then, every lead has to win on a set that did not produce it before going any further. Two other ideas from the same campaign went through the same test on the same day: only one passed it, and it is the only one that was kept.
- 18Intuition refuted25 September 2026
It is said everywhere that mean reversion works better on ETFs than on stocks. Is that true with swing exits?
Protocol
76 liquid ETFs: industries, countries, sectors, factors, broad indices, commodities and bonds. Daily prices since 2017, closing years 2018 to 2026, data stopped on 13 August 2026. Portfolio of 3 shared slots, sized for a universe of this dimension, production exits, structural stop, neutral ranking, expectancy in capped R. Each signal family played on its own, from 60 to 315 trades per arm. First measured on 13 August 2026, replayed on 25 September with the production exits as they run today.
Result
The short-term mean-reversion signal, a two-day RSI in oversold territory, returns −0.019 R per trade over 315 trades, with a profit factor of 0.97 and 3 winning years out of 9. It is the family that trades the most on this universe, and it gains nothing. In August, with the exits of the time, it returned exactly 0.000 R: between the two measurements, the verdict stayed flat. What does hold on ETFs are the families that buy strength. The volume surge returns +0.199 R over 195 trades, with a profit factor of 1.36 and 7 winning years out of 9. The yearly high breakout returns +0.106 R over 150 trades, the moving-average crossover +0.100 R over 244 trades, each with 6 winning years out of 9. The trend pullback, weak in August at +0.022 R, comes back to +0.082 R over 291 trades with the current exits: positive, but below the three strength families. Proximity to the yearly high does not pass (−0.013 R over 212 trades). The momentum family, restored to production in early September, stays flat: +0.010 R over 212 trades, 3 winning years out of 9. A recent family, too rare here to conclude on with 60 trades, is negative as well.
Link to this verdictWhat it changes
The intuition is not absurd: it comes from measurements made on indices with fast exits, at the first bounce. Our exits let gains run behind a trailing stop. They are built for trends, and the most likely reading is that on ETFs they hand back to the market what the bounce had given. The verdict therefore bears on the combination, not on the idea alone: mean reversion on ETFs with an exit at the bounce was not tested here. On the product side, this measurement justified widening the scan's ETF universe from 30 to 78 lines, covering distinct exposures rather than copies of the same index, because three families have a real edge there. No filter by market was put in place: the ETF scan runs every family, including the three this verdict finds weak. That is the point that stays open.
- 19Figure abandoned6 October 2026
How much did your stop save you from losing?
Protocol
A paired control on our real positions, not on a test universe. The 80 positions that ended on a stop, replayed with no stop and no target, each one its own control. The horizon of the no-stop version was fixed BEFORE looking at the result: the median duration of our winning positions, that is 23 sessions, measured over 48 winners. One position with a partial exit was excluded in advance, its result not being comparable, and 7 had no usable series: 72 positions measured. The null draws the SIGN of each position's difference at random, 5,000 times, which is the correct null when each position is its own control.
Result
With the stop, those 72 positions return −0.967 R on average, that is −81 € per position. With no stop, held 23 sessions, they return −1.206 R, that is −100 €. The average difference, 0.239 R or 19 € per stopped position, is the figure we were about to publish. It survives no check. Removing the three positions where the stock fell furthest brings it down to 0.021 R, at percentile 44 of the null: nothing is left. The median already says the opposite of the mean: on the typical position, truncated series aside, the stop COST 0.043 R. And trade by trade it is a coin toss, the stop having protected 37 times out of 72 and cost 35 times. No cut of the sample clears two standard deviations in the right direction. The family-by-family reading, on the other hand, is clear: the family that buys volume surges pays 1.320 R for the stop (6 positions), the one that buys trend pullbacks does not pay for it at all, 0.016 R over 17 positions, and mean reversion pays in reverse, the stop costing it 0.579 R (4 positions).
Link to this verdictWhat it changes
The figure is dropped. It was meant to serve as an argument, and it rested on three trades out of 72: publishing it would have been exactly what this page holds against everyone else. Nothing changed in production, where the stop stays in place on every family. What the measurement does say, and it is more useful than the figure it fails to give, is that the stop is not average insurance but a tail cut: it pays on the families that buy strength and turns against the ones that buy the pullback, which matches the verdict on price structure. A stop differentiated by family is the lead, and it is waiting to be investigated on a window that did not produce it. Stated reservation: 72 stopped positions are not enough to settle a difference of this size, and the measurement will be redone once the sample has doubled.
One verdict a week
The corpus runs to about thirty studies. The first 19 are here; the rest appear at a rate of one a week, including the weeks where the result is that there is nothing to keep. And what the strategies we kept have returned can be read, losses included, on the public track record.