Backtesting Discipline: The Easiest Number to Fake
The determining question is when the strategy logic was fixed relative to the performance period being shown.
Macro Context: Overfitting Is Curve-Fitting, Not Genuine Predictive Signal
A strategy backtest iterated against the same historical dataset dozens of times will eventually appear excellent, not because the underlying logic is sound, but because sufficient parameters have been tuned to fit the noise in that specific data. The resulting performance statistics reflect the fitting process as much as the strategy itself — a well-known failure mode in quantitative research, and one that produces an impressive-looking Sharpe ratio regardless of whether the logic generalizes.
The Structural Challenge: A Number That Looks Identical Whether Earned or Fitted
The resulting performance statistics reflect the fitting process as much as the strategy itself, and a backtest report alone provides no way to distinguish between the two — an institution evaluating the headline number cannot tell, from the number alone, whether it reflects genuine predictive logic or successful curve-fitting.
The Methodology: Locking Logic Before Out-of-Sample Testing
Genuine out-of-sample discipline means locking the strategy logic before testing it against data the model has never encountered, and discarding a strategy that performs well in-sample but fails out-of-sample — the discipline is in resisting the temptation to re-tune after a disappointing out-of-sample result, which reintroduces the same fitting problem in a different form.
can the team demonstrate precisely when the strategy logic was locked, and show performance on data that postdates that lock. A provider unable or unwilling to demonstrate this distinction is asking an institution to trust a number that may describe only the fitting process.
The Deterministic Outcome
An institution that evaluates the procedural discipline behind a backtest — not merely its headline statistic — can distinguish a strategy with genuine out-of-sample validation from one whose performance number reflects successful curve-fitting against historical noise.
Strategic Takeaways
- Lock strategy logic before testing against out-of-sample data, resisting the impulse to re-tune after a disappointing result
- Demand a clear record of when strategy logic was finalized relative to the performance period being shown
- Weight the discipline of the testing process at least as heavily as the resulting performance statistic
Discuss this with our team.
Tell us what you are working through, and we will route you to the right partner.
Algorithmic Trading Compliance Amid Shifting Markets
A compliance framework calibrated to the prior market cycle governs a market structure that has already moved on.
Risk Budgeting Across a Multi-Strategy Systematic Portfolio
Equal-capital allocation and equal-risk allocation are structurally distinct exercises, and that distinction is where most portfolios quietly go wrong.
The Case for Architectural Diversification Across Regimes
Portfolios diversified by asset class alone stay exposed to one dependency: the market condition every included strategy quietly requires to perform.




