Koryu · 黄龍
Koryu
← Research library
Research · 研究 · 19 · Integrity

Half the audit came back clean. We published the other half too.

17 Jul 20266 min readMethodologyKoryu Research

The audit described below was run against our own equity sibling's flagship backtest, and its unflattering findings were published in full at the time. Numbers are quoted from that audit. Nothing here is investment advice. Decisions are yours.

Before this site published a single number, our house pointed the tools institutions use for catching self-deception, probability-of-overfitting tests, deflated Sharpe ratios, survivorship bounds, at its own flagship equity backtest. Half the audit came back clean. The other half found a flaw large enough to bound the headline result down by as much as ninety percent in the worst scenario. We published both halves, and the damage report became the constitution this crypto site was built under.

Why audit yourself at all

Every quantitative shop faces the same temptation: your backtest is your marketing, and auditing it can only make it smaller. The reasons to do it anyway are practical rather than noble. Reality arrives eventually, and a live track record converges on the truth whatever the backtest said, with the gap between promise and delivery paid in reputation at the worst possible moment. An audit you run yourself, on your own terms, published on your own site, is worth more than the same flaw found later by a critic. And the decisive one: you cannot demand verification from an industry, the way our seven checks do, while exempting your own numbers from the knife.

The setup: two questions, two different tools

The subject was a five-year, four-engine systematic equity strategy whose locked backtest showed a 75x multiple, a Sharpe ratio near 2.4, and a maximum drawdown under 16%. Numbers good enough to deserve suspicion. Every strong backtest has to answer two questions, and they are not the same question. Is the result an artifact of selection, the winner plucked from so many tried variants that something was bound to look this good by chance? Or an artifact of data, a universe that quietly excludes the failures the strategy would have traded? Different tools answer each. They came back with opposite verdicts.

The selection audit came back clean

Selection has proper statistics behind it. The probability of backtest overfitting, computed by splitting the sample into in-sample and out-of-sample halves thousands of ways and checking whether the in-sample winner keeps winning, came back at 2.6% against a danger threshold of fifty. The variant that won in-sample stayed near the top out-of-sample in essentially every split. The deflated Sharpe ratio, which asks whether the headline survives once you account for how many variants were ever tried, held up even assuming three thousand historical trials, several times more than the program's documented history.

One detail from that pass is worth keeping. The single highest-Sharpe variant in the entire research archive was one the house had rejected before the audit ever ran, for being a lone spike rather than a stable plateau. The discipline of preferring robust parameter regions over lucky points turns out to be measurable, and it measured well.

The data audit did not

Then the second question, and it took one crude query. The strategy's universe held 6,134 symbols with five years of history, assembled the industry-standard way from a broker's current listings. How many of those symbols stopped trading during the window?

Zero. Not one death in five years, in an asset class where several percent of small companies delist annually. The universe was survivor-only, the backtest had never been forced to trade through a single corporate failure, and every result it produced inherited that flattery. Why momentum strategies suffer this worst is the graveyard problem.

Bounding the damage

A flaw you cannot remove, you bound. The audit modeled the missing graveyard adversarially: assume dead names would have taken candidate slots in proportion to an annual delisting rate, with a propensity boost because dying momentum names over-represent on breakout boards, and assume each such slot earned the year's average losing return while displacing a real trade. At a 4% annual delisting rate the 75x multiple bounds down to about 30x. At 8%, to 15x. At a pessimistic 12%, to 9x.

That bound is deliberately harsh. It ignores that short holding periods blunt the bias, and that dead names still had to clear quality gates on the way in. But its direction is not negotiable, so the summary we published was the blunt one: the true multiple lives somewhere well below the headline, and the exact floor is unknowable without better data.

What the audit changed

Findings without consequences are theater, so here is what the damage report actually bought. On the equity side the published marketing already leaned on risk-adjusted quality rather than the headline multiple, and the audit went out alongside the track record it criticized. On this side the consequences are structural. The crypto research universe has to be point-in-time and graveyard-inclusive from day one, and a data vendor that cannot produce a dead coin's final week is disqualified. Every backtest we publish states its universe construction and its trial count. The gates a strategy must pass, including the two tests the equity work passed and the survivorship standard it failed, are pre-registered in measurement vs advice. Our founding crypto experiment inherited the humility directly: the transplant's results were labelled upper bounds on a survivor-biased universe in the same breath as their publication.

What this means for you

Two portable lessons. The tools exist and are not house secrets: probability of backtest overfitting, deflated Sharpe ratios, and survivorship bounds are published methods, and any shop with real results can run them in a day. So when you meet a spectacular backtest, ask whether its owner has run them, and whether they will show you the report. That question sorts this industry faster than any returns figure.

The second lesson is that the two failure modes are independent. A strategy can be honestly selected on rotten data, or luckily selected on clean data, and only testing both questions closes the account. Our own report card, selection clean and data flawed and bounded, is public because that is exactly what we would demand of anyone else. The knife cuts both ways here, permanently. Decisions are yours.

Related reading
MethodologyA 365-day market: how 24/7 changes annualized numbers5 min readMethodologyWhat is overfitting? The crypto edition5 min readMethodologyWalk-forward validation: tuning on one era, proving on another5 min read
Frequently asked

What did the self-audit test?

Two independent failure modes: selection (is the result the lucky best of many tried variants?) via probability-of-backtest-overfitting and deflated Sharpe tests, and data (does the universe exclude failures?) via a census of dead symbols and an adversarial survivorship bound.

What were the audit's results?

Selection came back clean: PBO of 2.6% and a Sharpe surviving an assumed three thousand historical trials. Data did not: zero of 6,134 symbols died in five years of history, a survivor-only universe, and the headline 75x multiple bounds down to roughly 30x, 15x, or 9x under 4%, 8%, or 12% annual delisting assumptions.

Why publish an audit that shrinks your own headline?

Because reality arrives anyway: live results converge on the truth regardless of the backtest, and a flaw you publish yourself, bounded and dated, costs less than the same flaw discovered by a critic. It also earns the right to demand verification from everyone else.

What is the probability of backtest overfitting?

A test that splits history into in-sample and out-of-sample halves thousands of combinatorial ways, picks each split's in-sample winner, and measures how often that winner underperforms out-of-sample. High values mean the strategy selection was fitting noise; low values mean in-sample winners kept winning.

How did the audit change this site?

Structurally: graveyard-inclusive point-in-time universes are mandatory for research, vendors that cannot produce dead assets' final data are disqualified, trial counts are disclosed, and the same statistical gates the equity work passed are pre-registered pass-fail conditions for any strategy shipped here.