We built 19 Borsa Istanbul research systems but still cannot prove a real trading edge — what are we missing? #867
Replies: 1 comment
|
The number you already have and are not using is 19. I would turn it into a bar before running anything prospective, because the best of 19 architectures tested on one history looks good even when none of them has skill, and the more you test, the better the luckiest one looks. So the question is not whether the current system beats zero, but whether it beats what the luckiest of 19 no-skill trials would show on the same data. Bailey and López de Prado give that bar in closed form in The Deflated Sharpe Ratio (2014). The expected maximum Sharpe of N trials with no skill is the mean Sharpe across the trials plus their standard deviation times a term that grows with N, roughly the inverse normal of 1 − 1/N. The deflated Sharpe is then the probabilistic Sharpe of your chosen architecture measured against that bar instead of against zero, corrected for the length of the record and for the skew and kurtosis of its returns. The output is a probability; anything under 0.95 means the result is consistent with luck. It needs only five numbers you already have: the count of trials, the variance of the Sharpe across them, the length of the record in observations, and the skew and kurtosis of the chosen one. The count should include every parameter variant inside each architecture, not just the 19 headline designs, because the trial count is what sets the bar and 19 is a lower bound. The second test uses the daily after-cost PnL of every variant as one matrix, one column per variant. It is the probability of backtest overfitting (PBO) from the same authors with Borwein and Zhu (2015): split the rows into 16 equal time blocks, take every one of the 12,870 ways to call half of them in-sample, pick the in-sample winner by your own after-cost metric, and record whether it ranks below the median out-of-sample. The share of splits where it does is the probability that your selection process itself overfits; a value above 0.5 says the way you choose among architectures loses information rather than adds it. It needs no model and no distributional assumption, and it works directly on the history you already have. Both tests are cheap, and they answer a question the prospective shadow portfolio cannot: whether the selection you are about to freeze is worth freezing at all. They also speak to your second possibility, because you can build the matrix at each stage of the pipeline (raw candidates, then ranked, then allocated, then after exits) and see at which stage the in-sample winner stops beating the median. One caution from the DSR paper: in mean-reverting series overfitting does not fade, it reverses, because a strategy fitted to the most extreme past pattern is betting on exactly what is about to be undone, and thin emerging-market equities revert more than liquid ones. If your 19 systems share a ranking on recent strength, that is the first thing I would look at. And whatever you decide, write down the statistic, the threshold and the stopping rule before the first run. A test whose pass mark is chosen after seeing the result is a twentieth architecture. |
Uh oh!
There was an error while loading. Please reload this page.
We have been developing an independent decision-support system for Borsa Istanbul for approximately 22 months.
The system uses only delayed market information that was available at the exact decision time. It does not use future data, real-time private feeds, other stock markets, or automated order execution. It does not manage client money. Its purpose is to examine the market, identify candidates, compare their relative strength and risk, and produce a manual daily decision.
How the system works
Over time, we built and tested 19 different research architectures. These were not simple parameter changes; they examined different combinations of price behaviour, intraday movement, market and sector conditions, historical similarities, candidate ranking, downside protection, capital allocation, holding periods, and exit decisions.
The current AVCI architecture contains:
Historical daily and real one-minute BIST price and trading data have also been examined. Information created after the decision time is not supposed to enter the decision process.
The unresolved problem
Despite thousands of hypotheses, simulations, tests, and several complete architecture changes, we have not been able to prove a repeatable and executable after-cost edge.
Promising historical results often weaken or disappear when:
A further problem is that much of the historical period has already been examined during research. After thousands of experiments, even an apparently excellent historical result may simply be a false discovery caused by overfitting and repeated testing.
There are also unresolved differences between a paper result and actual capital growth. A correct candidate does not automatically mean a profitable trade. Entry price, liquidity, slippage, tradable quantity, transaction costs, corporate actions, holding time, and exit timing can all change the result.
Our historical records also do not provide a complete real-money ledger containing every decision, executed quantity, entry, exit, cost, and daily capital change. For this reason, some old results cannot be reconstructed as genuine executable performance.
At present, we cannot confidently distinguish between three possibilities:
What we are looking for
We are not looking for:
We are looking for experienced researchers, graduate students, quantitative developers, market-microstructure specialists, or independent practitioners who are willing to examine this problem carefully and patiently.
The central question is:
We are especially interested in people with experience in:
A negative conclusion is acceptable. The objective is not to make an unsuccessful system appear successful. The objective is to determine, with defensible evidence, whether a real edge exists, where it disappears, or why it cannot be extracted under the current constraints.
This is not a quick question that can be solved with one indicator or a few comments. We are looking for serious contributors who are willing to understand the architecture and help define a small number of decisive experiments.
A concise anonymized technical summary can first be shared with serious contributors. Proprietary selection rules, the complete source code, and raw data that we do not have the right to redistribute will not be posted publicly.
Our core question is:
Why, although this large research infrastructure appears able to identify strong candidates, can we not convert that ability into repeatable, executable, after-cost capital growth across different market periods?
All reactions