Blog/Case study · Jul 30, 2026

The clock that lied for ten years

OverfitCheck Research · 7 min read

One of the strategies in our research program is an intraday breakout system on a Nasdaq-100 CFD, hosted at a major CFD broker. It trades one narrow window each morning. We believed that window was anchored to the US cash open. Its backtest covers fourteen years and 867 trades, and it is profitable live. For roughly ten of those fourteen years, half of its historical trades did not happen at the time of day we thought they did.

This is how a broker’s server clock quietly corrupted a decade of session analysis, how we found it, and what it means for anyone whose backtest references clock time. None of it took unusual tools. It took distrusting a timestamp.

Server time is a claim, not a fact

MetaTrader-style platforms stamp every bar and every trade in server time. That is whatever clock the broker chooses to run. The common convention is UTC+2 in winter and UTC+3 in summer, tracking Eastern European daylight-saving, because that alignment keeps the daily candle closing with the New York session. Our broker’s clock behaved exactly that way in 2024, when we checked it. Our mistake was extrapolating. We assumed it had always behaved that way, back through the start of our data in 2012.

There was no documentation either way. Historical broker clock policy is not something a data export advertises. The backtest inherited the assumption, and every analysis built on the backtest inherited it too.

The forensics

The error surfaced while we were doing something else. We were running a census of intraday price discontinuities, with two analysts working independently. Both hit the same anomaly. Pre-2024 event signatures were not where the clock said they should be. Both then reconstructed the server’s actual UTC offset, date by date, from the price data itself.

The method is called joint multi-feature timestamp forensics. The idea underneath the name is simple. Certain market events are pinned to external clocks, so they can act as witnesses. The 8:30 a.m. ET macro-release cluster (CPI, payrolls) produces a volatility spike at a known instant. The US cash open at 9:30 a.m. ET has a distinctive auction signature. The Eurex pre-open leaves its own mark in the early European morning. Any one witness is circumstantial. Together they let you solve for the server’s offset on every trading day in the sample, and flag the ambiguous days for exclusion instead of guessing.

Both independent reconstructions landed on the same answer. Before April 2024, the server ran at a fixed UTC+2 all year round. It never sprang forward. Not for Europe’s daylight-saving, and not for America’s.

What half-wrong looks like

US daylight-saving covers roughly March through November, about eight months of the year. On every US summer-time day before April 2024, the assumed clock mapping was off by exactly one hour. Counted against the trade log:

867
Trades in the 14-year backtest
436
Trades mis-anchored · 50.3%
1 hour
Actual shift within the session

436 of 867 trades executed a full hour later in the US session than every piece of analysis had assumed. A trade we had filed as taken shortly after the open, in the aftermath of the opening auction, had actually been taken deep in the mid-morning. Different phase of liquidity, different participants on the other side.

Nothing on the surface flagged it. The equity curve was continuous. Live performance matched expectations, because the live era started after April 2024, the one era in which the clock assumption was true. The corruption lived entirely in the historical record, and only in eight months of each year. That is close to a perfect hiding place.

The damage to inference

Here is where it gets expensive. The strategy’s worst stretch was 2012 to 2015. Split those years by anchoring, and the mis-anchored trades average −$85 per trade (n = 127) while the correctly anchored trades average +$68 per trade (n = 57). The obvious reading is that the shifted window is the problem. That is exactly the conclusion the data cannot support. Every mis-anchored trade is a summer trade by construction, and every correctly anchored one is a winter trade, so window and season are perfectly confounded. A decade of session analysis, and the one comparison you most want to make is unidentifiable. Sample size is never just how many rows are in the file. It is how many of them actually bear on the claim you are making, which is lesson three of our short course.

We re-ran the full history on the corrected per-date clock. The headline totals roughly washed. The risk profile inverted. Without anyone designing it that way, the system had been running two different session exposures for a decade. Accidental diversification. Meanwhile every session-conditioned statistic we had ever computed was built on partly false labels. Time-of-day splits, seasonal effects, any regime label that referenced a session boundary.

One coda, because it cuts both ways. The accidental hour-late window became a pre-registered hypothesis of its own, tested on 187 winter-season trades that no optimization had ever touched. It validated, at a cost-corrected profit of about +$9,900. That was our program’s first out-of-sample mechanism prediction to survive contact with clean data. We retired it anyway. Its second confirmation arm missed a frozen profit-factor threshold of 1.10. It printed 1.096. Kill rules are only worth having if you apply them when they hurt. Writing the threshold down before the run, and sealing data you have not looked at, are the first two items on the checklist we would give anyone starting out.

The general lesson

Session-anchoring artifacts are a class of backtest corruption that staring at the equity curve cannot catch. If your strategy references clock time in any way, whether that is session opens, news windows, time-of-day filters, or daily bar boundaries, your backtest silently inherits every assumption your broker and your data vendor make about time. Three checks would have caught ours years earlier:

There is a statistical takeaway underneath all this too. When half your sample turns out to belong to a different session than you believed, the effective sample behind your actual claim is half what you thought. And any edge that genuinely is a session phenomenon will show up as concentration once the labels are fixed. Those are a sample-size problem and a regime-dependency problem, and both are testable without knowing anything about clocks. OverfitCheck’s regime-dependency and sample-size tests exist for exactly this class of failure. They cannot find a broker’s clock bug for you. They will tell you when a backtest’s profits are concentrated in a slice of data too small, or too strange, to trust.

Read next

Run your backtest through eight peer-reviewed statistical tests. Free, no credit card required.

Try OverfitCheck free →