Equity curves of an ORB trading strategy from 2016 to 2026 across Polygon, Alpaca, IQFeed, and IBKR showing large performance differences despite identical strategy parameters.
Backtest results for the ORB strategy using identical parameters across four data providers (2016–2026).

First published on Substack

This article was shared there two weeks ago. New research is published on Substack before it appears on the site.

Introduction

In quantitative trading, backtesting data quality can materially change results, yet it is often overlooked. Most attention goes to strategy design, execution, and risk management, yet the reliability of the underlying data provider can materially affect results.

While expanding our Opening Range Breakout (ORB) research across multiple data sources, we expected consistent results, but that is not what we found. Using identical code, parameters, and time periods, the same strategy produced materially different outcomes depending on the data provider.

This raises a fundamental question: how much of a backtest is driven by the strategy itself, and how much by the data? This article presents the analysis and its key findings, with full reproducibility details to be published separately.

Strategy Overview: Opening Range Breakout

Before analyzing the data, we define the strategy used throughout this study: the Opening Range Breakout, as described in Can Day Trading Really Be Profitable? (2023), co-authored with Andrew Aziz.

The formulation is intentionally simple and fully deterministic, making it suitable for cross-provider comparison.

Rules

Stops

Exit

Sizing

Scope and limitations

This study evaluates the ORB strategy under a specific set of assumptions, instruments, and market conditions. Results are not universal and depend on data quality, execution assumptions, and implementation details. Small changes in these inputs can lead to materially different outcomes. Independent validation is recommended.

Disclosure

This analysis is conducted independently for research and educational purposes. We are not affiliated with, sponsored by, or compensated by any data provider referenced in this study.

The Discovery: Same Strategy, Different Results

The key question is not whether differences exist, but how large they are.

To measure this, we ran the same backtest over the last 10 years across four different data providers:

All tests used identical dates, code, and parameters, as well as a starting capital of $25,000.

Two stop configurations were evaluated:

  1. H/L Stop – An opening range high/low (H/L) stop based on intraday extremes,
  2. ATR Stop – An ATR-based stop derived from daily volatility.

H/L Stop Results

ATR Stop Results

Under the H/L configuration, results diverge significantly, with final portfolio values ranging from $226k to $726k under identical strategy logic.

In contrast, the ATR-based configuration produces far more consistent results across providers. While performance differences remain, the overall behavior converges and the dispersion is significantly reduced.

This divergence reflects how sensitive the strategy is to the underlying data. When stops rely on intraday highs and lows, small differences in price can compound into large performance gaps. When stops are based on daily volatility, sensitivity to intraday variation is reduced.

This raises a more specific question: what differences in the data drive these outcomes?

Before analyzing each issue in detail, we summarize the main sources of discrepancy identified in the data:

We now examine each of these issues in turn.

Data Quality Issues

Issue 1: Phantom Highs and Lows

When comparing minute-level High and Low values across providers, we observe that most series align closely, but isolated spikes appear in a single dataset and are not confirmed by the others (e.g., Polygon in the examples above).

Although these discrepancies are small, they directly affect strategies that rely on intraday extremes. A slightly higher reported high or lower low can shift stop levels, alter position sizing under fixed risk, and ultimately change trade outcomes. For H/L-based stops, even minor differences can materially impact backtest results.

These discrepancies are typically not driven by actual market activity, but by differences in how providers construct their datasets. U.S. equities trade across multiple exchanges and off-exchange venues such as dark pools, where trades may be reported to the tape with a delay. As a result, some transactions can appear out of sequence relative to the prevailing market. Depending on the provider’s filtering and consolidation rules, these prints may be included, excluded, or adjusted, leading to isolated highs and lows that are not consistently observable across datasets.

Issue 2: Stale Bars

A second issue observed in the data is the presence of stale bars, defined as 1-minute candles where Open = High = Low = Close. While occasional occurrences are expected, extended sequences of such bars in a highly liquid instrument like TQQQ are unlikely to reflect actual market activity and instead point to missing or improperly aggregated data.

For intraday strategies, this has direct consequences. During the opening range, stale bars can collapse the range to a single price, reducing stop distance and destabilizing position sizing. Outside the opening window, they suppress volatility and can prevent stop or target conditions from being detected, leading to delayed execution and price gaps. In both cases, the time series no longer reflects continuous trading, and backtest results become unreliable.

A Case Study on Stale Bars: IBKR Data Used in the Original Paper

The original ORB study was written in early 2023 and used intraday data downloaded from Interactive Brokers (IBKR). The figure below shows the equity curve reported in that study.

Equity curve from the original ORB study using IBKR data available at the time of the analysis (February 2023).

At first glance, these results appear inconsistent with the IBKR backtests presented earlier in this article. This raises a natural question: if the same data provider is used, why do the results differ?

To investigate this, we compare the dataset used in the original study with the same trading days, re-downloaded from IBKR on March 24, 2026.

The difference is not subtle. The earlier dataset shows continuous intraday price variation consistent with active trading, while the more recent download contains extended sequences of flat, unchanging prices. In some sessions, more than 350 out of 390 one-minute bars show no price movement.

These differences are structural. The same provider, queried at different points in time, can produce materially different datasets, leading to different backtest outcomes.

Backtest results using IBKR data downloaded in 2023 versus 2026. Despite identical strategy logic, performance diverges due to differences in the underlying data.

For this reason, a trading strategy that appeared profitable in a backtest conducted in 2023 may lead to completely different conclusions when evaluated using data downloaded three years later. The issue is not necessarily the strategy itself, but the underlying dataset, which has changed over time. This highlights the critical importance of validating data integrity before drawing any conclusions from historical simulations.

Issue 3: Early-Close Day Leakage

On certain trading days, U.S. markets close early (typically at 13:00 ET). However, some data providers continue to report intraday bars beyond the official session close. These bars do not necessarily reflect actual tradable conditions during the regular session, but rather extend the time series into periods that fall outside standard market hours.

For backtesting, this introduces a structural inconsistency. The strategy implicitly assumes a different trading regime, effectively allowing execution outside regular trading hours. As a result, key implementation assumptions break down: market-on-close orders can no longer be used as intended, and the strategy may be exposed to liquidity conditions that are not representative of the regular session.

This issue also creates cross-provider inconsistencies. Some providers truncate the session correctly on early-close days, while others include additional bars beyond the official close. As a result, identical strategies can produce materially different outcomes depending on how each dataset handles these sessions.

Although early-close days are relatively infrequent, the impact is systematic and can bias results if not explicitly addressed. In practice, this requires enforcing a consistent session definition across the dataset. Fortunately, several Python libraries and data-cleaning workflows allow researchers to detect and correct these discrepancies, ensuring alignment with actual tradable hours.

Issue 4: Tick-to-Bar Assignment

A more subtle, yet highly impactful, source of discrepancy arises from how data providers assign individual trades (ticks) to time bars. In particular, trades occurring exactly at time boundaries (e.g., 09:35:00) may be included in either the preceding bar or the subsequent one, depending on the provider’s aggregation convention.

While this difference may appear negligible, it has direct consequences for strategies like ORB. The signal depends on the relationship between the opening price and the close of the 5-minute bar. A single trade allocated differently can slightly shift the bar close, potentially flipping the signal from long to short, or vice versa.

To quantify this effect, we extended the analysis by adding a fifth provider, Databento, and restricted all datasets to the same overlapping window of approximately 750 trading days.

Across this common sample, four providers produce broadly consistent results, while Databento exhibits significantly higher performance.

H/L Stops

ATR Stops

The source of this divergence is highly concentrated.

On just 16 out of 753 days (2.1%), Databento generates a different trade direction compared to the other providers. In most cases, the difference is extremely small, often as little as one cent in the 5-minute close. However, given the binary nature of the signal, this is sufficient to alter the trade.

Below you find some examples…

December 5, 2023:
December 24, 2024 (early close):
March 5, 2025:

When these discrepancies occur on strongly trending days, their impact is amplified. A small number of signal flips, concentrated in high-return environments, can meaningfully reshape the equity curve and create the appearance of superior performance.

The implication is not that one provider is correct and another is wrong, but that differences in tick-to-bar assignment can propagate into materially different outcomes for threshold-based strategies.

Issue 5: Venue Coverage and Trade Aggregation

A further, often overlooked source of discrepancy lies in how the underlying trades are sourced before being aggregated into time bars. U.S. equities trade across multiple venues, including exchanges such as NYSE, Nasdaq, and Cboe, as well as off-exchange facilities. Since no single venue captures all transactions, different data providers rely on different inputs, such as consolidated SIP feeds, proprietary direct exchange feeds, or internally aggregated subsets of venues.

Even when focusing only on executed trades, these choices lead to different datasets. SIP-based data aims to capture a broad cross-venue view, while direct-feed-based datasets may include only a subset of venues. For example, the Databento US Equities Mini package we have used is built by aggregating, through a proprietary methodology, a subset of trading venues and therefore represents only a synthetic “mini” market view rather than a fully consolidated one. Databento also offers direct feeds from multiple venues, but these require significantly more complex aggregation to reconstruct a comprehensive view of the market.

These differences directly affect the construction of 1-minute bars. Missing or additional trades within a time interval can change the observed high, low, close, and volume of a bar. For strategies sensitive to intraday extremes, these discrepancies can translate into materially different signals, even when each dataset is internally consistent.

Conclusion

This article set out to answer a simple question:

if the same strategy is tested over the same historical period using different data providers, should the results be the same?

The answer is NO.

Using the Opening Range Breakout (ORB) strategy from our paper Can Day Trading Really Be Profitable? (2023), and applying identical code, parameters, and trading days across multiple providers, including Massive (Polygon), Alpaca, IQFeed, Interactive Brokers, and Databento, we find that results can diverge materially depending on the data source.

In fact, identical code applied to different datasets can produce significantly different outcomes, with performance dispersion exceeding a threefold difference under certain configurations. These discrepancies are driven by a combination of data issues, including phantom highs and lows, stale bars, early-close leakage, and differences in tick-to-bar assignment, where small variations in how trades are aggregated can flip threshold-based signals.

Once these data issues are identified and addressed, results across providers converge to a consistent conclusion: independent data sources confirm the validity of the backtest results presented in our original paper.

The broader implication is that data is not a neutral input. Differences in how it is constructed, aggregated, filtered, or updated can materially alter conclusions about a strategy, even when the logic is unchanged.

For this reason, backtest results should never be taken at face value. Proper validation requires not only robust strategy design, but also a systematic approach to data quality. Knowing where to look, and how to detect and correct these issues, is an essential part of any serious quantitative research process.

Get research like this before it’s public.

Enter your email to receive our next data-driven analysis.

Live Experiment

Can You Beat a Systematic Strategy?

We’re running a research experiment to test whether day trading skill can improve the performance of a fully systematic intraday strategy.

No trade generation. No guessing.
Just managing exposure using price action — and we are measuring the result.

Join the Experiment →