A reproducible Python companion to “Can You Trust Your Intraday Database?“

DF_Intraday_Databases

First published on Substack

This article introduces concretum_tests, a reproducible Python tool for detecting and correcting intraday market data issues including phantom highs/lows, early-close leakage, stale bars, and missing bars. It also explains the methodology behind each detection and why some issues should be fixed while others should only be flagged.

Introduction

Intraday market data is one of the most overlooked variables in quantitative trading research.

In the previous article, Can You Trust Your Intraday Database?, we showed that running the same Opening Range Breakout (ORB) backtest across multiple data providers produced materially different outcomes. In some configurations the spread between the best and worst result exceeded a threefold difference, despite identical strategy logic. We identified five main sources of divergence: phantom highs and lows, stale bars, early-close leakage, tick-to-bar assignment, and venue coverage / trade aggregation.

Identifying that data quality matters is one thing. Knowing how to detect and correct the issues, and where it is genuinely safe to do so, is the next step. This article focuses on three of the five issues, ranked by how clean a fix is possible:

  1. Early-Close Day Leakage : fully deterministic, fix is trivial once the official NYSE half-day calendar is loaded.
  2. Phantom Highs and Lows : probabilistic detection, with a clearly parameterized trade-off between false positives and false negatives. A correction is possible, but the result is necessarily a reconstructed value, not a price reported by the provider.
  3. Stale Bars and Missing Bars : easy to detect, but we recommend against fabricating prices to fill them. We explain why.

We accompany this article with concretum_tests, a Python CLI tool that runs all four detections (the three above plus missing bars) on any 1-minute OHLCV CSV and, at the user’s discretion, produces corrected output files.

Research Tool
concretum_tests is currently an internal tool used in our own research workflow. It is not yet a fully packaged library with a stable API; think of it as a working prototype that already performs its intended job well, but may continue evolving as the methodology matures.

For now, the single-file script attached to this article is the canonical version. We plan to eventually publish a proper GitHub repository with versioned releases and expanded documentation as the project develops further.
Google Drive
Download concretum_tests
Download the single-file Python CLI used throughout this article. The tool detects early-close leakage, phantom highs/lows, stale bars, and missing intraday bars in 1-minute OHLCV datasets.

Install the required dependencies locally and run it directly from the terminal on your own CSV files.
Download →

Intraday Data Quality Testing with concretum_tests

concretum_tests is a single-file Python script. You run it from the terminal, it reads your data, tests it, tells you what is wrong, and optionally fixes what can be fixed cleanly.

Installation

pip install polars rich exchange-calendars

No heavy frameworks, no database connections, no API keys. The exchange_calendars package provides the authoritative NYSE early-close calendar directly.

Expected Data Format

The tool expects 1-minute OHLCV CSVs placed in a Data/Raw/ folder with the following layout:

datetime,open,high,low,close,volume
2016-04-04 09:30:00,105.65,105.70,105.32,105.48,27391.0
2016-04-04 09:31:00,105.44,105.46,105.33,105.44,5066.0
...
ColumnDescription
datetimeBar timestamp (YYYY-MM-DD HH:MM:SS)
openOpening price of the 1-minute bar
highHighest price within the minute
lowLowest price within the minute
closeClosing price of the 1-minute bar
volumeNumber of shares traded

Why the provider matters

Different providers use different bar-labeling conventions. Most providers timestamp a bar at the start of the minute: a bar labeled 13:00 covers the interval 13:00:00 to 13:00:59. IQFeed timestamps bars at the end of the minute: a bar labeled 13:00 covers 12:59:00 to 13:00:00. This affects which bar is the last valid bar before an early close. If the tool does not know which convention your data uses, it may remove one bar too many or too few.

The tool auto-detects the provider from the filename. If your file is named TQQQ_Polygon_1min.csv, it will detect “POLYGON” and apply the start-of-minute convention. Recognized providers: ALPACA, POLYGON, MASSIVE (treated as Polygon), IBKR, IQFEED, DATABENTO. You can also force it with --provider.

If your provider is not in the list, use --bar-label start or --bar-label end to specify the convention manually.

Usage

python concretum_tests.py                              # all CSVs in Data/Raw/
python concretum_tests.py Data/Raw/TQQQ_Polygon_1min.csv  # specific file
python concretum_tests.py --zb 3 --zw 3 --zv 4        # custom phantom thresholds
python concretum_tests.py --no-fix                     # skip the fix prompt
python concretum_tests.py --provider MyFeed --bar-label end  # custom provider

Issue 1: Early-Close Day Leakage

What it is

On a small but recurring number of trading days each year, the NYSE closes early, typically at 13:00 ET. Black Friday, the day before Independence Day, and Christmas Eve are the three regular cases.

Any 1-minute bar reported after the early close on these days is structurally outside the regime a strategy assumes. Some providers (e.g., Polygon, IBKR) truncate the session correctly. Others (e.g., Alpaca, IQFeed) continue to report bars all the way to 16:00 ET on half-days.

Why it is easy to fix

The tool uses the exchange_calendars package to retrieve the authoritative NYSE session schedule, which includes every early-close day and its exact close time. This covers all standard cases as well as any one-off early closes.

What the tool does

Detection: Flags every bar on an NYSE early-close day whose timestamp falls at or after the market close (adjusted for the provider’s labeling convention).

Correction: Removes all flagged bars. This is a clean, lossless operation: no price is fabricated, the leaked bars simply should not have been there.

Issue 2: Phantom Highs and Lows

What they are

A phantom high is an isolated 1-minute upper-wick spike that is not confirmed by the same minute in any other provider’s data. Phantom lows are the symmetric case. They typically last exactly one bar and leave essentially no footprint in the volume series. Visually, they look like a thin needle pointing away from the prevailing price action.

Why they happen

U.S. equities trade across multiple exchanges and off-exchange venues, including dark pools, where trades can be reported to the consolidated tape with a delay. Depending on the provider’s filtering and consolidation rules, these late-reported prints may be included, excluded, or adjusted, leading to isolated highs or lows that are not consistently observable across datasets.

Different providers also rely on different feeds. SIP-based feeds aim for broad cross-venue coverage; direct-feed-based or “mini” packages may include only a subset of venues. The choice of feed influences which prints make it into a 1-minute bar.

How we detect them

The detection uses three conditions, all of which must hold simultaneously for a bar to be flagged. We describe the high-side rule; the low side is symmetric.

Condition 1 (body baseline): The bar’s upper wick exceeds z_b times the day’s largest body. The body of a candle is |close - open|, representing the largest genuine intrabar price movement on the day. If a wick dwarfs the biggest body observed that session, it is geometrically anomalous.

Condition 2 (wick baseline): The bar’s upper wick exceeds z_w times the day’s median upper wick. This compares the spike to what is typical on the same side for that day. A wick that is several multiples of the day’s median wick is an outlier by its own standard.

Condition 3 (volume coherence): The bar’s volume is below z_v times the day’s median volume. This is the key coherence check. A real price move large enough to produce such a spike should attract trading activity. If the bar that hosts the spike traded at a fraction of normal volume, the move is unlikely to reflect genuine market interest.

The first two conditions are geometric and independent of each other: the wick is anomalous on two separate yardsticks. The third condition filters out volatile-but-legitimate bars (which would have elevated volume) from suspicious ones (which do not). The AND logic means all three must trigger simultaneously.

Our default thresholds are z_b = 2, z_w = 2, z_v = 5.

The threshold trade-off

There is no single correct answer for these parameters. Tightening them (higher z_b, z_w; lower z_v) reduces false positives (legitimate volatile bars wrongly flagged) at the cost of more false negatives (real phantoms missed). Loosening them does the opposite. Since there is no ground-truth label to optimize against, the right operating point depends on what the user is willing to trade off.

We recommend starting with the defaults, auditing the flagged bars on a representative sample of days, and adjusting if the false-positive / false-negative balance is unsatisfactory. The tool accepts --zb, --zw, and --zv flags for this purpose.

An optional correction

For applications that cannot tolerate phantom-driven outliers (strategies whose signals depend on intraday extremes: stops based on high/low, ATR-based volatility estimates, breakout triggers), the tool can replace the phantom value with a reconstructed price.

The reconstruction is intentionally simple:

The medians are the day’s typical upper/lower wick. The intuition: we replace an anomalous wick with a wick of typical size for that session, anchored on the body of the bar that hosted the phantom. The result preserves the bar’s body and sign, but removes the spike.

The corrected value is an artificial reconstruction, not data reported by the provider. It is fine to use it as a defense against pathological outliers, but it should always be flagged as such in any audit trail.

Issue 3: Stale Bars and Missing Bars

What they are

Stale bars and missing bars are two faces of the same phenomenon: the absence of trading activity during a 1-minute interval that the provider was meant to cover.

A strong stale bar is a 1-minute bar where open == high == low == close. A single price is reported for the whole minute. Some providers synthesize a stale bar from the previous close when no trade is observed; others omit the minute entirely. The latter shows up as a missing bar: a gap in the 1-minute grid.

In a highly liquid instrument such as TQQQ, isolated stale bars or single-minute gaps are plausible. Long sequences of either indicate one of two things:

  1. The market was halted and no actual trading occurred. This is legitimate.
  2. The data is structurally incomplete: the provider failed to record trades or aggregated them in a way that lost minute-level detail.

Distinguishing case (1) from case (2) without external information is hard. But cross-provider comparison is often enough.

Halts vs. data quality

The clearest illustration is the COVID-shock days of March 2020, which contained multiple LULD halts. On those days a strategy querying IBKR sees a sequence of stale bars with volume = 0 across each halt; a strategy querying Alpaca sees a gap in the 1-minute grid. The underlying market state is identical but the two providers represent it differently.

The lesson is operational: when assessing database quality, one must look at stale and missing counts together, and be aware that some of those counts are legitimate representations of halts.

Detection

Strong stale bars: check whether all four OHLC values are identical on a bar. Trivial.

Missing bars: compute the elapsed time between consecutive rows on the same trading day. Any gap greater than 1 minute indicates missing data. The tool uses each day’s observed first and last bar as session bounds, so early-close days and overnight gaps are handled naturally.

Why we do not correct these

Unlike a phantom, where we have too much information (an obvious anomaly we can subtract), a missing or stale region is a place where we have too little information. Any filling is a fabrication.

The cost of a wrong fill is also asymmetric. Backtests on filled data can produce trades that would have been impossible in reality (market-on-close orders inside halts, stops triggered against synthesized volatility). A backtest that appears tradeable but is not is far worse than one that flags days as “untrustworthy.”

The pragmatic recommendation: flag days with significant stale or missing content and handle them explicitly. Exclude them, report separate statistics on the “clean” subset, or cross-validate any signal that fires on those days.

Putting It All Together

All four detections are wrapped into the single concretum_tests CLI. It takes one or more CSVs and produces:

  1. A structured terminal report with pass/fail/warn for each check.
  2. Corrected CSVs (optional) with early-close bars removed and phantom highs/lows reconstructed. Stale and missing bars are reported but never modified.

Example output

── Data Integrity ─────────────────────────────────────────────────────────────

  ✓ Column schema valid                          passed
  ✓ No NaN in OHLC                               passed
  ✓ No NaN in volume                             passed
  ✓ No zero prices (OHLC)                        passed
  ✓ No negative prices                           passed
  ✓ Low ≤ High                                   passed
  ✓ Open within [Low, High]                      passed
  ✓ Close within [Low, High]                     passed
  ✓ No negative volume                           passed
  ✓ Chronological order                          passed

── Early-Close Leakage ────────────────────────────────────────────────────────

  ✗ No bars after early-close cutoff             FAILED  21 bars on 21 half-day(s)

── Phantom Detection  zb=2.0  zw=2.0  zv=5.0 ─────────────────────────────────

  ! Phantom highs                                warn    14 events
  ! Phantom lows                                 warn    24 events

── Structural Quality ─────────────────────────────────────────────────────────

  ! Strong stale bars (O=H=L=C)                  warn    10,374 bars (1.07%)
  ! Missing bars (intraday gaps)                 warn    4,383 gaps  ~5,079 missing min

────────────────────────────────────────────────────────────────────────────────
  Summary   10 passed · 1 failed · 4 warnings   0.6s

Multi-provider workflow

Place your files in Data/Raw/ and run once:

Data/
└── Raw/
    ├── TQQQ_Alpaca_1min.csv
    ├── TQQQ_Polygon_1min.csv
    ├── TQQQ_IBKR_1min.csv
    ├── TQQQ_IQFeed_1min.csv
    └── TQQQ_Databento_1min.csv

The tool processes each file independently and shows a grand summary at the end.

Conclusion

Reliable intraday market data is a foundational requirement for any serious quantitative trading workflow.

The previous article asked: If you run the same backtest over the same period using different data sources, do you get the same results?

The answer was no. This follow-up asked: knowing where to look, can we detect and fix the issues?

The answer is yes for two of them, with one important caveat. Early-close leakage is fully deterministic and can be removed cleanly. Phantom highs and lows can be detected with a parameterized rule and replaced with a reconstructed value, but the user must remain aware that the corrected value is not data reported by the provider. Stale and missing bars are straightforward to detect, but we strongly recommend against fabricating prices to fill them.

Once these checks are in place, the backtest results across providers converge meaningfully. The broader implication remains: data is not a neutral input.

The tool is attached to this article as a single Python file. It is an internal tool we are sharing in its current working state. As we refine and extend it, we will publish a versioned repository. In the meantime, we hope this is a useful starting point for your own data integrity workflow.

Get research like this before it’s public.

Enter your email to receive our next data-driven analysis.

Live Experiment

Can You Beat a Systematic Strategy?

We’re running a research experiment to test whether day trading skill can improve the performance of a fully systematic intraday strategy.

No trade generation. No guessing.
Just managing exposure using price action — and we are measuring the result.

Join the Experiment →