Does 25-delta skew mark intraday tops and bottoms?

506 sessions, 209,364 entries, a matched-random control and a matched-minute control

Published 2026-08-24 · Sample SPX 0DTE chain, 506 archived sessions, 209,364 entries across three horizons · Data per-minute 25-delta risk reversal from the measured book, licensed tick tape · Status measurement, not a signal

A widely repeated practitioner claim — it appears verbatim in the glossary of a widely followed dealer-positioning publisher — is that the 25-delta risk reversal marks the edges of a near-term range. When put demand rises, the metric moves and that "can often mark the bottom of a near-term range"; when calls are bid instead, it "can mark a near-term top". Fear at the lows, greed at the highs, and the option surface says which one you are looking at before price does.

It is a good claim to pick a fight with, because it is fully testable. The 25-delta risk reversal is computed every minute in our book: put implied vol minus call implied vol at 25 delta on the 0DTE chain, in vol points. "Extreme" needs no judgement either — rank each minute against everything the session has printed so far and take the top and bottom deciles. Expanding window, never the whole day: ranking a 10:05 reading against the full session uses the afternoon to classify the morning, which is lookahead and is exactly how a previous study's one "profitable" cut turned out to be fake.

The sign, because it is the whole result

Our number is put IV minus call IV, so HIGH means puts are bid, which is fear, which by the claim marks a bottom and is scored as a buy. LOW means calls are bid, which marks a top and is scored as a sell. The published metric moves the other way round, so the translation matters: throughout this post a positive number means the claim worked. The first draft of the test scored it inverted and would have reported a strong confirmation of a claim the data contradicts.

1. The naive result

Forward returns in basis points at 5, 15 and 30 minutes, sign-adjusted as above, t-statistics clustered by day — 196k minutes inside 506 sessions are not 196k independent samples, and treating them as such inflates t by roughly √390. Each signal is set against a matched-random control: the same number of entries, on the same day, at the same horizon, drawn uniformly from the same eligible minutes. Whatever the claim shares with plain index drift appears in both columns and cancels.

sidehorizonsignal bpcontrol bpedge bpt (day-clustered)prior bpdays
puts bid → buy5m−0.57+0.14−0.71−2.75−0.39478
puts bid → buy15m−1.35−0.22−1.14−2.90+0.43474
puts bid → buy30m−2.47+0.25−2.72−3.46+2.75473
calls bid → sell5m−0.13−0.06−0.07−1.60+0.60506
calls bid → sell15m−0.23−0.05−0.18−1.10+0.93506
calls bid → sell30m−0.54−0.11−0.43−1.28+0.55506

Tails are the top and bottom 10% of the session-to-date distribution. The t column is on the signal column, day-clustered; edge is signal minus control. prior is the raw return over the h minutes before entry, not sign-adjusted.

Every cell is negative. The claim's own scoring, applied to its own metric, loses. The puts-bid half — the "fear marks a bottom" half, the one that gets quoted — is the half with a real magnitude behind it: buying the fear tail is −2.72 bp against its control at 30 minutes, with the signal mean at t −3.46. The calls-bid half is not wrong so much as absent: −0.43 bp, t −1.28, nothing that would survive a second look. Split chronologically at 2025-12-23, the first 70% gives −2.05 bp at 30 minutes (t −2.03) and the last 30% gives −4.45 bp (t −3.24). The sign holds out of sample and the magnitude roughly doubles.

2. The confound that decides how to read it

Look at the prior column. Entries in the fear tail sit on a positive preceding 30-minute return of +2.75 bp, and out of sample +5.10 bp. The sequence is not "fear, then a fall". It is: price rises, puts get bid, the rise gives some back. Extreme skew arrives after a move, which makes the entire forward return suspect — if the index simply gives back short-horizon moves, the table above is what you would print with skew contributing nothing at all.

The matched-random control cannot catch this. It corrects for drift, because it samples the average minute; the confound here is the average minute that follows the same move. Different question, different control.

3. The matched-minute test

So each signal minute is paired with a comparison minute drawn from the same session whose skew percentile is plainly not extreme (25–75), choosing the one whose prior return over the same horizon is closest. Matching is without replacement, so one convenient neighbour cannot serve many signals, and any pair whose prior returns differ by more than 3.0 bp is thrown away rather than used: a bad match is worse than no match, because it silently reintroduces the confound it was built to remove.

horizonsignal bpmatched bpdifference bpt (day-clustered)prior bpmean match gappairsdays
5m−0.63+0.23−0.86−3.12−0.270.31 bp13,229476
15m−1.27+0.45−1.72−3.26+0.500.44 bp12,150473
30m−2.57+0.22−2.79−4.42+2.770.52 bp11,492471

The mean match gap is the diagnostic: the paired minutes' prior returns differ by 0.31–0.52 bp on average, against prior moves of a few bp, so the matching did what it claims. The t here is on the difference, day-clustered.

It does not collapse. At 30 minutes the fear tail is followed by a price 2.79 bp below a minute of the same session after an all-but-identical preceding move, t −4.42, on 11,492 pairs over 471 days. Out of sample (after 2025-12-05) the difference is −5.97 bp, t −4.44. So the skew, not the rally, is doing the work — but read what the work is. The matched minute went up (+0.22 bp) and the signal minute went up less, or stalled. This is an advance that stops, not a reversal, and "puts bid marks a bottom" is the opposite of both.

4. The regime cuts, and the bar they have to clear

The obvious next question is whether the effect concentrates somewhere. Same pairs, same matching, 30-minute horizon; only the split is new, so every row reads against the base. The base for this pass is −2.63 bp, t −4.14, 11,443 pairs over 471 sessions — it differs in the third digit from the table above because the pairing is greedy and the two passes walk the pool in a different order, which is itself a useful reminder of how much precision these numbers carry.

Four splits with two or three cells each is roughly ten comparisons. At the usual ±2 threshold, one of ten comes up "significant" by luck about 40% of the time — so the bar column is the two-sided t required after a Bonferroni correction for the number of cuts made here, which is the honest threshold once you admit you looked more than once. A row that clears ±2 but not its bar is not a weak finding; it is what a lucky subgroup looks like, and the only reason it can be told apart from a real one is that the bar was computed before reading the column.

splitcelldifference bpt (by day)barverdictpairsdays
allevery pair−2.63−4.143.29REAL?11,443471
gamma signdealer long gamma−4.06−5.722.83REAL?8,961420
gamma signdealer short gamma−0.71−0.632.842,482250
vs flipspot above zero-gamma−4.04−5.662.83REAL?8,959420
vs flipspot below zero-gamma−0.71−0.632.842,482250
ROOMspot near a band edge−1.19−0.812.841,372230
ROOMspot deep inside the band−2.56−1.812.851,691203
capacityhigh absorbing capacity−0.95−0.882.843,282317
capacitylow absorbing capacity−3.00−3.782.84REAL?5,764378
sessionfirst third−3.25−4.242.83REAL?5,543418
sessionmiddle third−0.08−0.052.844,367330
sessionlast third+1.59+1.202.851,533205

"REAL?" means the row cleared its adjusted bar, and keeps its question mark on purpose. The bar varies by row because it is a t-quantile and each cell has its own number of days.

As it happens, no row in this table landed in the trap zone: every cut either cleared its adjusted bar or failed to reach ±2 at all. That is luck, not virtue — the bar was there to catch the −2.4s, and had ROOM's "deep inside the band" row (t −1.81) come in at −2.3 instead, it would have been reported as noise wearing a suit rather than a discovery.

Three things about the surviving cuts, in the order that matters.

The two strongest splits are one variable, not two. Dealer gamma positive at spot and spot above the zero-gamma flip agree on 100.0% of pairs — the flip is by definition where volume-convention gamma changes sign. Reporting both as separate confirmations would have been double-counting one cut, which is why they carry nearly identical numbers.

It holds out of sample and strengthens. Inside the long-gamma regime: −3.27 bp (t −4.19, 296 sessions, 6,785 pairs) in sample, −5.97 bp (t −3.94, 124 sessions, 2,176 pairs) out. And it is not a rare corner — long gamma covers 78% of all pairs (8,961 of 11,443) on 420 of 471 sessions, which is the ordinary state of the book rather than a curiosity.

Time of day is a separate, weaker cut. Within long gamma the first third still carries most of the effect (−4.01 bp, t −4.10, 373 sessions) while the middle third is nothing (−1.39 bp, t −0.79). So session phase is not merely a restatement of the gamma regime, but on its own it is the lesser of the two.

The long-gamma result is worth slightly more than its t-statistic, for one reason that has nothing to do with statistics: it is the direction the mechanism predicts. Long-gamma hedging leans against moves, so a skew spike in that regime being followed by an advance that stops is consistent with the story; short gamma, where the same hedging amplifies, shows −0.71 bp, t −0.63. That split was not chosen to fit.

The verdict, in one sentence

The 25-delta risk reversal does not mark bottoms — the fear tail is followed by an advance that stops, 2.79 bp below a matched minute of the same session at 30 minutes (t −4.42, 11,492 pairs, 471 days), and the "calls bid marks a top" half of the claim is empty (−0.43 bp, t −1.28).

What it is not

Not costed, and not costable from this study. Nothing above touches a bid-ask spread, a fee or a fill. The largest effect measured anywhere in it, the −4 bp long-gamma cut, is still under a tenth of the single-episode spread — the dispersion of individual outcomes swamps the average by more than an order of magnitude, so a 2.79 bp mean is a statement about a large pile of minutes and never about the next one. No SPX round-trip cost is measured here, so no post-cost number is quoted; the honest phrasing is that the effect is smaller than the frictions it has not yet been charged.

Not a reversal. The signal minute's forward return is negative in the table, but its matched minute's is positive. What is being measured is a stalled advance relative to a comparable minute, not a turn. Any reading of it as "short the fear spike" is reading a difference as a level.

Not a mechanism. The regime split is observational. Dealer gamma sign correlates with a great many other things about a session — realised vol, time of day, the calendar, the composition of the tape — none of which were held constant. Consistency with the hedging story is not evidence for it.

Not a refutation of everything the claim implies. Two deciles of one metric on one index at three horizons were tested. A slower version, a different tail width, or an underlying with a different flow mix are separate questions with separate answers.

Not a full-archive result. 506 sessions, not the full published history: the 25-delta reading is returned only when both wings bracket 0.25 delta, because an extrapolated wing is a fabricated number and this series feeds a chart where a flat line reads as information.

Nothing on this page is investment advice; see the Terms.

Reproduce it

The book behind it is licensed. The 25-delta risk reversal here is computed from per-print SPXW trades and NBBO quotes from ThetaData — the SPXW trade + NBBO tick feed, a paid subscription. Reproducible does not mean free, and we cannot redistribute the tape.

What is free is the output series, which is all this test actually consumes. Every finished session's JSON at https://gex.live/snapshots/YYYY-MM-DD.json carries, aligned one-for-one with minutes (09:30 to 15:59 ET, 390 values) and spot:

Two series and a price column rebuild the whole study. The dates are the ones listed at /sessions.

pythonreproduce
import numpy as np, json
d = json.load(open("2026-08-11.json"))
s  = np.array(d["spot"], float)          # 390 one-minute prints
sk = np.array(d["skewpct"], float)       # 25-delta put/call ratio, percent

# session-to-date percentile, EXPANDING -- never the whole day
rk = np.full(len(sk), np.nan)
for i in range(30, len(sk)):
    rk[i] = (sk[:i] < sk[i]).mean() * 100

H = 30
fwd = np.full(len(s), np.nan); fwd[:-H] = (s[H:] / s[:-H] - 1) * 1e4   # bp
pri = np.full(len(s), np.nan); pri[H:]  = (s[H:] / s[:-H] - 1) * 1e4

ok   = np.isfinite(rk) & np.isfinite(fwd) & np.isfinite(pri)
sig  = np.flatnonzero(ok & (rk >= 90))              # the fear tail
pool = np.flatnonzero(ok & (rk >= 25) & (rk <= 75))  # plainly not extreme

# greedy nearest neighbour on the PRIOR return, no replacement, reject beyond 3 bp
avail, diffs = list(pool), []
for i in sig:
    if not avail: break
    a = np.array(avail); j = int(a[np.argmin(np.abs(pri[a] - pri[i]))])
    if abs(pri[j] - pri[i]) > 3.0: continue     # drop rather than fake a match
    avail.remove(j); diffs.append(fwd[i] - fwd[j])
# per-session mean of diffs; then t across sessions -- CLUSTER BY DAY, never by pair

Run that over every date in the archive and the 30-minute row of the matched table falls out. To reach the regime cuts you also need flip, hold_lo and hold_hi from the same JSON — the gamma-sign split is spot against the flip, and ROOM is the distance to the nearer band edge as a share of band width. The terminal marks these minutes live as they occur; the layer exists to show where the state happened and how each one turned out, not to tell anyone to do anything.

Part of gex.live research. Measured on the free session archive; every session is free to replay.