Does 25-delta skew mark intraday tops and bottoms?
506 sessions, 209,364 entries, a matched-random control and a matched-minute control
A widely repeated practitioner claim — it appears verbatim in the glossary of a widely followed dealer-positioning publisher — is that the 25-delta risk reversal marks the edges of a near-term range. When put demand rises, the metric moves and that "can often mark the bottom of a near-term range"; when calls are bid instead, it "can mark a near-term top". Fear at the lows, greed at the highs, and the option surface says which one you are looking at before price does.
It is a good claim to pick a fight with, because it is fully testable. The 25-delta risk reversal is computed every minute in our book: put implied vol minus call implied vol at 25 delta on the 0DTE chain, in vol points. "Extreme" needs no judgement either — rank each minute against everything the session has printed so far and take the top and bottom deciles. Expanding window, never the whole day: ranking a 10:05 reading against the full session uses the afternoon to classify the morning, which is lookahead and is exactly how a previous study's one "profitable" cut turned out to be fake.
The sign, because it is the whole result
Our number is put IV minus call IV, so HIGH means puts are bid, which is fear, which by the claim marks a bottom and is scored as a buy. LOW means calls are bid, which marks a top and is scored as a sell. The published metric moves the other way round, so the translation matters: throughout this post a positive number means the claim worked. The first draft of the test scored it inverted and would have reported a strong confirmation of a claim the data contradicts.
1. The naive result
Forward returns in basis points at 5, 15 and 30 minutes, sign-adjusted as above, t-statistics clustered by day — 196k minutes inside 506 sessions are not 196k independent samples, and treating them as such inflates t by roughly √390. Each signal is set against a matched-random control: the same number of entries, on the same day, at the same horizon, drawn uniformly from the same eligible minutes. Whatever the claim shares with plain index drift appears in both columns and cancels.
| side | horizon | signal bp | control bp | edge bp | t (day-clustered) | prior bp | days |
|---|---|---|---|---|---|---|---|
| puts bid → buy | 5m | −0.57 | +0.14 | −0.71 | −2.75 | −0.39 | 478 |
| puts bid → buy | 15m | −1.35 | −0.22 | −1.14 | −2.90 | +0.43 | 474 |
| puts bid → buy | 30m | −2.47 | +0.25 | −2.72 | −3.46 | +2.75 | 473 |
| calls bid → sell | 5m | −0.13 | −0.06 | −0.07 | −1.60 | +0.60 | 506 |
| calls bid → sell | 15m | −0.23 | −0.05 | −0.18 | −1.10 | +0.93 | 506 |
| calls bid → sell | 30m | −0.54 | −0.11 | −0.43 | −1.28 | +0.55 | 506 |
Tails are the top and bottom 10% of the session-to-date distribution. The t column is on the signal column, day-clustered; edge is signal minus control. prior is the raw return over the h minutes before entry, not sign-adjusted.
Every cell is negative. The claim's own scoring, applied to its own metric, loses. The puts-bid half — the "fear marks a bottom" half, the one that gets quoted — is the half with a real magnitude behind it: buying the fear tail is −2.72 bp against its control at 30 minutes, with the signal mean at t −3.46. The calls-bid half is not wrong so much as absent: −0.43 bp, t −1.28, nothing that would survive a second look. Split chronologically at 2025-12-23, the first 70% gives −2.05 bp at 30 minutes (t −2.03) and the last 30% gives −4.45 bp (t −3.24). The sign holds out of sample and the magnitude roughly doubles.
2. The confound that decides how to read it
Look at the prior column. Entries in the fear tail sit on a positive preceding 30-minute return of +2.75 bp, and out of sample +5.10 bp. The sequence is not "fear, then a fall". It is: price rises, puts get bid, the rise gives some back. Extreme skew arrives after a move, which makes the entire forward return suspect — if the index simply gives back short-horizon moves, the table above is what you would print with skew contributing nothing at all.
The matched-random control cannot catch this. It corrects for drift, because it samples the average minute; the confound here is the average minute that follows the same move. Different question, different control.
3. The matched-minute test
So each signal minute is paired with a comparison minute drawn from the same session whose skew percentile is plainly not extreme (25–75), choosing the one whose prior return over the same horizon is closest. Matching is without replacement, so one convenient neighbour cannot serve many signals, and any pair whose prior returns differ by more than 3.0 bp is thrown away rather than used: a bad match is worse than no match, because it silently reintroduces the confound it was built to remove.
| horizon | signal bp | matched bp | difference bp | t (day-clustered) | prior bp | mean match gap | pairs | days |
|---|---|---|---|---|---|---|---|---|
| 5m | −0.63 | +0.23 | −0.86 | −3.12 | −0.27 | 0.31 bp | 13,229 | 476 |
| 15m | −1.27 | +0.45 | −1.72 | −3.26 | +0.50 | 0.44 bp | 12,150 | 473 |
| 30m | −2.57 | +0.22 | −2.79 | −4.42 | +2.77 | 0.52 bp | 11,492 | 471 |
The mean match gap is the diagnostic: the paired minutes' prior returns differ by 0.31–0.52 bp on average, against prior moves of a few bp, so the matching did what it claims. The t here is on the difference, day-clustered.
It does not collapse. At 30 minutes the fear tail is followed by a price 2.79 bp below a minute of the same session after an all-but-identical preceding move, t −4.42, on 11,492 pairs over 471 days. Out of sample (after 2025-12-05) the difference is −5.97 bp, t −4.44. So the skew, not the rally, is doing the work — but read what the work is. The matched minute went up (+0.22 bp) and the signal minute went up less, or stalled. This is an advance that stops, not a reversal, and "puts bid marks a bottom" is the opposite of both.
4. The regime cuts, and the bar they have to clear
The obvious next question is whether the effect concentrates somewhere. Same pairs, same matching, 30-minute horizon; only the split is new, so every row reads against the base. The base for this pass is −2.63 bp, t −4.14, 11,443 pairs over 471 sessions — it differs in the third digit from the table above because the pairing is greedy and the two passes walk the pool in a different order, which is itself a useful reminder of how much precision these numbers carry.
Four splits with two or three cells each is roughly ten comparisons. At the usual ±2 threshold, one of ten comes up "significant" by luck about 40% of the time — so the bar column is the two-sided t required after a Bonferroni correction for the number of cuts made here, which is the honest threshold once you admit you looked more than once. A row that clears ±2 but not its bar is not a weak finding; it is what a lucky subgroup looks like, and the only reason it can be told apart from a real one is that the bar was computed before reading the column.
| split | cell | difference bp | t (by day) | bar | verdict | pairs | days |
|---|---|---|---|---|---|---|---|
| all | every pair | −2.63 | −4.14 | 3.29 | REAL? | 11,443 | 471 |
| gamma sign | dealer long gamma | −4.06 | −5.72 | 2.83 | REAL? | 8,961 | 420 |
| gamma sign | dealer short gamma | −0.71 | −0.63 | 2.84 | — | 2,482 | 250 |
| vs flip | spot above zero-gamma | −4.04 | −5.66 | 2.83 | REAL? | 8,959 | 420 |
| vs flip | spot below zero-gamma | −0.71 | −0.63 | 2.84 | — | 2,482 | 250 |
| ROOM | spot near a band edge | −1.19 | −0.81 | 2.84 | — | 1,372 | 230 |
| ROOM | spot deep inside the band | −2.56 | −1.81 | 2.85 | — | 1,691 | 203 |
| capacity | high absorbing capacity | −0.95 | −0.88 | 2.84 | — | 3,282 | 317 |
| capacity | low absorbing capacity | −3.00 | −3.78 | 2.84 | REAL? | 5,764 | 378 |
| session | first third | −3.25 | −4.24 | 2.83 | REAL? | 5,543 | 418 |
| session | middle third | −0.08 | −0.05 | 2.84 | — | 4,367 | 330 |
| session | last third | +1.59 | +1.20 | 2.85 | — | 1,533 | 205 |
"REAL?" means the row cleared its adjusted bar, and keeps its question mark on purpose. The bar varies by row because it is a t-quantile and each cell has its own number of days.
As it happens, no row in this table landed in the trap zone: every cut either cleared its adjusted bar or failed to reach ±2 at all. That is luck, not virtue — the bar was there to catch the −2.4s, and had ROOM's "deep inside the band" row (t −1.81) come in at −2.3 instead, it would have been reported as noise wearing a suit rather than a discovery.
Three things about the surviving cuts, in the order that matters.
The two strongest splits are one variable, not two. Dealer gamma positive at spot and spot above the zero-gamma flip agree on 100.0% of pairs — the flip is by definition where volume-convention gamma changes sign. Reporting both as separate confirmations would have been double-counting one cut, which is why they carry nearly identical numbers.
It holds out of sample and strengthens. Inside the long-gamma regime: −3.27 bp (t −4.19, 296 sessions, 6,785 pairs) in sample, −5.97 bp (t −3.94, 124 sessions, 2,176 pairs) out. And it is not a rare corner — long gamma covers 78% of all pairs (8,961 of 11,443) on 420 of 471 sessions, which is the ordinary state of the book rather than a curiosity.
Time of day is a separate, weaker cut. Within long gamma the first third still carries most of the effect (−4.01 bp, t −4.10, 373 sessions) while the middle third is nothing (−1.39 bp, t −0.79). So session phase is not merely a restatement of the gamma regime, but on its own it is the lesser of the two.
The long-gamma result is worth slightly more than its t-statistic, for one reason that has nothing to do with statistics: it is the direction the mechanism predicts. Long-gamma hedging leans against moves, so a skew spike in that regime being followed by an advance that stops is consistent with the story; short gamma, where the same hedging amplifies, shows −0.71 bp, t −0.63. That split was not chosen to fit.
The verdict, in one sentence
The 25-delta risk reversal does not mark bottoms — the fear tail is followed by an advance that stops, 2.79 bp below a matched minute of the same session at 30 minutes (t −4.42, 11,492 pairs, 471 days), and the "calls bid marks a top" half of the claim is empty (−0.43 bp, t −1.28).
What it is not
Not costed, and not costable from this study. Nothing above touches a bid-ask spread, a fee or a fill. The largest effect measured anywhere in it, the −4 bp long-gamma cut, is still under a tenth of the single-episode spread — the dispersion of individual outcomes swamps the average by more than an order of magnitude, so a 2.79 bp mean is a statement about a large pile of minutes and never about the next one. No SPX round-trip cost is measured here, so no post-cost number is quoted; the honest phrasing is that the effect is smaller than the frictions it has not yet been charged.
Not a reversal. The signal minute's forward return is negative in the table, but its matched minute's is positive. What is being measured is a stalled advance relative to a comparable minute, not a turn. Any reading of it as "short the fear spike" is reading a difference as a level.
Not a mechanism. The regime split is observational. Dealer gamma sign correlates with a great many other things about a session — realised vol, time of day, the calendar, the composition of the tape — none of which were held constant. Consistency with the hedging story is not evidence for it.
Not a refutation of everything the claim implies. Two deciles of one metric on one index at three horizons were tested. A slower version, a different tail width, or an underlying with a different flow mix are separate questions with separate answers.
Not a full-archive result. 506 sessions, not the full published history: the 25-delta reading is returned only when both wings bracket 0.25 delta, because an extrapolated wing is a fabricated number and this series feeds a chart where a flat line reads as information.
Nothing on this page is investment advice; see the Terms.
Reproduce it
The book behind it is licensed. The 25-delta risk reversal here is computed from per-print SPXW trades and NBBO quotes from ThetaData — the SPXW trade + NBBO tick feed, a paid subscription. Reproducible does not mean free, and we cannot redistribute the tape.
What is free is the output series, which is all this test actually consumes.
Every finished session's JSON at https://gex.live/snapshots/YYYY-MM-DD.json
carries, aligned one-for-one with minutes (09:30 to 15:59 ET, 390 values) and
spot:
skew— the 25-delta risk reversal in vol points, put IV minus call IV on the 0DTE chain.skewpct— the same skew as a ratio: how much richer the 25-delta put is than the equidistant call, in percent. This is the field the test ranks, because it survives a change in the vol level — one vol point of skew means something different on a 12-vol day than on a 40-vol one.
Two series and a price column rebuild the whole study. The dates are the ones listed at /sessions.
import numpy as np, json d = json.load(open("2026-08-11.json")) s = np.array(d["spot"], float) # 390 one-minute prints sk = np.array(d["skewpct"], float) # 25-delta put/call ratio, percent # session-to-date percentile, EXPANDING -- never the whole day rk = np.full(len(sk), np.nan) for i in range(30, len(sk)): rk[i] = (sk[:i] < sk[i]).mean() * 100 H = 30 fwd = np.full(len(s), np.nan); fwd[:-H] = (s[H:] / s[:-H] - 1) * 1e4 # bp pri = np.full(len(s), np.nan); pri[H:] = (s[H:] / s[:-H] - 1) * 1e4 ok = np.isfinite(rk) & np.isfinite(fwd) & np.isfinite(pri) sig = np.flatnonzero(ok & (rk >= 90)) # the fear tail pool = np.flatnonzero(ok & (rk >= 25) & (rk <= 75)) # plainly not extreme # greedy nearest neighbour on the PRIOR return, no replacement, reject beyond 3 bp avail, diffs = list(pool), [] for i in sig: if not avail: break a = np.array(avail); j = int(a[np.argmin(np.abs(pri[a] - pri[i]))]) if abs(pri[j] - pri[i]) > 3.0: continue # drop rather than fake a match avail.remove(j); diffs.append(fwd[i] - fwd[j]) # per-session mean of diffs; then t across sessions -- CLUSTER BY DAY, never by pair
Run that over every date in the archive and the 30-minute row
of the matched table falls out. To reach the regime cuts you also need flip,
hold_lo and hold_hi from the same JSON — the gamma-sign split is
spot against the flip, and ROOM is the distance to the nearer band edge as a share of band
width. The terminal marks these minutes live as they occur; the layer exists to show where
the state happened and how each one turned out, not to tell anyone to do anything.
Part of gex.live research. Measured on the free session archive; every session is free to replay.