Does price at the cumulative-gamma trough predict a larger move?
63 sessions, 22,680 minute-observations, 1,002 visits: −0.08 pt
The idea is specific enough to implement, which is why it is worth testing. Take the per-strike gamma ladder, sort the strikes ascending, take the running sum, and find the strike at which that running sum is at its minimum. Call it the trough. The claim: when price reaches the trough, the move that follows is larger than usual.
It is not true on this sample, and the interesting part is not the null. It is that the pilot which motivated the test — six sessions, a clean +6.58 pt median spread — was six sessions chosen the way everyone chooses them.
The definition tested
Exactly as stated, with no room left for interpretation:
- Ladder. Per-strike gamma summed over expiries, convention-signed
(calls positive, puts negative) × open interest — the book the terminal's payload serves
as
go. Open interest is the 06:30 ET pre-open book of the same session, so the ladder at minute t uses nothing from after t. - Trough. Strikes sorted ascending,
cumsum,argmin. Recomputed every minute. - Visit. |spot − trough| ≤ 3 points. Three points is the operator's threshold, not a fitted one; the whole surface from 1 to 30 points is below.
- Outcome. The range of spot over minutes t+1 … t+30, strictly forward. The last 30 minutes of a session carry no outcome, leaving 360 observations per session.
The number
On 63 sessions and 22,680 minutes, price was within 3 points of the trough on 1,002 minutes across 34 sessions. The forward 30-minute range on those minutes minus the forward 30-minute range on the other 21,678:
| statistic | at the trough | elsewhere | difference |
|---|---|---|---|
| median forward 30m range | 13.12 pt | 12.74 pt | +0.38 pt [−2.17, +3.76] |
| mean forward 30m range | — | −0.08 pt [−2.77, +2.60] | |
| day-clustered t | — | −0.06 | |
| n | 1,002 visits | 21,678 minutes | 34 of 63 sessions |
Every interval on this page is day-clustered (CR1, 63 clusters); median differences carry a day-block bootstrap CI, 1,000 draws, days resampled with replacement.
The Welch t on the medians is −0.26. There is no effect, and the interval is not wide enough to be hiding a large one — more on what it can and cannot exclude at the end.
The pilot, and why six recent days is a specific way to fool yourself
The test exists because a pilot found the effect. On six sessions the median forward range at the trough was 21.51 against 16.05 elsewhere, Welch t +2.01, on 119 visits — a median spread of +6.58 pt. That reproduces. The visit count reproduces exactly (n = 119 on the same definition), and two control numbers that are properties of the non-event rows reproduce to the decimal: trailing 30-minute return elsewhere +1.87, and the quiet-tercile forward range elsewhere 10.99.
The six sessions were 2026-07-24 to 2026-07-31: the last six in the archive at the time. Not sampled, not held out. The six that were on screen.
| book / scope | window | n visits | med at trough | med elsewhere | Welch t |
|---|---|---|---|---|---|
| convention×OI, all expiries | pilot 6 sessions | 119 | 21.51 | 16.05 | +2.01 |
| convention×OI, all expiries | all 63 sessions | 1,002 | 13.12 | 12.74 | −0.26 |
| convention×OI, 0DTE only | pilot 6 sessions | 132 | 23.65 | 15.90 | +4.85 |
| convention×OI, 0DTE only | all 63 sessions | 901 | 17.07 | 12.69 | +8.66 (naive) |
The general point is worth stating carefully, because it is not the ordinary multiple-comparisons warning and it is not survivorship bias. It is this: the sessions you reach for first are the ones you have already watched.
An idea about intraday structure does not arrive from nowhere. It arrives because something was noticed — a level that price hung around on a day that moved, a shape that seemed to matter while it was on the screen. The most recent sessions are exactly the ones that supplied the noticing. When they are then used to check the idea, the check is run on the data that generated the hypothesis. The window is small enough that a handful of memorable minutes dominate it, and those minutes are memorable because the market was moving, which is the outcome variable.
Recency does not merely make the sample small. It makes it correlated with the thing being measured, in the direction of the claim, and it does so silently — nothing about "the last six days" looks like a choice. A random six-session draw from the same 63 would be honest and underpowered. Six recent sessions are underpowered and biased, and the bias has no error bar attached because it never entered the arithmetic.
The tell is available in advance, without more data: ask how the window was picked. "The most recent N" and "the ones I had open" are the same answer. Here, the pilot's +6.58 pt median spread becomes +0.38 pt at depth on its own definition, with a day-bootstrap 95% CI of [−2.17, +3.76].
The pilot's hourly detail dissolves the same way. Its strongest cell, +30.70 at 14:00, rested on 7 observations; at n = 113 that hour is +0.38 with t +0.14. Its weakest cell, −2.03 at 10:00, is now +2.48. Neither number was ever about the market.
Neighbouring definitions, so the null is not one specification
Three books were built — the measured/signed book, the convention×OI book, and the convention×volume book that a widely followed public dealer-gamma feed publishes — each in two scopes (all expiries in the ladder, and 0DTE only) and three strike bands. Eighteen combinations, all printed. The strike band does not matter at all: full chain, ±2.5% and ±1.0% agree to two decimals. The scope and the book do.
| book / scope | n visits | days w/ visit | med at trough | med elsewhere | mean diff | day-clust t | naive t |
|---|---|---|---|---|---|---|---|
| measured, all expiries | 796 | 33 | 9.52 | 12.90 | −4.19 | −2.97 | −10.64 |
| measured, 0DTE | 819 | 35 | 9.81 | 12.90 | −4.02 | −3.03 | −10.34 |
| convention×OI, all expiries (the claim) | 1,002 | 34 | 13.12 | 12.74 | −0.08 | −0.06 | −0.24 |
| convention×OI, 0DTE | 901 | 35 | 17.07 | 12.69 | +2.83 | +1.90 | +7.62 |
| convention×volume, all expiries | 2,780 | 63 | 13.41 | 12.70 | +0.17 | +0.20 | +0.76 |
| convention×volume, 0DTE | 2,776 | 63 | 13.40 | 12.70 | +0.11 | +0.14 | +0.52 |
The sign flips across books. In the measured book the trough predicts a smaller forward move, significantly so. In the volume book — the one that is actually published live, and the only one with a visit on all 63 sessions — it is flat zero. One of six families is positive, at t +1.90, which does not clear 2. That is what a multiple-comparison artefact looks like: a single positive cell among sign-inconsistent neighbours.
The rest of this post carries that best cell as well as the claim's own definition, because a null is only worth reading if the most favourable variant was pushed hardest.
The best cell, taken seriously and then taken apart
The 0DTE convention×OI trough gives +2.83 pt, day-clustered t +1.90, 95% CI [−0.10, +5.76] — median 17.07 against 12.69. Its naive t is +7.62. The gap is the whole story of overlapping windows: consecutive minutes share 29 of the 30 minutes of their outcome, the design effect is 16.2, and the effective n is 1,402 of 22,680. De-overlapped there are 307 distinct visits on 35 sessions, median visit length 2 minutes.
Every gate after that:
| gate | result |
|---|---|
| day-clustered SE | t +1.90, CI [−0.10, +5.76]; effective n 1,402 of 22,680 |
| residualisation | +1.87 raw on the common sample → −0.06 (t −0.06) once trailing 30m and 120m realised vol enter; trailing 30m RV alone takes it to −0.09 |
| day fixed effects | +0.70 (t 0.89) — and day FE absorb the whole day-level vol regime by construction |
| permutation on residuals | 82.4th percentile (free shuffle), 68.0th (block shift) of its own within-day null |
| chronological split | first 60% +5.44 (t 3.35), last 40% +1.43 (t 0.85); May+June +5.70 (t 3.54), July +1.30 (t 0.78) |
| drop 5 most helpful days | +1.61 (t 1.16) |
| direction | none: 48.2% up against a 52.4% baseline, t +0.58 |
One control does the damage: the trailing 30-minute realised vol. The pilot believed it had already handled vol clustering, because it split on terciles of trailing range and found the effect strongest in the quiet tercile. Trailing range is a coarse statistic. The sum of squared minute returns over the same window is not, and it is the one that absorbs the signal.
Threshold and horizon, unfitted
The 3-point threshold is not a choice that rescues or ruins anything, which is the point of printing the whole surface. For the 0DTE cell the effect is smooth in the threshold and peaks at 10 points (t +2.47) — a 20-point-wide window around the level, which is a neighbourhood, not a level. For the claim's own all-expiry definition the surface is flat at zero throughout: max |t| = 0.52.
| threshold | n visits | mean diff | day-clust t | 95% CI |
|---|---|---|---|---|
| 1 pt | 294 | +3.10 | +1.69 | [−0.51, +6.71] |
| 2 pt | 578 | +2.83 | +1.88 | [−0.12, +5.78] |
| 3 pt (as claimed) | 901 | +2.83 | +1.90 | [−0.10, +5.76] |
| 5 pt | 1,484 | +2.93 | +1.98 | [+0.03, +5.84] |
| 10 pt | 2,921 | +3.19 | +2.47 | [+0.66, +5.72] |
| 15 pt | 4,594 | +2.54 | +2.14 | [+0.22, +4.87] |
| 30 pt | 9,452 | +1.27 | +1.05 | [−1.10, +3.65] |
0DTE convention×OI trough. The claim's own all-expiry definition never leaves zero on this surface.
And the outcome definition is not load-bearing either. These are all raw variants — the ones that look best, before section 6's controls — and none of them changes the conclusion:
| variant | 0DTE convention×OI: mean diff (t) | all-expiry (the claim): mean diff (t) |
|---|---|---|
| per-second forward range | +3.427 (+2.14) | −0.016 (−0.01) |
| absolute forward 30m return | +1.920 (+1.48) | −0.060 (−0.06) |
| log forward range | +0.235 (+2.46) | +0.034 (+0.37) |
| drop 15:00–15:29 | +3.184 (+2.17) | +0.055 (+0.04) |
| drop first 30 minutes | +2.569 (+1.59) | −0.689 (−0.54) |
| May + June only | +5.696 (+3.54) | — |
| July only | +1.302 (+0.78) | — |
The placebo that settles it
Take the same trough and slide it a fixed distance. The geometry survives; the meaning does not.
| shift | n visits | mean diff | day-clust t | 95% CI |
|---|---|---|---|---|
| −50 pt | 343 | +5.52 | +2.58 | [+1.33, +9.72] |
| −25 pt | 816 | +1.59 | +1.44 | [−0.57, +3.74] |
| −10 pt | 834 | +3.84 | +2.85 | [+1.20, +6.48] |
| −5 pt | 897 | +3.69 | +2.48 | [+0.77, +6.62] |
| 0 — the trough itself | 901 | +2.83 | +1.90 | [−0.10, +5.76] |
| +5 pt | 843 | +2.30 | +1.47 | [−0.76, +5.36] |
| +10 pt | 1,033 | +0.66 | +0.44 | [−2.33, +3.65] |
| +50 pt | 699 | −5.26 | −4.40 | [−7.60, −2.91] |
A level 5, 10 or 50 points below the trough produces a larger and more significant effect than the trough. Whatever is being detected is regional — spot is somewhere in the zone at or under the crossover, which is where it tends to be when it has just rallied into a moving market — and not a property of the strike where the cumulative curve bottoms. Other levels in the same book behave the same way, in both directions: the largest-absolute-gamma strike gives −3.38 (t −4.39), the argmax of the same cumulative curve gives −2.62 (t −2.11), and the nearest 25-point round number, which uses no options input whatsoever, gives +0.73 (t +1.51).
What is actually happening at the trough
The mechanism can be named rather than guessed, because the state variables at the moment a visit begins are observable:
| state variable | mean at arrival | mean over all minutes | ratio |
|---|---|---|---|
| trailing 30m range | 22.33 | 16.12 | 1.39 |
| trailing 120m RV | 29.99 | 24.02 | 1.25 |
| implied 30m range (0DTE straddle) | 23.45 | 21.95 | 1.07 |
| forward 30m range | 20.25 | 15.56 | 1.30 |
| minutes from open | 116.62 | 179.50 | 0.65 |
The market is already moving when price reaches the trough — 39% busier than normal — and the next 30 minutes are 30% busier, which is less elevated than the preceding 30. That is vol clustering with mild mean reversion. The trough is a lagging indicator of the vol state with a 3% duty cycle, and it under-performs the state variable it proxies.
One result from the exercise is worth keeping, and it is not about gamma. Controlling only for what the 0DTE at-the-money straddle implies over the next 30 minutes, the trough is incremental: +2.90 pt, clustered t +3.05, ΔR² +0.0027. A 0DTE straddle prices the whole remaining session, so it is a badly-averaged 30-minute forecast, and a short-horizon state variable can beat it. But trailing realised vol beats it by far more — ΔR² +0.0698 over the straddle, against the trough's +0.0027 — and once trailing RV is in, the trough adds +0.0002 (t +0.96). Everything the trough knows that the straddle does not, trailing RV also knows, on every minute instead of 3% of them, with no gamma data at all.
Even taken at face value, the level is not usable
The 0DTE convention×OI trough has an intraday standard deviation of 114.4 points — the argmin of that cumulative curve teleports between distant local minima as the gamma kernel narrows through the session. On 28 of 63 sessions price never comes within 3 points of it at all, and when it does the visit lasts a median of 2 minutes. Only the convention×volume book produces a stable, always-visited level, and that book shows +0.17 pt with t +0.20.
And the tradeable framing closes it. Realised 30-minute range divided by what the straddle implied is 0.84 at the trough and 0.69 elsewhere — both far below 1. A 30-minute straddle bought at the trough still loses; the trough merely loses less. To monetise that you would have to be selling, and the signal would be telling you to sell less, which is the least valuable form of the information.
What it does not say
A null of this size rules out a large effect on this definition, not every effect. Be precise about which:
What is excluded. On the claim's own definition the day-clustered 95% interval is [−2.77, +2.60] points against a base of 12.74 points elsewhere. An effect larger than about +2.6 points — roughly a fifth of the typical forward 30-minute range — would have been visible in this sample and is not there. For the most favourable neighbouring definition the interval is [−0.10, +5.76], which is a weaker statement: an effect up to about +5.8 points is not excluded by that interval alone, and it is the controls, the placebos and the out-of-sample halving that argue against it, not the width of the CI. There is no formal power calculation behind either sentence; the intervals are the intervals.
What is not excluded. Anything smaller than the interval. A different threshold rule, a different horizon than 30 minutes, a different signing convention for the book, a level built from a different ladder, or the same idea in a different regime — this is 63 sessions, 2026-05-01 to 2026-07-31, one market. The measured book's negative result (−4.19 pt, t −2.97: forward moves are smaller at that book's trough) is a finding of the same sample and carries the same caveat.
What this is not. It is not a claim that dealer positioning is irrelevant, and it is not a critique of anyone's product. It is one precisely stated level tested on its own terms. Nothing on this page is investment advice; see the Terms.
Reproduce it
The per-strike ladder is published with every finished session, so the trough can be
rebuilt and the whole test re-run on price alone. Each session's JSON at
https://gex.live/snapshots/YYYY-MM-DD.json carries:
minutesandspot— the one-minute series, 09:30 to 15:59 ET (390 points).frames— one ladder frame everystepminutes (stepis 5). Each frame hasstrikes(ascending),go— convention-signed × open-interest gamma summed over all expiries, in $M per point — andgo0, the same thing on the 0DTE slice only.idxis the frame's index intominutesandspot, andtis its clock time.
The trough is one line on go; the visit test and the outcome are three
more:
import json, urllib.request, numpy as np d = json.load(urllib.request.urlopen("https://gex.live/snapshots/2026-08-11.json")) spot = np.array(d["spot"], float) # 390 one-minute prints for f in d["frames"]: # one frame every d["step"] = 5 minutes k = np.array(f["strikes"], float) # ascending g = np.array(f["go"], float) # convention-signed OI gamma, all expiries trough = k[np.argmin(np.cumsum(g))] # f["go0"] for the 0DTE-only slice i = f["idx"] # index into d["minutes"] / d["spot"] if abs(spot[i] - trough) <= 3 and i + 30 < len(spot): fwd = spot[i+1:i+31] # strictly forward, 30 minutes print(d["day"], f["t"], trough, fwd.max() - fwd.min()) # Then: the same forward range on every NON-visit minute is the comparison set. # Cluster by session, not by minute — consecutive minutes share 29 of the 30 # minutes of their outcome, so the design effect on this panel is 15.0 for the # all-expiry book and 16.2 for the 0DTE one.
The dates are the ones listed at /sessions. Two differences from the study, both stated so the reader is not surprised: the payload serves the ladder on a 5-minute frame grid, while the panel here recomputes it every minute, and the study's spot is the pipeline's put-call-parity spot rather than the published index series. Neither changes any conclusion, but visit sets will differ by a handful of minutes — on the pilot window a sweep of all 18 ladder definitions × six thresholds × both outcome samplings matched the pilot's visit count exactly (n = 119) without matching all four of its headline numbers at once.
What is not reproducible for free is the minute-by-minute rebuild and the implied-move control. Those come from licensed tick data — the SPXW trade and NBBO tick feed from ThetaData, which a reader would have to subscribe to in order to redo that part. Reproducible does not mean free. The published 5-minute ladder and the one-minute price series are enough to rebuild the trough, the visits and the outcome, which is the claim itself.
Part of gex.live research. Measured on the free session archive; every session is free to replay.