How fast does a dealer-gamma level go stale?

Drift, inversion and the value of a fresh read — 1,077 sessions, 2022-04-14 to 2026-08-21

Published 2026-08-24 · Sample SPX regular session, 2022-04-14 to 2026-08-21, 1,077 trading days · Levels zero-gamma flip, hold-band edges, band width (measured dealer book) · Status measurement, not a signal

The published hold-rate study reads each level once, at 10:00 ET, and judges it against the rest of the day. That is the right way to avoid lookahead, and it leaves a question standing behind it: the terminal recomputes these levels every minute, so how far does a level actually travel between the read and the close — and if a reader waits and reads it later, does the later read hold better for any reason other than being further from spot?

Three measurements, each with the session as the unit of observation. All of them are judged on minute closes, not the per-second tape, so every hold rate on this page is an upper bound — a spike that pierced a level and came back inside the minute is invisible here. The bias is identical in the levels and in their baselines, so the comparisons survive it even though no single rate does.

The sample

1,088 archived session files, 1,086 after the contiguity walk drops two isolated seed days, 1,077 after nine shortened sessions are dropped for not having the full 390 one-minute rows: 2022-11-25, 2023-07-03, 2023-11-24, 2024-07-03, 2024-11-29, 2024-12-24, 2025-07-03, 2025-11-28, 2025-12-24. Nothing else is filtered, and every further exclusion below is counted where it happens.

1. The flip is not a stable object

Read the flip at 10:00, then read it again later the same day and measure the distance between the two, in index points and in units of the move the tape was already making at 10:00 — σ = rv30(10:00) × √(minutes elapsed), the same scaler as the touch surface. That second unit is the one that matters: it asks whether the level moved more or less than price itself did over the same stretch.

flip, read atsessionsmedian driftmean driftmedian drift / σmean drift / σno flip at T
11:001,02114.22 pts23.31 pts0.847σ1.355σ22
12:0099517.29 pts28.59 pts0.752σ1.166σ48
13:0099520.76 pts34.49 pts0.761σ1.120σ48
14:0098324.88 pts38.99 pts0.789σ1.108σ60
15:001,00029.57 pts43.27 pts0.817σ1.115σ43
close1,03850.86 pts65.96 pts1.225σ1.494σ5

Drift is |flip(T) − flip(10:00)|, absolute, so it cannot cancel. "No flip at T" counts sessions where the gamma ladder had no root at that minute; the 10:00 base exists on 1,043 sessions.

By the close the flip sits a median 50.9 points from where it was at 10:00 — 1.23σ, which is to say the level moved further than the tape did. Measured against the session's own high-low range the drift is 133.2% on average (se 3.4) and 110.9% at the median: on more than half of sessions the flip travelled further between 10:00 and the close than SPX travelled from its own high to its own low that day.

And it does not merely move, it changes sides. On 23.5% ± 1.3% of sessions (n 1,038) the closing flip sat on the opposite side of spot(10:00) from where the 10:00 flip sat. On nearly a quarter of days the level's directional meaning inverted within the session.

Which piece is steadiest

The same measurement for the hold band's two edges and — separately — for the band's width, which is what the band actually claims: the corridor in which the measured book stays long gamma.

median drift from 10:0011:0012:0013:0014:0015:00closeclose, in σ
hold band — upper edge9.0210.9812.4615.0119.3833.59 pts0.858σ
zero-gamma flip14.2217.2920.7624.8829.5750.86 pts1.225σ
hold band — lower edge13.5116.6919.4023.5028.2947.81 pts1.177σ
hold band — width16.1121.1527.6434.5241.8284.33 pts2.067σ

Points, median over sessions; n falls from 971 at 11:00 to 886 at the close as the band stops existing on some minutes (exclusions: 55, 92, 119, 139, 123, 140). The band is defined on 1,026 sessions at 10:00 and brackets spot on all 1,026 of them, so at most one of its two edges can invert.

The band's width is the least steady quantity of the four: a median 84.3 points of change by the close, 2.07σ, and 184.3% of the day's high-low range at the median (mean 214.5%, se 5.2). Against a median opening width of 140.3 points that is a corridor being redrawn, not adjusted. The two edges invert their side of spot(10:00) on 17.6% ± 1.3% and 17.7% ± 1.3% of sessions.

The steadiest piece is the band's upper edge: 33.6 points, 0.858σ, 77.1% of the day's range at the median — a third less drift than the flip in σ terms. That is the same edge, and the only one, that the published hold-rate study found beating its distance baseline (+3.0 pp, ±0.7, against every other level landing within one standard error of its baseline). The level that moves least is the level that outperforms. Two independent measurements pointing at the same edge is a reason to keep testing it; it is not evidence of a mechanism, and neither study was designed to find one.

The shape of the drift

Most of the migration is bought in the first hour. In σ units the flip is already 0.847σ away by 11:00, then 0.752σ, 0.761σ, 0.789σ, 0.817σ at noon, 13:00, 14:00 and 15:00 — flat, or very slightly falling, for five hours. In points the drift keeps growing; normalised by the move price itself was making, it does not. Then it jumps at the close: 1.225σ. The band width does the same thing more sharply, 0.966σ at 11:00, flat through the afternoon, 2.067σ at the last minute. A level's day is a fast repricing in the first hour, a long stretch in which it drifts no faster than the tape, and a settlement-hour jump.

Which side of spot the flip sits on

Before any hold rate is quoted, one fact has to be on the table: at 10:00 the flip is below spot on 97.6% of sessions, a mean of 86.8 points below (median 74.9 below, n 1,043). It is a downside level almost always, and drifts toward being one less reliably as the day goes on — below spot on 94.5% at 11:00, 93.2% at noon, 91.0% at 13:00, 89.6% at 14:00, 85.4% at the close.

So an aggregate "the flip holds 88.8% of the time" is almost entirely a statement about downside levels, and hiding the split would misreport it:

flip read atabove spot: nheldbelow spot: nheld
10:002429.2%1,01690.2%
11:005446.3%98191.3%
12:006546.2%94293.6%
13:008949.4%91892.9%
14:0010365.0%89695.1%

A flip above spot is a rare configuration and a weak level: on the 24 sessions where it happened at 10:00 it held on seven of them.

2. A later read holds better — and it is entirely distance

Take the flip at time T, and judge it only over the minutes after T: it held if price never closed on the other side of it, "other side" fixed by spot(T). A later read gets a shorter window and a different distance to spot, so the raw rates are not comparable across T on their own. Both are reported: dist_sig is |flip(T) − spot(T)| divided by rv30(T) × √(minutes left), and the placebo is described in the next section.

flip read atsessionsminutes judgedheldsedistance, mean σmedian σplaceboexcess95% CI
10:001,04035988.8%1.02.069σ1.914σ88.8%−0.03 pp[−1.31, +1.29]
11:001,03529989.0%1.02.720σ2.568σ88.8%+0.15 pp[−1.02, +1.39]
12:001,00723990.6%0.93.526σ3.241σ89.6%+0.92 pp[−0.28, +2.15]
13:001,00717989.1%1.04.253σ3.889σ89.1%+0.01 pp[−1.19, +1.26]
14:0099911992.0%0.95.064σ4.539σ91.0%+0.96 pp[−0.13, +2.05]

Excess is hold − placebo; the interval is a 2,000-resample bootstrap over sessions (bootstrap se 0.67, 0.61, 0.62, 0.63, 0.56 pp). Exclusions at 10:00: 34 sessions with no flip (24 censored up, 10 censored down), 0 with no rv30, 3 with spot within 0.02% of the level; at 11:00 through 14:00 the no-flip counts are 39, 64, 65, 75.

The raw hold rate does rise with a later read: 88.8% at 10:00 to 92.0% at 14:00. So does the distance — 2.069σ to 5.064σ, because the window shrinks faster than the level moves toward spot. And so does the placebo, 88.8% to 91.0%, in step. Every excess confidence interval contains zero. A level read at 14:00 holds better than one read at 10:00 for exactly the reason that any mark 5σ away holds better than one 2σ away. Reading later does not buy a better level; it buys a shorter day.

A placebo that cannot fail is not a placebo

This is worth stating carefully, because the obvious construction is vacuous and the study measured it explicitly in order to say so.

The natural placebo is: put a fake level at the same signed σ-distance from spot(T) as the real flip, and score it on the same session's forward path. But spot(T) + sign × dist_sig × σ is the flip — the same number, arrived at by rearranging its own definition. Its hold indicator is identical on every session, so the excess is exactly 0.000 at every read time, with no sampling error at all. That is reported in the study output as an identity, not a result. A comparison that cannot come out any other way teaches nothing about the level.

The informative version keeps the session's own distance and its own direction, and takes the forward path from somewhere else: for session i, the placebo rate is the share of the other sessions j read at the same minute whose forward path, measured from their own spot(T) in σ units, never reached dist_sig(i) in direction sign(i). Leave-one-out, so a session never grades itself; split by direction, because selloffs travel further in σ than rallies; bootstrapped 2,000 times over sessions, with the placebo ECDF rebuilt inside each resample so the interval carries the placebo's own sampling error rather than treating it as known. What is left over — hold minus placebo — is what the flip gets for being the flip rather than a mark at that distance and side. On this sample, at every read time, that is zero.

3. Re-reading a level buys nothing measurable

The two results above concern different windows. This one fixes the window and varies only the level's age. Take the window [10:00+k, close]. Score it twice: once with the flip as it stood at 10:00 (age k minutes, stale) and once with the flip read at 10:00+k (age zero, fresh), with "other side" defined for both by spot at the window start. Identical minutes judged, identical sample, one difference — how old the number is.

k (minutes stale)window startsminutes judgedsessionsmean |fresh − stale|held, staleheld, freshfresh − stalesesessions disagreeing
010:003591,0400.00 pts88.8%88.8%0.00 pp0.000
1510:153441,02614.12 pts89.2%90.0%+0.78 pp0.6848
3010:303291,02318.57 pts89.4%90.7%+1.27 pp0.7559
6011:002991,01423.29 pts90.2%89.3%−0.99 pp0.8778
12012:0023998628.53 pts91.4%91.0%−0.41 pp0.8672
24014:0011997639.00 pts94.4%92.4%−1.95 pp0.9281

Exclusions per k (no flip at 10:00 / no flip at 10:00+k / spot within 0.02% of either level): 34/0/3, 34/10/7, 34/16/4, 34/22/7, 34/48/9, 34/60/7. "Sessions disagreeing" counts days where the two ages reached opposite verdicts.

The raw difference flips sign with k. Freshness looks worth +1.27 pp at k=30 (1.7 se) and worth −1.95 pp at k=240 (2.1 se) — the only difference past two standard errors runs the wrong way, with the four-hour-old level holding better. That is the giveaway: a stale level has drifted further from spot (mean 39.00 points of drift at k=240; distance 7.16σ stale against 5.10σ fresh), and a far level breaks less often. The raw column flatters age, not freshness.

So score both ages against the same cross-session distance-matched placebo:

kplacebo, staleplacebo, freshexcess, staleexcess, freshexcess differencese
088.8%88.8%−0.03 pp−0.03 pp0.00 pp0.00
1588.9%89.1%+0.24 pp+0.85 pp+0.60 pp0.60
3089.2%89.5%+0.20 pp+1.26 pp+1.06 pp0.67
6089.9%89.0%+0.31 pp+0.20 pp−0.11 pp0.72
12091.2%89.9%+0.15 pp+1.12 pp+0.97 pp0.70
24094.0%91.3%+0.36 pp+1.15 pp+0.79 pp0.77

Once distance is taken out, the sign stops flipping — freshness is worth between −0.11 and +1.06 pp — and nothing reaches 1.6 standard errors at any k. Plainly: a level 39 points of drift out of date holds the remainder of the session as well as one computed that minute. Whatever the value of recomputing a dealer book every minute is, on this sample it is not in the hold rate.

Where this agrees with the published hold-rate study

The hold-rate post scored the 10:00 flip against the empirical touch surface — pooled minute-observations across sessions, matched on side, σ-distance and minutes left — and found an excess hold of −0.1 pp (±1.0). This study scores the same level at the same minute against a different baseline: a leave-one-out, direction-split, bootstrapped cross-session ECDF of forward reach, with the session as the unit rather than the minute. It finds −0.03 pp, 95% CI [−1.31, +1.29]. Two constructions, two error models, the same answer to two decimal places of a percentage point: the flip's hold rate at 10:00 is what its distance predicts and nothing more. A reader who took the first result should read this as it is meant — an independent confirmation, not a restatement.

What it does not say

It does not say the flip is useless, and it does not say re-computation is wasted. The staleness test measures exactly one property — does price close on the other side of this number — and it holds the two things a re-read is mostly for fixed by construction: the fresh level's distance and its side both enter the placebo, so what the test reports is what remains after they have been removed. A trader who re-reads the level at noon is reading a different distance and, on 23.5% of days, a different direction; this page says that once you know the distance and the side, the age of the number adds nothing measurable to the crossing question. It does not say the distance and the side were worthless information.

It does not measure profitability, position sizing, or anything about the per-second path. Hold rates are upper bounds throughout. The sample is a single regime-spanning window — 2022's bear market through mid-2026 — and the measured hold band exists only for as long as the measured book has; every number here is a statement about these 1,077 sessions, not a law. Nothing on this page is investment advice; see the Terms.

One structural caveat on section A: drift is absolute, so a level that wandered and came back scores the same as one that walked away. That choice is deliberate — a reader acting on a level at 11:00 is exposed to where it is at 11:00, not to its round trip — but it means the migration table overstates net repositioning and understates nothing.

Reproduce it

No licensed data is needed for this one. The dealer book behind the levels is rebuilt from a licensed SPXW trade + NBBO tick feed (ThetaData), but the outputs this study reads — the per-minute flip, hold_lo, hold_hi and spot arrays — are published free in every finished session's JSON at https://gex.live/snapshots/YYYY-MM-DD.json, aligned with minutes. Index 30 is 10:00 ET and index −1 is the close. One session, one line, to see the shape of it:

pythonreproduce
import json, urllib.request
j = json.load(urllib.request.urlopen("https://gex.live/snapshots/2026-08-11.json"))
print(j["minutes"][30], j["spot"][30], j["flip"][30], j["flip"][-1])
# 10:00 7762.89 7652.63 7720.75      -> the flip moved 68.1 points after 10:00

The median migration of section A, in full. The date list is the one at /sessions, restricted to 2022-04-14 through 2026-08-21 with the nine short sessions named above removed; files are about 7 MB each, so fetch them sequentially:

pythonreproduce
import json, statistics, urllib.request
SHORT = {"2022-11-25", "2023-07-03", "2023-11-24", "2024-07-03", "2024-11-29",
         "2024-12-24", "2025-07-03", "2025-11-28", "2025-12-24"}
d = []
for day in dates:                          # 2022-04-14 .. 2026-08-21, from /sessions
    if day in SHORT: continue              # shortened sessions: fewer than 390 minutes
    u = "https://gex.live/snapshots/" + day + ".json"
    j = json.load(urllib.request.urlopen(u))
    a, b = j["flip"][30], j["flip"][-1]     # 10:00 and the close
    if a is None or b is None: continue     # 5 sessions have no flip at one of the two
    d.append(abs(b - a))
print(len(d), round(statistics.median(d), 2), round(sum(d) / len(d), 2))
# 1038 50.86 65.96
# hold_lo, hold_hi and the width (hold_hi - hold_lo) are the same three lines.
# sigma units: rv = trailing-30-minute std of diff(log spot) * spot, at minute 30
#   drift / (rv[30] * sqrt(minutes elapsed))
# inversion: sign(flip[30] - spot[30]) != sign(flip[-1] - spot[30])

Sections B and C need one more array each — the forward running max and min of spot from each minute, which is two lines of numpy — and the placebo is the leave-one-out ECDF described above. Each session page at /sessions states that day's levels in prose, and the MCP server returns the same summaries, which is enough to spot-check any single row against a day you remember.

Part of gex.live research. Measured on the free session archive; every session is free to replay.