Does open interest pin the index?
Two published claims, seven pre-specified tests, 1,072 SPX sessions
Two papers ask the same question with opposite answers. Golez & Jackwerth (2012), "Pinning in the S&P 500 Futures", Journal of Financial Economics 106(3), 566–585, found that S&P 500 futures are pulled toward the nearest option strike on serial-expiration days: 13.56% of futures settle within $0.25 of a strike against 10% expected under independent draws, about 16% in a tighter 1998–2009 subsample. Elms (2026), "From Pinning to Amplification" / "Does Options Open Interest Pin the Underlying? Five Tests on 2,294 Trading Days" (Zenodo DOI 10.5281/zenodo.19540116, also SSRN 6564078), ran five tests on 2016–2025 and found the pinning gone. What survived was the opposite: daily range by ATM open-interest tercile of 0.889% / 0.999% / 1.033%, monotonically increasing, high-OI days 16.2% wider, t = −3.67, p = 0.0003 — his only significant cell out of five.
Attraction or amplification: both are claims about what the open interest sitting at a strike does to the index. Elms used SPY open interest as a proxy for SPX dealer positioning. This replication uses SPX open interest directly, on the 0DTE book, strike by strike. Removing that proxy step is the reason the work is worth redoing, and it is where the two studies part company.
First, the lookahead check
Every result below is a statement about open interest, so before any of them: the open interest used is the book carried into the session, never the book left after it traded. The daily open-interest file is stamped for date D, published around 06:30 ET on D, and reflects the close of D−1 trading; the production surface builder labels that column "OI (start of session, constant intraday)". A separate column holds the D+1 book — open interest after the session — and this study never reads it. That is the only way a pinning test can mean anything: if the OI were measured at the end of the day it would already know where price went, and every test would confirm itself.
Golez & Jackwerth: the clustering is not there
The test is theirs, simplified to what the data supports: how often does SPX close within $0.25 of a listed strike? The strike grid is derived rather than assumed — the modal gap between listed 0DTE strikes within ±1% of the close is 5 points on 1,069 of the 1,072 sessions and 10 points on the other 3. Under a close uniformly distributed over a 5-point grid, the chance of landing within $0.25 of a strike is 2 × 0.25 / 5 = 10%.
| window | sessions | closes in window | share | expected under a uniform close | ratio | binomial p |
|---|---|---|---|---|---|---|
| within $0.25 of a strike | 1,072 | 110 | 10.26% | 10.00% | 1.026 | 0.76 |
| within $1.00 of a strike | 1,072 | 435 | 40.58% | 40.00% | 1.014 | 0.71 |
| Golez & Jackwerth, futures settle | — | — | 13.56% | 10.00% | 1.36 | — |
The mean distance from the close to the nearest listed strike is 1.2479 points against a uniform expectation of 1.2500 — a match to two thousandths of a point. Whatever the SPX cash index does into the close, it does not know where the strikes are. A ratio of 1.026 with p = 0.76 is not a weak effect; it is the absence of one, measured precisely enough that a 36% excess would have been impossible to miss.
The deviation that bounds this. Golez & Jackwerth measured a futures settlement price. This study holds no futures and measures the SPX cash index at 15:59 ET, from put-call parity. SPX index options are cash-settled, so there is no futures settle anywhere in the data. Any pinning that lives specifically in the futures basis is invisible here, and this table cannot refute it. What it can say is that the cash index, in 2022–2026, closes as if the strike grid did not exist.
Elms's four nulls replicate
Four of his five tests found nothing. All four find nothing here too, and the agreement is close enough to be worth putting side by side. His column is 2,294 days of SPY open interest; ours is 1,072 SPX sessions of SPX open interest.
| test | Elms (2026), SPY OI | his p | this study, SPX OI | our p |
|---|---|---|---|---|
| close nearer the max-OI strike than the open | 44.9% toward | 0.40 | 48.2% toward | 0.26 |
| …and does it strengthen with OI (high vs low half) | 47.5% vs 42.2% (+5.3 pp) | 0.19 | 50.2% vs 46.3% (+3.9 pp) | 0.199 |
| minutes spent within 0.10% of the max-OI strike | no difference | — | 0.60% vs 1.20% of minutes | 0.083 |
| final-hour convergence toward the max-OI strike | 43.0% vs 45.0% | 0.78 | 50.7% vs 47.6% | 0.30 |
Two details are worth pulling out. On the first test, restricting the max-OI strike to within ±5% of the open — the legacy open interest parked at far round levels is otherwise the "max" on many days — moves the toward share to 47.7%, p = 0.13. Still nothing.
On the third, the mean shares are small and their difference has the wrong sign for pinning, but the vivid number is the median: the median session in both the high-OI and the low-OI half spends 0% of its minutes within 0.10% of the max-OI strike. The typical session does not visit that strike at all. Whatever gravity is supposed to hold price near it, the tape shows a majority of days spending no time there.
The one significant result reverses
Elms's surviving finding was amplification: high ATM open interest, wider daily range, monotonic across terciles. On SPX open interest the sign flips, and it flips hard.
| ATM-OI tercile | sessions | ATM OI range (contracts) | Elms, daily range % | this study, daily range % | median |
|---|---|---|---|---|---|
| low | 357 | 0 – 1,889 | 0.889 | 1.360 | 1.193 |
| mid | 357 | 1,899 – 3,249 | 0.999 | 0.962 | 0.811 |
| high | 358 | 3,252 – 70,607 | 1.033 | 0.977 | 0.785 |
| high vs low | — | — | +16.2% wider, t = −3.67, p = 0.0003 | −28.1% narrower, t = +6.49, p = 1.7e−10 | — |
High ATM-OI sessions are 28.1% narrower, not 16.2% wider, at a t-statistic roughly ten times the size of his and with the opposite sign. The Mann-Whitney version gives p = 6.3e−20, so it is not a tail artefact. The pattern is not monotonic on our side either — almost all of the effect is the low tercile standing apart, with mid and high nearly equal.
Why the sign flips: low ATM OI marks a gap
A contradiction is worth little; a mechanism is worth something. The diagnostics point at one, and it is a property of the measure rather than of the market.
"ATM open interest" here — and in Elms — is the open interest at the single listed strike nearest today's open. But that strike carries yesterday's open interest. On a gap day, the strike nearest this morning's open was not near the money yesterday, so nobody had built a position there and it holds very little OI. And gap days are wide-range days. Low ATM OI therefore does not measure thin positioning; it marks a gap, and the range result follows from that.
The three rank correlations say exactly this:
| Spearman correlation | value |
|---|---|
| ATM open interest vs daily range % | −0.318 |
| |overnight gap| vs ATM open interest | −0.198 |
| |overnight gap| vs daily range % | +0.331 |
If the reversal were an artefact of a knife-edge single-strike definition on our side, then widening the definition should weaken it. It does the opposite — every widening makes the negative relation stronger:
| definition of "ATM open interest" | range %, low / mid / high tercile | high vs low | t | p |
|---|---|---|---|---|
| single strike nearest the open (primary) | 1.360 / 0.962 / 0.977 | −28.1% | 6.49 | 1.7e−10 |
| terciles formed within each calendar year | 1.238 / 1.044 / 1.017 | −17.9% | 3.77 | 0.00018 |
| three strikes around the ATM strike | 1.379 / 0.998 / 0.922 | −33.1% | 8.20 | 1.4e−15 |
| all 0DTE OI within ±1% of the open | 1.508 / 0.974 / 0.818 | −45.8% | 12.56 | 1.3e−31 |
De-trending by forming the terciles inside each calendar year — the book grows over the sample — keeps the sign and about two thirds of the size. The two genuinely wider OI definitions push the gap from −28% to −33% and −46%. So the negative sign is robust on SPX, and the likely reading of Elms's positive sign is not that SPY behaves differently but that a knife-edge OI measure is a noisy proxy for a gap, and which way the noise points depends on the underlying and the OI source.
Divide by the volatility already on the tape
The honest objection to everything above is that a raw daily range is mostly a volatility level. So the same test again, with the range measured from 10:00 to the close and divided by what the tape had already realised by 10:00: the trailing 30-minute realised move per minute, projected over the minutes remaining. A value of 1.0 means the session travelled exactly as far as its own early volatility implied.
| ATM-OI tercile | sessions | raw range % | range ÷ own volatility | SE |
|---|---|---|---|---|
| low | 357 | 1.360 | 1.117 | 0.028 |
| mid | 357 | 0.962 | 1.180 | 0.035 |
| high | 358 | 0.977 | 1.251 | 0.040 |
Scaled, the residual points weakly Elms's way: high-OI sessions realise +12.0% more range than their own volatility predicts, +0.206 sigma, t = −2.75, p = 0.0061 — clearing the Bonferroni threshold below by a hair. It should not be leaned on. It survives one alternative OI definition (three strikes: +11.8%, p = 0.0040) and evaporates on the other (±1% band: −2.3%, p = 0.55, sign flipped). One of three variants disagreeing is not a finding; it is a hint.
The general lesson is the one worth keeping: the conditioning variable decides the sign. The raw range says high OI means narrower by 28%; the volatility-scaled range says high OI means wider by 12%; they are the same 1,072 sessions and the same open interest. Any raw-range result conditioned on something correlated with the volatility level is a volatility result wearing a different label — which is what the ATM-OI measure turns out to be, through the overnight gap.
An independent reproduction, on a different conditioning variable
The same volatility-scaled range, cut by day-of-expiry instead of by open interest: the 49 monthly third-Friday sessions realise 14.5% less range than their own volatility predicts (scaled 1.018 vs 1.190 on the other 1,023 sessions, t = −2.52, p = 0.014, −0.303 sigma). A companion measurement on a different code path put the same effect at −15% and −0.180 sigma, t = −2.74. Two independent constructions landing on the same sign and roughly the same size is the sort of agreement that is worth more than either number alone — and note that it is a compression effect on expiration days, sitting next to a null on pinning. Range and direction are separate questions.
Seven tests, one threshold
Seven tests were pre-specified. Bonferroni at α = 0.05 / 7 = 0.00714. Raw p-values, and which cleared:
| test | raw p | p < 0.05 | p < 0.00714 |
|---|---|---|---|
| T1 pinning, close vs open | 0.2584 | no | no |
| T2 pinning by open interest | 0.1993 | no | no |
| T3 daily range by ATM-OI tercile | 1.7e−10 | yes | yes |
| T4 minutes near the max-OI strike | 0.0831 | no | no |
| T5 final-hour convergence | 0.2990 | no | no |
| T6 volatility-scaled range by ATM-OI | 0.0061 | yes | yes |
| T7 close within $0.25 of the strike grid | 0.7600 | no | no |
Two of seven clear, and they are the two that point in opposite directions on the same variable. Everything else in this post — the within-year terciles, the alternative OI definitions, the ±5% band on the max-OI strike, the monthly split — is labelled a decomposition of one of those seven, not a new hypothesis, and is deliberately excluded from the Bonferroni family. No threshold anywhere in the study was tuned, and nothing was selected after seeing a result.
What it does not say
The sample is 1,072 sessions, roughly half the length of Elms's 2,294 days, so his standard errors are about √2 tighter than ours; where we agree with him, we agree less precisely. The instrument is the SPX cash index at 15:59 ET, not a futures settle, not the 16:15 index print, and not the AM settlement used by the third-Friday AM-settled contracts — so Golez & Jackwerth's futures-basis pinning is out of reach here rather than refuted. It is one index: nothing here says anything about single names, where the original pinning literature found its strongest effects. The open interest is a single end-of-previous-day stamp and does not distinguish dealer from customer positioning — neither did either paper. Nine sessions were dropped for having a short or gapped minute series, and one for missing its open-interest file.
And a null on pinning is not a claim that no expiry effect exists anywhere. This study finds one: monthly expirations compress the volatility-scaled range. What it rejects is that the index is drawn toward a strike because open interest sits there. Nothing on this page is investment advice; see the Terms.
Reproduce it
This one splits cleanly in two, and the split is worth being blunt about.
The price half is free. Every finished session's JSON at
https://gex.live/snapshots/YYYY-MM-DD.json carries minutes and
spot — 390 aligned one-minute values, 09:30 to 15:59 ET. That is enough for every
range, every volatility scaling, the overnight gap, the monthly split, and all of T7. The
dates are the ones listed at /sessions.
The open-interest half is not in those files. The published session files carry the flow ladder, not per-strike open interest, so a reader has to bring their own OI source: OCC's series-level data, or an options data vendor. The open interest used here is the ThetaData per-strike OI book (thetadata.net), which is a paid subscription — reproducible does not mean free for T1 through T6.
T7 — the test that refutes Golez & Jackwerth — needs no open interest at all, only the close and the strike grid, so it is fully free:
import json, urllib.request, numpy as np GRID = 5.0 # modal gap between listed 0DTE strikes near the close, # on 1,069 of 1,072 sessions; 10.0 on the other 3 d = [] for day in dates: # the dates listed at /sessions u = "https://gex.live/snapshots/%s.json" % day s = json.load(urllib.request.urlopen(u))["spot"] # 390 prints, 09:30..15:59 ET c = float(s[-1]) # the 15:59 close d.append(abs(((c + GRID / 2) % GRID) - GRID / 2)) # distance to the nearest strike d = np.array(d) # under a close uniform over the grid, P(dist <= x) = min(2x/GRID, 1) for x in (0.25, 1.0): print(x, (d <= x).mean(), 2 * x / GRID) # actual share vs expectation print(d.mean(), GRID / 4) # mean distance vs 1.25 under uniform # Golez & Jackwerth report 13.56% inside $0.25 against 10% expected. # A ratio near 1.00 here is the refutation; fetch sequentially, ~7 MB per day.
The session pages at /sessions state each day's levels in prose, so any individual session in the panel can be spot-checked by eye against a day you remember.
Part of gex.live research. Measured on the free session archive; every session is free to replay.