What the 0DTE literature actually claims

25 papers, sorted by whether a tick tape with a signed dealer book can test them

Published 2026-08-24 · Sample a literature sweep run 2026-08-02, 25 entries across five topics · Scope empirical papers on US index options and index intraday dynamics · Status a map of other people's claims, not a measurement of ours

Two things make the published record on 0DTE options hard to use. Most of the popular claims about dealer gamma are not in the literature at all — they circulate as charts with no sample attached. And most of the claims that are in the literature cannot be checked by a reader, because checking them needs data that costs money, or because the paper itself is behind a paywall that will not open. This post is the sorting exercise: every paper the sweep found, its central claim in one clause, and a verdict on whether the claim can be run against a tick tape — with the reason, and with the identifier so anyone can go and read the thing.

The sweep was run on 2026-08-02. It has no numbers of its own; every number below is the paper's, reported as the paper reports it.

What "testable" assumes

The verdict column is not "is this true". It is "could this claim be run, end to end, against the following data": a tick SPXW tape with the prevailing NBBO at every print; 5-minute full-chain quotes; per-strike open interest; a Black-Scholes surface carrying gamma, vanna and charm per contract per minute; a signed dealer-inventory book; per-second dealer-gamma levels; and the index recovered by put-call parity. That is what this terminal builds each session — the finished ones are at /sessions — and the tick layer under it is licensed, not free (see Reproduce it). A paper marked not testable is usually not wrong; it is about a different instrument, or it reports no effect size to clear.

1. Conditional variance risk premium

paperthe claimverdictwhy
Vilkov (2026) — "0DTE Trading Rules: Tail Risk, Implementation, and Tactical Timing"
SSRN 4641356; code at github.com/vilkovgr/0dte-strategies
A positive 0DTE variance risk premium exists but is too small to monetise unconditionally; conditional timing works only as directional classification, and only for some structures once costs are charged. testable Same instrument, same vendor, same 09/2016–01/2026 window, MIT-licensed code released. The closest thing to a pre-built benchmark in the sweep. Unconditional net: straddle SR −0.51, iron butterfly/condor −0.96.
Almeida, Freire & Hizmeri (2025) — "0DTE Asset Pricing"
SSRN 4701401; free PDF at fma.org
0DTE investors are compensated mainly for upside moves, producing a large positive 0DTE VRP that negatively predicts intraday market returns; ~70% of 0DTEs violate second-order-stochastic-dominance price bounds. testable Every input is on the surface: risk-neutral variance from the chain, realised variance from the index, bid/ask fills from the tape. Annualised 0DTE VRP 1.54%–2.95% by entry time, vs 0.55% at 1DTE.
Muravyev & Ni (2020) — "Why do option returns change sign from day to night?"
JFE 136(1), DOI 10.1016/j.jfineco.2018.12.006
The entire short-volatility premium in S&P 500 options is earned overnight; intraday, delta-hedged option returns are positive. testable The decomposition — −0.7%/day overall, close-to-open −1.0%/day, open-to-close +0.3%/day — reproduces exactly on SPXW at real bid/ask and extends a decade past their sample. [mid-fill, unconfirmed] [pre-2020]
Cheng (2019) — "The VIX Premium"
RFS 32(1), 180–227; free post-print at utoronto.scholaris.ca
The ex-ante volatility premium falls or stays flat when ex-ante risk rises — selling volatility is worse compensated in high-risk states, the opposite of the folk model. partly The instrument is VIX futures, which this terminal does not hold. The claim survives as an analogue on our own chain: does short-dated VRP per unit of risk shrink when 10:00 ET implied variance is high? [pre-2020] — sample ends May 2016.
Aït-Sahalia, Karaman & Mancini (2020) — "The Term Structure of Variance Swaps and Risk Premia"
Journal of Econometrics 219(2); SSRN 2136820
The variance risk premium has a downward-sloping term structure; the jump component inverts to upward-sloping in turbulent times. partly No variance swaps here, but the synthetic swap rate is buildable from the chain at every 5-minute snapshot. The obstacle is that the paper reports no effect size to clear. [mid-fill] [pre-2020]
Dew-Becker, Giglio, Le & Rodriguez (2016) — "The price of variance risk"
JFE 123(2); no open version found
Only realised variance carries a risk premium; forward variance is essentially unpriced, and the premium is concentrated at the very short end. partly Directionally supportive of working at the short end, but testing it needs a variance-swap term structure we would have to synthesise. It says where to look, not what to trade. [mid-fill] [pre-2020]
Hu & Jacobs (2020) — "Volatility and Expected Option Returns"
JFQA 55(3), DOI 10.1017/S0022109019000310
Expected call returns decrease and expected put returns increase with the volatility of the underlying — a mechanical leverage/moneyness effect. not Wrong instrument: the cross-section of single-name equity options. Listed as the standard caution against reading "high vol ⇒ rich options" off raw returns.

2. Intraday realised-volatility forecasting

This is the thinnest area in the sweep, and the thinness is itself the finding. No published paper was found that forecasts intraday realised volatility at a 30-minute horizon from option state and benchmarks it against trailing RV. The literature splits into next-day forecasting from short-dated implied vol, and high-frequency measurement of spot volatility from options, which is estimation rather than forecasting. An internal result in this shape therefore has no published competitor to beat and no published corroboration. That is a gap, not a validation.

paperthe claimverdictwhy
Albers (2025) — "A New Star Is Born: Does the VIX1D Render Common Volatility Forecasting Models for the US Equity Market Obsolete?"
Journal of Futures Markets, DOI 10.1002/fut.70023
Cboe's 1-day VIX1D overestimates next-day S&P 500 volatility; after a risk-premium adjustment the adjusted VIX1D beats HAR-family models. testable VIX1D is a VIX-methodology integral over the 0DTE strip, rebuildable from our chain at every 5-minute snapshot rather than only at Cboe's publication times. Numbers unverified — the publisher PDF refused automated retrieval. Companion: Albers & Kestner (2024) on VIX1D's overnight bias.
Chong & Todorov (2024) — "Volatility of Volatility and Leverage Effect from Options"
arXiv 2305.04137v2
Model-free spot volatility, vol-of-vol and leverage-effect estimators can be built from high-frequency observations of short-dated options. partly Yes as machinery, no as a trade: it is the right way to extract spot volatility from 5-minute chain snapshots without inverting Black-Scholes contract by contract. An econometrics paper — rates of convergence and a CLT, no effect size, by design. Companion: Todorov (2019), Annals of Applied Probability 29(6).
Andersen, Fusari & Todorov (2017) — "Short-Term Market Risks Implied by Weekly Options"
Journal of Finance 72(3), 1335–1386; free author PDF
Short-dated options isolate a negative jump tail risk factor unspanned by market volatility, and that factor predicts future equity returns where volatility does not. testable The tail-shape estimator is computable from the chain. The 1-week coefficient, t = 4.507, is the strongest t-statistic in the sweep — on 1,105 trading days of weekly options in 2011–2015, with overlapping returns. [mid-fill] [pre-2020]
Alexiou, Bevilacqua & Hizmeri (2026) — "Uncovering the Asymmetric Information Content of High-Frequency Options"
Journal of Banking & Finance, DOI 10.1016/j.jbankfin.2026.107720
Option realised semivariances and signed jumps, built from the sign of high-frequency option returns, carry directional information incremental to aggregate option realised measures. testable One of the few papers whose inputs map one-to-one onto a tick tape with NBBO. Paywalled; the abstract quotes no magnitudes, so nothing is extracted here.
Bollerslev, Patton & Quaedvlieg (2016) — "Exploiting the errors"
Journal of Econometrics 192(1)
HAR-Q: adjust HAR loadings by the realised quarticity — the measurement error — of the RV estimate. not Not a candidate at all. It is here to set the bar: HAR-Q, not plain trailing RV, is the honest benchmark any RV signal has to clear.

3. 0DTE specifically

The two substantive 0DTE papers — Vilkov, and Almeida/Freire/Hizmeri — sit in section 1, because their central claim is about conditional variance risk premium rather than about 0DTE as such.

paperthe claimverdictwhy
Beckmeyer, Branger & Gayda (2023) — "Retail Traders Love 0DTE Options... But Should They?"
SSRN 4404704 [SSRN-gated]
0DTE options are popular with retail traders who mainly lose money in them; dealers are structurally short 0DTE gamma. not Already absorbed. The load-bearing line — dealers short 0DTE gamma — is a claim about flow mix, not a tradable signal, so there is nothing new to run. Note the sizing correction Almeida et al. add: 0DTE is over 75% of retail SPX option trading, but roughly 94% of all 0DTE SPX volume is institutional.
Muravyev & Pearson (2020) — "Options Trading Costs Are Lower than You Think"
RFS 33(11), 4973–5014, DOI 10.1093/rfs/hhaa010
Option prices are predictable at high frequency, so traders who time executions pay far less than quoted or effective spread measures imply. testable Tick prints with the prevailing NBBO are precisely what is needed to measure how far actual SPXW prints land from the midpoint. Their headline: execution-timing traders pay less than 40% of conventional effective spread, and the overall average is one-quarter smaller. [pre-2020]
Bandi, Fusari & Renò (2023) — "0DTE Option Pricing"; Bandi, Fusari, Gazzani & Renò (2026) — "Ultra-short-term volatility surfaces"
SSRN 4503344 [SSRN-gated]; arXiv 2603.29430 (free)
Closed-form pricing that fits the 0DTE implied-volatility surface; pronounced oscillations in the ATM implied-vol term structure across ultra-short tenors that classical models cannot fit jointly. partly No as a strategy, yes as a diagnostic. The oscillating ATM term structure is a measurable stylised fact — and a warning: if a Black-Scholes-inverted surface shows the same oscillation, any cross-tenor VRP comparison is partly measuring model misfit. Fit papers report RMSE, not Sharpe. [mid-fill]
Adams, Fontaine & Ornthanalai (2024) — "The Market for 0-Days-to-Expiration: The Role of Liquidity Providers in Volatility Attenuation"
SSRN 4881008 [SSRN-gated]
0DTE liquidity providers attenuate rather than amplify volatility — dealer hedging in 0DTE dampens the underlying. partly Note the live disagreement: Dim, Eraker & Vilkov (2024) and Adams et al. reach attenuation from net-open-interest measures; Brogaard, Han & Won (2023/2026) reach the opposite from a different measure. Effect sizes not obtainable — SSRN blocked, no working-paper mirror found. Related and also [SSRN-gated]: Vasquez, Amaya, Pearson & Garcia-Ares (2025), "0DTE Index Options and Market Volatility: How Large is Their Impact?", SSRN 5113405.

4. Dealer inventory and positioning

paperthe claimverdictwhy
Muravyev (2016) — "Order Flow and Expected Option Returns"
Journal of Finance 71(2), 673–708, DOI 10.1111/jofi.12380
Market-maker inventory risk has a first-order effect on option prices; the inventory component of price impact is larger than the asymmetric-information component, and past order imbalances predict option returns better than standard predictors. testable A signed dealer-inventory book is a reconstruction of exactly the quantity this paper says matters, and the claim is about option returns rather than index direction. The abstract quotes no t-statistic; the headline is that inventory-driven imbalances have five times the price impact previously thought. Caveat: measured on a daily panel of mostly equity options, where inventory is genuinely scarce. [pre-2020]
Gârleanu, Pedersen & Poteshman (2009) — "Demand-Based Option Pricing"
RFS 22(10), 4259–4299
When options cannot be perfectly hedged, end-user demand pressure raises a contract's price in proportion to the variance of its unhedgeable part. partly The theory gives the right shape for a dealer-inventory signal: impact should scale with unhedgeable variance, so ATM 0DTE near expiry is where demand pressure should bite hardest. The obstacle is their "end-user demand", which comes from an exchange customer/firm/market-maker breakdown that a tape-based book approximates rather than observes. [pre-2020]
Chen, Joslin & Ni (2019) — "Demand for Crash Insurance, Intermediary Constraints, and Risk Premia in Financial Markets"
RFS 32(1), 228–265; NBER w25573
Public net buying of deep-OTM index puts measures intermediary constraint tightness; a tightening comes with rising option expensiveness and deteriorating funding liquidity. partly Their horizon is monthly to quarterly. The deep-OTM put segment of a signed book is the direct analogue and the intraday version is unpublished, but confidence that it survives at intraday horizons is low. [pre-2020]
Chordia, Kurov, Muravyev & Subrahmanyam (2021) — "Index Option Trading Activity and Market Returns"
Management Science 67(3), DOI 10.1287/mnsc.2019.3529
Weekly index put order flow positively and robustly predicts weekly S&P 500 returns, driven by net put buying, and more strongly in high-VIX periods. partly Testable — signed flow and the VIX conditioning are both buildable — but this is index direction from option flow at a weekly horizon, the exact shape of result that has repeatedly failed here. Rated below the option-return tests for that reason. [pre-2020]
Barbon & Buraschi (2021) — "Gamma Fragility"
SSRN 3725454; free PDF at abarbon.com
Large negative aggregate dealer gamma imbalance produces higher volatility and intraday momentum; positive imbalance produces lower volatility and mean reversion — but the channel requires limited liquidity in the underlying. partly The direction half is the momentum claim that failed to replicate here, and the paper's own mechanism needs underlying illiquidity, which SPX does not have. The amplitude half — gamma imbalance moves the volatility level, t up to −5.07 — is the part that is alive. No strategy is backtested anywhere in the paper. [pre-2020]
Christoffersen, Goyenko, Jacobs & Karoui (2018) — "Illiquidity Premia in the Equity Options Market"
RFS 31(3), 811–851; free post-print at utoronto.scholaris.ca
Option illiquidity is priced in the cross-section of option returns. not Wrong instrument, and a cross-sectional premium cannot be traded with one index underlying. Listed for completeness. [pre-2020]

5. Expiry-day and pinning effects

paperthe claimverdictwhy
Golez & Jackwerth (2012) — "Pinning in the S&P 500 Futures"
JFE 106(3), 566–585; free working paper at kops.uni-konstanz.de
S&P 500 futures are pulled toward the ATM strike on serial-expiration days, driven by market makers rebalancing delta hedges as decay changes their deltas. testable 13.56% of futures settle within $0.25 of a strike on serial expiration, against 10% expected under independent draws. Testable, but expect to reject — the direct modern replication below already finds it gone. [pre-2020], by sixteen years.
Elms (2026) — "From Pinning to Amplification: Evidence of a Regime Shift in S&P 500 Options Expiration Dynamics, 2016–2025"
Zenodo, DOI 10.5281/zenodo.19540116; also SSRN 6564078
Golez & Jackwerth's pinning is gone in 2016–2025; what remains is amplification. testable Five tests on 2,294 matched trading days; four null. The one significant result: daily range rises monotonically across ATM-OI terciles, 16.2% wider ranges on high-OI days, p = 0.0003. Cheap to replicate — but weigh the source: a self-published preprint, not peer reviewed, one significant cell in five with no multiple-testing adjustment, using SPY open interest as a proxy for SPX dealer positioning.
Ni, Pearson & Poteshman (2005) — "Stock price clustering on option expiration dates"
JFE 78(1)
Pinning in US equity options, with an aggregate market-cap effect of about $9 billion per expiration date. not Wrong instrument — single-name equities — and the modern index version is already answered by Elms. Cited as the origin of the literature. [pre-2020]

The count

25 entries. Ten are testable as stated, ten are testable only in part — usually because the instrument is wrong but the claim survives as an analogue, or because the paper reports no effect size to clear — and five are not, because they are cross-sectional equity-option papers, a benchmark rather than a candidate, or a claim about flow mix with nothing to run. Ten carry a [pre-2020] flag, meaning the sample ends before daily SPXW listings began on 2022-05-11. Only Vilkov, Almeida et al., Elms, Albers, Alexiou et al., Chong-Todorov and the gated 0DTE-impact papers use samples that include the daily-expiry regime at all.

What to test first, and why

The sweep ranks its own candidates. Reproduced in order:

  1. Muravyev & Ni (2020) — the day/night decomposition. A −1.0%/day overnight versus +0.3%/day intraday split is a large, mechanical, unambiguous claim, priceable at real bid/ask and extendable a decade past their sample. If it holds, it says in one line why intraday short-vol variants keep dying, and that 0DTE — which has no overnight leg — is the worst possible vehicle for the premium. Highest information per unit of work.
  2. Vilkov (2026) — conditional out-of-sample 0DTE timing. Same instrument, same vendor, same window, published replication code and reference tables, and a cost model close to ours. A rare chance to check a pipeline against someone else's published numbers before trusting either.
  3. Almeida, Freire & Hizmeri (2025) — conditional after-cost Sharpe. Their Low-RV versus High-RV split (after-cost Sharpe 0.404–0.470 against 0.004–0.082) is the largest state-conditioning effect anyone reports for index option selling, and their own claim that the edge died after 2022-05-11 is directly falsifiable on the window where modern data is densest. Decisive either way.
  4. Muravyev & Pearson (2020) — real execution costs. Not a strategy, but it targets the parameter that kills results. Measuring where actual SPXW prints sit relative to the prevailing NBBO is a day of work, and a 25–60% cut in the assumed cost would reopen several closed files.
  5. Muravyev (2016) — inventory-driven imbalance to option returns. The best mapping onto a signed dealer-inventory book, and it predicts option returns rather than index direction.
  6. Elms (2026) plus the Barbon-Buraschi amplitude half. Cheap, and the two agree that expiry and gamma concentration widen rather than compress ranges.
  7. Albers (2025) on VIX1D, and the Chong-Todorov / Todorov spot-vol machinery. Feature generators, not strategies.
  8. Andersen-Fusari-Todorov negative-tail factor. Strong t-statistics, but 2011–2015 and mid-priced.
  9. Chordia-Kurov-Muravyev-Subrahmanyam put order flow. Testable, but it is index direction from option flow.
  10. Gârleanu-Pedersen-Poteshman, Chen-Joslin-Ni, Aït-Sahalia et al., Dew-Becker et al. Useful for the shape of a signal; none gives an effect size to clear, and all are pre-2020.

The gated rows, and why you should not trust their numbers

SSRN returns 403 to every automated fetch. Four entries in the map were therefore read from title and abstract metadata plus whatever the citing literature says about them, not from the PDF: Beckmeyer, Branger & Gayda (2023) (SSRN 4404704), Bandi, Fusari & Renò (2023) (SSRN 4503344), Adams, Fontaine & Ornthanalai (2024) (SSRN 4881008) and Vasquez, Amaya, Pearson & Garcia-Ares (2025) (SSRN 5113405). Their claims are stated here as the citing literature states them, and their effect sizes are either absent or second-hand. Nothing in those four rows should be trusted until someone opens the PDF by hand. Two more entries are unverified for the same practical reason from a different publisher: the Albers (2025) PDF and the Gârleanu et al. and Chen et al. PDFs refused automated retrieval, so no magnitudes were extracted from them.

Vilkov's SSRN PDF is gated too, but its full text and code are mirrored on GitHub under an MIT licence, which is why that row is the best-documented one in the map.

Where we have already run one

Honesty about linkage: the sweep does not pair any paper with a published post on this site. Where it points to prior work it points at internal results — a dealer-gamma to next-30-minute realised-volatility regression (β = −0.047, p = 0.0013), a killed net-GEX direction study, a charm signal that flipped sign in 2026 — and none of those has been published here. The amplitude side of Barbon & Buraschi, the attenuation side of Adams et al., and Elms's widening-range result all land on the same sign as that gamma-to-RV regression, which is the strongest convergent evidence in the sweep; the regression itself is not on this site, so no link is offered for it. The published measurements here — the touch surface, level hold rates, the OCC dealer-to-dealer floor, daily book turnover and the pre-registered dealer-state test — are separate work, and inventing a correspondence between one of them and a paper in this map would be the exact error this post exists to avoid.

What it does not say

This page contains no measurement. "Testable" is a judgement about data availability, not about whether a claim is correct, well-identified or likely to replicate; several papers marked testable are expected to be rejected on modern data, and the map says so in the row rather than pretending otherwise. The sweep is one pass by one reader on one day and is certainly incomplete — it applied standing filters that exclude a set of already-dead ideas, so the map is not a neutral census of the field. The claims are summarised in one clause each, which is not enough to argue with; the identifiers are given so you can go and read the argument. Where a paper is paywalled, the summary is second-hand and flagged as such. Nothing on this page is investment advice; see the Terms.

Reproduce it

This post has no numbers of its own, so what there is to reproduce is the search. The sweep used, in this order: the arXiv API, OpenAlex, Semantic Scholar, DuckDuckGo HTML results, and direct PDF pulls from author and repository pages. All four indexes are free and take no key. The route:

pythonreproduce
# arXiv (full text, free) — e.g. Chong & Todorov
http://export.arxiv.org/api/query?search_query=all:%220DTE%22&max_results=50
https://arxiv.org/pdf/2305.04137

# OpenAlex — metadata, DOI, open-access location, citing works
https://api.openalex.org/works?filter=title.search:0DTE%20options

# Semantic Scholar — abstract verification when the publisher is paywalled
https://api.semanticscholar.org/graph/v1/paper/DOI:10.1016/j.jfineco.2018.12.006?fields=title,abstract,year

# Zenodo — the Elms preprint, open PDF
https://doi.org/10.5281/zenodo.19540116

# GitHub — Vilkov's replication code and annotated paper, MIT licence
https://github.com/vilkovgr/0dte-strategies

What refused automated access. SSRN returns 403 to every automated fetch, with no exception found — that is what gates the four flagged rows. The Wiley publisher PDF for Albers (2025) refused automated retrieval, though its DOI landing page is open. Oxford University Press PDFs blocked the fetch for Gârleanu, Pedersen & Poteshman (2009) and Chen, Joslin & Ni (2019), so both were verified from abstract only. Free author or repository copies did work for Cheng (2019), Andersen-Fusari-Todorov (2017), Barbon & Buraschi (2021), Golez & Jackwerth (2012), Christoffersen et al. (2018) and Almeida et al. (2025).

What running the tests would cost. Every "testable" verdict assumes the tick layer described at the top, and that layer is licensed, not free: the SPXW trade and NBBO tick feed from ThetaData (https://thetadata.net) is the product a reader would have to subscribe to in order to redo any of these on prints rather than on midpoints. Vilkov names ThetaData or Massive as his own vendor, which is why his row is the one directly comparable benchmark. The derived per-session layer — levels, the band, the per-minute series — is free: each finished session is published as https://gex.live/snapshots/YYYY-MM-DD.json, and the dates are listed at /sessions. That is enough for the amplitude and pinning rows (Elms, Golez & Jackwerth, the Barbon-Buraschi amplitude half); it is not enough for anything that needs a fill price.

Part of gex.live research. Measured on the free session archive; every session is free to replay.