Polymarket L2 orderbook depth visualization

Polymarket Historical Data: A Quant Guide to L2 Orderbook Snapshots

A technical guide to accessing, storing, and analyzing Polymarket historical orderbook data for quantitative research on crypto prediction markets.

Written by Convex Lake Research
· 8 min read
#orderbook#polymarket#historical-data#quant

Why historical Polymarket data matters

Polymarket is the largest crypto-native prediction market venue by open interest, and its CLOB (central limit orderbook) produces per-second L2 snapshots that behave much more like a traditional exchange feed than a typical AMM. For quantitative teams, that history is the foundation for backtesting event-driven strategies, calibrating implied probability models, and studying microstructure phenomena — spread dynamics, quote fade, and reaction latency around real-world events.

Unlike price-only feeds, historical L2 orderbook data preserves the full depth of the book at every snapshot, letting you reconstruct realistic fills, slippage, and venue-specific edge without look-ahead bias.

What "historical L2 orderbook" means here

Each snapshot is a point-in-time picture of the visible order book for a given market (e.g. btc-updown-5m-1781049600) and contains:

  • UTC timestamp (millisecond precision)
  • Best bid and best ask (BBO)
  • Depth ladder — price and size for each level on both sides
  • Market/outcome identifier and resolution timestamp

Convex Lake stores these snapshots as gzipped, date-partitioned files in object storage — one file per market, per interval, per UTC day. Files are typed by md_type (for example orderbook) and grouped by venue and symbol so a single download corresponds to a single, self-describing time series.

Data schema

Typical BBO / L2 record (JSON lines, gzipped):

{
  "ts": 1781049605123,
  "market": "btc-updown-5m-1781049600",
  "outcome": "up",
  "bids": [[0.52, 1200], [0.51, 3400], [0.50, 8000]],
  "asks": [[0.54, 900],  [0.55, 2100], [0.56, 6000]],
  "mid": 0.53,
  "spread_bps": 380
}

Every record is self-contained: no cross-file joins are required to reconstruct the book at time ts. That is deliberate — it keeps backtests reproducible and lets you shard workloads across markets trivially.

Downloading Polymarket historical data

Convex Lake exposes a small, S3-backed REST API. Two endpoints: /api/info for progressive discovery and /orderbook for downloads. They accept the same filter parameters: exchange, ticker, timeframe, and date.

# List available orderbook files for a Polymarket BTC 5m market on 2026-06-10
curl -H "x-api-key: $CONVEXLAKE_API_KEY" \
  "https://api.convexlake.com/info?md_type=orderbook\
&exchange=polymarket&ticker=btc&timeframe=5m&date=2026-06-10"

# Download a specific file
curl -H "x-api-key: $CONVEXLAKE_API_KEY" \
  "https://api.convexlake.com/orderbook\
?exchange=polymarket&ticker=btc&timeframe=5m&date=2026-06-10\
&slug=btc-updown-5m-1781049600-bbo.gz" -o snapshot.gz

The API is tier-gated: sample orderbook data is available on the free tier; full historical Polymarket depth requires a Pro key. See the API docs for the full schema.

Processing snapshots at scale

A day of 5-minute BTC-updown markets is typically a few hundred KB gzipped. A full year of every listed market fits comfortably on a single workstation. For larger universes, the recommended pattern is:

  1. List available files via /api/info.
  2. Download in parallel by date × symbol.
  3. Stream-parse the JSONL directly into Polars or DuckDB — no intermediate CSV.
  4. Persist derived features (mid, microprice, weighted depth) as Parquet.
import polars as pl, gzip, json

with gzip.open("snapshot.gz", "rt") as f:
    rows = [json.loads(l) for l in f]

df = pl.from_dicts(rows).with_columns([
    pl.col("ts").cast(pl.Datetime("ms")),
    (pl.col("mid") * 1.0).alias("mid"),
])
print(df.head())

Common research use cases

  • Event-driven backtests — measure how BBO and depth react to macro prints, on-chain events, or oracle updates.
  • Implied probability calibration — convert mid-price into a probability estimate and compare to realized outcomes.
  • Microstructure studies — quote lifetime, spread persistence, and adverse-selection cost per venue.
  • Cross-venue arbitrage research — align Polymarket snapshots with Kalshi and Limitless feeds on a common UTC clock.

Best practices

  • Always work in UTC. Prediction market resolution timestamps are UTC-native; local time will silently corrupt event alignment.
  • Treat the visible book as an upper bound on liquidity. Hidden or reserve orders are not represented; model fill probability accordingly.
  • For strategy research, resample to a consistent grid (e.g. 1s) before computing features — raw snapshot cadence varies by market activity.
  • Cache raw files locally and derive features on top. Re-downloading is bandwidth you do not need to spend.

Getting started

Create a free Convex Lake account to generate an API key, or explore the API documentation first. If you're evaluating full historical Polymarket coverage for a research team, reach out via the support page and we'll set up an evaluation window.

Related reading

Convex Lake

A comprehensive financial technology platform for prediction market data and quantitative analytics

Resources

Company

© 2026 Convex Lake. All rights reserved.