Why historical Polymarket data matters
Polymarket is the largest crypto-native prediction market venue by open interest, and its CLOB (central limit orderbook) produces per-second L2 snapshots that behave much more like a traditional exchange feed than a typical AMM. For quantitative teams, that history is the foundation for backtesting event-driven strategies, calibrating implied probability models, and studying microstructure phenomena — spread dynamics, quote fade, and reaction latency around real-world events.
Unlike price-only feeds, historical L2 orderbook data preserves the full depth of the book at every snapshot, letting you reconstruct realistic fills, slippage, and venue-specific edge without look-ahead bias.
What "historical L2 orderbook" means here
Each snapshot is a point-in-time picture of the visible order book for a given market (e.g. btc-updown-5m-1781049600) and contains:
- UTC timestamp (millisecond precision)
- Best bid and best ask (BBO)
- Depth ladder — price and size for each level on both sides
- Market/outcome identifier and resolution timestamp
Convex Lake stores these snapshots as gzipped, date-partitioned files in object storage — one file per market, per interval, per UTC day. Files are typed by md_type (for example orderbook) and grouped by venue and symbol so a single download corresponds to a single, self-describing time series.
Data schema
Typical BBO / L2 record (JSON lines, gzipped):
{
"ts": 1781049605123,
"market": "btc-updown-5m-1781049600",
"outcome": "up",
"bids": [[0.52, 1200], [0.51, 3400], [0.50, 8000]],
"asks": [[0.54, 900], [0.55, 2100], [0.56, 6000]],
"mid": 0.53,
"spread_bps": 380
}Every record is self-contained: no cross-file joins are required to reconstruct the book at time ts. That is deliberate — it keeps backtests reproducible and lets you shard workloads across markets trivially.
Downloading Polymarket historical data
Convex Lake exposes a small, S3-backed REST API. Two endpoints: /api/info for progressive discovery and /orderbook for downloads. They accept the same filter parameters: exchange, ticker, timeframe, and date.
# List available orderbook files for a Polymarket BTC 5m market on 2026-06-10 curl -H "x-api-key: $CONVEXLAKE_API_KEY" \ "https://api.convexlake.com/info?md_type=orderbook\ &exchange=polymarket&ticker=btc&timeframe=5m&date=2026-06-10" # Download a specific file curl -H "x-api-key: $CONVEXLAKE_API_KEY" \ "https://api.convexlake.com/orderbook\ ?exchange=polymarket&ticker=btc&timeframe=5m&date=2026-06-10\ &slug=btc-updown-5m-1781049600-bbo.gz" -o snapshot.gz
The API is tier-gated: sample orderbook data is available on the free tier; full historical Polymarket depth requires a Pro key. See the API docs for the full schema.
Processing snapshots at scale
A day of 5-minute BTC-updown markets is typically a few hundred KB gzipped. A full year of every listed market fits comfortably on a single workstation. For larger universes, the recommended pattern is:
- List available files via
/api/info. - Download in parallel by
date × symbol. - Stream-parse the JSONL directly into Polars or DuckDB — no intermediate CSV.
- Persist derived features (mid, microprice, weighted depth) as Parquet.
import polars as pl, gzip, json
with gzip.open("snapshot.gz", "rt") as f:
rows = [json.loads(l) for l in f]
df = pl.from_dicts(rows).with_columns([
pl.col("ts").cast(pl.Datetime("ms")),
(pl.col("mid") * 1.0).alias("mid"),
])
print(df.head())Common research use cases
- Event-driven backtests — measure how BBO and depth react to macro prints, on-chain events, or oracle updates.
- Implied probability calibration — convert mid-price into a probability estimate and compare to realized outcomes.
- Microstructure studies — quote lifetime, spread persistence, and adverse-selection cost per venue.
- Cross-venue arbitrage research — align Polymarket snapshots with Kalshi and Limitless feeds on a common UTC clock.
Best practices
- Always work in UTC. Prediction market resolution timestamps are UTC-native; local time will silently corrupt event alignment.
- Treat the visible book as an upper bound on liquidity. Hidden or reserve orders are not represented; model fill probability accordingly.
- For strategy research, resample to a consistent grid (e.g. 1s) before computing features — raw snapshot cadence varies by market activity.
- Cache raw files locally and derive features on top. Re-downloading is bandwidth you do not need to spend.
Getting started
Create a free Convex Lake account to generate an API key, or explore the API documentation first. If you're evaluating full historical Polymarket coverage for a research team, reach out via the support page and we'll set up an evaluation window.
Related reading
- Kalshi historical data: trades, orderbook depth and market history — the equivalent guide for Kalshi, including where its official API falls short.
- Maker rebate programs compared: Polymarket, Kalshi, Limitless, predict.fun — how Polymarket's rebate mechanics shape the liquidity in these snapshots.

