Abstract monochrome orderbook depth ladder illustrating Kalshi historical data

Kalshi Historical Data: A Quant's Guide to Trades, Orderbook Depth & Market History

What Kalshi historical data actually contains, how it's captured, where to get it, and the mistakes that quietly wreck backtests built on it.

Written by Convex Lake Research
· 10 min read
#kalshi#historical-data#orderbook#trades#quant

Kalshi's live API only gets you so far when you're building a strategy, a calibration study, or a research pipeline. And if you're outside the US, it may not get you anywhere at all (more on that below). The exchange splits its data into a live tier and a historical tier, and the historical side is where the interesting research questions live: full trade history, orderbook depth, and price history that runs through settlement.

Written for quants, researchers, and developers who want to work with the data programmatically, not for traders looking for a dashboard.

What Kalshi historical data actually includes

"Kalshi historical data" isn't one dataset. It's several related ones, usually sold or served together.

Markets carry the metadata: tickers, titles, categories, open and close times, market status. Candles give you OHLC-style price and volume bars, and each file covers a market's full life from open through settlement. Trades are individual executions, timestamped, with YES/NO prices, size, and taker side. Orderbook history is the bid/ask depth at each price level over time, not just whatever the last trade happened to print. And category and event metadata tells you which series and event a market belongs to: weather, politics, sports, crypto, economics, culture.

A dataset that only gives you daily closing prices is missing most of that list. If your use case involves execution (modeling slippage, fill probability, spread cost), you need orderbook history. A price series alone won't get you there.

EndpointKey fieldsTypical use
GET /candlestimestamp, yes_bid, yes_ask, price, volume. One file per market, open to settlement.Price history, lifecycle analysis
GET /tradescreated_time, yes_price, no_price, count, taker_side, taker_outcome_side, is_block_tradeVolume analysis, execution research
GET /orderbooktype, sid, seq, market_ticker, yes_dollars_fp, no_dollars_fp, price_dollars, delta_fp, side, ts_ms. Gzip, one snapshot or delta per line.Slippage modeling, liquidity research, market making

Field names above reflect the current API docs at time of writing. Always check the API docs for the live reference before integrating; endpoint schemas can change.

Last price vs. full order-book depth: why it matters

A last-traded price tells you what happened, once. It doesn't tell you what it would have cost to actually get filled: the size resting at each level, how far the price would move to fill a larger order, how wide the spread was at that moment. For any strategy involving real execution, rather than a hypothetical fill at the mid, that gap is the whole problem.

Full order-book depth is the complete bid/ask ladder, both sides, every level. It's what lets you measure realistic slippage instead of assuming you always get the mid price. Between a backtest that looks profitable on paper and one that survives contact with real fills, this is usually the difference.

Kalshi's official historical API vs. third-party providers

Kalshi's own API splits data into live and historical tiers at a cutoff timestamp you can query directly. The official historical endpoints hand you raw market, trade, and settlement data, but no bulk orderbook history, no pre-cleaned exports, and no cross-venue normalization. Pagination, storage, joins: all on you.

There's a bigger obstacle before any of that, though. Getting access at all. Kalshi's API is tied to a Kalshi trading account, and opening one requires identity verification as part of onboarding to a CFTC-regulated exchange. That verification isn't available to residents of most countries outside the US. For a large share of the researchers, quants, and developers who want Kalshi data, the official API simply isn't an option, regardless of what it covers. It's an access problem, not a completeness one, and it's the main reason third-party providers exist for Kalshi data at all.

Past that access question, providers differ on coverage too. Some focus on no-code browser access and analysis tooling. Others capture raw orderbook data at high frequency. Others normalize Kalshi alongside other prediction-market venues under one schema. None of these is strictly better than the rest; they trade off differently on coverage, orderbook granularity, and how much integration work is left for you to do.

Source typeAccess requirementOrderbook depthSetup effortBest for
Kalshi official APIVerified Kalshi account (identity check) — not available in most countries outside the USLive book only; historical trades/settlement, no bulk depth archiveYou build storage, pagination, joinsUS-eligible teams with existing data infrastructure
No-code analysis platformsAccount with the provider, no Kalshi verification neededVaries by providerMinimal — browser-based queryingResearchers who want charts and exports without writing code
Raw-capture / tick-level providersAccount with the provider, no Kalshi verification neededFull L2 depth, high frequencyAPI integration requiredBacktesting execution-sensitive strategies
Cross-venue normalized APIsAccount with the provider, no Kalshi verification neededVaries by providerAPI integration requiredComparing Kalshi against Polymarket, Deribit, and other venues in one schema
Convex LakeAPI key, self-service. No Kalshi account or verification needed.Batch and real-time orderbook snapshots, trades, candlesREST API, documented endpointsQuant research spanning Kalshi plus other prediction and crypto-options venues, including from outside the US

How historical data is captured and stored

Two broad capture strategies produce meaningfully different datasets. Continuous polling hits the exchange's public REST orderbook on a loop and records what it sees; realized frequency depends on request budget and how many markets are being tracked at once. Event-driven capture listens for state changes and only writes a snapshot when something actually moves. That's more efficient, but it depends on not missing events.

Either way, order-book depth only runs forward. If nobody was recording when a price level changed, that state is gone for good. You can sometimes backfill a daily close after the fact; you can't backfill order-book history. So "how far back does your orderbook archive go" is a more useful question to ask a provider than "how far back does your data go" in general. Trade and settlement history is often available further back than full-depth orderbook capture.

Storage format matters once you're at scale, too. CSV and JSON are fine for exploring, but get unwieldy past a few million rows. Columnar formats like Parquet compress better and query faster, which is usually what bulk historical research needs.

Downloading and querying Kalshi historical data via API

  1. Get an API key. With Convex Lake this is self-service: no Kalshi account, no verification. Generate one on the API Keys page and send it as x-api-key on every request.
  2. Pick a venue and market. Kalshi requests are scoped with exchange=kalshi plus a ticker and date or timeframe.
  3. Pull the file. Download endpoints return the requested trades, orderbook, or candle file for that market and date.
  4. Store and join. Land the files in your own storage and join markets, trades, candles, and orderbook data on the ticker.
curl -O -J -H "x-api-key: do_YOUR_KEY" \
  "https://api.convexlake.com/trades?exchange=kalshi&ticker=<market-ticker>&date=2026-06-10"

Field names and exact query parameters differ slightly by endpoint and venue. Check the API docs for the current /trades, /orderbook, and /candles reference before building against it.

Kalshi market categories in the historical dataset

Kalshi runs markets across several category types, and how far back historical coverage goes can vary by category depending on the provider.

  • Weather: temperature, precipitation, and other climate contracts
  • Politics: elections, approval ratings, legislative and policy outcomes
  • Sports: game and season outcomes
  • Crypto: short-dated BTC/ETH/SOL up-down and threshold markets
  • Economics: inflation, employment, and other macro data releases
  • Culture: entertainment, awards, and similar events

Weather and crypto markets tend to be short-dated and high-frequency, often 15-minute or hourly windows, which piles up a lot of completed market history fast. That's useful if you want a large sample for calibration work without waiting years for enough political markets to settle.

Common research and backtesting use cases

Volume and liquidity analysis looks at which markets and categories concentrate trading activity, and when. Calibration research uses candle price history through settlement to check whether prices were well-calibrated against outcomes. Backtesting simulates a strategy against recorded book states to estimate fills, slippage, and P&L instead of assuming a perfect fill at the mid. Lifecycle analysis studies how a market's implied probability evolves from listing through settlement. Cross-venue arbitrage research compares Kalshi pricing against Polymarket or other venues for the same or correlated events.

The starting point tends to differ by role. Quant researchers usually start from orderbook depth and trade history, to model execution. Market makers care most about historical spread and depth, to size quotes. Journalists and analysts tend to want settlement-price trends and category-level activity rather than tick data. Students and independent researchers often start with a single category, politics or weather, and a manageable date range before scaling up.

Data quality pitfalls to watch for

Ticker reuse. Series tickers can get reused across different market instances over time. Join purely on ticker without also checking the event or series ID, and you can silently merge unrelated markets.

Timezone and settlement-time handling. Settlement timestamps and market close times need consistent timezone handling, especially for short-dated hourly or 15-minute markets, where an off-by-one-hour bug changes which snapshot a backtest scores against.

Survivorship bias in "settled only" datasets. If a dataset only includes markets that fully settled, it may be excluding cancelled or voided markets in a way that biases calibration studies.

Live vs. historical cutoff gaps. Live and historical data are often served by different endpoints with different latency, which can leave a short window where neither feed has the most recent data yet.

Verification-gated access shaping what datasets even exist. Because Kalshi's own API requires a verified account, most third-party Kalshi datasets are built by a provider re-serving data they collected themselves, rather than a direct exchange feed. Worth asking any provider how their data was captured, not only what it contains.

Glossary: key terms in Kalshi historical data

Orderbook depth is the full set of resting bid and ask sizes at each price level, rather than only the best price. YES/NO ladder refers to how Kalshi markets are binary: the orderbook has a YES side and a NO side, where a NO bid at price p is equivalent to a YES offer at 1 − p, and markets settle to YES, NO, or void at expiry. Calibration measures how closely a market's implied probability matches the actual settlement rate across many markets. Brier score is a standard scoring rule for probabilistic forecasts, where lower is better. VWAP is the volume-weighted average price, used to smooth noisy tick data into a steadier reference price. Slippage is the gap between the price you expected and the price you'd actually have been filled at, given real order-book depth.

Historical data formats and delivery

Delivery format depends on how the data gets used. CSV and XLSX are easiest for ad-hoc analysis in spreadsheets or pandas, but inefficient for very large pulls. JSON is the default for REST API responses, convenient for scripting and less so for bulk storage. Parquet is columnar and compressed, and the practical choice once you're working with millions of orderbook snapshots. REST gives request-response access to markets, trades, orderbook snapshots, and candles. WebSocket is for streaming live updates; historical backfills are typically pulled over REST even on platforms where live data comes through a socket.

FAQ

Do I need a verified Kalshi account to get Kalshi historical data?

Only if you're using Kalshi's own official API. It's tied to a Kalshi trading account, which requires identity verification as part of onboarding to a CFTC-regulated exchange, and that verification isn't available to residents of most countries outside the US. Third-party providers, Convex Lake included, don't require a Kalshi account or verification; access comes from a provider API key instead.

Does Kalshi provide historical data for free?

Kalshi's official API exposes historical trades and settlement data without a paid tier, for accounts that can pass its verification. It doesn't include a bulk order-book depth archive either way. Several third-party providers offer free tiers with limited history or row counts, with paid plans for full-depth or full-archive access.

Can I download Kalshi historical data as CSV?

Yes. Most providers, including Kalshi's own export tools and third-party APIs, support CSV export for trades and market data. Full order-book history more often comes as JSON or Parquet, given its size.

Does Kalshi's API include historical order book data?

Kalshi's official API gives you the live order book plus historical trades and settlement, but not a bulk historical order-book archive. That gap is what most third-party providers, Convex Lake included, exist to fill.

How far back does Kalshi historical data go?

It depends entirely on when a given provider started recording. Kalshi as an exchange has operated since 2021, but any specific orderbook archive only covers the period after that provider began capturing it. Order-book depth can't be backfilled retroactively.

What’s the difference between "historical" and "settled" markets on Kalshi?

"Historical" generally covers any data before the live/historical cutoff, including markets that are still open. "Settled" specifically means the market has reached a final YES/NO outcome. A historical dataset can include both closed-then-archived markets and markets still awaiting settlement.

Is there a rate limit on Kalshi’s historical endpoints?

Yes, on both the official API and third-party providers, usually tiered by plan. Check the specific provider's documentation for current limits before designing a bulk pull.

Getting started with Convex Lake's Kalshi historical data

Convex Lake covers Kalshi alongside Polymarket, predict.fun, Limitless, and Deribit under one API, batch and real-time: orderbook snapshots, trade history, candle data. One schema instead of stitching together five. See the API docs for the current endpoint reference, or create an account to get access.

Related reading

Convex Lake

A comprehensive financial technology platform for prediction market data and quantitative analytics

Resources

Company

© 2026 Convex Lake. All rights reserved.