Monochrome line-art illustration of an open book surrounded by isometric city buildings, representing prediction market venues and data

Prediction Market Data: The Complete Guide

What prediction market data covers across Kalshi, Polymarket, Predict.fun and Limitless — APIs, orderbook and trade data, calibration, and how the unified-API landscape actually compares.

Written by Convex Lake Research Team
· 10 min read
#prediction-markets#data#kalshi#polymarket#calibration#api

"Prediction market data" covers three distinct things that get lumped together: order book depth (what it would actually cost to trade a contract right now), trade history (what actually executed), and price history through settlement (how a market's implied probability evolved). Most people asking about this are looking for one of the three, not all of them, and the venue they need it from shapes what's actually available.

This covers the four venues Convex Lake tracks, how their data differs, where the existing "unified API" tools stand, and what calibration actually means when someone claims a prediction market is accurate.

The venues

Kalshi is a CFTC-regulated US exchange. That regulatory status is also its main access friction: its official API is tied to a verified trading account, and identity verification isn't available to residents of most countries outside the US. Categories span politics, sports, economics, weather, crypto, and culture. Full schema and access details: Kalshi historical data.

Polymarket runs on-chain, with no verification gate on its data. Markets span an even broader range of topics than Kalshi's, including a large volume of crypto and current-events contracts. Full detail: Polymarket historical data.

Predict.fun runs on BNB Chain and is the official prediction-market provider inside Binance Wallet, giving it direct distribution most competitors don't have. It's reported over $1.8 billion in cumulative volume and more than 130,000 users since launch.

Limitless runs on Base, built around short-duration markets, hourly and daily windows rather than long-running contracts, with a no-liquidation-risk structure. It's reported more than $1 billion traded.

Kalshi vs. Polymarket

The two largest venues split on structure more than on what they let you trade. Kalshi is a regulated exchange with identity-verified accounts and CFTC oversight; Polymarket is a decentralized, on-chain venue with no such gate. That difference cuts both ways: Kalshi's regulation gives it legal clarity in the US that an offshore, permissionless venue doesn't have, while Polymarket's lack of a verification gate makes its data and its markets accessible to anyone, anywhere, immediately.

On coverage, both run central limit order books per contract, and both span a wide range of topics. Where they diverge sharply is category depth. A snapshot of markets actually live on each venue, across the categories both share: Polymarket runs 578 crypto markets to Kalshi's 43, 417 economics to 18, 441 esports to 15, 1,063 sports to 10, 568 weather to 3. Polymarket leads every single shared category, sometimes by two orders of magnitude, a direct consequence of Kalshi's slower, regulated contract-approval process versus Polymarket's more permissionless listing model. Neither is strictly better: Kalshi's tighter catalog trades breadth for a more curated, regulated contract set, which matters if regulatory clarity is part of what a research or product use case actually needs.

What prediction market data actually contains

Data typeWhat it tells youTypical use
OrderbookResting bid/ask depth at each price levelSlippage modeling, liquidity research, market making
TradesIndividual executions: price, size, side, timestampVolume analysis, execution research
CandlesPrice and volume bars, open through settlementCalibration studies, lifecycle analysis

Exact field names, file formats, and access requirements differ by venue, covered in the dedicated guides linked above and in the API docs.

The unified-API landscape

Convex Lake isn't the only product trying to put multiple prediction-market venues behind one interface. PMXT is an open-source SDK (Python and TypeScript) covering Polymarket, Kalshi, and Limitless, built more for execution, placing orders, checking balances, than for historical data archives. FinFeedAPI, from the team behind CoinAPI, normalizes Polymarket, Kalshi, Myriad, and Manifold into one schema. PolyRouter aggregates Kalshi, Polymarket, Limitless and others behind a single key. Tatum offers one API key and one response format across Polymarket and Kalshi specifically.

All four solve real fragmentation, one schema instead of four. None of them cross over into crypto derivatives. That's the specific gap Convex Lake fills: the same unified-schema approach, but spanning Kalshi, Polymarket, Predict.fun, and Limitless on the prediction-market side, plus Deribit and Binance options on the crypto-derivatives side, under one API and one key. If your research or product only ever touches prediction markets, any of the four above is worth evaluating on its own terms. If it touches both categories, that's where the gap actually shows up. (Full comparison against crypto-only providers Tardis, Amberdata, and CoinAPI, and prediction-market-only providers Lychee and PolymarketData.co: data provider alternatives.)

Common cross-venue research use cases

Cross-venue arbitrage research compares pricing on the same or correlated events across two or more venues, real price gaps have been documented and studied directly, though execution against them is bounded by how much size each venue's book can absorb before the gap closes (more on this in the research roundup). Calibration comparison asks whether the same category, political markets are a well-studied example, shows the same accuracy pattern independent of which venue hosted it; recent research found persistent underconfidence in political markets on multiple platforms independently, suggesting it's a property of the category rather than any one venue's pricing. Category-coverage research, like the breakdown above, is its own use case when the question is which venue actually has enough simultaneously-active markets in a given category to build a real sample from, rather than which venue is faster or cheaper to query.

Data delivery and formats

Across venues, REST is the standard for request-response access to markets, trades, orderbook snapshots, and candles; historical backfills are typically pulled this way even on venues where live data streams over WebSocket. CSV and JSON suit exploration and scripting; Parquet, columnar and compressed, is the practical choice once a pull runs past a few million order-book rows, which full-depth history reaches quickly on any actively-traded venue.

Calibration and the Brier score

A prediction market being "accurate" specifically means it's well-calibrated: when it prices something at 70%, that thing happens roughly 70% of the time across many such markets, not that any single call was right or wrong. The standard way to measure this is the Brier score, the mean squared difference between a predicted probability and the actual binary outcome, scored 0 to 1, where 0 is a perfect forecast and 1 is the worst possible one.

Commonly cited reference ranges: human superforecasters typically score around 0.15–0.20; aggregate prediction-market pricing often falls in the 0.12–0.18 range; sports betting lines average roughly 0.18–0.22. Treat these as general benchmarks rather than a single hard number, they vary by study, by domain, and by how the sample is constructed, and calibration itself varies by category rather than being one fixed property of a platform. Political markets, for instance, have shown persistent underconfidence in recent research, independent of which platform hosted them. What a price means depends on what's being traded, not just where.

FAQ

Is there a resolution-specific dataset separate from price history?

Not on Convex Lake. Candle data covers a market’s full lifecycle from open through settlement, which includes the resolution point, but there’s no separate "resolution" endpoint distinct from that. If a use case specifically needs settlement outcomes in isolation, that’s derivable from the last candle in a market’s series rather than a dedicated feed.

Is prediction-market sports data the same as sports betting odds data?

No, and they’re often confused. Sportsbook odds APIs cover traditional bookmaker lines, spreads and moneylines from operators like the major sportsbooks. Prediction-market sports data, what Convex Lake covers via Kalshi’s sports category, is trading data on event-contract markets: order book depth, trades, and price history for contracts like "will Team X win," priced by an open market rather than set by a bookmaker. Related concepts, different data structure and different source.

What’s a "unified prediction market API," and do I need one?

A single interface covering multiple venues under one schema, instead of integrating each venue’s API separately. Worth it as soon as a research question or product needs to compare or combine more than one venue; not necessary if the work is scoped to a single platform.

Does Kalshi vs. Polymarket data show real arbitrage opportunities?

Price gaps between the two venues on the same or correlated events do show up in the data and have been studied directly; execution against them is constrained by how much size each venue’s book can actually absorb before the gap closes. See the prediction market research roundup for specific findings on this. See the prediction market research roundup

Which venue should I start with for historical data?

Depends on the research question more than on which venue is "better." Kalshi for US-regulated, politics- and economics-heavy analysis; Polymarket for broader category coverage and no access gate; Predict.fun or Limitless if the question is specifically about newer, faster-moving venue structures. All four are covered under one schema through Convex Lake’s API. API docs

Convex Lake

A comprehensive financial technology platform for prediction market data and quantitative analytics

Resources

Company

© 2026 Convex Lake. All rights reserved.