Prediction markets went from a niche academic curiosity to a multi-billion-dollar market between 2024 and 2026, and prediction market research followed the money. Most of the rigorous empirical work on Kalshi and Polymarket specifically is barely a year or two old, because there wasn't enough trading history to study before that. What follows is a survey of twelve papers published between November 2024 and July 2026, organized around four questions that come up constantly in prediction market microstructure work: how liquidity actually gets provided, whether an arbitrage strategy is real money or a theoretical curiosity, how well-calibrated the prices actually are, and whether markets can be pushed around by whoever has the deepest pockets.
Every paper below links to its primary source. Where a finding is a direct quote from an abstract, it's marked as one.
Papers at a glance
| Paper | Authors | Published | Core finding |
|---|---|---|---|
| Makers and Takers: The Economics of the Kalshi Prediction Market | Bürgi, Deng, Whelan | Sep 2025 | Kalshi prices show a favorite-longshot bias; makers and takers show distinct patterns |
| Prediction Market Accuracy: Crowd Wisdom or Informed Minority? | Gómez-Cram, Guo, Jensen, Kung | 2026 | ~3% of traders drive most price discovery; the rest fund their profits |
| Unravelling the Probabilistic Forest: Arbitrage in Prediction Markets | Saguillo, Ghafouri, Kiffer, Suarez-Tangil | Aug 2025 | ~$40M in realized arbitrage profit extracted from Polymarket |
| Arbitrage Analysis in Polymarket NBA Markets | Cheng, Yang, Zou | Apr 2026 | Arbitrage exists but is liquidity-bounded to retail-scale size |
| The Anatomy of a Blockchain Prediction Market | Tsang, Yang | Mar 2026 | Naive volume overstates trading by more than 2x; market quality improved sharply |
| Decomposing Crowd Wisdom: Domain-Specific Calibration Dynamics | Nam Anh Le | Feb 2026 | Political markets show persistent underconfidence; calibration is domain-specific |
| Manipulation in Prediction Markets: An Agent-based Modeling Experiment | Smart, Mark, Bastian, Waugh | Jan 2026 | Whale agents distort prices temporarily; markets self-correct across most parameters |
| Optimal Market Making in Prediction Markets | Feil, Nendel | Jul 2026 | A stochastic control framework for bid/ask quoting under settlement risk |
| Adaptive Liquidity in Prediction Markets via Online Learning | Nueve, Nguyen, Frongillo, Waggoner | May 2026 | Treats liquidity depth itself as a learnable, adaptive parameter |
| Toward Black Scholes for Prediction Markets | Dalen | Oct 2025 | Proposes a logit jump-diffusion kernel as an IV analogue for quoting and hedging |
| Bootstrapping Liquidity in BTC-Denominated Prediction Markets | Shabashev | Sep 2025 | Compares three liquidity-bootstrapping methods for BTC-settled markets |
| Designing AMMs for Combinatorial Securities | Hossain, Wang, Yu | Nov 2024 | Connects AMM pricing to computational geometry for efficient combinatorial markets |
How liquidity actually gets provided
This is where most of the recent prediction market liquidity research has concentrated. Liquidity in a prediction market has a structural problem that ordinary markets don't: every contract settles to exactly $0 or $1, so a market maker's inventory risk doesn't fade with diversification the way it does in equities, it concentrates right up until resolution. Five recent papers approach this from different angles.
Feil and Nendel build a stochastic control framework specifically for this settlement-risk problem. Their model treats the market price as a conditional probability generated by a latent belief process, and derives optimal bid and ask quotes by solving a Hamilton-Jacobi-Bellman equation that accounts for both ordinary inventory risk and the binary settlement risk that vanishes at resolution. In their numerical analysis, the resulting quoting strategy "substantially improves downside protection while preserving most of its expected profit relative to a myopic benchmark," which is a fairly direct answer to the practical question of whether it's worth building a real risk model instead of quoting a fixed spread around the mid.
Nueve, Nguyen, Frongillo, and Waggoner attack a different piece of the same problem: most cost-function market makers fix their liquidity parameter in advance, which forces a static tradeoff between how responsive prices are and how much the market maker can lose in the worst case. Their paper reframes liquidity selection as an online learning problem, mixing a family of cost-function markets via learnable weights so the effective liquidity adapts to order flow and inventory conditions as they change, while still preserving no-arbitrage and bounded worst-case loss.
Dalen's paper is the most ambitious of the group conceptually: an attempt to give prediction markets something like the Black-Scholes framework options markets have had for fifty years. The proposed model treats the traded probability as a martingale under a logit jump-diffusion process, which separates ordinary "belief volatility" from discrete jumps (news events) and yields quotable risk factors a market maker could actually hedge against, along with a family of derivative instruments analogous to variance and correlation products in options markets.
Two papers look at liquidity from a market-design rather than a single-market-maker angle. Hossain, Wang, and Yu tackle combinatorial prediction markets, where the number of possible outcome combinations can explode, by connecting AMM pricing and trade updates to the range query problem from computational geometry. That connection lets them build market makers with provably sublinear-time pricing when the outcome structure has bounded complexity, and show that no such efficient algorithm exists once that complexity is unbounded. Shabashev's paper on BTC-denominated markets is a more applied comparison: it evaluates cross-market making, automated market making, and DeFi-based liquidity redirection as three ways to bootstrap a brand-new market denominated in Bitcoin rather than stablecoins, finding that cross-market making offers the best risk profile for users but needs active professional participation, while pure AMM liquidity is simple to deploy but capital-inefficient and exposes liquidity providers to permanent loss.
Is arbitrage in prediction markets real money, or a rounding error
Two papers put actual dollar figures on this question, and the numbers are bigger and smaller than intuition might suggest, depending on which slice of the market you're looking at.
Saguillo, Ghafouri, Kiffer, and Suarez-Tangil ran a systematic empirical arbitrage analysis across all of Polymarket, distinguishing between two structurally different kinds of mispricing: rebalancing arbitrage within a single market (where related outcome prices don't sum to $1) and combinatorial arbitrage that spans multiple related markets. Using on-chain historical order-book data, they estimate a realized total of roughly $40 million in arbitrage profit extracted from the platform, and confirm that both categories have actually been exploited by traders, not just theoretically available.
Cheng, Yang, and Zou narrow the lens considerably, focusing only on Polymarket's NBA game markets. Reconstructing continuous market state from more than 75 million limit order book snapshots across 173 games, they find single-market arbitrage anomalies are "exceedingly rare," producing only 7 executable in-game episodes with a median duration of 3.6 seconds. Combinatorial arbitrage across related markets is more common (290 active episodes, concentrated in the final minutes of live play) and yields a statistically meaningful median return of 101 basis points, but execution is bottlenecked hard by order-book depth: 76.9% of those opportunities were constrained to an average executable size of just 14.8 shares. Their conclusion is blunt: risk-free extraction is real, but structurally bounded to retail scale by how thin the book actually is.
That liquidity constraint connects directly to who's actually making money in these markets day to day. Gómez-Cram, Guo, Jensen, and Kung studied roughly $13.76 billion in trading volume across 1.72 million accounts and nearly 99,000 events, and their answer to "is it crowd wisdom or informed trading" is neither, cleanly. Prices, per the paper, get accurate because of "a tiny, persistently skilled minority roughly 3% of traders," who react to news within moments of it breaking and systematically trade against known biases like the favorite-longshot effect, while the bulk of retail volume contributes little information and effectively funds that minority's profits.
Bürgi, Deng, and Whelan land on a closely related pattern from a different angle: using transaction-level data on over 300,000 contracts, they document a clear favorite-longshot bias in Kalshi pricing, where low-probability contracts win less often than their price implies and high-probability contracts win slightly more often than theirs does, and they attribute distinct pieces of this pattern to market makers versus the takers who accept their quotes.
How well-calibrated are the prices, really
A price isn't automatically a probability, which is the whole point of a prediction market accuracy study: two papers here dig into what it takes for that reading to actually hold.
Nam Anh Le's calibration study is the largest empirical exercise in this list by data volume: 353 million trades across 429,000 binary contracts on both Kalshi and Polymarket. The central finding is that calibration is domain-specific rather than a single platform-wide property, and the most robust pattern across the whole dataset is persistent underconfidence in political markets specifically, where prices compress toward 50% relative to what actually happens, a pattern that replicates on both platforms independently. On Kalshi, large political trades are associated with even more compression, though that effect doesn't hold up as robustly on Polymarket. A descriptive model of these calibration slopes explains 87.3% of in-sample variance on Kalshi, dropping to 71.5% out-of-sample, and a Bayesian error model suggests roughly half of the raw slope variation is just estimation noise once conservative clustering is applied. The paper's own framing is worth keeping: "a price's meaning depends on what, when and how much is traded."
Tsang and Yang's study of Polymarket's 2024 presidential election market is nominally about volume, but it doubles as a market-quality study, because they show how badly naive volume figures can mislead anyone judging market quality from headline numbers. Using full on-chain Polygon data, they build a transaction-level accounting framework that separates genuine exchange-equivalent turnover from token minting and burning, and the gap is large: naive aggregation reports $958 million in October Trump-market volume, versus $391 million once properly decomposed, less than half. On the more encouraging side, market quality improved substantially over that period: arbitrage-deviation half-lives fell from hours to under a minute, and Kyle's lambda, a standard measure of price impact per unit of order flow, dropped from 0.53 to 0.01. During the large-account trading episode in October that drew public scrutiny, they find capital flowed into both sides of the market simultaneously, which is more consistent with traders holding genuinely different beliefs than with one-sided manipulation.
Can a well-funded trader actually move the market
Smart, Mark, Bastian, and Waugh built an agent-based simulation specifically to test this, combining it with an analytical characterization of price dynamics rather than relying on the simulation alone. Their model includes bettors with heterogeneous expertise, noisy private information, and different learning rates and budgets, all observing public opinion to inform their trades. Across a broad range of parameters, the market exhibits what the authors describe as self-regulatory price discovery, meaning it tends to correct itself. But a well-resourced "whale" agent with a biased valuation can still temporarily shift prices away from the informationally efficient level, and the magnitude and duration of that distortion grow specifically when ordinary bettors herd or learn slowly. The practical reading is that manipulation resistance in these markets isn't a fixed property of the mechanism, it's conditional on how the rest of the crowd actually behaves.
What this means if you're building on top of this data
A few patterns hold across most of these twelve papers regardless of which specific question they set out to answer. Naive metrics mislead: naive volume overstates real turnover by more than 2x in the Polymarket election study, and naive "is it accurate" framing misses that accuracy and calibration are conditional on domain, trade size, and timing rather than a single number. Liquidity is genuinely scarce at the depth that matters: the NBA arbitrage study found over three-quarters of exploitable mispricings constrained to under 15 shares of executable size. And the field is moving from "does this work" toward "how do you actually build and trade on it," reflected in the shift from earlier mechanism-design papers toward the stochastic-control and online-learning market-making frameworks published in just the last several months of this list.
Doing any of this kind of research firsthand, reconstructing order-book state, measuring realized arbitrage, or building a calibration study, depends on having the underlying tick-level historical data these papers are built on: full trade history and order-book depth, not just daily closes. That's the same data problem covered in more product-specific detail in the Kalshi and Deribit historical data guides on this site.
A theme worth naming directly: several of these results are less flattering to the platforms than the marketing narrative around prediction markets usually is. A 2x gap between naive and decomposed volume, a favorite-longshot bias that persists on a regulated exchange, calibration that only holds up domain-by-domain rather than platform-wide, these aren't reasons to distrust the mechanism, but they are reasons to distrust any single summary statistic pulled from a dashboard without knowing how it was constructed. That's arguably the single most consistent finding across all twelve papers: the mechanism works better than the raw numbers on any given day suggest, once you correct for how those numbers are built.
Scope and methodology
This survey covers twelve papers found through targeted search across arXiv, SSRN, and university working-paper series, filtered to those published between November 2024 and July 2026 that deal directly with prediction market liquidity, market making, arbitrage, calibration, or manipulation. Each paper's existence and core claims were checked against its own abstract page or PDF directly rather than taken from a news summary or a single secondary source.
What this survey doesn't claim to be is exhaustive. Search-based discovery misses working papers that haven't been indexed yet, non-English-language research, and results published in venues that don't surface well through general search. It also doesn't cover the older, pre-2024 literature on prediction markets used for corporate forecasting and internal decision markets, a genuinely different research tradition from the retail-facing, high-volume platforms this survey focuses on. Treat this as a snapshot of what's verifiable as of writing, not a comprehensive literature review.
FAQ
Are prediction markets accurate?
It's conditional rather than a flat yes or no. Recent research finds accuracy is concentrated in a small minority of skilled traders (Gómez-Cram et al. estimate roughly 3%) and that calibration varies by domain, with political markets showing persistent underconfidence relative to how often events actually occur (Nam Anh Le, 2026).
Is there real money in prediction market arbitrage, or is it theoretical?
Both, depending on scope. Saguillo et al. estimate around $40 million in realized arbitrage profit across all of Polymarket, while a narrower study of just NBA markets (Cheng, Yang, and Zou, 2026) found real but structurally small opportunities, most constrained to under 15 shares of executable size by thin order books.
How is liquidity typically provided in prediction markets?
Through a mix of automated market makers, limit order books, and increasingly, adaptive or model-based quoting strategies that account for the settlement risk unique to binary markets. Recent academic proposals include stochastic-control quoting (Feil and Nendel, 2026) and treating the liquidity parameter itself as something to be learned online (Nueve et al., 2026), rather than fixed in advance.
Can large traders manipulate prediction market prices?
Agent-based modeling suggests a well-funded "whale" can temporarily distort prices, with the size and duration of the distortion depending heavily on whether the rest of the market herds or reacts slowly (Smart et al., 2026). Across most simulated conditions, the market tends to self-correct.
Where can I find the primary sources for this kind of research?
Most of the papers referenced here are preprints hosted on arXiv, with a few distributed through SSRN and university working-paper series (CESifo, CEPR) rather than as a formal research report from either exchange itself. Every citation in this article links directly to the primary source rather than a secondary summary.
What kind of data do these studies typically rely on?
Full trade-level and order-book data, not daily summaries. The NBA arbitrage study reconstructed market state from over 75 million limit order book snapshots; the calibration study drew on 353 million individual trades. Studies working from anything coarser than tick-level history tend to miss exactly the short-lived mispricings and microstructure effects most of this research is trying to measure.
Running this research on your own data
Convex Lake covers tick-level trades and order-book depth for Kalshi, Polymarket, predict.fun, Limitless, and Deribit options under one API. See the API docs for the current endpoint reference, or create an account to get access.

