Monochrome isometric trading venues, clocks and fee symbols

Trading Fees and Execution Latency Across Six Venues: What's Published, and What Isn't

Fee formulas for Kalshi, Polymarket, Predict.fun, Limitless, Deribit and Binance options, plus what public sources actually say about execution latency, especially Polymarket fills.

Written by Convex Lake Research Team
· 8 min read
#trading-fees#execution-latency#polymarket#kalshi#market-structure

Fees decide whether a spread is an edge. Latency decides whether you get the fill. Fees are published, but in six different formulas. Latency mostly isn't published at all, and the numbers that circulate measure something other than what people assume.

This covers Kalshi, Polymarket, Predict.fun, Limitless, Deribit, and Binance options: what each charges, what each says about speed, and where the public record stops.

Fees: six formulas, one comparable example

Prediction markets, taker fees.

VenueMakerTakerSource
KalshiZero on most series; a 1.75% multiplier on series flagged for maker fees7% × contracts × P × (1 − P), rounded up to the cent. Max 1.75¢ per contract at 50¢. Multiplier varies by series: fee-free series at 0, some sports at 0.5Kalshi help center, plus a third-party summary citing the schedule effective 2026-07-07
PolymarketNever chargedContracts × rate × P × (1 − P). Rate by category: crypto 0.07; sports, economics, culture, weather, other 0.05; finance, politics, mentions, tech 0.04; geopolitics 0Polymarket docs
Predict.funNoneBase rate × min(P, 1 − P) × shares. Ranges from 0.018% to 2% depending on price. 10% discount availablePredict.fun docs
LimitlessNoneCLOB buys 0.40%–3.00% (paid in outcome tokens); sells 0.42%–1.50% (paid in USDC). AMM markets flat 0.40%Limitless docs

Options venues.

VenueTrading feeCapExercise / delivery
Deribit0.03% of underlying, maker and taker12.5% of option price0.015% (third-party summaries; Deribit's own page blocked fetching)
Binance options0.024% of spot index price × contract unit10% of option price0.015%, same 10% cap. Liquidation 0.19%, capped at 25%

A worked example makes the prediction-market rows comparable. Take 100 contracts at 50¢, the price where the quadratic fees peak:

  • Kalshi: 0.07 × 100 × 0.5 × 0.5 = $1.75
  • Polymarket, crypto markets: the same formula, same $1.75. At the 0.05 rate, $1.25. At 0.04, $1.00. Geopolitics, $0.
  • Predict.fun: 2% × 0.5 × 100 = $1.00
  • Limitless: the published curve peaks at 3.00% on buys and 1.50% on sells, but the docs say the curve isn't published as a closed formula, so a dollar figure would be a guess. The exact fee comes back as effectiveFeeBps on execution.

That's arithmetic from published formulas at one price, not an all-in cost. But it shows something the headline rates hide: Kalshi and Polymarket's crypto fee are the same shape and the same size. The gap between venues is mostly which category you're trading, not which platform.

On the options side, both caps bind for cheap options. Deribit's fee is the lesser of 0.03% of the underlying and 12.5% of the premium. Binance's is the lesser of 0.024% and 10%. Both flip at the same place: when the option premium is under 0.24% of the underlying. Below that, the percentage-of-premium cap sets the fee. Above it, the flat rate does.

Latency: four layers, and only some are published

"Latency" gets used for four different things, and public numbers rarely say which one they mean.

  1. Network round trip. Your machine to the venue's endpoint.
  2. Deliberate delay. A hold the venue adds on purpose.
  3. Engine processing. Time inside the matching engine.
  4. Settlement finality. When the trade is final, for venues that settle on-chain.

What each venue actually publishes:

VenueNetworkDeliberate delayEngineSettlement
DeribitNot publishedNone documentedPublished: p50 76 µs, p99 234 µsn/a
PolymarketThird-party only, and it measures the CDN edge250 ms on selected up/down markets; cut to 50 ms on 5m/15m cryptoNot publishedDocs list the statuses, no timeline
KalshiThird-party onlyNone documentedNot publishedn/a
Predict.funNot foundNot foundNot foundOn-chain, BNB Chain
LimitlessNot foundNot foundOn-chain, no figuresOn-chain, Base
Binance optionsNot foundNone documentedNot foundn/a

Deribit is the exception, and a recent one.

Deribit: the only venue with real engine numbers

In its own September 21, 2026 write-up, Deribit reports its new matching engine, called Starbase, at a median round trip of 76 microseconds, with p90 at 106 µs and p99 at 234 µs. The classic API it replaced measured 4.7 ms at the median, 26.4 ms at p90, and 124 ms at p99. That's roughly 61 times lower at the median and 531 times lower at p99, compared across six days before rollout and six days after over 90% of order-entry messages moved to the new system. In the busiest single second, August 25, with 23,836 requests, median latency rose to 609 µs against a 79 µs baseline, and the worst round trip stayed just under 0.8 ms.

One caveat on reading those numbers: they're the round trip through the matching system. You still pay your own network distance to get there. A trader in another region isn't seeing 76 µs. The engine's share of the total is now small.

Kalshi: the number you see is the edge, not the engine

Third-party VPS vendors report that Kalshi's matching engine runs in AWS us-east-2 (Ohio), that a FIX connection from Chicago sees roughly 10 ms round trip with very low jitter, and that WebSocket updates arrive about 25 ms after the event. The catch is the number that gets quoted most: ping times near 1 ms to Kalshi's public REST and WebSocket hosts. Those hosts sit behind Amazon CloudFront, which terminates your connection at the nearest edge. A 1 ms ping measures how close you are to a CDN node. It says nothing about the engine. Every one of these figures comes from vendors selling low-latency hosting, so treat them as estimates. Independent probes, like Glassnode's public prediction-market latency monitor, describe the same setup and say their own numbers are "directional baselines."

Polymarket: what the public record actually contains

Polymarket fill speed is the most-asked case, colocation included. Here is the whole public record.

The matching engine location. Polymarket's off-chain order book runs in AWS eu-west-2 (London). Orders are signed intents, matched off-chain by an operator, then submitted to an Exchange contract on Polygon.

What "colocation" measurements measure. A QuantVPS post reports 0.83 ms from a Dublin VPS to clob.polymarket.com. A DEV Community author's own 2026 test (100 requests per region, plain HTTPS GET, not orders) found time-to-first-byte at the median of about 6 ms from Amsterdam, 16 ms Frankfurt, 19 ms London, 110 ms US-East, and 195 ms Singapore. Notice that Amsterdam beats London, though the engine is reported to sit in London. That pattern points to the CLOB API being CDN-fronted, so a fast probe from Amsterdam or Dublin is hitting an edge node. The author says of his own numbers only that they "will drift" and should be re-measured. Neither source measures an order being submitted and filled.

The delay that dwarfs the network. Polymarket's docs say that on selected crypto and finance up/down markets, "the order is held for 250 ms, then validation runs again and the order is matched or placed on the book." During that hold, the order is pending and can't be cancelled. If validation fails when the delay ends, the order is rejected. Sports markets apply their own delay windows around live game conditions.

A third-party report, citing a Polymarket Developers announcement, says the delay on taker orders in 5-minute and 15-minute crypto markets was cut from 250 ms to 50 ms effective August 17, 2026. The official order-lifecycle page still says 250 ms as of this writing, so check which applies to the market you're trading. The itode flag on the public CLOB endpoint marks which markets carry a delay. The delay exists to protect resting quotes: a market maker quoting BTC outcomes gets a short window to update stale prices before a taker can hit them.

What that means for colocation. On the markets where speed matters most, a taker is waiting 50 to 250 ms by design. Colocation trims single-digit milliseconds off the network leg of that. It doesn't touch the hold. For a maker, the picture inverts: the delay shrinks the time to pull a stale quote before a taker can match it, which is why the shortened window puts more weight on fast cancels.

Settlement. Matched trades go MATCHED → MINED → CONFIRMED, or RETRYING → FAILED. The docs specify no timeline beyond finality on Polygon.

Rate limits add a hidden layer. POST /order allows bursts of 5,000 requests per 10 seconds and 120,000 per 10 minutes. When limits are exceeded, requests are throttled, delayed or queued, rather than rejected. That means an over-limit bot doesn't see errors. It sees slower fills.

Bottom line for Polymarket: no public source gives a measured submit-to-fill latency distribution. What exists is network round trip to a CDN-fronted API and a documented, intentional hold of 50 to 250 ms on the fast markets. If you need a real fill-time number, you have to measure your own.

Execution caveats worth knowing

  • Fees peak at 50¢. Kalshi, Polymarket, and Predict.fun all use price-curved fees that fall toward zero at the extremes. Limitless's buy curve is the opposite shape: highest at low prices.
  • Fee currency differs. Limitless charges buys in outcome tokens and sells in USDC. That changes how you reconcile cost.
  • Rounding differs. Kalshi rounds each fee up to the next cent, which weighs more on small orders. Polymarket rounds to five decimals with a floor of 0.00001 USDC.
  • Zero maker fees aren't free liquidity. A resting order still faces adverse selection, which is what Polymarket's taker delay is designed to soften on its fastest markets.
  • Fee schedules change. Kalshi's schedule is dated and carries per-series multipliers. Polymarket's rates differ by category. Check the current page for the exact contract.
  • Ping isn't fill time. CDN-fronted endpoints make a low ping look better than the path to the engine.

What we couldn't verify

Kalshi's official fee PDF and Deribit's own fee page both blocked automated access, so those two rows rest on third-party summaries of the official schedules. Nothing public gave Binance options, Limitless, or Predict.fun a latency figure. No public source measured Polymarket or Kalshi from order submission to fill.

Historical order-book and trade data across Kalshi, Polymarket, Predict.fun, Limitless, Deribit, and Binance options, in one schema, is available through the API docs. The order-book data is what you'd use to estimate realized slippage per venue instead of assuming it. For how fee drag fits alongside the settlement-definition risk that hits cross-venue positions harder, see the arbitrage piece. For where each engine sits geographically, see the infrastructure comparison.

FAQ

Which prediction market has the lowest taker fee?

It depends on category and price. At 50¢, Predict.fun's formula gives about $1.00 per 100 shares, against $1.75 on Kalshi and on Polymarket crypto markets. Polymarket's finance and politics markets come in at $1.00, and its geopolitics markets charge nothing. Limitless's buy curve reaches 3.00%.

Do makers pay fees on Polymarket?

No. Polymarket says makers are never charged, and only takers pay.

Is Deribit or Binance cheaper for options?

On the flat rate, Binance: 0.024% against Deribit's 0.03%. Both cap at a percentage of option price, 10% for Binance and 12.5% for Deribit, and both caps start to bind below a premium of 0.24% of the underlying.

How fast is Deribit's matching engine?

Deribit reports a median round trip of 76 µs and a p99 of 234 µs on its new engine, measured through the matching system. Your own latency also includes the network path to it.

What is Polymarket's actual fill latency?

No public source measures it end to end. Network round trip to the API is measurable, but the endpoint sits behind a CDN. On selected fast crypto markets there's also an intentional taker delay, documented at 250 ms and reported cut to 50 ms for 5m and 15m crypto markets from August 17, 2026.

Does colocation help on Polymarket?

It shortens the network leg, by single-digit milliseconds in the measurements published. It doesn't shorten the taker delay, which is 50 to 250 ms on the markets where speed matters most.

Convex Lake

A comprehensive financial technology platform for prediction market data and quantitative analytics

Resources

Company

© 2026 Convex Lake. All rights reserved.