
How to use AI to find mispriced bets on Kalshi and Polymarket
Learn how to build a simple AI-powered edge-finding workflow for prediction markets using free tools like Claude, Kalshi, and Polymarket data.
Ask ChatGPT which stock to buy this week and it'll happily give you an answer. Ask it who wins Sunday's NFL game and it'll pick a team. The problem isn't that the AI refuses to answer — it's that a confident-sounding answer and a profitable one are two completely different things. Most people stop at "the AI said X" and never ask the question that actually matters: is X priced into the market already?
That's the gap this guide closes. You're going to learn how to build a repeatable, step-by-step workflow that uses an AI model to find the difference between what a market is pricing and what's actually likely to happen — and why, most of the time, the market is going to win anyway.
What are prediction markets, exactly?
Prediction markets are places where people trade contracts on real-world outcomes — will it rain in New York tomorrow, will a bill pass Congress, will the Chiefs cover the spread. The price of a contract reflects the crowd's collective belief about the probability of that outcome. If a "Yes" contract trades at 62 cents, the market is saying there's roughly a 62% chance it happens.
Two platforms dominate this space right now. Kalshi is a CFTC-regulated U.S. exchange covering everything from weather to elections to economic data, and it has a full API documentation site for anyone who wants to pull data programmatically. Polymarket runs on-chain and covers a huge range of political, crypto, sports, and culture markets, with its own API reference for developers.
The core idea behind using AI here isn't "ask the robot who wins." It's building a model that estimates a fair probability independently, then comparing that number to the market's price to see if there's a gap worth exploiting.
Why doesn't ChatGPT just tell you the right answer?
Because large language models are trained to sound confident, not to be calibrated. When you ask an LLM for a prediction without giving it a structured process, it pattern-matches to what sounds plausible — team momentum, recent headlines, vibes. It has no real mechanism for turning that into a probability, let alone comparing that probability to a live market price.
The fix isn't a smarter model. It's a smarter process. That means breaking the task into discrete steps: set your rules first, establish what "fair price" even means, gather clean data, build and backtest a model, generate a prediction, size a position only if there's genuine edge, and then — critically — challenge your own output before you trust it.
A public prompt chain built around exactly this structure is available on GitHub, and it's a useful reference for how to sequence these steps. The repo lays out 8 prompts that take you from your rules to a model-backed position call, meant to be copied in order, one prompt per step, where each step checks the one before it.
How do you set up the rules before you even look at a market?
This is the step almost everyone skips, and it's the one that keeps you from blowing up a bankroll on a "sure thing." Before touching any market, decide three things: how much you're willing to risk total, what minimum edge you require (expressed as expected value per dollar after fees), and how aggressively you'll size positions relative to that edge — a fraction of the Kelly criterion is the standard approach, since full Kelly sizing is brutal on real bankrolls.
The idea is to state your bankroll, your minimum edge threshold, and your Kelly fraction up front, then confirm the formulas you'll use for edge and stake before sizing anything. Locking this in before you see a specific game or market stops you from rationalizing a bet after you've already fallen in love with an outcome.
How do you figure out the market's "fair price"?
Every market price already has fees or a "vig" baked in — the house's cut. To know if you actually have an edge, you need to strip that out and compare apples to apples across different platforms, since a sportsbook's implied odds and a prediction market's price don't always mean the same thing once fees are removed.
This step involves removing the vig from a sportsbook price using both proportional and power methods, then comparing it to the prediction-market price including its fees, using only the actual prices you've captured as the market reference. Screenshotting the actual odds rather than letting the model guess at current prices matters a lot here — AI models don't have live, accurate market data by default, so feeding in your own screenshots keeps the comparison grounded in reality.
What data do you actually need before building a model?
Garbage in, garbage out applies doubly hard to forecasting. For a sports example, that means pulling verified data sources rather than trusting an LLM's internal "knowledge," which is often stale or just wrong about recent rosters and injuries.
A solid data step pulls sources like nflverse play-by-play data, injury reports, depth charts, and weather forecasts for the specific game, verifies starting quarterbacks from play-by-play data, lists current injuries for both teams, and flags any data that contradicts other sources. This sanity-check habit — cross-referencing instead of trusting a single source — is the difference between a model built on reality and one built on hallucination.
How do you build and test the model without fooling yourself?
This is where most amateur forecasting falls apart. People build a model, run it once on the data they already have, and declare victory. The honest way to do it is to tune the model only on historical seasons, then run it against seasons it has genuinely never seen, so you can measure its real-world accuracy rather than its ability to memorize the past.
The approach is to build a win-probability model from the gathered data, tuning it only on older seasons and testing it on the most recent seasons it never saw. That backtest is also where you get an honest reality check on how good your model actually is compared to the market itself — and the results are humbling. Running this kind of backtest across more than a decade of games typically shows the market's own implied probabilities outperforming both a coin flip and a custom model on pure accuracy (measured by a lower Brier score). That's not a sign the exercise is pointless — it's the whole point. The market already aggregates the opinions of thousands of sharp, well-capitalized traders. Your job isn't to beat that collective wisdom across the board; it's to find the narrow pockets where your model disagrees with the market and has good reason to.
How do you decide when a disagreement is actually an edge?
Once you have a fair price and a market price, the gap between them tells you whether there's an opportunity. The rule from your very first step — minimum edge after fees — gets applied here mechanically, not emotionally.
For example, if your model's fair price puts a team's win probability higher than what the market implies, you calculate exactly what price you'd need to get on that side of the trade to clear your required edge threshold, something like needing a specific cents-on-the-dollar price or better to hit a 2% expected value target. If the market's current price doesn't clear that bar, you pass. No edge, no trade — full stop, regardless of how confident the model "feels."
Why should you always challenge your own prediction?
Because confirmation bias is the single biggest leak in any forecasting system, human or AI. After you've got a prediction and a position sized, the most valuable step is asking the model to argue against itself: why might the market be right and your model be wrong?
This step matters more than people expect. In many backtests, this challenge step is exactly where you discover the market "knows" something your data pipeline missed — a late scratch, a line move driven by insider information, a factor your model simply doesn't account for. Building this adversarial check into your process, every single time, is what separates a disciplined forecaster from someone who just likes being right.
Which tools do you actually need to get started?
You don't need a hedge fund budget. Here's a minimal, practical stack:
- An AI assistant with strong reasoning — Claude, ChatGPT, or similar, used with a structured prompt chain like the one in the beat-the-market GitHub repo.
- Kalshi for regulated, U.S.-based prediction markets, with public market data endpoints that don't require API keys, allowing direct access to market data from production servers immediately.
- Polymarket for a broader range of political, crypto, and sports markets, with official API documentation covering market data and order management.
- Free public data sources relevant to your market category — nflverse for NFL stats, weather APIs for weather markets, official economic releases for macro markets.
Start on paper. Run the full eight-step chain on a market that's already resolved, compare your model's call to what actually happened, and only move real money once you've seen the process hold up across dozens of backtested examples.
What's the real takeaway here?
The market isn't something to be casually outsmarted with a clever prompt. It's millions of dollars of sharp, informed opinion constantly correcting itself. Most of the time, it's going to be right, and your model is going to be wrong. The value in building this kind of AI-assisted process isn't that it hands you easy money — it's that it forces discipline: defined rules, honest backtesting, and a built-in habit of doubting yourself before you doubt the market.
That discipline is worth more than any single winning bet. Build the habit first, and the edge — when it genuinely shows up — will actually mean something.