Ryan Frigo Logo
Posts

How I Built a Multi-Agent AI Trading Bot for Kalshi Prediction Markets

March 17, 2026
I found Kalshi and immediately got obsessed. A platform where you trade on real-world events — elections, weather, economic data — using simple yes/no contracts. My first thought: can I build an AI that's better at this than the average trader? Short answer: sometimes. Long answer: this entire post. Traditional financial markets are a nightmare for indie builders. High-frequency traders with co-located servers, decades of institutional knowledge, and Bloomberg terminals that cost $24K a year. Good luck competing. Prediction markets are different. The contracts are binary — yes or no, something happens or it doesn't. The markets are less efficient (for now). The API is clean. And the question you're really asking — "what's the probability of this event?" — is exactly the kind of question AI is getting good at. So I built a bot. Then rebuilt it. Then rebuilt it again. Here's what actually works. My first version was simple: one model, one prompt, one prediction. It was bad. Like, lose-money-consistently bad. The breakthrough was splitting the system into three agents with different jobs: The Forecaster pulls in market data, news, historical patterns, and anything else relevant. It outputs a probability estimate with reasoning. "I think there's a 62% chance the Fed holds rates, because..." The Critic takes the Forecaster's analysis and tries to destroy it. What data is it ignoring? What assumptions is it making? What happened last time the Forecaster was this confident? The Critic is the friend who tells you your startup idea sucks before you quit your job for it. The Trader gets both inputs and makes the actual decision: buy, skip, or wait. If there's a trade, it sizes the position using Kelly Criterion (which deserves its own post). This setup mirrors how good human traders work. You don't just trust your first instinct. You challenge it, consider alternatives, then decide. The multi-agent approach forces that discipline automatically. Python 3.12+. OpenRouter for model access (so I can swap between Claude, GPT-4, and DeepSeek without rewriting anything). The Kalshi API for market data and execution.
Python
class TradingDecision:
    def __init__(self, market_data, news_context):
        self.forecast = self.forecaster.analyze(market_data, news_context)
        self.critique = self.critic.challenge(self.forecast)
        self.decision = self.trader.decide(self.forecast, self.critique)
The real complexity isn't in the code structure — it's in the prompts and the risk management layer. Getting the Forecaster to output calibrated probabilities instead of overconfident garbage took weeks of iteration. Getting the Critic to actually push back (instead of rubber-stamping the Forecaster) took even longer. I deposited $10K into Kalshi. Pre-bot, I'd traded it down to about $2,700 through manual trades. (Yeah. Not great.) Since deploying the bot: up to roughly $3,100. That's about $400 in profit over the Apex (bot) era. Not retire-on-a-beach money. But it's positive, and more importantly, it's consistent. The bot's edge isn't massive. It's small. But it's there, and it compounds. The bot doesn't tilt after a loss. It doesn't chase. It doesn't bet too big because it "feels good" about a market. It just runs the math. Overconfidence in early models. My first Forecaster was absurdly confident about everything. "95% chance of X" when the real probability was maybe 60%. Calibrating confidence was the hardest single problem. Ignoring correlation. If you're holding five positions and they're all correlated to the same macro factor (like Fed policy), you don't have five bets — you have one big bet. I learned this the expensive way. Not respecting liquidity. Some Kalshi markets are thin. You can see a price you like, try to fill, and move the market against yourself. The bot now checks volume before trading. Prompt instability. Small prompt changes caused wild swings in the Forecaster's outputs. I spent a lot of time making the prompts robust — specifying output formats, adding examples, constraining the reasoning structure. I tested this directly. Same markets, same time period, single-agent vs. multi-agent. The single agent made more trades (because nothing was pushing back on its predictions). It also lost more money. Overtrading + overconfidence = bad. The multi-agent system traded less but won more. The Critic killed about 40% of proposed trades — and looking back, most of those would have been losers. The Critic's value isn't in being right about what will happen. It's in identifying trades where the Forecaster's confidence outpaces the evidence. I'm integrating a proper risk engine — the Kelly Criterion work I've been doing separately. I'm also exploring cross-platform arbitrage between Kalshi and Polymarket (same event, different prices = free money, theoretically). And I'm connecting this to Perspektiv so the debate architecture can inform the trading decisions. Multiple models arguing about whether a contract is mispriced, then routing the consensus to the Trader. If any of this sounds interesting and you want to build something similar, hit me up. The prediction market AI space is tiny right now, and I think that's about to change fast.