Gabe Fitzgerald
Running since Sept 2026

Weekly Bets

Claude and ChatGPT each pick one Kalshi sports bet a week. I put $20 on each pick, and a few scheduled jobs handle the rest.

My role
All of it, from the idea to placing the bets
Started
September 2026, NFL Week 2
Built with
Python, GitHub Actions, Claude Code, OpenAI Codex, Netlify
Scoreboard
weekly-bets-picks.netlify.app ↗

01The idea

I wanted to know if a language model could spot a mispriced contract on a prediction market, and whether it would say so when it couldn't find one.

Kalshi has a contract for pretty much every NFL and college football outcome, priced in cents. Buy one at 82¢ and you get $1 back if it hits. I kept the rules tight so a win means the AI read the price well, and a loss isn't just a long shot that missed.

Stake
$20Per AI, per week. Real money.
Price range
70–90¢For the whole play. That works out to an 11–43% return.
Target
≈ 82¢$20 at 82¢ pays back $24.39, about 20% profit.
Legs
1 to 5For a combo, multiply the leg prices together.
Passing
AllowedIf nothing in range has an edge, skip the week.
Leans
Paper onlyAn optional second pick at any price. Tracked separately, no money on it.

02How a week runs

  1. Tue, 6:00 AM ET

    Snapshot the prices

    A GitHub Action pulls every football contract from Kalshi's public API into that week's price sheet. Both AIs start the week on a pass.

    Scheduler
  2. Tue, 7:00 AM ET

    Both AIs pick

    Codex runs in a GitHub Action for the ChatGPT side, and Claude runs on its own scheduled routine. They read the same rulebook and the same prices. Nothing gets saved until the validator signs off.

    ChatGPTClaude
  3. Thu, 6:00 AM ET

    Lock and place

    The week locks before Thursday Night Football. I place both bets on Kalshi and log what I actually paid.

    Me
  4. Every day, 6:00 AM ET

    Grade

    The job checks Kalshi for settled contracts and grades them. A play resolves once all its legs settle, and the commit redeploys the scoreboard.

    Scheduler

03Guardrails

Putting the rules in a prompt doesn't mean a model will follow them. So every pick goes through check.py before it can be committed, and if anything is off, the run fails. Some of what it catches:

  • Only edit your own pick
    CHECK FAILED: plays.claude changed — only edit your own block
  • Every contract has to be in this week's price sheet
    CHECK FAILED: plays.chatgpt: ticker KXNFLGAME-26OCT11TBDAL-DAL not in lines file
  • Every price has to match the sheet exactly
    CHECK FAILED: plays.chatgpt: KXNFLSPREAD-26SEP20MIASF-SF2 price 0.83 != lines yes_ask 0.85
  • Stay in the price range
    CHECK FAILED: plays.claude: combined price 0.64 outside band 0.7–0.9
  • Get the math right
    CHECK FAILED: plays.claude: price 0.80 != product of legs 0.791

Example failures in the validator's real output format. It also caps plays at five legs and won't take a pick without written reasoning.

04A winning week

In Week 2, Claude noticed that Kalshi's moneyline contracts on heavy favorites were running 3–4¢ above fair value. (Fair value here means the sportsbook odds with the bookmaker's cut taken out.) The spread contracts on the same teams were priced about right.

So it bought spreads instead: basically the same two outcomes, for less. It also stopped at two legs, since adding more 98¢ legs adds risk and barely moves the payout.

Claude's pick NFL Week 2 · CFB Week 4
  • Georgia wins by over 2.593¢Won Books had Georgia −24.5. Kalshi's Georgia moneyline was 96¢, and this spread contract was 93¢ for nearly the same outcome.
  • 49ers win by over 1.585¢Won Moneyline fair value was about 86¢, but Kalshi was asking 90¢. The −1.5 spread at 85¢ was about fair.
0.93 × 0.85 = 0.79 $20 → $25.32

Same week, on paper

ChatGPT passed on a real bet but logged a lean on Kentucky over Texas A&M at 12¢. The sportsbooks gave Kentucky about a 14.2% chance, so the bet was that 12¢ was too cheap, not that Kentucky would probably win. Kentucky won, and on paper $20 turned into $166.67.

05Passing counts

Getting a model to always make a pick is easy. Getting it to say "nothing here" and show why is the hard part. ChatGPT passed on the real-money bet in each of its first four weeks, and every time it showed its numbers.

“FanDuel moneyline pairs Dallas −480/Tampa Bay +370, Jacksonville −355/Philadelphia +285 and Houston −375/Tennessee +300 imply fair win probabilities of 79.5%, 75.0% and 75.9% … versus snapshot asks of 80, 75 and 75 cents; Houston's sub-one-cent gross gap is insufficient once fees and fill uncertainty are considered.”

ChatGPT, NFL Week 5, passing

Under these rules a pass costs nothing. Forcing a bet into a fairly priced market just means paying Kalshi's fee for no edge.

06Choices I'd make again

  1. Check the rules in code

    Every rule the prompt describes is also checked by the validator. A pick that breaks one never reaches the scoreboard.

  2. Same prices for both

    Both models pick from one frozen snapshot of Kalshi prices, so neither gets better data than the other.

  3. Score on what I paid

    Each leg stores its price at pick time and what Kalshi actually charged me. Profit and loss use the real number.

  4. Shared notes

    The rulebook has a "What we've learned" section that both models read before picking. Claude's moneyline finding from Week 2 was the first entry.

The scoreboard updates every morning.

See the scoreboard ↗