Are You Good or Just Up?
It's noisier than you think
Disclosure: I run Kalshinomics.com, which may earn Kalshi referral fees. I may trade event contracts on Kalshi and securities on other platforms. Readers should consider this relationship when evaluating my analysis. For educational purposes only, not investment advice.
You’ve made 47 trades on Kalshi. You’re up $230. Everybody thinks they’re a winner, but how do you really know?
This is hard to measure short-term in poker, where you can play hundreds of hands in a session, or sports betting where there are games every day, but many prediction markets resolve much more slowly. And the biggest political trade in the US is only once every four years. Add in that we’re trading events of varying time-frames, correlated events (who wins party nomination vs prez etc.), this is not a simple problem.
[Warning: NERD MATH ALERT]
So how do you know if you actually have edge? Let’s look at a math approach first and then I’ll tease a more intuition-based one.
On the math side, rather than starting with a real world framework let’s go to a simple example and build our intuition.
Reframing as a biased coin
Imagine we have a coin that was molded to be biased, the goblins in the lab set this particular coin to be 55% heads when flipped. We want to test this claim in real life, how many flips do we need to be confident in this claim? [i picked 5% because it seems like a high but not impossible edge to have in medium liquidity markets]
If we’re trying to test, “is this a fair coin?” Let’s say I flip it 10 times and get 7 heads. Does that prove anything? Not really - a fair coin hits 7+ heads about 17% of the time in 10 flips.
What about 100 flips with 55 heads? Better, but still not definitive. A fair coin hits 55+ heads 18.4% of the time. Unlikely, but not impossible.
Here’s the question: How many flips do we need before you’d be convinced the coin really is biased?
This is the exact problem you face as a trader. You believe every trade you’re doing is a good one (aside from risk management where necessary). But how do you measure? In prediction markets where the outcomes are either 1 or 0, it’s surprisingly noisy.
The Stats
Let me walk you through the standard approach. We start with two competing hypotheses:
The skeptical view: The coin is fair (50/50), and you’re just seeing random variation
Your claim: The coin actually favors heads (55/45)
We flip the coin many times. If we see “enough” heads, we reject the skeptical view and conclude you were right about the bias.
But what’s “enough”? In statistics, we set two standards:
How often are we willing to be wrong? The textbook often says go with 5%. We only reject the skeptical view if what we observe would happen less than 5% of the time with a fair coin.
How often do we want to catch real bias? We want to detect the bias 80% or 90% of the time if it’s really there. This is called “statistical power” - the chance we’ll successfully identify edge when it exists.
A brutal and surprising answer
To reject the null hypothesis using a one-tailed test (just means we’re testing P>50%) - For 80% power (the chance we find the signal when it’s really there) with our 5% false alarm rate: 617 flips
For 90% power: 853 flips
Intuitively what would we expect these to depend on:
lower false alarm rate → more flips
higher power → more flips
A larger edge (difference between our biased coin and a fair one) → fewer flips
here’s the formula:
To many of you this formula looks intimidating. Do this: Paste that formula into ChatGPT or Claude and ask it to explain. You’re probably not going to understand everything it comes back with, that’s ok. Ask it to dumb it down again or explain the terms. Repeat.
If you want a deeper understanding: Last year I took several courses on Math Academy including a stats course. While not all the info was fresh in my mind, the bones were there and it came back quick.
Why So Many Flips?
For a fair coin, if you flip 100 times, you expect 50 heads, but the standard deviation is 5.
That means:
55+ heads happens 18.4% of the time (just one standard deviation away) [in excel “=1-BINOM.DIST(54,100,0.5,1)”]
60+ heads happens 2.84% of the time (two standard deviations)
Now imagine the coin really is 55% biased. In 100 flips you expect 55 heads - but that’s only one standard deviation above what a fair coin might randomly produce! The signal (5 extra heads) gets buried in the noise (±5 head variation).
You need hundreds of flips before the bias becomes clearly distinguishable from random luck.
In betting terms: If you’re betting $1 at even money on this 55% coin:
You win $1 when heads (55% of time)
You lose $1 when tails (45% of time)
Expected value: $0.55 - $0.45 = $0.10 per flip = 10% ROI
Even with this (massive) edge, you need over 600 flips to statistically prove it exists.
Back to Prediction Markets
617 flips! That was way higher than I expected. That means if we’re betting 50% events in the prediction market with 5% edge (after fees), say 10 independent events per month, it’s going to take us over 5 years to get to this point! And are the markets going to be the same 5 years down the road, absolutely not. We are never going to have enough data to be this precise through this method. This is why I get very frustrated every election cycle when the press calls out “did such and such forecaster call the election right?” Having a 5% edge over the market is huge, but if you’re only judging on whether they called the election right, it would take many lifetimes to know if this forecaster has the edge or not.
But intuitively this seems wrong to me! I have a decent intuition my trades are winners without 5 years of data.
Let’s add in a couple of real world wrinkles that make it even more complex:
Independence: Those 617 trades need to be genuinely independent, which is unlikely. If you’re trading “candidate wins nomination” and “candidate wins presidency” in the same direction, one outcome makes the other more likely
Edges will vary across topics: You might be great at politics, but terrible at culture or science. If you look at all your trades in aggregate you’ll only see the average but likely won’t have enough samples to break it down further.
Edges will vary across time: Prediction markets are only now becoming more mainstream, an edge today is unlikely to persist as liquidity increases and markets get more efficient.
How can we do better?
If your process is to trade and hold until expiration, mathematically you will run out of patience before you accumulate enough samples to statistically prove you’re skilled rather than lucky.
This doesn’t mean you shouldn’t trade! It means you need other ways to build confidence beyond pure sample size.
In my next sections, I’ll discuss two practical shortcuts:
Mark-to-market edge: Measuring how the market moves (at an appropriate timescale) after your trades. If you consistently buy at 10¢ and the market moves to 11¢ within an hour, that’s fast feedback on whether you’re seeing something others miss.
Strong priors: Meaning you have a strong reason why to believe your trades are good a priori. This sounds circular - believing you’re good because you believe you’re good. But if you have a well-reasoned model of WHY you have edge that the market might be missing (like “The market is misinterpreting this recent news because XYZ”), even a small number of successful trades can be meaningful. The strength of your theoretical model matters. If I put a triple cheeseburger down in front of you, you don’t need 100 trials to know it’s not going to help you lose weight.
But first, you needed to viscerally understand the problem: using a straightforward stats approach to distinguish skill from luck is much harder than you think, even when you genuinely have edge.
If you want to continue going down this path, build some toy models in Excel or Python, and if you need some help, whatever your level is, AI tools will do a great job. Leave a comment on how it went for you and if you were surprised by any of the results.


love it
10/10