A leaderboard sorted by win rate is noise: a quarter of our wallet store shows exactly 100%, most on a single market. What has to be measured first.
- Auteur
- EdgeMarket
- Publié
- Temps de lecture
- 7 min de lecture
"Smart money" is a phrase that does an enormous amount of work it has never earned. In equities it usually means institutional. On-chain it usually means large. In prediction markets it usually means a wallet somebody screenshotted after it was right, which is not a category at all.
The phrase is worth keeping, but only if it is attached to something measurable. This article is about what that something has to be, and about what we found when we audited the obvious version of it on our own data.
The obvious ranking does not survive contact
The natural thing to build is a leaderboard: every wallet, its historical hit rate, its profit and loss, a composite score, sorted descending. We had one. Here is what an audit of it found, on a store of 145,141 wallets — every figure below measured on 17 August 2026:
- Every wallet carries a
win_rate. The field is never empty, which means it never tells you whether there is a sample behind the number. - 36,354 wallets — 25.05% of the store — show exactly 100%. That is not a distribution of skill. 18,914 of them have a single market counted, and 28,710 have three or fewer.
- The highest composite score anywhere in the store is 86. Six wallets score above 80, and their market counts run from 5 to 56 — the same score describes an address with five markets behind it and one with fifty-six.
- Sorting by volume returns junk at the top: 11 of the top 20 addresses by recorded volume have at most one market counted, and 2,082 wallets carry a non-zero volume with no trade at all in the trade store.
- 2,209 trades, from 1,415 addresses and two markets, carry an amount never divided down from wei. The largest reads as 8.5 × 10⁷⁰ dollars. They were copied into a quarantine table on 15 August 2026 at 23:52 UTC and are still present in the live trade table.
Publishing that leaderboard would have meant presenting statistical noise as expertise. Any trader who looked at it for thirty seconds would have seen "100% on one market" and correctly stopped trusting everything else on the page.
This is the general shape of the problem, and it is not specific to us. A ranking of anonymous addresses by realised outcomes is dominated, at every level of the table, by wallets whose sample is too small to say anything. The top of the list is a lottery result, presented as a hierarchy.
What would have to be true
For a wallet to be worth watching, three things have to hold simultaneously, and dropping any one of them produces the table above.
Enough distinct markets. Not enough trades — enough markets. Six hundred orders in one market that happened to go up is one observation, not six hundred. Treating them as independent is the single error that generates absurd significance figures.
A statistic, not a percentage. A hit rate with no denominator is unreadable. What matters is whether the result would survive being a coincidence, which is a question about the ratio of the effect to its own variability.
A measurement that does not depend on self-reporting. Declared profit and loss on-chain is reconstructable but fragile — it depends on cost basis assumptions, on which transfers you count, and on whether you can see every leg. Something more robust is needed.
What we measure instead
Rather than outcomes, we measure the price movement that follows a wallet's trades. For each trade, the price of that market is compared with where it goes by the time of the next trade in the same market, in the direction of the position taken. If an address is consistently entering just before the market reprices its way, that shows up here regardless of what its profit and loss says, and regardless of whether the market has resolved yet.
The method, exactly as applied:
- Window: trades from 1 April 2026 to 24 June 2026, at prices between 0.05 and 0.95. The extremes are excluded because they are mechanically resolved most of the time and would swamp the measurement with contracts that were never in doubt.
- For every trade, the subsequent price movement in that market, in the direction taken.
- Aggregate by market first, then by wallet. This is the step that makes the rest honest. Without it, one address in this window posted 621 observations across 13 markets and led the ranking with a t-statistic of 5.02 — and 601 of those 621 observations sat in a single market. Aggregated by market, the same address has 13 observations, a mean of −4.49 points, and a t-statistic of −0.87. The sign flips.
- Filters: at least 25 distinct markets, a t-statistic above 2.5, and the corrupted-amount rows excluded.
Here is the part the first version of this article got wrong, and it is worth stating flatly. Across 16,551 wallets and 10,940 markets, the window yields 307,540 observations. 604 wallets reach 25 distinct markets. Four of them clear a t-statistic of 2.5.
Four out of 604 is roughly what a one-sided threshold at 2.5 produces from noise alone. And of those same 604 wallets, 277 — 45.9% — have a positive mean at all, which is a coin flip. So the filter does not isolate a population of skilled addresses. On this window it isolates four addresses that are not distinguishable from the right-hand tail of chance.
Drop the market-aggregation step, keeping the same filters, and 12 wallets pass instead of four, which is the entire reason the step exists. The number of names a method returns is mostly a statement about the method.
What it does not mean
Those four are within the sample. They describe the window that was measured. They do not promise anything about the next one, and no wallet-level out-of-sample validation exists.
What does exist is validation at the level of deciles, and it is one-sided. Ranking the 727 wallets that traded at least ten distinct markets in February and March, then measuring the same statistic from 1 April to 24 June: the bottom decile ran at −1.27 points of subsequent movement, and the top decile at −0.06. The three worst deciles came in at −1.27, −0.37 and −0.45; every other decile sits between −0.11 and +0.27.
Read that honestly. The ordering carries information about who keeps losing. It carries none about who keeps winning — and that is the half everybody wants.
There is one more limit worth stating plainly, because it is easy to imply otherwise: these are wallets, not entities. No address clustering exists in our data — no clustering table, no clustering column, no entity tag. No column anywhere in the schema is named for a cluster, an entity, an owner or a group, and the tag array is empty on all 145,000-plus wallets. The readable alias an address carries is a generated display label, not a grouping: the same label lands on several unrelated addresses. The only classification that exists is a size bucket: more than 130,000 fish, over 9,000 shark, over 3,000 whale, which is a statement about position size and not about identity. An address is a pseudonym for a stream of transactions, and nothing more. What that costs you when you try to act on it is covered in seven ways copying a wallet goes wrong.
Size is not skill, and neither is confidence
The folk version of smart money conflates three unrelated things: capital, conviction and accuracy.
Capital is the easiest to observe and the least informative. A large position tells you someone had a large balance, which is a fact about their bank account. Our own store says as much: among wallets with a non-zero reconstructed profit and loss, about 52% of the whale bucket is down against about 44% of the fish bucket — the large addresses are, if anything, the slightly worse half. (More than 65,000 wallets carry a profit and loss of exactly zero and are outside that comparison, which is itself a reminder of how partial reconstructed P&L is.)
Conviction is not observable at all. A single large print looks like certainty and is equally consistent with a hedge, a wind-down, a market-making inventory adjustment, or a mistake. Five ways a large position gets built, and what each leaves on the tape goes through the shapes and, more importantly, through what each shape fails to distinguish.
Accuracy is the only one that means anything, and it is the only one that requires a denominator, a window, and a stated method. Which is why it is the one nobody puts in the headline.
What the phrase is actually good for
Used carefully, "smart money" is a shorthand for a subset of addresses whose subsequent price movement, measured over a stated window with a stated filter, is distinguishable from the rest. That is a mouthful, it is the only version that survives being checked, and on the window measured here it comes to four addresses out of 16,551 — which is to say it may come to nothing at all.
Used as a reason to enter a position, it is a way of outsourcing judgement to a stranger whose constraints, hedges, horizon and remaining balance are all invisible to you. The measurement is a prior on whose activity is worth reading. It is not a trade.
The same discipline applied to price rather than to wallets is on the public register — including the price bands where the market is right and our own signal is not. If you want the general version of the argument, forecast calibration covers how to tell a good forecast from a lucky one, and the tools a prediction-market trader actually needs covers where this kind of data sits in a stack that mostly consists of a spreadsheet.