Last month we published the grading-arbitrage math — the probability-weighted equation that tells you whether buying a raw card, grading it, and selling the slab actually makes money. It ended with a caveat that bothered us: the whole trade lives inside two numbers, your true gem rate and the 10-to-raw multiple, and most people guess at both.
So we built the thing that stops you from guessing. It's live at worththerip.com/grade, and like everything else on this site, it's free.
What it actually does
The Grading Arbitrage Scanner has two halves.
The first is a calculator that runs the full equation from the blog post on any card you type in: raw price, slab comps, grading fees, selling friction, and your honest read on the grade distribution. It gives you expected profit, ROI, and — the number we find most useful — the break-even gem rate. If a card needs a 45% gem rate to break even and it's a modern print with known centering problems, you have your answer before any money moves.
The second half is the part we're either proud of or slightly embarrassed by, depending on the day. On a schedule, a scanner walks a watchlist of cards where the 10-to-raw multiple is still wide, pulls the cheapest raw listings off eBay, and hands the seller's photos to a vision model that's been prompted to think like a grader: centering first, then surface under whatever light the photos offer, then edges and corners. It produces a probability distribution over the grade the card would actually receive — not "looks NM to me!" but 38% ten, 44% nine, and here's the bottom-edge whitening that's doing it. That distribution goes through the same EV math, and anything that clears the bar gets published to the page.
Most scans publish nothing. That's not a bug. That's the entire point of screening.
The hard part was teaching it pessimism
Anyone who has bought raw cards online knows the fundamental problem: listing photos are a marketing document. Glare hides print lines. Sleeves hide edges. A flattering angle turns 60/40 centering into "looks pretty good." A model that takes photos at face value will hand out gem rates like participation trophies, and an over-optimistic gem rate is the exact failure mode that loses real money.
So the grader is built to lean the other way. It knows the population-wide gem rate on modern Pokémon is roughly a coin flip, and that a photo merely failing to show defects is not evidence of their absence. When photo quality limits what it can verify, it shifts probability mass down toward the 9 and 8, never up. And when the photos are genuinely useless — stock images, back-of-card-only, potato resolution — it refuses to score the listing at all.
Then we check its work. Every card with a known real-world grade becomes a labeled example: the model predicts from the photos, reality supplies the answer, and the gap between the two gets baked into a calibration layer that adjusts every future prediction. If the model's "70% gem" calls only gem 55% of the time, future 70s get haircut accordingly. It's the least glamorous kind of machine learning — no training run, just an accumulating record of being wrong and a correction for it — but it's the honest kind, and every card we or anyone else actually submits makes it a little sharper.
What it can't do, stated plainly
Three limitations, because this site doesn't do fine print:
The comps are asks, not sales. Sold-price data requires API access we don't have yet, so slab values come from the median of the lowest live asks with a 10% haircut. Asks run optimistic. Treat every comp on the page as a ceiling, and check the real sold history before buying anything.
A screen is not a grader. The model sees four compressed JPEGs. PSA sees the card under halogen with a loupe. The scanner exists to tell you which listings are worth your eyeballs — the centering measurement, the zoomed surface check — not to replace them. Never buy on the model's read alone. We mean it enough that it says so on the page.
The boring risks are still yours. A 45–90 day turnaround during which prices move, shipping damage, and the fee stack on exit are all outside the model. They were outside the blog post's model too. They will eat you regardless of what any scanner says.
Why bother, then?
Because the edge in grading arbitrage — the narrow, real edge that survived 2021 — has always belonged to people who screen ruthlessly and say no a hundred times for every yes. The math was never the hard part; the discipline was. A tool that runs the math on every candidate, haircuts its own optimism, and mostly returns "not this one" is discipline you don't have to summon.
Go run your own card through it. Odds are the answer is no. Now you know for free.