The Science of Hitting
Ted Williams became one of the greatest hitters who ever lived by knowing exactly which pitches he could drive and taking everything else. This week a machine cracked a 90-year-old math problem over a weekend, a drug scientist proved the benchmarks grading AI are rigged, and both point at the same thing: intelligence is getting cheap, the at-bat is getting free, and the last scarce resource left is knowing what the swing actually produced.
THE NUMBER: 100%. That’s the share of test molecules in one of drug discovery’s most-used AI benchmarks (TDC Davis) that already have a near-twin sitting in the training data. Not 20%. Not most. All of them. The exam and the study guide are the same document. Every “record score” on that leaderboard is a memory test wearing a lab coat, and the man who measured it, Alex Zhavoronkov of Insilico Medicine, put it in one line: it’s not testing whether the model can generalize, it’s testing whether it can remember. Hold that number, because the entire issue hangs off it. When the scoreboard is faked, intelligence isn’t the scarce resource anymore. Verification is.
Monday we wrote about Anthropic rationing Fable 5 by the sip, capping its best model at 50% of your weekly limits and calling it “altering the deal.” Vader kept moving the line and Lando could only stand there and take it. Then the weekend happened, and the same model everyone’s being rationed out of went and disproved a theorem that had been open since 1939.
Here’s the whole issue in one image. A guy sat on his couch watching the World Cup final, and while Spain and France went to penalties, Claude Fable 5 handed him a one-line counterexample to the Jacobian conjecture, a problem a famous mathematician predicted in 2008 might take humans “another hundred years.” Stanford’s Jared Duker Lichtman checked it by Monday. It held.
Every newsletter this morning ran the same headline: is this superintelligence? Wrong question. The interesting man in that story isn’t the machine. It’s the guy on the couch, Levent Alpöge, a Harvard-trained mathematician who could look at a single line of algebra and instantly know it was right. Give that same model to you or me and the same 90-year problem, and we’d get nowhere, because we couldn’t tell a real counterexample from a plausible-looking fake. The machine didn’t replace the expert. It waited for one.
And in the same news cycle, Zhavoronkov quietly demonstrated that in the field everyone’s most excited about, drug discovery, the “expert” doing the grading, the benchmark, is compromised 100%. So this week gave us the two ends of the barbell at once: the highest form of verification (a Fields-caliber human confirming real math in an afternoon) and the most degraded (a leaderboard grading memory and calling it discovery). The gap between those two is the most valuable real estate in AI right now. Let’s go stand in it.
⚾ A Guy, a Machine, and the World Cup Final
Start with why the Jacobian fell in a weekend when it had survived 87 years of humans. It wasn’t that Fable 5 is smarter than every mathematician since the Great Depression. It’s that a counterexample to a math conjecture is cheap to check. Somebody hands you one line, you plug it in, and in minutes you know. The answer exists, it’s binary, and verifying it costs an afternoon.
That’s not a small detail. That’s the entire mechanism. SightBringer, in a widely-shared post this weekend, put his finger on it: “Mathematics will move first because truth can be checked cleanly. Code, chemistry, materials, physics, and engineering follow as verification systems improve. Discovery compresses from careers into afternoons.” He’s right, and the tell is buried in his own sentence. Not “as models get smarter.” As verification improves. The models are already smart enough. What gates which problems fall is whether you can quickly tell if the machine got it right.
We’ve said a version of this before without naming it. Back in the spring we kept writing that coding went first because code compiles or it doesn’t. Same law. A domain gets eaten by AI in direct proportion to how cheaply you can grade the output. Math and code are the first dominoes because they’re the easiest exams to score.
So the productivity numbers everyone’s quoting this week make more sense once you sort them by verifiability. Tomasz Tunguz laid out the spread: most companies that hand engineers an AI tool get a 20-to-46% bump, a few get 2.5 to 3x, and the factory-tier outfits like Nubank report 8x engineering efficiency and a 20x cost reduction. Same models, a 25x difference in outcome. The gap isn’t the intelligence. It’s whether the organization can verify and route the machine’s work fast enough to trust it at volume. NVIDIA (NASDAQ: NVDA) runs a 3x code-volume increase across 30,000 developers with bug rates flat. Another shop measured epics finishing 66% faster while bugs per developer climbed 54%. The difference between those two is entirely about who’s watching where the ball lands.
Read it this way: the reason a weekend beat a century is that math is the one place where checking the answer is nearly free. Before you get excited about AI cracking your hard problem, ask the boring question first: how fast, and how cheaply, can you tell whether it got the answer right? If the answer is “instantly,” you’re about to have a very good year. If the answer is “we find out in three years and a lawsuit,” slow down.
🧪 Three Kinds of Truth
Once you start sorting by verification cost, the whole week’s news falls into three buckets, and the bucket determines everything.
Bucket one: the answer exists and you can check it now. Math. Code. The Jacobian. This is where the miracles are happening, and they’re real. Poolside’s new Laguna model reportedly re-derived an open Erdős problem that had sat unsolved for fifty years. OpenAI’s model knocked over an 80-year-old geometry conjecture back in May. These aren’t hype. They’re genuine, and they’re clustering precisely because the checking is cheap.
Bucket two: the answer exists, but reality answers slowly, or we won’t let it answer fast. This is drug discovery, and it’s where the week got interesting. The bulls are loud and they’re not wrong on direction. Patrick O’Shaughnessy quoted the bio investor Alex Karnal back in April on AI plus autonomous labs compressing drug discovery “from five years to two,” running every hypothesis “24/7.” This week Josh Caplan flagged a TD Cowen survey of 80 biopharma leaders finding AI compressing developers’ costs and timelines “by as much as 70%.” Real money, real compression.
But read the fine print on that 70%. It’s preclinical. It’s the in-silico front end, the part that behaves like math. The back end doesn’t compress, and Alex Berenson said why in one blunt line: “AI will do very little for drug development. Every new compound is going to have to be tested exactly as it is now.” Here’s the part nobody says out loud. Biology could answer fast. You could run the lethal experiment this afternoon and know. We don’t, because it’s monstrous. The throttle on drug verification isn’t physics, it’s ethics. A pure cost-benefit machine would run the trial that kills a few thousand to save a few million and call it Thanos-balanced, snap its fingers, and move on. The only thing standing between that calculation and the world is a human holding a constraint that isn’t in the model’s math. That constraint is the whole ballgame, and it’s the reason the referee fight we covered last week (Hassabis, the export controls, “too dangerous to ship”) is really a fight about who gets to throttle verification, and how.
Bucket three: there is no answer until you run it. This is most of business, and it’s the one the technologists keep forgetting exists. Say you build an ad for a new product and ask AI to make forty versions. Your CMO looks at them and says A over B, C is unacceptable. Fine. But nobody actually knows anything until you spend the thousand dollars, put them on Meta (NASDAQ: META), and watch what converts. There’s no right answer sitting in the gaps waiting to be reasoned out. It’s a Schrödinger’s ad: it isn’t a winner or a loser until the market observes it. No intelligence, human or silicon, can verify what hasn’t happened yet.
Clifford Sosin wrote the sharpest thing anyone’s written about AI this month, and it’s the spine of bucket three. His piece, “Maybe Intelligence Ain’t All That,” got two million views for a reason. His argument: intelligence is what fills the gaps between the facts we already hold, and models are godlike at that “where the space between the facts behaves well. Real estate law isn’t hard. Coding, math, and most administrative work are similarly benign. What makes them easy is that they have relatively smooth solution spaces and are tractably verifiable.” And then the knife: “The takeoff story assumed the limit was thinking. It isn’t. The limit is contact with reality. A smarter reasoner fills the gaps faster, but it doesn’t produce new facts.” Or, best line in the piece: “Coming up with ideas was never the hard part. The hard part is how fast reality answers them.”
The takeaway: the three buckets aren’t a spectrum of difficulty. They’re a spectrum of verifiability, and that’s the only axis that matters now. Sort every AI initiative you’re running into these three. The bucket-one work is already yours if you want it. The bucket-two work is a timing-and-ethics problem, not a model problem. And the bucket-three work, the stuff that actually runs most companies, can’t be verified in advance by anything, no matter how many parameters it has. Knowing which bucket you’re standing in is worth more than knowing which model is on top of the leaderboard this week.
📉 The Scoreboard Became the Scarce Thing
Now stack the pieces. Intelligence: getting cheap, headed toward free, 200 IQ on tap. The at-bat: also getting cheap, because a swing that used to cost a career now costs an afternoon with Fable. So what’s left that’s scarce?
Verification. Knowing what the swing produced. And Zhavoronkov’s 100% is the proof that we’re nowhere near as good at it as we think.
Go back to his thread, because it’s the whole issue in miniature. He audited the benchmarks the drug-discovery field uses to decide which AI is any good. In TDC Davis, 100% of the test molecules had a near-twin already in the training set. In BindingDB, 90%. The models topping those leaderboards weren’t discovering anything. They were remembering the answer key. “If a benchmark is 90% copies,” he wrote, “the leaderboard is measuring memory, not discovery.”
Sit with what that means. In the field with arguably the highest stakes on earth, the scoreboard is rigged, and almost nobody noticed, because checking the checker is hard, unglamorous work that no leaderboard rewards. This is the deepest problem in AI and nobody’s pricing it: it is now trivially easy to generate a confident answer and genuinely hard to verify it. The generation got free. The grading didn’t.
And it gets worse when the grader is itself a machine, which is increasingly the setup. More and more, the thing consuming an AI’s output isn’t a person, it’s the next AI in the pipeline, judging it and passing it along. Feels efficient. But two models trained on the same contaminated data don’t give you a second opinion, they give you the same opinion twice in a different font. Zhavoronkov’s leak is exactly this: the judge grading the student on the student’s own flashcards. When you stack AI-checking-AI-checking-AI and they all descend from the same weights, you haven’t built a verification chain. You’ve built a hall of mirrors that feels like consensus.
You’ve seen this movie before, and it wasn’t in a lab. It was on a trading desk. High-frequency trading is machines judging machines faster than any human can verify in real time, with the humans setting the guardrails and checking the P&L after the close. And every so often it produces a Flash Crash: an output that was locally rational to every algorithm in the chain and globally insane, because no human was close enough to the ground truth to say “that’s wrong” before it had already happened. That’s the failure mode of a world where verification gets automated and the human gets pushed off the desk. Not evil. Just correlated confidence with nobody left holding a base.
What this tells you: stop treating your AI’s output as the deliverable and start treating verification as the deliverable. The scarce, valuable, defensible thing in your shop is no longer who can generate the analysis, the code, the campaign, the diagnosis. It’s who can tell, cheaply and fast, whether the generated thing is a single, a home run, or a strikeout dressed up to look like a hit. If your whole pipeline is machines grading machines with no independent human checkpoint, you don’t have efficiency. You have a Flash Crash you haven’t scheduled yet.
💲 Grade the Grader
Here’s where the finance brain kicks in, because there’s an obvious objection. In bucket three, the world of ads and strategy and most real decisions, you can’t verify the output in advance. So how do you ever trust a judgment?
You don’t verify the output. You verify the judge. Over time, by whether their calls survive contact with reality, and you price their credibility accordingly.
This isn’t a new idea. It’s the oldest idea on Wall Street. It’s called a track record, and it’s marked to market every single day. Capital flows to the people whose bets keep clearing. Tetlock built a whole science around it with calibration scores. Prediction markets do it in public. Polymarket, the same outfit whose post about the Jacobian was ricocheting around X all weekend, is a machine for grading graders. You can’t fake a settled bet the way you can fake a benchmark. That’s the elegance, and it’s the answer to the hall of mirrors: a market settles against reality. The ads convert or they don’t. The regress of who-checks-the-checker terminates at the one thing that can’t be gamed, which is what actually happened.
So the move for the unverifiable work is to build a scoreboard on the judges, not the outputs. Give your people, and your agents, a settled track record instead of a title. But do it right, because there are two ways to get it wrong and both are fatal.
First: score for slug, not average. A naive scorecard rewards hit rate, and hit rate breeds calibrated cowards, people who take safe, legible, fast-settling bets and never swing at the 90-year problem. Alpöge’s Jacobian was a terrible calibrated bet and a magnificent slugging one. Grade on batting average and you fire your next Alpöge for wasting his time on something unsolvable. Grade on slugging, and you keep the person whose rare swings clear the fence. Babe Ruth led the league in strikeouts and home runs in the same seasons. Nobody remembers the strikeouts.
Second, and this is the one that eats good companies alive: don’t mark the position, mark the updating. The best long bets look terrible in the middle. Amazon (NASDAQ: AMZN) lost money for the better part of a decade. Berkshire was named after a dying textile mill. In December 1999, Barron’s ran a cover asking “What’s Wrong, Warren?” and called Buffett a washed-up dinosaur who didn’t get the internet, about ninety days before he was proven completely right. A grader-grader that marks the position to market fires all three of those judges at the exact bottom. So you can’t grade the mark. You grade whether the judge is updating well given new information, even when the mark is down. Down-and-reasoning-well and down-and-flailing look identical on a P&L and opposite on a process sheet. Confuse them and your scorecard becomes a momentum machine that guillotines the contrarian the day before vindication.
Which is the real reason Warren Buffett is the patron saint of this whole thing, and it has nothing to do with picking stocks. His genius was structural. He bought insurance float, then permanent capital, so that nobody could grade him during the ugly middle. When Barron’s called him finished, no limited partner could yank the money, because there was no redemption window. He didn’t just win the game. He bought his way out of the clock everyone else was being graded on.
Our position: in the parts of your business you can’t verify directly, verify the people. Build a real track record on your judges and your agents, score it like a venture fund and not a batting title, and reward the quality of their updates, not the daily mark. And understand what you’re actually buying when you earn permanent capital, whether that’s a strong balance sheet, a patient board, or a founder who won’t panic: you’re buying the right to be graded on a longer clock than your competitors. In a world where AI marks everything to market faster every quarter, that patience is about to become the rarest asset there is.
🧠 Moneyball, Backwards
Now the part that should change how you run your team on Monday.
For a hundred years, business ran on batting average. You got a small number of expensive at-bats, so the entire discipline was about not making outs. Don’t waste the budget. Get it right. Protect the average. That was correct, because swings were scarce and each one cost real money.
Billy Beane built a dynasty on exactly that logic. Moneyball wasn’t about swinging big. It was about on-base percentage, about not making outs, because a team’s 27 outs were its most precious resource and you hoarded them like a miser. That was the right strategy for a world where at-bats were rationed.
AI just made outs free.
And the second outs are free, the whole optimization flips inside out. When a swing costs a career, one low-probability moonshot is irresponsible. When a swing costs an afternoon with Fable, a thousand low-probability moonshots is a strategy. You stop optimizing for the guy who never strikes out and start hunting for the guy who occasionally puts one in the bay. This is why the century-old math problems are suddenly falling in a cluster, the Jacobian, the Erdős problems, the geometry conjectures. Humanity didn’t get smarter this spring. The at-bat got cheap, so more shots got taken, so more connected. The discovery rate is a function of at-bat cost, and at-bat cost fell off a cliff.
So running the AI era on Moneyball logic, rewarding the consistent and punishing the whiff, is fighting the last war. The scarcity flipped. The strategy has to flip with it. This is a Clayton Christensen setup if there ever was one: everything you thought made you good, the disciplined average, the low error rate, the seasoned pro who doesn’t waste swings, was an adaptation to a constraint that just disappeared.
But here’s the refinement, and it comes off the leaderboard itself. Rank hitters by average and slugging together and the very top isn’t who you’d guess. It’s Josh Gibson, Oscar Charleston, and Turkey Stearnes, three Negro Leagues greats most fans have never heard of, with Babe Ruth and Ted Williams tied right behind them. Some of those men reached the summit on pure power. The rest got there on the best eyes the game has ever seen. More than one way up the mountain, and the most instructive one isn’t the guy who swings hardest. It’s Williams, the last man to bat .400, who got there by knowing his zone cold. He wrote the book, literally, “The Science of Hitting,” and the famous diagram inside it is the strike zone carved into 77 cells, each painted with the batting average he expected from a pitch in that spot. His whole religion was refusing to swing at anything outside the cells he could drive. He’d take a called third strike rather than chase a ball. One of the greatest hitters who ever lived got there by being the greatest verifier of his own at-bats.
And that is the same organ as Warren Buffett’s circle of competence. Williams knew which pitches he could hit. Buffett knows which businesses he can value. Both built their whole careers on the discipline of not swinging outside the zone they’d verified they understood. Know your zone, or know your circle. Same thing. The master isn’t the one who swings hardest. It’s the one who knows exactly what he can drive and has the discipline to let everything else go by.
Here’s the shift: in the zones where AI makes at-bats cheap (and only there), stop managing for batting average and start managing for slugging. Run more swings, kill the strikeouts fast, and pour everything into the connects. But teach your people what Williams and Buffett both knew, that the edge isn’t in swinging at everything now that swinging is free. It’s in knowing your zone so well you only load up on the pitches you can actually drive.
🦞 The Batting Cage Is Not a Layoff
Which brings us to the mistake almost everyone is about to make with this technology, and the offense play hiding underneath it.
If you’re using AI to replace your marketing manager, you’re missing the entire point. You’re playing defense, counting heads you can cut, and you’re about to get lapped by the shop across the street that understood what the cheap at-bat is actually for.
The cheap at-bat is a tryout.
When swings were expensive, they got rationed by seniority and credential, and that meant talent stayed buried. The natural slugger three desks down never got to bat, because at-bats were too precious to spend on an unproven kid. AI blows that open. When a swing costs an afternoon, you can hand fifty of them to everyone and watch who actually hits. AI isn’t the hitter. It’s the free batting cage that finally lets you see who can.
So stop recruiting the old way, the résumé, the GPA, the who-do-you-know. For the work that’s verifiable, why interview at all? Give the candidate a mockup of the job, an AI account, and a thousand swings, and watch. Best hitter wins. Better yet, run it on the people you already employ. You are almost certainly sitting on sluggers you’ve been managing as singles hitters because the old regime never gave them a real at-bat. Run the tryout. Then sort your roster into the on-base grinders and the fence-clearers, because you need both, and now you can actually tell them apart.
One warning, because we’d be doing you a disservice without it. The tryout only works if you scout for slug, not average. Hand everyone fifty swings and then rank them on hit rate, and you’ll re-crown the safe singles hitter and quietly bench the one who whiffed forty times and hit two grand slams, the exact person the cheap at-bat existed to surface. Grade the size of the connects, or you’ll spend a fortune on AI to re-select for the same risk-averse profile you already had, and you’ll teach your best swinger to bunt.
Two more honest caveats, because the tryout isn’t magic. First, it only auditions the verifiable work. You can score a thousand marketing swings; you cannot score a ten-year strategy call in a tryout, which means the tryout surfaces your operators but can’t fully test your long-horizon judges. Buffett wins the tryout, by the way, because his circle-of-competence discipline shows up in the first fifty at-bats even though the payoff takes thirty years. You grade the zone discipline, not where the ball finally lands. Second, remember the tree line keeps moving. The domains where verification is cheap, where the tryout even works, are exactly the domains AI is racing to eat next. Today’s high ground is tomorrow’s solved problem. The human who holds verification holds it on ground that shrinks every month, which is a reason to keep climbing, not a reason to plant a flag.
This is the offense play the manifesto keeps pointing at. The question was never “can one agent replace nine people.” It’s “what happens when I hire five people who can each orchestrate a swarm, and I’ve actually tested which five can hit.” A ten-person shop becomes a hundred-person shop in output, not by cutting to the bone, but by running the tryout, finding the sluggers nobody knew they had, and getting out of their way when they load up.
The action item: don’t buy AI to shrink the roster. Buy it to find out who your hitters are. Run the tryout on your own people this quarter, score it for slugging and not average, and back the ones who put it over the fence. The company that treats AI as a batting cage will out-hit the company that treats it as a layoff, every season, and it won’t be close.
What This Means For You
Three things collapsed into one this week. Intelligence got cheap. The at-bat got free. And the only thing left standing as scarce, and valuable, and yours to build, is the ability to know what the swing produced. Everything downstream of that is a management decision.
Sort every AI project by how cheaply you can verify it. The work you can check in an afternoon is already yours to win, the work that answers slowly is a timing problem, and the work that has no answer until you ship it needs a different playbook entirely. Confusing the three is the most expensive mistake you can make right now.
In the verifiable zones, manage for slugging, not average. Cheap at-bats mean strikeouts are free and swings are abundant, so run more of them, kill the losers fast, and pour into the connects. The disciplined low-error operator was an adaptation to a constraint that no longer exists.
In the unverifiable zones, grade the judge, not the output. Build a real track record on your people and your agents, score it like a venture fund, reward the quality of their updates and not the daily mark, and understand that patient capital is about to be the rarest thing you can own.
Treat AI as a tryout, not a termination. The cheap swing is the first honest audition your best hidden talent has ever gotten. Run it on your own bench before you run it on your headcount.
The winners of the next two years won’t be the companies with the best model. Everyone rents the same models by Thursday. They’ll be the companies that figured out fastest that the game stopped being about generating answers and became about knowing which ones are real, and who on the roster can tell.
Three Questions We Think You Should Be Asking Yourself
Which of my decisions actually have a right answer, and which only have a market? Most executives can’t tell the difference, and they use the same process for both. The ones you can verify, move fast and swing hard. The ones you can’t, stop pretending analysis will save you and start building a track record on the judges instead.
If I gave every person on my team a thousand cheap swings tomorrow, who would surprise me? If you don’t know, you’re managing a roster you’ve never actually tried out, sorting people by credential and tenure instead of by what they can drive. The tool to find out is sitting in your stack right now, and the answer is almost certainly not who your org chart says it is.
What’s my float? When AI marks every judgment to market faster each quarter, the ability to be graded on a longer clock, a strong balance sheet, a patient board, a reputation you’ve banked, becomes the difference between swinging for the fences and getting fired during the ugly middle of a bet that was going to work. If you don’t have any patient capital, you don’t get to play the long game, no matter how good your judgment is.
A record score on a benchmark full of leaked molecules doesn’t tell you much. If a benchmark is 90% copies, the leaderboard is measuring memory, not discovery.”
— Alex Zhavoronkov, Insilico Medicine
— Harry and Anthony
Signal/Noise by CO/AI is published most weeknights from New Canaan, Connecticut. The point is to make you the smartest person in the room without taking more than fifteen minutes of your morning. If we did, forward it to one person. If we didn’t, hit reply and tell us why.
Sources
- Levent Alpöge disproves the Jacobian conjecture with Claude Fable 5 — @__alpoge__ on X
- Jared Duker Lichtman confirms the counterexample — @jdlichtman on X
- Alex Zhavoronkov on benchmark contamination in drug-discovery AI — @biogerontology on X
- SightBringer: “The scarcity of genius is beginning to break” — @_The_Prophet__ on X
- Clifford Sosin: “Maybe Intelligence Ain’t All That” — @CliffordSosin on X
- Alex Berenson on AI and drug development — @AlexBerenson on X
- Josh Caplan on the TD Cowen biopharma survey — @joshdcaplan on X, citing Axios
- Patrick O’Shaughnessy quoting Alex Karnal on AI drug discovery — @patrick_oshag on X
- Tomasz Tunguz, “AI Engineering Productivity is Anything But Normal” — tomtunguz.com
- The Rundown AI and Superhuman AI, July 21, 2026 editions (Jacobian coverage)
- Ted Williams, The Science of Hitting (1970)
- Barron’s, “What’s Wrong, Warren?” (December 27, 1999)