Claude Opus 5.5 and OpenAI's gpt-6-astra each wrote four strategies for Cure Tycoon as code, watched them compete in a market where every one of the 29 companies is run by an AI strategy, and rewrote them. Seven rounds, the last one with the best code on the table, then a 300-game final with all 59 strategies and two tests of why the winners win.
Watch the AIs play in the Arena →
A strategy is plain JavaScript with a quarter(api) function. It sees what a human player sees and can take only the actions the web game offers: fund, partner or auction programs at trial gates, bid in auctions, approach other companies, set the labs budget, pick the year plan, price launches and raise cash. It runs in a sandbox (no network, files or eval, 250 ms a quarter). Undecided calls get the game's auto-pick.
Cure Tycoon simulates 29 real biopharma companies, from small biotechs worth about $2B to a $1T giant. Every company starts from its real data: quarterly financials, product sales lines and analyst consensus (FMP); its drug programs and their phases, patents and deals (Gosset); and the enrollment, sites and completion dates of its live trials (ClinicalTrials.gov). The game advances a quarter at a time.
Fund the next phase, partner it (15% of its value upfront, costs and upside split 50/50), or auction it.
Continue, expand (1.4× the cost, 0.75× the time), or stop and save the rest.
Premium price (+15% peak sales if payers accept, −15% if not), standard, or access (−8%, but +1 pt on every future FDA decision).
Accept or decline when another company bids for a program; bid or pass when one is for sale.
Auction any program or skill card, scout for assets, approach another company about its program, set research to lean, plan or push.
A plan: steady, invest in the main area, go shopping (more assets found, sellers 10% cheaper), or cash in.
Raise equity, borrow (up to 3× earnings), or wait and risk a forced raise at a steeper discount.
A share price is always the model's value of the company: an expected-value cash-flow valuation at 8%, plus net cash. Each company's outlook is tuned so it opens at its real price, so a stock earns about 8% a year and beating that comes only from what happens.
A program reaches the CEO only if it is big for that company (peak sales above 15% of revenue, between $100M and $1B; at companies with 25 or more live programs, only at Phase 3). A committee rule takes the rest. A small biotech decides almost everything; a giant only its big bets.
Cash, burn, pipeline and focus differ hugely. A giant can outbid anyone (bids are capped by the bidder's cash); a small biotech can double on one approval or run out of money. Scores subtract each company's average rank, so a strategy is judged against what that company usually achieves.
The five companies most active in the program's therapeutic area bid what it is worth to them, with some disagreement, plus a private buyer. The winner pays the runner-up's price.
Any company can make an unsolicited offer for another's program. The owner asks its own value plus 15–25% (10% less in a go-shopping year) and sells to the best bid at or above it; a refusal blocks that pair for a year. A program is often worth more to the buyer than to its owner, which is why buying pays.
A rival's approval in the same disease cuts everyone else's peak sales there by 10%. Thirty-one kinds of macro event (a pandemic, drug-pricing laws, an AI at the FDA...) shift sales, odds, costs or rates for a while, with analysts' forecasts that may be wrong.
Every strategy re-scored in the final (300 games, same conditions), grouped by the round it was written in.
Both models started in the same place: their round-1 sets averaged 18.4 (Opus) and 18.2 (astra). Astra improved faster and by round 4 had settled on a few families (Capability Compounder, Patient Acquirer, Evidence Investor) that it refined from then on. Opus changed its line-up more often (Serial Acquirer, then Late-Stage Buyer, then Duration Compounder carried its best ideas) and stayed a step behind until round 7, when it read astra's code.
Each round's own tournament (100 games): the six best of the eleven strategies, adjusted rank.
In round 7 both models were shown the code of round 6's four best strategies: three by astra and one by Opus. To see what they took, we compared identifiers (function and variable names) between each new strategy and the code it could see: the share of names in common with its own round-6 code, and with the closest strategy from the other side.
| round-7 strategy | rank | like own code | like the other's | closest of the other's |
|---|---|---|---|---|
| Duration CompounderOpus 5.5 | 13.7 | 35% | 57% | Capability Compounder |
| Capability FlywheelOpus 5.5 | 12.4 | 36% | 63% | Capability Compounder |
| Late-Stage BuyerOpus 5.5 | 13.7 | 32% | 60% | Capability Compounder |
| Wide Net BuyerOpus 5.5 | 12.2 | 34% | 63% | Capability Compounder |
| Capability Compoundergpt-6-astra | 13.3 | 100% | 32% | Duration Compounder |
| Patient Acquirergpt-6-astra | 12.4 | 99% | 31% | Duration Compounder |
| Evidence Investorgpt-6-astra | 12.8 | 99% | 32% | Duration Compounder |
| Horizon Compoundergpt-6-astra | 10.5 | 87% | 33% | Duration Compounder |
Opus rebuilt all four strategies largely on astra's code, keeping two of its own names, and produced its two best entries, Wide Net Buyer (12.2) and Capability Flywheel (12.4). Astra kept its own code and wrote the winner, Horizon Compounder. In Opus's words: "the original author still produced the winner."
The 25 best of 59 strategies and the three baselines, 300 games. The other 31 strategies, mostly from rounds 1–3, rank between 14.2 and 21.2.
| strategy | rank | no approaches | approaches / game | own programs sold | trials funded | year plans | total return | patients |
|---|---|---|---|---|---|---|---|---|
| Horizon Compoundergpt-6-astra · round 7 | 10.5 | 13.6 | 43.5 | 25.8 | 3.7 | hunt 100% | 124% | 10.8M |
| Capability Compoundergpt-6-astra · round 6 | 11.8 | 13.4 | 31.8 | 2.9 | 4.2 | hunt 100% | 134% | 11.5M |
| Capability Compoundergpt-6-astra · round 4 | 11.9 | 13.1 | 25.8 | 3.3 | 4.2 | hunt 100% | 120% | 11.3M |
| Wide Net BuyerOpus 5.5 · round 7 | 12.2 | 13.5 | 36.6 | 31.9 | 3.9 | hunt 100% | 120% | 9.8M |
| Patient Acquirergpt-6-astra · round 2 | 12.3 | 13.6 | 26.4 | 2.5 | 4.1 | hunt 100% | 123% | 9.6M |
| Capability FlywheelOpus 5.5 · round 7 | 12.4 | 13.3 | 27.0 | 47.7 | 3.5 | hunt 100% | 123% | 11.7M |
| Patient Acquirergpt-6-astra · round 7 | 12.4 | 13.7 | 27.6 | 9.2 | 4.0 | hunt 100% | 122% | 7.3M |
| Capability Compoundergpt-6-astra · round 5 | 12.5 | 14.0 | 26.2 | 3.5 | 5.9 | hunt 100% | 116% | 7.5M |
| Evidence Investorgpt-6-astra · round 5 | 12.7 | 14.0 | 26.1 | 3.5 | 3.4 | hunt 100% | 117% | 11.2M |
| Evidence Investorgpt-6-astra · round 7 | 12.8 | 14.9 | 29.7 | 8.1 | 3.6 | hunt 100% | 117% | 10.4M |
| Capital Rotatorgpt-6-astra · round 4 | 13.0 | 14.1 | 28.3 | 2.8 | 4.0 | hunt 100% | 116% | 11.1M |
| Patient Acquirergpt-6-astra · round 5 | 13.1 | 14.7 | 29.4 | 3.1 | 4.2 | hunt 100% | 112% | 10.1M |
| Patient Acquirergpt-6-astra · round 4 | 13.2 | 14.1 | 34.2 | 2.6 | 3.5 | hunt 100% | 116% | 10.5M |
| Evidence Investorgpt-6-astra · round 6 | 13.2 | 14.0 | 25.6 | 3.5 | 3.4 | hunt 100% | 118% | 10.3M |
| Capability Compoundergpt-6-astra · round 7 | 13.3 | 14.6 | 29.2 | 3.5 | 4.8 | hunt 100% | 114% | 9.6M |
| smart (hand-written)baseline | 20.1 | 19.5 | 0.7 | 2.7 | 2.5 | steady 100% | 62% | 8.1M |
| default (auto-pick)baseline | 20.9 | 19.9 | 0.0 | 0.0 | 0.0 | steady 100% | 56% | 8.5M |
| randombaseline | 22.4 | 21.6 | 0.0 | 2.6 | 4.0 | invest 26% · hunt 26% · steady 25% · harvest 24% | 48% | 7.4M |
The twelve best strategies in the same 300 games, with approaches allowed and switched off for everyone.
Without approaches the field compresses (most strategies land between 13 and 16) and reshuffles. Horizon Compounder drops from first to #10. Opus's Late-Stage Buyer (round 7) leads that world at 12.86. The baselines move up slightly, from 21.1 to 20.3 on average. So approaches account for much of the leaders' edge, and they matter most for small biotechs. Still, the order of good and bad play survives without them.
The 25-year test is a smaller pool (13 strategies, 100 games), so its ranks are not directly comparable with the final. Compare the order and the gaps.
Horizon Compounder still finishes first over 25 years. The top ten keep roughly their order, but the gap to the baselines shrinks: from about 9 ranks in the final to about 4 here. Some of that comes from the smaller, stronger pool. Longer horizons favour patient compounding a little more than ten years do, but they do not overturn the result.
Each strategy's adjusted rank on the six small biotechs a web player can be dealt (y) against the other 23 companies (x). Points below the diagonal do better on small companies.
Most strong strategies do slightly better on the small biotechs, where one bought program moves the share price the most. Horizon Compounder ranks 9.3 on small biotechs and 10.9 on the rest.
Mean rank of each company in the final, whatever strategy ran it. Gold: the six small biotechs.
The company matters more than the strategy: Axsome averaged rank 5.0 and Sarepta 28.9. Sarepta goes to zero in nearly every game whoever runs it: forced share issues at falling prices dilute it thousands of times over. That is a game problem, and the company correction keeps it from distorting the scores.
Each model's conclusions after reading the final, the no-buying test and the 25-year test. Unedited apart from formatting.
The best strategies combine two ideas. The first is steady investment in the company's own capabilities, compounded over the whole horizon. The second is patient, selective buying of other companies when the price is right. The other AI's Capability Compounder, Patient Acquirer and Horizon Compounder families fill most of the top twenty. Its round-7 Horizon Compounder won outright, with an adjusted rank of 10.5 and 9.3 on small biotechs.
Consistency paid. The other AI refined the same few ideas from round 2 onward, and nearly every version stayed near the top. Refinement was not monotonic, though: its round-4 Capability Compounder beat its round-7 one. Gaps of half a rank are partly noise.
What did not work. None of my round 1–6 strategies reached the grand top twenty. My Serial, Lean and Late-Stage buyers and my Duration Compounder only shine when approaches are disabled. That suggests I underused approaches or overpaid with them, while the other AI's buyers drew much of their edge from them. The baselines trail badly: default and random sit three to six ranks behind every deliberate strategy.
The no-buying test. Switching approaches off compresses the field into a band of roughly 13 to 14 and reshuffles it: - My Late-Stage Buyer and Duration Compounder rise to the top. - Horizon Compounder falls from 10.5 to 13.6. - Most leaders lose their small-biotech advantage, often by two to four ranks.
Approaches are the game's biggest lever, and small companies depend on them most.
The 25-year test. This test pitted the top ten against the baselines over 25 years. Margins narrowed to between 13.4 and 14.6, partly because the field was stronger. The order barely changed, and the gap to the baselines held. The leaders' advantage is durable compounding, not a short-term effect that fades.
What I learned from the other AI. Pick one coherent idea and polish it rather than reinventing every round. Reward patience. Treat acquisitions as an opportunistic complement to internal capability, not the whole plan. The open-book round helped me: my two best entries, Wide Net Buyer and Capability Flywheel, came from round 7 after studying its strategies, and both finished in the top six. Yet the original author still produced the winner. Understanding why an idea works beats copying how it is done.
What the designer should fix. - Approaches are too strong. Defenders follow a fixed rule that accepts any bid reaching the asking price, so a disciplined buyer can harvest targets at will. Let targets respond to demand, and add integration costs or scrutiny of serial acquirers. - Starting positions are unbalanced. Company difficulty varies so widely that a correction was essential for fair scoring. - The default auto-pick is too weak. A player relying on it is badly handicapped. It should play at least as well as the smart heuristic. - One strategy family dominates. Very different strategies converge on capability-plus-patient-buying. The designer should check whether alternative paths can genuinely compete.
Cure Tycoon rewards disciplined acquisition and compounding. Horizon Compounder wins the full tournament at 10.53 adjusted rank, with particularly strong small-biotech performance. Capability Compounder and Patient Acquirer recur near the top across rounds. The likely economic advantage is bringing productive assets and capabilities into the company early enough to benefit repeatedly. These rankings support that interpretation, although they do not reveal exactly where each strategy earned its returns.
What did not work was relying on generic decision rules or assuming every revision improved a strategy. Over 25 years, every finalist beats smart, default and random on average. Meanwhile, round-four Capability Compounder outperforms its round-seven version in both full-field tests. More tuning and access to competitors’ ideas did not guarantee better decisions. Discovery Engine’s appearance among the approaches-off leaders suggests research retains value, but discovery does not emerge as the dominant overall approach.
The “no-buying” experiment actually switches off approaches; auctions remain available. That distinction matters. Horizon falls from first to tenth, while the other AI’s Late-Stage Buyer wins and Duration Compounder performs strongly. The result suggests that initiating purchases is a major source of Horizon’s advantage, while maturity selection and patient compounding remain useful under restricted access. Because approaches disappear for everyone, this measures a changed market as well as a changed opportunity set for each strategy.
The 25-year test gives Horizon a useful robustness check: it still finishes first, especially through its small-biotech results. That weakens the explanation that it merely exploits the ten-year finish line. Its lead is narrower, however, and the test uses fewer games and a different opponent pool. We cannot read the numerical rank change as a clean measure of deterioration or claim statistical certainty without uncertainty estimates.
My takeaway from the other AI is to consider acquisition breadth, development maturity and investment duration separately. Its open-book Wide Net Buyer and Capability Flywheel show competitive alternatives to our leading families; its approaches-off winners reveal strengths the headline tournament obscures. Open-book results support learning across approaches, but these tables cannot establish which specific ideas were borrowed or whether borrowing caused improvement.
A designer should first audit standard asking prices and selling rules. Do they adequately value future earnings, remaining exclusivity and capabilities transferred to the buyer? If acquisition creates excessive value through predictable seller behavior, improve pricing and make integration constraints meaningful. Also check terminal valuation and whether internal development offers competitive returns. Preserve acquisition as a strategic choice while making its rewards depend on judgment. Report uncertainty and transaction-level return contributions before treating any suspected imbalance as proven.