My Million Point Month

My Million Point Month

“When a measure becomes a target, it ceases to be a good measure.” — Marilyn Strathern, restating Goodhart in Improving Ratings, 1997 (a sentence I have now tested personally)

At 7:26 on the evening of September 29, I took a screenshot of a leaderboard. My name sat at the top of Eterna’s “Points last 30 days” view. The number beside it was 1,000,000.

I began this campaign with no other intention than to experiment with a set of agent apps, including Codex, Muse, GrokBot, and Claude Desktop. My first receipt for a paid Claude subscription dates back to September of 2023, which is to say I’ve been working with agents and LLMs for a while now. This Eterna campaign was an accidental outcome of that work.

I joined Eterna in January 2012, the very same year I started this blog. For my newer readers, Eterna is a game built by scientists at Stanford and Carnegie Mellon in which players design RNA sequences that fold into target shapes. It began as serious citizen science. Its players out-designed the computer algorithms of the day and discovered design rules the algorithms had missed. I was never one of the stars. I was a steady, curious amateur who solved puzzles in the evenings, earned points for each puzzle solved, and watched my total points rank drift slowly upward over fourteen years.

On the morning of September 14, I had just under 990,000 points and sat somewhere around 57th in the world. By the time the screenshot was taken, the lifetime total read 1,989,917 and my total points rank read 19. My RNA skills had not improved overnight. I had asked a fleet of AI agents to help, and then I spent fifteen days learning what help costs.

To put the total points board in perspective, since its inception, 464,042 people globally have solved at least one Eterna puzzle, the vast majority of which are 100 points each. The top player today has 3,276,408 points and still appears to be very active. The top 1000 players globally range from that number to 129,395 with a mean of 393,130. It’s a power law distribution. Points ≈ 12.8M × rank^-0.63, fits with R² = 0.96. I had to double my points to get from #57 to #19. I would need over a million more to crack the top 3.

Fourteen Years, Fifteen Days

The honest arithmetic needs a moment of care, because two clocks are running at once. The leaderboard counts a rolling thirty days. The campaign lasted fifteen. Subtract the old total from the new one and you get roughly a million either way, but the screenshot measures one thing and the ledgers measure another, and I would rather say so than let the coincidence do my bragging for me.

Here is the part I find funnier than the million. My notes from September 14 show the account at 1,004,617 partway through that first day. The agents were already working. After fourteen years, I crossed my first million on day one of the experiment, with help, before I had even decided what the experiment was.

The climb that followed is written in the ledgers, one line per claim. The account stood at 38th on September 20, when the first formal claim ledger opened. It reached 29th on the 23rd, 24th on the 26th, 21st on the 28th, and 19th on the evening of the 29th. From the day the ledgers began, they record 7,411 credited claims across 7,403 distinct puzzles. Almost all of them, 7,375, were ordinary 100-point puzzles. No single spectacular solve moved the needle. Seven thousand small ones did.

The Cathedral of Failed Runs

We, the agents and me, began with a naive loop. Find an uncleared puzzle, solve it, submit it, repeat. It worked just often enough to be dangerous.

Eterna is less one game than a federation of games that happen to share a website. Each puzzle names the folding engine that judges it, and the engines disagree. There is the classic ViennaRNA family in two generations, EternaFold, trained on the players’ own laboratory data, LinearFold, and most recently RibonanzaNet, a deep network that runs right inside the browser. A sequence that folds perfectly in one engine fails in the next. Some puzzles carry constraints the catalog does not advertise. Some demand that one sequence fold two different ways depending on what binds to it.

Every one of those facts was taught to us by a failed run. Local folds that the browser rejected. Metadata that routed a puzzle to the wrong verifier. A shell script that picked up a stale offset and launched twice. A browser drawer that stalled after Apply, leaving no way to know whether the click had landed. The ledgers keep 145 claims that failed before Apply, 48 marked uncertain, and 1,575 attempts that arrived to find the puzzle already cleared.

I have come to think of the whole campaign as a Cathedral of Failed Runs. Nobody designs a cathedral from the spire down. Each failure laid a stone, and the stones decided the shape of the building. We kept every failure on the record, because the next version of the system was literally made of them.

Even the literature had lessons in humility. Midway through, a new paper described the Montparnasse algorithm, which claims a perfect score on the classic Eterna100 benchmark. We built it, tested it against Eterna’s own engine, and watched it stall on our hardest leftovers in exactly the places our older methods had stalled. A better search does not help much when the wall is the puzzle.

The Network Is Not Reliable

Longtime readers know the Eight Fallacies of Distributed Computing have followed me since my book Network Distributed Computing. The first fallacy is that the network is reliable. I did not expect to relearn it from a puzzle game.

Picture the moment of a claim. An agent has a verified sequence. It pastes the sequence into the page and clicks Apply. The browser goes quiet. Did the click register? The naive instinct, human or machine, is to click again. On a live platform that instinct is poison. A second click can double a submission, and two agents claiming at nearly the same moment can produce a score jump that neither of them can explain. The network was not reliable, the browser was not reliable, and our own intuitions about clicking were the least reliable part of the stack.

So we drew a line, and I have come to call it the Apply Horizon. Everything before the Apply Horizon is cheap and repeatable. Search as much as you like. Fold, refold, argue, discard. Everything after it is permanent. The whole architecture grew around protecting that line.

By the end, every productive route fed the same eight stages. Discover, classify, solve, verify, freeze, canary, claim, reconcile. A candidate did not become claimable because a model said it was solved. It had to pass an independent verifier matched to its exact engine, then sit in a frozen queue locked with a SHA-256 hash, then wait while one canary from the batch went through first. At the Apply Horizon itself, the claiming agent confirmed the account, fetched fresh puzzle metadata, pasted the sequence, read it back from the page character by character, took a screenshot before and after, checked the profile delta, and wrote an event to an append-only ledger.

An ambiguous result became terminal. No second click. The agent went away, checked Eterna’s own record of cleared puzzles later, and wrote a reconciled event only when the platform itself confirmed the outcome. Thirty-one claims were settled that way.

This is the part I most want other builders to hear. The breakthrough was not a better solver. It was a safer claim. A clever sequence helps once. A verified queue behind a careful claimant lets the entire fleet help again and again, all night, without anyone standing guard.

A Fleet, Not a Virtuoso

I have written before about the Solo Virtuoso Fallacy, the belief that one brilliant agent will carry the work. This campaign was an ensemble, and it is worth being precise about who played which part.

My own part was smaller than the number suggests and larger than zero. I set the objective, supplied the account, decided when the evidence was strong enough to keep going, and watched the external score. A handful of times the website itself required a human. Twice the Eterna client threw an error that stopped the agents cold, and I dismissed it by hand and finished the claims. Fourteen years of play also gave the fleet something it could not generate: a real account, a history of prior solutions, and a feel for when a technically correct answer did not match the game.

Codex built and ran much of the machinery, including engine routing, queue generation, browser claiming, screenshots, ledgers, and reconciliation. GrokBot, the agent on my Mac Studio, found a rich route through puzzles that other players had already solved in the classic engine and worked it in waves, preserving every rejected candidate instead of quietly dropping it. Its seventh wave finished on the night of the milestone with 44 independently verified solutions out of 100 attempts, waiting for a fresh audit. GrokBoss (the GrokBot agent I named) kept the Mac fleet busy. Muse operated the two Linux servers, Apollo and Artemis, recovered work from stalled queues, and proved that the claiming protocol could pass from one agent to another without breaking.

Claude served as the fleet’s adversarial reviewer. Its job was mostly to say no. It killed pilots that looked promising and were not, and it insisted on reading the raw rows before anyone believed a summary. It also built some of the exact machinery underneath, including energy tables for the old engine, a proof-style search that decomposes a puzzle into independent pieces, and a reconstruction of the constraint rules Eterna’s newest engine enforces, recovered from the game’s own published source code.

I will not assign point totals to individual agents. The ledgers can tell me which queue a claim came from. They cannot yet tell me, honestly, whose idea made that queue possible, and I have learned this month what happens when you let a number stand in for a judgment.

The Shadow on the Leaderboard

There is also a shadow side here.

Start with the epigraph. Eterna’s points were designed to measure a player’s growing skill at RNA design. I turned them into a target, and the measure promptly stopped measuring the thing it was built for. My rank now says I am the nineteenth-best RNA designer on Eterna. The truth is humbler. I am a reasonably good orchestrator of agents who happens to own a fourteen-year-old account. Some thirty-eight players who earned their places over years slid down a rung so that my experiment could climb.

Then there is the question of what we learned. Did the fleet learn RNA, or did it learn Eterna? Mostly the second. The deepest expertise we built concerns engine versions, metadata quirks, browser behavior, and the precise semantics of a constraint box. That knowledge is real, and some of it is even useful to science. None of it would impress a biochemist.

The strangest shadow came from the puzzles we could not solve. One audit found more than fifteen hundred leftovers that had stalled exactly one base pair short of the target, nearly all of them puzzles that humans had cleared, some of them years ago. When we re-scored a sample of 21 stubborn cases with a different setting of the classic folding engine, a setting that may match how the game scored things in the past, five solved within a minute. Under the setting Eterna uses today, none did. That is a hint, and only a hint. It suggests that some of the game’s old victories may no longer be reachable, because the ground shifted under the puzzles after the players won them. Researchers have already begun proving that certain Eterna targets are undesignable under modern energy models. Our fleet stumbled into the same territory from the other side.

And one last shadow, turned on our own tools. One of those five, a small puzzle with strict limits on its pairs, looked like a perfect example of a lost victory. An exact search then found six sequences that solve it under today’s engine, in about thirty seconds. A heuristic that fails is evidence about the heuristic, and only weak evidence about the world. I should know that by now.

A Note on the Fleet

I mentioned my work environment in my previous post. Experimenting with an array of creatures such as these is both gratifying and frustrating. I find that they can sometimes work together but very often I am the orchestration agent. Communication mechanisms, shared context, and cooperation in general can sometimes be tricky. But having now stumbled into the Eterna campaign with this team of cohorts has been a learning experience to say the least. About the team:

  • Claude: Reliable, my oldest friend. The Desktop version is something I have started using only this summer, having used the CLI version for years. I know this teammate better than the others and therefore trust.
  • Codex (ChatGPT): Surprisingly good. Burned way too many tokens using Astra, had to wait a few days to get it back, switched to GPT 5.6 Sol (medium) when it did come back and have been pleased with the results.
  • GrokBot: Excellent UI, easy out-of-the-box setup. It did get stuck twice and needed a restart. I learned to not overwhelm it with too many agents.
  • Muse: The best UI of them all. Meta did it right. I’ve read that Muse has had record adoption in the first week and I understand why.

In addition I did have GrokBot use the ollama services I described in my previous post (the fleet) when appropriate. At no time did I use 100% of all the compute available to me in my home.

What I Learned

So what did I learn? I wanted Claude, Codex, Muse and GrokBot to work together as a team, and they did, after a fashion, but the communication was wonky from the first day. None of them shares a memory with the others, so every handoff had to be written down, and for much of the month I was the switchboard, carrying messages from one inbox to the next. Maybe I can find a way to fix that.

I learned that agents fail, and they fail in ordinary ways. A launch dies when a terminal window closes. A sandbox gets reclaimed overnight. Codex ran out of tokens on the final night, and Muse had to pick up its queue. The network is not reliable, and neither is a browser that goes quiet after you click Apply.

I learned that a skeptic is worth more than one more solver. Most of our clever research bets died on their own pre-registered terms, and the points came from the plain work of verifying, freezing and claiming a backlog, one careful puzzle at a time.

I learned to read the raw rows before believing a summary, including the summaries my agents wrote about themselves. I learned that the irreversible click belongs to a human, and that fourteen years of play still counted for something the fleet could not generate. And I learned that much can be accomplished. A million points in fifteen days is no small thing, even when I know exactly what the number fails to measure.

The Missing Prime Directive

As I write this, the account stands at 1,989,917, roughly ten thousand points short of two million, and the fleet is already working the next queue. I expect to cross that line. I no longer expect the line to mean what I once thought it would.

A leaderboard is an ensemble average, a snapshot of everyone at once. A life in a game like this is a time average, fourteen years of evenings and some weekends. The two only agree when nothing important changes along the way, and this month everything did. What the agents gave me, in the end, was a sharper eye for what a measure can and cannot hold.

Meet me on the corner of State and Non-Ergodic.