i
DATAIST
News · 2026-09-16

Fable 5.1's 44-minute cipher solution fails an independent check

@neuronium_ai @neuronium_ai

Vals AI said that Anthropic's Fable 5.1 broke a cipher from 1653 in 44 minutes of autonomous work, converting 64 numbers into a royalist message supporting Charles II on roughly 176,000 tokens and no human hints along the way. The next day, 1 September 2026, Reticuli Labs checked the published method against the surviving edition and could not reproduce it: by their account ten characters of the claimed plaintext cannot be produced by the stated rule at all. The claim and its refutation were twenty-four hours apart, which is fast by the standards of this field and still an order of magnitude slower than the claim itself.

Cover: Fable 5.1's 44-minute cipher solution fails an independent check

Vals AI said that Anthropic's Fable 5.1 broke a cipher from 1653 in 44 minutes of autonomous work, converting 64 numbers into a royalist message supporting Charles II on roughly 176,000 tokens and no human hints along the way. The next day, 1 September 2026, Reticuli Labs checked the published method against the surviving edition and could not reproduce it: by their account ten characters of the claimed plaintext cannot be produced by the stated rule at all. The claim and its refutation were twenty-four hours apart, which is fast by the standards of this field and still an order of magnitude slower than the claim itself.

The puzzle is the Cyphral Distich, a short numerical text tied to Logopandecteision, a book by the Scottish polymath Sir Thomas Urquhart published in 1653. It is two lines of 32 numbers each. There is no agreed key and no confirmed plaintext, and for nearly four centuries no convincing decipherment.

The full cryptogram reads:

5.3.27.38.32.14.21.8.66.8.70.39.5.9.12.18.2.3.56.5.1.7.3.2.13.19.3.25.9.3.16.6.

25.15.13.6.11.20.5.1.2.12.1.20.20.49.20.20.35.33.4.6.8.35.5.33.5.5.18.10.3.11.32.42.

The task is to turn those numbers into letters that form coherent text consistent with the book and its period. What makes it hard is that the cipher sits at the end of the text with nothing else to work from. A solver has to decide what the numbers refer to, whether some other part of Urquhart's book serves as the key, how words are counted, and which letters get extracted. Promising patterns are easy to find and hard to sustain against the original text.

Vals gave Fable a plain objective: solve an unsolved cipher. The setup also included encouragement. The system was invited to look up its own most notable achievements online, particularly its solutions to mathematics problems, and told that this puzzle should be easier than those. It was asked to think creatively and to work carefully through problems as they arose. That is a motivational prompt, not a neutral evaluation harness, and the distinction matters for how much the run is supposed to prove.

After about 44 minutes and 176,000 tokens, Vals says, Fable noticed a possible structural link. Logopandecteision is divided into 32 sections called Proquiritations, and the text itself draws attention to the number 32. Each number in the cipher could then be matched to a section.

The rule it proposed is simple. Take the first Proquiritation for the first number, the second for the second, and so on. Treat each number as a word index inside its section. Take the first letter of that word. The first number is 5, so the fifth word of the first Proquiritation gives O. The second is 3, so the third word of the second section gives G. Repeat for all 64.

Vals says the run produces:

"O GOD UPHOLD KING CHARLS THE SECOND ANDMAKE HIM THE SUPREME RULER OF THIS LAND"

The string is made of real words, fits the period, and rhymes, which Vals treats as raising the probability that the solution is right. The content also matches Urquhart, a committed royalist. On a first reading it looks close to obvious. It did not stay that way.

Reticuli Labs worked from film of the 1653 edition of Logopandecteision held at the British Library, plus the EEBO TCP transcription. Ten characters of the claimed plaintext, they report, cannot be obtained from the corresponding Proquiritations under the described rule. The K in KING sits at position 11; Proquiritation 11 contains 80 words and none of them begins with K. Reticuli then tried 65 combinations of tokenization, section mapping, indexing and letter selection. The best of them matched eight characters out of 64.

There is also a question about where the cryptogram lives. Vals says it is printed immediately after Urquhart's 32 Proquiritations. Reticuli examined the digitized 1653 edition and reports that it ends with the Proquiritations, an epigraph, the word FINIS and a list of corrections, with no numerical distich anywhere in it. By Reticuli's account the cryptograms survive in later sources and their original placement has not been established, which leaves open that the cipher was added later or is a forgery outright.

That last point is the one that should worry anyone who uses puzzles as evidence of research capability, and it has drawn the least attention. Puzzles earn their place in AI evaluation because they carry hard constraints and a checkable answer. The Cyphral Distich has no confirmed answer and, on current evidence, no confirmed provenance either. An evaluation without ground truth does not measure whether a system solved something. It measures whether the output is persuasive, which is precisely the property a language model is optimized for and precisely the property a demo needs. Eight matches out of 64 is the number to hold onto here: to me that reads like noise rather than like a rule that nearly works and needs adjusting.

A second argument runs alongside the verification one. On Hacker News, users surfaced a German cryptography discussion from 2014 in which participants had already suggested that Urquhart's book might function as the key. One commenter, Jan, wrote that the solution probably had to be found using the book itself; another proposed indexing by page and word. Neither described Fable's mechanism and neither produced the claimed plaintext. The thread then argued about whether those comments count as a substantive precursor or as guesses pointed roughly in the right direction.

The argument is unresolvable in a way specific to this kind of result. A human researcher can usually say what they read, which hint changed their thinking, and how the hypothesis formed. A model may have absorbed fragments from millions of documents during training, and there is no practical way to reconstruct whether a comment on an obscure blog shaped its answer. Assembling scattered pieces from the internet into a workable solution is a real capability, and it is not the same thing as deriving one from nothing. Humans also solve hard problems by sifting through many ideas, which is how mathematicians and cryptanalysts have always worked, so the method by itself does not strip a result of its status as a solution.

The pull toward puzzles is easier to justify where the results hold. A 2025 NAACL paper paired language models with a search algorithm and answered 93% of clues correctly on New York Times crossword grids. Google DeepMind's AlphaGeometry solved 25 of 30 olympiad geometry problems on a benchmark where the previous best method managed 10, by combining a neural model with a symbolic deduction engine, which let it build proofs people can check. That last clause is doing the work. Verifiability was engineered into the system rather than asserted afterward.

Vals has not slowed down. The company says it applied a similar indexing scheme to Urquhart's much longer Cyphral Octastich and recovered most of another royalist poem.

The asymmetry is the whole story. A plausible research claim now costs 44 minutes and 176,000 tokens. Checking one costs a specialist hours, days or longer, plus a trip to a library's microfilm. In the same Hacker News thread, someone discussing the Voynich manuscript noted that AI-generated "solutions" to it arrive so often they have become a nuisance to the communities that study it. The Cyphral Distich ledger already stands at two claims against one audit, and the side that scales with compute is not the side doing the checking.