
Reliable Machines From Unreliable Parts
Why LLM errors need to be designed around rather than waited out, what John von Neumann worked out about unreliable components in 1952, and why a decision-only model like Jev fits his framework better than a chat model does.
If you’ve built anything on an LLM, you know most of the effort goes into making its output reliable, meaning there’s a high probability that it does its job, under stated conditions, for a stated time or number of steps.
The problem of getting reliable results from unreliable parts is a problem solved at the root of creation; every time one of your cells divides, it copies about six billion letters of DNA with an enzyme that gets about one in every hundred thousand wrong, then proofreads and repairs the copy until only about one in a billion is still wrong. Nonetheless, it’s a problem we still struggle with in the machines we build, and have since the first computers.
In January 1952, John von Neumann formalized this idea in a series of lectures at Caltech, with vacuum tubes and neurons as the parts. They were published in 1956 as Probabilistic Logics and the Synthesis of Reliable Organisms from Unreliable Components, and the introduction states his premise:
Error is viewed, therefore, not as an extraneous and misdirected or misdirecting accident, but as an essential part of the process under consideration…
He showed mathematically that a machine can be made as reliable as you like out of parts that fail at random, as long as five things hold:
- Each part is right more often than not (there’s a ground truth)
- The parts fail independently (rarely the case in the real world)
- Errors aren’t fed back in
- Someone pays for the redundancy
- The signal is discrete
You can quickly infer (🥁) why LLMs struggle with reliability today, but what’s nice about these principles is that they can be, and have been, applied well beyond vacuum tubes. The Space Shuttle flew on four computers running the same software, cross-checking each other about 440 times a second and voting out any that disagreed, plus a fifth programmed by a different company in case all four shared a bug.

But today, let’s focus on the conditions and how we can apply them to LLMs.
Errors compound
While individual errors may be insignificant, they accumulate:
In a complicated network, with long stimulus-response chains, the probability of errors in the basic organs makes the response of the final outputs unreliable, i.e., irrelevant, unless some control mechanism prevents the accumulation of these basic errors.
In a multi-step LLM job, if each step is right 99% of the time, a 100-step job finishes with no errors 36.6% of the time, and a 1,000-step job 0.004% of the time. At 99.9% per step, the 1,000-step job succeeds about a third of the time. In general, a job with 1/(error rate) steps finishes clean about 37% of the time.
A more reliable model pushes these curves to the right, but a long enough job still fails. If the components can’t be made reliable enough, the system has to correct their errors. Von Neumann’s first step was redundancy:
Instead of running the incoming data into a single machine, the same information is simultaneously fed into a number of identical machines, and the result that comes out of a majority of these machines is assumed to be true.
He called the voting component a majority organ. With three copies and a two-out-of-three vote, one faulty copy gets outvoted. Anyone who’s worked on distributed systems will know it as a quorum.
Voting once, on three copies of the whole machine, just moves the problem: the final vote becomes a single point of failure. Voting after every step fixes that, but multiplies the part count by three at every step. He calculated that a computation 160 steps deep would need about 2 × 1076 components, “somewhat above the putative order of the number of electrons in the universe”. At 200 steps it came to 2.5 × 1095, and he wrote only that “this requires no comment”.
His workable version is called multiplexing. Each signal is carried on a bundle of many wires instead of one, and the bundle’s value is whatever nearly all of its wires say. After every computing stage comes a restoring organ, which takes a bundle where most wires are right and returns one where more of them are right, and before each one the wires are shuffled so that errors which occurred together don’t stay together.
The bundles have to be large. With components that fail once in 200 operations, a bundle of 1,000 wires fails more often than a single component does. For a 2,500-tube computer meant to run eight hours between errors, he estimated about 14,000 wires per bundle, and concluded that the method was impractical with 1950s hardware. He didn’t rule it out for the brain, though, where size wasn’t the obstacle: “The nerves are bundles of fibres,” he noted, “like our bundles.”
Each component has to be right more often than not
Von Neumann’s design needed parts that were wrong about 1% of the time or less. The simplest form of this condition is Condorcet’s jury theorem from 1785: a majority vote among independent voters gets more accurate as voters are added, but only if each voter is right more than half the time. Below that, adding voters makes the majority more likely to be wrong.
The problem is that LLMs are wrong, often. Sampling a model repeatedly and taking the majority answer can raise accuracy and then lower it as calls are added, because extra samples help on questions the model usually gets right and hurt on questions it usually gets wrong. When most questions are the first kind, the gains are large. Majority voting over samples raised Llama 2 13B from 35% to 59% on a grade-school math benchmark, above a single Llama 2 70B at 54%.
So before you vote on anything, find out which kind of question you’re asking. Run the model on a labelled sample, see where it’s usually right, and only vote there. Everywhere else, voting just makes it confidently wrong.
Components have to fail independently
Von Neumann assumed each component fails “statistically independently of the general state of the network and of the occurrence of other malfunctions”, and noted right away that dependent failures are “a good deal more realistic”.
It’s very difficult to produce genuine independence, given how interconnected reality is. In 1986, a study had 27 students at two universities write the same program from one specification, then ran each version on a million inputs. Each program was very reliable on its own, but inputs where two or more versions failed together were far more common than independence predicts. A 2026 repeat of the experiment with coding agents, on the same specification, found 429 coincident failures where about 115 were expected. Whether it’s because people copied each other’s homework, or they all learned from the same textbook and examples, these kinds of situations often show up; the students in that study mostly named the same parts of the problem as the hardest.
Distributed systems ran into the same wall. Raft only commits what a majority of its servers have, but it’s built for servers that fail by going quiet. One that keeps answering, wrongly and convincingly, is called a Byzantine fault, which is how LLMs usually fail, and tolerating those takes more than two-thirds of the voters to be sound (while only a simple majority is needed for servers that just go quiet). Even that assumes the bad ones fail independently, which is why Ethereum wants its validators spread across different clients: a bug in software run by two-thirds of them could finalize the wrong chain.
Much like the students’ bugs, LLM errors are correlated given models are trained on largely the same data, so they make many of the same mistakes. In a study of more than 350 models, when two of them were both wrong on one leaderboard, they gave the same wrong answer about 60% of the time, compared with about 33% by chance. More accurate models had more correlated errors, including across providers. A panel of nine frontier models from seven families, used as judges, carried the information of about 2.18 independent judges, given how correlated their errors were. When all nine agreed, they were wrong 9.1% of the time; nine independent judges with the same individual error rates would be wrong about 0.02% of the time. Unintuitively, asking the judges to reason step by step made their errors more correlated.
Voting across copies of one model therefore buys little. The voters need to differ in ways that affect their errors, such as model family, prompt and the evidence they’re given, and their correlation should be measured rather than assumed.
It’s also a pretty strong argument for why we need open, diverse models (and why regulating them away is a bad idea), lest our models get hit with an exploit like the Panama disease, which wiped out the Gros Michel banana. A diverse ecosystem is a healthy ecosystem.
As for measuring, it’s pretty cheap: on a labelled sample, count how often two voters are wrong on the same input, and compare that with the product of their separate error rates, which is what independence would give you. You can also borrow von Neumann’s shuffle to break up errors that share a cause, for instance by shuffling the order of the options for each voter, since models favour some positions over others. And the most independent second opinion is often not a model at all: a rule, a lookup or a test fails in completely different ways.
Errors can’t be fed back in
In circuits with feedback, von Neumann warned “it is possible, that the machine remembers its mistakes, so to speak, and thereafter perpetuates them”.
Making such a prediction ~75 years ago is borderline prophetic, given that LLMs today do exactly this: on long, simple tasks with the necessary knowledge and plan provided, their per-step error rate rises as the task goes on, partly because a model is more likely to make a mistake when its context contains its own earlier mistakes. Larger models did it more; reasoning models mostly didn’t. An agent that appends its own output to its context is exposed to this at every step.
So check each step before its output goes back in, and when one fails, retry it from a clean context rather than asking the model to fix it in the one that holds the mistake. Better still, don’t carry the whole history forward at all: give each step the verified state it needs and nothing else.
Redundancy costs something
Multiplexing multiplied the number of components by about three times the bundle size, and later work proved that some computations can’t be made reliable without a logarithmic overhead. MAKER shows a similar cost for LLMs. It solved a 20-disk Towers of Hanoi puzzle, 1,048,575 moves, with no errors, using gpt-4.1-mini. Each move was a separate decision given only the context it needed, each decision was sampled until one answer was ahead by three votes, and responses that were too long or malformed were discarded. The authors show that the required vote margin grows with the logarithm of the number of steps, so total cost grows as steps × log(steps). The model was small and inexpensive; the reliability came from the decomposition and the voting.
Most inputs are easy, and asking a model the same easy question over and over is money down the drain. MAKER’s rule already saves some of it: if the first three answers agree, it stops there, and it only keeps asking when they disagree. Cheaper still is to double-check only the inputs the model is likely to get wrong, but that needs the model to tell you which ones those are. Most models don’t.
The signal has to be discrete
Multiplexing works because each wire is either on or off, so a restoring stage can push a bundle back toward all on or all off. Von Neumann also analyzed a continuous version, where a number is represented by the fraction of wires that are on, and found that noise accumulates until “the excitation levels are more likely to resemble a random sampling of numbers than mathematics”.
Free text behaves like the continuous version. You can’t take a majority vote over three paragraphs, or check that a paragraph still says what the original said. Self-consistency, which samples several reasoning paths and takes the most common final answer, raised PaLM 540B’s accuracy on grade-school math by 17.9 points, and it works because the final answer is a single number. The closest thing for open-ended text, Universal Self-Consistency, hands the answers to an LLM and has it pick the most consistent one. The catch is that the votes are now counted by another model, which can get the count wrong too.
So wherever a decision gets made, have the model answer from a closed set: yes or no, one of five labels, a score out of ten. Structured output keeps the format dependable, and plain code can then compare, count and vote. Free text can be pulled into that shape too, by breaking it into claims and asking a yes-or-no question about each one (“does the summary say the refund was approved?”), and those you can vote on. There’s now a kind of model that does nothing but this. The first one, Jev, had the category to itself for about a day before the open-source clones showed up, and it’s up next.
Where Jev fits
Jev is TypeSafe’s System One model, and it doesn’t generate text at all. You give it an input and a set of typed questions (yes or no, pick from a list, give a score), and it returns a probability for each answer, evaluating all the questions in parallel. It’s named after Jevons, of the paradox I discussed last time.
It isn’t the most accurate model available, even on TypeSafe’s own benchmark, but the models that beat it cost dozens to hundreds of times as much per case. What makes it interesting is how closely it fits von Neumann’s conditions:
- Its outputs are discrete. Every answer is one of the values you declared. TypeSafe says it can’t hallucinate, which is true only in the sense that it can’t return a value outside the schema; a valid answer can still be wrong.
- It tells you which inputs it’s likely to get wrong. Von Neumann gave each component one fixed error probability and called that “an unrealistic assumption”, since real components fail more on some inputs than others. Jev returns a probability with every answer, and TypeSafe trains it to make those honest, with a method it calls Reinforcement Learning for Calibrated Decisions: when it says 0.9, it should be right about 90% of the time. TypeSafe hasn’t said how the method works, but a published method with the same goal scores each answer with a rule that punishes confident wrong answers and timid right ones, so the model’s best strategy is to say how sure it actually is.
- It’s cheap enough to use redundantly. Extra questions about the same input add tokens but almost no latency, so checking every field of every record is affordable.
Calibrated probabilities make a cascade possible: act on the answers it’s confident about, whichever way they go, and send the rest to something more expensive, such as a vote across different models, a larger model or a person. The vote has to be across different models, though; asking Jev the same question three ways is still one model voting three times.

Take TypeSafe’s calibration numbers with a grain of salt: they grade Jev against the averaged answers of two frontier models, not against the right answers. Checked against human labels in an independent study, its answers at 0.9 or above were right about 82% of the time on a typical task, and on one task its confidence meant nothing at all. The cascade still paid off: keeping the answers it was sure of and sending the rest to an LLM matched the LLM alone at a quarter to half of the cost. Just set the thresholds on your own labelled data.
Errors are a feature, not a bug
The ability to generate - that is, to create many different variations of an idea - goes hand in hand with the ability to err, seemingly a trade-off underpinning our world. Thus, a more “creative” model is necessarily a model that produces errors, and in many situations, eliminating all of them is both impractical and undesirable.
Still, a certain threshold of reliability needs to be met for practical purposes, and the final step in closing the loop is learning from the errors you do get. Nassim Nicholas Taleb goes further than von Neumann in Antifragile, treating errors not only as part of the system but as something the system can get better from. His example is flying: “Every plane crash brings us closer to safety”, because each one is investigated and its cause designed out, while in the global economy “errors spread and compound”. What separates the two is independence and feedback, the same conditions von Neumann needed. An LLM system can get better from its errors the same way, as long as the ones it catches go into the design, as test cases and retuned thresholds, and not back into the job that made them.
The beauty of von Neumann’s and Taleb’s principles is that they generalize to anything built from unreliable parts: vacuum tubes, cells, spacecraft, databases, and now LLMs. For an LLM system, they come down to a short list:
- Measure the error rate of each step, not only the end-to-end result.
- Only vote on questions the model usually gets right.
- Vote across models that fail differently, and measure how correlated they are.
- Check each step’s output before it goes back into the context.
- Spend redundancy on low-confidence inputs.
- Return a typed value wherever a decision is made, and count the votes with plain code.
None of this gets you to zero errors, and it isn’t meant to; the goal is a system that keeps working while its parts keep failing. The oldest one we know of is running inside you right now.
Your cells, after all that proofreading, still leave a few typos every time they divide, and you turned out (mostly) fine.



