Compound uncertainty: AI’s hidden risk in safety-critical development

Here’s a question your grandfather could have answered: Would you rather have a million dollars today or a penny that doubles every day for a month? Most people take the million. The penny reaches $5 million by day 30.
Human intuition is simply bad at exponential math. We think in straight lines, and compounding curves surprise us every time. Now run that intuition in reverse.
An AI coding agent that’s right 98% of the time sounds impressive. And 98% is a generous assumption, probably better than reality for most tasks. But apply that accuracy across 34 steps with no verification in the loop, and you’ve crossed the coin-flip line. More likely wrong than right. The math is 0.98^34 = 0.505.
The surprise is the same one your grandfather felt. And in a safety-critical development environment, the consequences are not a missed investment opportunity.
Sophisticated agentic systems don’t run open-loop. They compile, test, lint, and self-verify at each step, and the public record shows it works.
Andreas Kling ported Ladybird’s LibJS engine from C++ to Rust using AI agents across hundreds of human-directed prompts, producing 25,000 lines of Rust with zero regressions across 65,359 tests and byte-for-byte identical bytecode output. The human was in the loop at every decision point, which is precisely why it worked.
The Bun JavaScript runtime went further. AI Weekly highlighted that Claude agents rewrote roughly one million lines of Zig to Rust autonomously across 6,755 commits, passing 99.8% of its existing test suite. It also left 13,044 unsafe Rust blocks, where a comparable handwritten project would have 73. A passing test suite doesn’t surface this safety debt—it stops a safety-critical certification in its tracks.
Both of these projects succeeded because verification ran inside the loop at every step. They also illustrate exactly where the limits are. In most software development contexts, the floor is an efficiency problem. Verification catches it, the agent retries, and the process converges. Expensive in tokens and time, but recoverable.
In safety-critical development, the calculus is different. This is where functional correctness testing and safety-critical qualification part ways. Bun passed its own test suite. Ladybird produced byte-for-byte identical bytecode. Those are impressive results. But they are not safety cases. ISO 26262, DO-178C, and IEC 62304 don’t recognize self-generated test results as qualified verification evidence.
Your braking system software doesn’t get partial credit for passing tests it wrote for itself. Your insulin pump firmware isn’t certified on a curve. The standards assume deterministic tools producing verifiable evidence—qualified tools, documented configurations, and traceable outputs. An agentic workflow that self-verifies is better than one that doesn’t. But in safety-critical development, it still isn’t enough.
What safety-critical compliance actually requires isn’t vague.
ISO 26262 mandates a documented safety plan, requirements with bidirectional traceability from hazard analysis through to verified implementation, and evidence that coding guidelines—typically MISRA C or CERT C—were enforced by a qualified tool using a qualified configuration.
DO-178C adds structural coverage requirements. At the highest criticality levels, every statement, every branch, and every condition and its complement must be exercised by tests that are themselves traced to requirements.
IEC 62304 requires a software development lifecycle with documented verification activities at each phase. In every case, the evidence must be generated as the work happens rather than reconstructed afterward—and not self-certified by the tool that produced the artifact being evaluated.
The open-loop pipeline isn’t an edge case; it’s what every team promises to fix after the next release. A requirements review is handed to a code generator, a documentation tool, and a traceability updater with testing saved for the end. That’s not an agentic worst case. That’s a pipeline. At 98% per-step accuracy across 34 stages, you’ve crossed the coin-flip line before you’ve run a single test.
The answer isn’t a better model. It’s the same answer safety-critical engineers have always given to unreliable processes. You don’t improve your way to acceptable; you gate your way there.
Static analysis enforces expected coding patterns and flags dangerous anti-patterns like uninitialized memory, undefined behavior, and violations of MISRA or CERT rules that exist precisely because they’ve caused failures before.
Unit tests verify that individual components behave as specified under known conditions. And coverage in safety-critical development isn’t a spot-checking exercise. DO-178C requires 100% MC/DC coverage at DAL A, and ISO 26262 requires the same at ASIL D. Every line. Every branch. Every condition.
Each gate resets the accumulated uncertainty back toward zero before the next stage compounds it further. That’s not a new idea. It’s how you build software that people’s lives depend on.
The question AI raises isn’t whether to use gates. It’s whether the gates you already have are positioned to catch what an AI agent introduces and whether you’ve thought carefully about where in the workflow the uncertainty is actually accumulating.
The gates were designed for a world where code has an author who made deliberate choices. A human developer who writes an uninitialized variable made a mistake. A human developer who skips a boundary check made a tradeoff. Static analysis flags both—the developer understands the finding in context, and the correction is made by someone who knows what the code is supposed to do. The evidence trail is intact. The intent is recoverable.
An AI agent doesn’t make mistakes in that sense. It produces outputs that are statistically consistent with its training: plausible, often correct, and occasionally wrong in ways that look right.
The static analysis tool will still flag the MISRA violation. The unit test will still fail on the boundary condition. But the developer reviewing the finding is now one step removed from the original intent because there wasn’t original intent in the human sense. There was a probability distribution. And when you ask the agent why it made that choice, the answer is not recoverable in any form a certification auditor can use.
The gates catch the artifact. They don’t reconstruct the argument. In a safety case, you need both, and one of them must have been generated as the decisions were made, not reverse engineered from the output afterward.
The consumer technology press calls it “hallucination,” which means the AI confidently states something wrong. This term captures the symptom, but not the mechanism.
In safety-critical engineering the mechanism is what matters. ISO/PAS 8800, the emerging automotive standard for AI safety that the broader embedded industry is watching closely as a template, uses the term “functional insufficiency”: an unexpected error under specific conditions not adequately represented during development. As EDN noted, for engineers building software for medical devices, industrial automation, rail, aerospace, and defense, dismissing this document as “just for cars” would be a missed opportunity.
The distinction matters. Hallucination implies the system invented something from nothing. Functional insufficiency describes something more precise. The system performed exactly as its training data suggested it should, and the training data didn’t cover this case.
You can’t fix a hallucination by improving the model. You can’t fix a functional insufficiency that way either. What you can do is bound it, monitor it, and build an architecture that prevents it from propagating into a safety-critical decision unchecked.
None of this is an argument against AI in safety-critical development. These industries already have the architectural foundations to manage it responsibly. That argument is already lost, and it should be. AI tools are accelerating development, surfacing defects earlier, and handling the kind of repetitive verification work that exhausts engineers and introduces its own error rate.
The question was never whether AI would enter these industries. It’s here. The question is whether the engineering discipline surrounding it will keep pace.
Compound uncertainty doesn’t care about your intentions or your vendor’s benchmark scores. A 98% accurate agent in a 34-step open-loop workflow has already crossed the coin-flip line. Those numbers don’t improve because the use case is important or the schedule is tight.

Compound uncertainty in multi-step workflows. Even with 95% per-step accuracy, overall success rate declines sharply as the number of workflow steps (N) increases—not because model performance degrades, but because the workflow itself compounds error. Source: Parasoft
What does improve the outcome is treating AI in safety-critical development the way these industries have always treated unreliable components: with gates, evidence, and documented reasoning that survives an audit.
The standards that govern medical devices, aviation software, and automotive systems were written for a deterministic world. But the principles they encode—rigorous verification, traceable decisions, complete coverage, and structured safety arguments—turn out to be exactly the right response to a world where probabilistic behavior slipped into the development process before anyone checked its credentials.
ISO/PAS 8800 is the automotive industry’s first formal attempt to extend those principles into AI-specific territory. Other domains are watching. The framework outlined in the embedded world—manage uncertainty, bound it, argue it, and monitor it—applies whether you’re building firmware for a ventilator or a flight control system or an autonomous vehicle.
You will never eliminate functional insufficiency from an AI system. However, you can build an architecture that catches it before it becomes a safety event. That’s not a limitation of technology. It’s just engineering.
Arthur Hicken is a senior software evangelist at Parasoft.
Ricardo Camacho is director of product strategy for embedded and safety critical compliance at Parasoft.
Related Content
- AI Safety Moves to the Forefront
- Specifying Objectives is Key to AI Safety
- Can We Trust AI in Safety Critical Systems?
- Safe Automated Driving Starts with Architecture
- The impact of AI/ML on qualifying safety-critical software
The post Compound uncertainty: AI’s hidden risk in safety-critical development appeared first on EDN.


