AI & DataReview ArticlePublished 8/10/2026 · 31 views8 downloadsDOI 10.66308/air.e2026066

Why Large Language Models Cannot Be Certified for Safety-Critical Systems

Akbar SayakovBase80, AI Systems Architect, San Francisco, CA, USA
Received 7/16/2026Accepted 8/9/2026
large language modelshallucinationfunctional safetycertificationIEC 61508DO-178Csafety-critical systemsdeterministic execution
Download PDF
Cover: Why Large Language Models Cannot Be Certified for Safety-Critical Systems

Abstract

Large language models (LLMs) are moving rapidly into domains governed by functional-safety certification: aviation, road vehicles, medicine, and critical infrastructure. Certification regimes such as IEC 61508, DO-178C, and ISO 26262 combine system-level risk targets with process-, traceability-, configuration-, and evidence-based obligations, in mixes that differ by regime but share a demand for verifiable, bounded behavior. This review synthesizes three literatures that rarely meet: statistical learning theory on hallucination, empirical measurements of LLM error rates in high-stakes domains, and the standards and regulatory documents that define certification. The convergent finding is that the mismatch between the two worlds is structural rather than incidental. Impossibility results establish a nonzero error floor for calibrated probabilistic generators under stated conditions; reported benchmark metrics in legal, medical, and agentic tasks are not directly convertible to certification targets, and no published deployment supplies the system-level hazard and exposure model that a compliance demonstration would require; and every major mitigation family (retrieval augmentation, guardrails, formal verification, uncertainty quantification, and neurosymbolic hybrids) is documented to narrow, but not close, the gap. We formalize error compounding over execution horizons, tabulate the documented limits of each mitigation, and examine why plausibility cannot substitute for assurance. Two coherent exits emerge: statistical acceptance criteria in standards for bounded tasks, and architectures that confine the LLM to an untrusted-proposer role inside a deterministic, independently verifiable execution envelope. Certifying the envelope rather than the model is, on the current evidence, the most defensible path consistent with both the mathematics and the standards.

Keywords: large language models, hallucination, functional safety, certification, IEC 61508, DO-178C, safety-critical systems, deterministic execution

Cite asAkbar Sayakov (2026). Why Large Language Models Cannot Be Certified for Safety-Critical Systems. American Impact Review. https://doi.org/10.66308/air.e2026066Copy

Declarations

Data availability

All sources reviewed are publicly available at the locations cited in the References; no new data were generated.

Funding

The author received no external funding for this work.

Competing interests

The author declares no conflicts of interest.