Validating AI-Generated Automation Code Before It Touches a Machine
AI tools can now draft PLC routines, but a routine that compiles has not yet been shown to do what the machine needs. UTEC Industrial designs, engineers, machines, fabricates, and installs custom material handling systems for aerospace and heavy industry from its Spokane Valley, WA facility, integrating Allen-Bradley PLC and motion control with in-house CNC machining, heat treating, and stress relief. This article sets out how AI-drafted automation code is checked before it reaches a machine: what published studies found, what a compiler and a code review catch, how risk sets test depth, where simulation, FAT and SAT fit, and where safety-related code stays outside the process. On a heavy handling system the code sits in the controls link of one chain, design → engineering → parts machining → fabrication → assembly → weld fatigue → stress relief → drives → controls → tuning → monitoring, and it is validated against the physical machine the earlier links produced.
What does it mean to validate AI-generated automation code?
The FDA's General Principles of Software Validation guidance, written for medical device software, draws a distinction that is useful here by analogy. It says software verification "provides objective evidence that the design outputs of a particular phase of the software development life cycle meet all of the specified requirements for that phase," and that software testing is one of many verification activities; the others it lists include static and dynamic analyses, code and document inspections, and walkthroughs. For purposes of that guidance, FDA considers software validation to be "confirmation by examination and provision of objective evidence that software specifications conform to user needs and intended uses, and that the particular requirements implemented through software can be consistently fulfilled." The same section says the level of validation, verification and testing effort "will vary depending upon the safety risk (hazard) posed by the automated functions of the device."
The guidance's scope is medical device software, not machinery; its use here is an analogy and engineering reasoning, not a requirement on industrial equipment. The next two sentences are also engineering reasoning. Read that way, verification asks whether a drafted routine matches its specification, and validation asks whether the specification, as implemented, does what the machine's user needs. The failure mode that AI drafting adds is a routine that matches the prompt but not the requirement the prompt was meant to express. The article on what PLC copilots can and cannot draft covers the vendor tools themselves; this article covers the checks their output has to pass (FDA 2002, General Principles of Software Validation, §3.1.2 pp. 6–7).
Why is AI-drafted PLC code not trustworthy on sight?
NIST's Generative AI Profile, NIST AI 600-1, names two risks that bear directly on reviewing generated code. It defines confabulation as a phenomenon in which generative AI systems "generate and confidently present erroneous or false content in response to prompts," notes that these phenomena are colloquially also referred to as "hallucinations" or "fabrications," and says confabulations "are a natural result of the way generative models are designed." It also says humans "may over-rely on GAI systems or may unjustifiably perceive GAI content to be of higher quality than that produced by other sources," and calls this "an example of automation bias, or excessive deference to automated systems."
A study of control-logic generation reports a related observation. In an exploratory study of ChatGPT using GPT-4, researchers at ABB Research found that "many times, ChatGPT answers appear correct and well-formulated but can also contain subtle errors or omissions."
The NIST profile is cross-sectoral and does not mention PLCs or machinery; the next two sentences are engineering reasoning. The two risks compound in a controls review: output that reads well meets a reviewer inclined to accept it. A permissive with one inverted condition, or a timer preset in the wrong unit, can pass a quick read and still move a load at the wrong moment (Autio et al. 2024, NIST AI 600-1, §2.2 p. 6 and §2.7 p. 9; Koziolek, Gruener and Ashiwal 2023, §VI).
What have published studies found when large language models write PLC code?
Two conference papers give measured results, each under narrow conditions.
- Koziolek, Gruener and Ashiwal (ETFA 2023). This exploratory study created 100 prompts in 10 categories and generated answers with ChatGPT using GPT-4. Its abstract reports that the model "generated syntactically correct IEC 61131-3 Structured Text code in many cases." The authors also report that generating graphical notations, including Function Block Diagrams and Sequential Function Charts, in textual notations or ASCII art "largely failed," and that the model used language features and standard functions "which may not be supported by all PLC environments." Their threats-to-validity section states: "Many of the answers have not yet been compiled and thoroughly checked for functional correctness."
- Fakih et al., LLM4PLC (ICSE-SEIP 2024). The abstract states that LLMs "such as GPT-4 and LLaMa2 fail to produce valid programs for Industrial Control Systems (ICS) operated by Programmable Logic Controllers (PLCs)." The authors built a user-guided iterative pipeline with grammar checkers, compilers and SMV verifiers, validated it on a laboratory manufacturing testbed, and tested GPT-3.5, GPT-4, Code Llama-7B and -34B and fine-tuned versions of both Code Llama models. They report that the pipeline "improved the generation success rate from 47% to 72%, and the Survey-of-Experts code quality from 2.25/10 to 7.75/10."
The results carry conditions. Both studies targeted IEC 61131-3 Structured Text as the PLC language; neither tested ladder logic on a Rockwell Logix controller. The models tested were GPT-4 through ChatGPT in the 2023 study and the set listed above in the 2024 study, and the LLM4PLC figures come from one laboratory testbed. As engineering reasoning, the percentages describe those setups and do not transfer to another copilot, another language or another PLC platform (Koziolek, Gruener and Ashiwal 2023, abstract, §V, §VI and §VII; Fakih et al. 2024, abstract and §1).
What does a compile check catch, and what does it miss?
Koziolek and co-authors state the boundary plainly: "Syntactical problems can be easily identified by attempting to compile the generated code. Semantic problems may be harder to find." As engineering reasoning, a clean compile shows only that the compiler accepted the code; it does not show that the code does what the machine needs.
LLM4PLC adds a second, formal stage. Its pipeline checks syntax with an open-source IEC 61131-3 Structured Text compiler, then verifies compilable code with a symbolic model checker against formal properties, and feeds errors back to the model "with the option for human intervention." The authors state that on a verification-stage success "the code can be immediately deployed." Their properties are themselves produced by translating the plant's constraints from natural language into the model checker's specification language, using a feedback loop similar to the one used for the code.
The rest of this answer is engineering reasoning. A formal check proves the code against the properties that were written down, and no further. If the property set omits a requirement, for example that a transfer car must stop when its end-of-travel limit is made, the code can pass verification and still violate it. Where the same kind of model drafts both the code and the properties, a misreading of the specification can appear in both. A formal-verification pass is strong evidence about stated properties; the property list itself still needs a human review against the written requirements (Koziolek, Gruener and Ashiwal 2023, §VI; Fakih et al. 2024, §1 and §4).
How should a code review of AI-drafted logic be run?
NIST AI 600-1 includes, among its suggested actions to manage generative AI risks, "Review GAI system outputs for validity and safety: Review generated code to assess risks that may arise from unreliable downstream decision-making" (action MS-2.6-004). It does not say how to review PLC code.
The PLCopen Coding Guidelines, Version 1.0 (2016), supply a published rule set for PLC coding and code reviews. The document says IEC 61131-3 "does not describe how a programmer can increase the quality of the application program," and that it "consists of a set of rules that a programmer can use during coding and code reviews." It also states that it "contains no design rules," that "The set is certainly not complete," and that a programmer or organization can add or remove rules; the working group recommends tailoring a sub-set of rules to the application. Rules that its Annex 1 lists at High priority include:
- 5.3: All variables shall be initialized before being used
- 5.7: Error information shall be tested
- 5.8: Floating point comparison shall not be equality or inequality
- 5.11: Avoid multiple writes from multiple tasks
- 5.13: Physical outputs shall be written once per PLC cycle
Two conditions apply. The guidelines state that the working group's results should be based on the first and second editions of IEC 61131-3 and be easily extensible to the third, released in February 2013, and that the rules are written with the third edition in mind; they predate IEC 61131-3:2025, the fourth edition, which specifies the syntax and semantics of programming languages for programmable controllers as defined in IEC 61131-1. The guidelines also do not mention AI. The rest of this paragraph is engineering reasoning. Using them as a checklist for AI-drafted code is this article's application, and the reviewer is best someone other than the person who wrote the prompt. A copilot-added routine that writes an output the existing program already writes is one failure the multiple-writes and once-per-cycle rules address (PLCopen 2016, Coding Guidelines v1.0, §1 p. 9, §2 p. 10, §2.1 p. 11 and Annex 1; Autio et al. 2024, NIST AI 600-1, MS-2.6-004 p. 32; IEC 61131-3:2025, abstract).
How much testing does an AI-drafted routine need?
FDA's Computer Software Assurance for Production and Quality Management System Software guidance, issued February 3, 2026, offers a risk-based model by analogy. Every page is headed "Contains Nonbinding Recommendations," and its scope is computers and automated data processing systems used as part of medical device production or the quality management system, not industrial machinery in general. Its risk framework "can be applied, but is not limited, to automation tools (e.g., BOTS or automatic workflows), data analytic tools, artificial intelligence/machine learning tools, and cloud computing when used as part of production or the quality management system."
The guidance considers a software feature, function or operation to pose a high process risk "when its failure to perform as intended may result in a quality problem that foreseeably compromises safety, meaning a medical device risk." It defines unscripted testing, including scenario testing and experience-based testing such as error guessing and exploratory testing, and scripted testing, in which test cases are recorded and then executed manually or automatically. For high process risk features, manufacturers "may choose to consider more rigor such as the use of scripted testing or a hybrid approach of scripted testing and unscripted testing, scaled as appropriate"; for features that are not high process risk, they "may consider using unscripted testing methods." The guidance adds that these examples are not exclusive to those categories.
The rest of this answer maps the guidance to machinery; it is engineering reasoning, not FDA's text. On a handling machine, the analogue of high process risk is a standard-logic routine whose failure could move a load unexpectedly, drop a permissive, or mis-sequence a transfer. Those routines earn recorded, scripted test cases with expected results, including boundary values and fault reactions. Alarm text, status words and HMI data preparation can be covered by scenario and exploratory testing, provided the tester records what was run. Safety-related code is outside this scale entirely, as covered below (FDA-2022-D-0795, 2026, §I pp. 1–2, §V.A pp. 6 and 9, and §V.A(4) pp. 13–14).
Can the AI tool write its own test cases?
Some vendor tools are documented to generate test cases. The key-capabilities page of Siemens' Eigen Engineering Agent documentation, V1.5.0, lists "Generate and import LAD and SCL code" and "Generate executable test cases for SCL and LAD blocks" under code development and testing. Rockwell Automation's FactoryTalk Design Studio Copilot documentation states that the Copilot "can explain, but cannot modify, safety-related PLC code and device configuration." The Siemens documentation pages cited here do not state an equivalent restriction, and this article draws no conclusion from that about how the Siemens tool treats safety-related code.
Koziolek and co-authors note that testing control code without physical I/O "may require simulated input signal values, which could also be synthesized with ChatGPT."
The next three sentences are engineering reasoning. Generated tests can add coverage, but a test drafted by the same tool, from the same prompt, carries the same reading of the requirement as the code. If the prompt misstates a requirement, the code and its test agree and both are wrong. Acceptance test cases therefore trace to the written requirements, not to the prompt; the article on writing a user requirement specification covers how each requirement is made verifiable and assigned a verification method (Siemens AG 2026, Eigen Engineering Agent V1.5.0 documentation, key capabilities; Rockwell Automation 2026, FactoryTalk Design Studio Copilot, undated web documentation, accessed September 2026; Koziolek, Gruener and Ashiwal 2023, §VI).
Where do simulation and virtual commissioning fit before startup?
Virtual commissioning has its own guideline. The publisher's abstract for VDI/VDE 3693 Blatt 1:2025-05, Virtual commissioning — Model types, terms, and definitions, says it "provides a clear and systematic definition of virtual commissioning (VIBN) and classifies it in the life cycle of an automated plant or machine." It says the 2025 edition "has been revised and supplemented, particularly with regard to the representations of the basic test configurations, the test methods, the types of models required for this and simulation configurations in VIBN," and it addresses commissioning engineers, automation engineers and software developers among others. The abstract does not name the model types or test configurations, and this article does not name them either.
The order below is engineering reasoning, drawn together from the sources above:
- Compile, to catch syntax errors.
- Review against the tailored coding rules.
- Run the routine against simulated I/O or a plant model, including fault injections, before it reaches hardware; Koziolek and co-authors note that testing without physical I/O may require simulated input signal values.
- Repeat the functional tests on the real control panel and field devices at the factory acceptance test.
- Repeat the site-dependent tests at the site acceptance test.
The next two sentences are also engineering reasoning. A simulation proves the logic against the model's behavior, not against the machine's. Sensor noise, wiring errors, encoder scaling and direction, mechanical compliance and drive tuning belong to the physical machine and stay open until hardware tests close them (VDI/VDE 3693 Blatt 1:2025-05, publisher abstract; Koziolek, Gruener and Ashiwal 2023, §VI).
What must AI-drafted code prove at FAT and SAT?
The title of IEC 62381:2024 names three tests for automation systems in the process industry: the factory acceptance test (FAT), the site acceptance test (SAT) and the site integration test (SIT); its abstract also mentions the factory integration test (FIT). Its title names the process industry, and its clauses are not cited here. UTEC Industrial performs factory acceptance testing and on-site commissioning. The custom machinery project lifecycle article covers what a FAT proves and what has to wait for the site, and the URS article covers how requirements carry into factory and site acceptance testing and what regulated-industry frameworks add.
The rest of this answer is engineering reasoning. AI-drafted logic gets no lighter FAT than hand-written logic. Each acceptance test traces to a written requirement, and the test record shows which routines were AI-drafted, which lets a reviewer weight attention toward them. A failure mode to plan for is a FAT run on simulated field signals that passes, followed by a site failure when a real encoder counts in the opposite direction or a real load cell is scaled differently (IEC 62381:2024, title and abstract).
What changes when AI-drafted code is near a safety function?
The boundary itself is covered in the copilots article: Rockwell's Copilot documentation states that it can explain, but cannot modify, safety-related PLC code and device configuration. Two standards bear on the safety side.
- IEC 61508-3:2010, Software requirements, per its IEC abstract, applies "to any software forming part of a safety-related system or used to develop a safety-related system within the scope of IEC 61508-1 and IEC 61508-2." It "provides specific requirements applicable to support tools used to develop and configure a safety-related system within the scope of IEC 61508-1 and IEC 61508-2," and, with Parts 1 and 2, provides requirements for support tools "such as development and design tools, language translators, testing and debugging tools, configuration management tools." It establishes requirements for safety lifecycle phases and activities, including "measures and techniques, which are graded against the required systematic capability." The abstract does not mention AI or code generators. The IEC webstore page gives this edition a stability date of 2027 and shows a third edition as under development, with a forecast publication date of 22 October 2026.
- ISO 13849-1:2023 specifies a methodology and provides related requirements, recommendations and guidance for the design and integration of safety-related parts of control systems that perform safety functions, for high demand and continuous modes of operation; it does not apply to low demand mode of operation.
The rest of this answer is engineering reasoning. Whether a code-generating AI tool would count as a support tool under IEC 61508-3:2010 is not settled by the abstract, and a project that wants AI help near safety functions needs that question answered by its functional-safety lead before the tool is used. The simpler practice is to keep AI-drafted code on the standard side of the program. A failure mode that crosses the boundary anyway is a standard-side routine that writes a tag the safety logic reads as an input or permissive, which is one more reason to review every tag a drafted routine writes (IEC 61508-3:2010, Ed. 2.0, abstract; ISO 13849-1:2023, scope; Rockwell Automation 2026, FactoryTalk Design Studio Copilot, undated web documentation, accessed September 2026).
Which controls and sensing behaviors should the test plan exercise?
A drafted routine runs inside the controller's task model, not on its own. Rockwell's Logix 5000 design considerations manual says tasks can be configured as continuous, periodic, or event, and that "A periodic task performs a function at a specific time interval." The copilots article covers task priorities and the scan in detail. The PLCopen rules that bear on this layer include 5.11, avoid multiple writes from multiple tasks; 5.12, manage synchronization among tasks; 5.13, physical outputs shall be written once per PLC cycle; and 5.16, read a variable written by another task only once per cycle, all listed at High priority in Annex 1.
The test list below is engineering reasoning for a heavy handling machine:
- Position feedback: encoder loss, wrong count direction, and a position jump after a drive fault reset.
- Limit and zone sensing: end-of-travel and overtravel switches made in the wrong order, a stuck-made switch, and a zone interlock dropped mid-move.
- Load sensing: a load cell reading out of range, or zero with a load present.
- Drives: a VFD or servo fault word during motion, and the routine's response when the drive refuses a start.
- Networks: loss of the remote I/O or drive connection, and the state of every output the routine writes when it returns.
- Modes: manual, automatic and maintenance changes mid-sequence, and restart after a stop.
As engineering reasoning, that list follows the signature chain: the frame design, machining, fabrication and stress relief set the stiffness and alignment the sensors and drives see, the drives are tuned against that machine at commissioning, and monitoring then trends it in service. UTEC Industrial builds this layer as a Rockwell Automation Recognized System Integrator on Allen-Bradley ControlLogix and CompactLogix platforms, with VFD and servo drives and UL 508A panel building (Rockwell Automation 1756-RM094N-EN-P-2025, Ch. 5 pp. 39 and 41; PLCopen 2016, Coding Guidelines v1.0, §5.11–5.16 and Annex 1).
What records should follow AI-drafted code into service?
NIST AI 600-1 lists "code generation and review" among the applications in which AI actors interact with generative AI systems, and says that organizations' use of generative AI systems "may also warrant additional human review, tracking and documentation, and greater management oversight." The profile describes itself as a cross-sectoral profile of and companion resource for the AI Risk Management Framework, NIST AI 100-1, whose core is organized into four functions: Govern, Map, Measure and Manage. The copilots article applies those functions to a controls team. The LLM4PLC authors add a warning of their own in their section on industry challenges: "un-regulated use of such methods will inevitably lead to unverified and unmoderated deployments in critical systems."
The record below and the failure mode after it are engineering reasoning, one entry per accepted AI-drafted change:
- the tool, its version, and the requirement or prompt the change answered
- the person who accepted the change and the person who reviewed it
- the coding rules applied and any rule deliberately waived, with the reason
- the test cases run, their results, and whether each ran on simulation, at FAT or at SAT
- the controller program revision that first carried the change
A failure mode this record guards against is a later edit that reintroduces an error a reviewer already removed, because nobody could see why the earlier version was changed (Autio et al. 2024, NIST AI 600-1, Appendix A p. 47; Fakih et al. 2024, §8; NIST AI 100-1, 2023).
- Can Generative AI Write PLC Code? Copilots and Their Documented Limits — what copilots can and cannot draft
- Writing a User Requirement Specification (URS) for Custom Machinery — verifiable requirements that acceptance tests trace to
- Custom Machinery Project Lifecycle: Concept, Design, FAT, Install, Support — FAT and commissioning, where drafted logic meets hardware
- Connecting a FANUC Robot to an Allen-Bradley PLC over EtherNet/IP — handshake logic that drafted routines are tested against
- Lockout/Tagout for CNC Equipment: OSHA Requirements and Best Practices — why tested logic still does not replace energy isolation
References
- FDA: General Principles of Software Validation; Final Guidance for Industry and FDA Staff. U.S. Food and Drug Administration, 2002.
- Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., and Roberts, K. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. National Institute of Standards and Technology, 2024.
- Koziolek, H., Gruener, S., and Ashiwal, V. (2023). "ChatGPT for PLC/DCS Control Logic Generation." 2023 IEEE 28th International Conference on Emerging Technologies and Factory Automation (ETFA), 1-8.
- Fakih, M., Dharmaji, R., Moghaddas, Y., Quiros Araya, G., Ogundare, O., and Al Faruque, M. A. (2024). "LLM4PLC: Harnessing Large Language Models for Verifiable Programming of PLCs in Industrial Control Systems." Proc. 46th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP '24), ACM, 192-203.
- PLCopen: Coding Guidelines, Version 1.0 (Software Construction Guidelines). PLCopen, 2016.
- IEC 61131-3:2025: Programmable controllers — Part 3: Programming languages. IEC, 2025 (Ed.4).
- FDA-2022-D-0795: Computer Software Assurance for Production and Quality Management System Software: Guidance for Industry and Food and Drug Administration Staff. U.S. Food and Drug Administration, 2026.
- Siemens AG: Eigen Engineering Agent, V1.5.0 online documentation (DocV001). Siemens AG, 2026.
- Rockwell Automation (2026). FactoryTalk Design Studio Copilot. FactoryTalk Design Studio Online Documentation. Rockwell Automation, 2026 (undated web documentation, accessed September 2026).
- VDI/VDE 3693 Blatt 1:2025-05: Virtual commissioning — Model types, terms, and definitions. VDI Verein Deutscher Ingenieure e.V., 2025.
- IEC 62381:2024: Automation systems in the process industry — FAT, SAT, FIT and SIT. IEC, 2024 (Ed.3).
- IEC 61508-3:2010: Functional safety of electrical/electronic/programmable electronic safety-related systems — Part 3: Software requirements. IEC, 2010 (Ed.2.0).
- ISO 13849-1:2023: Safety of machinery — Safety-related parts of control systems — Part 1: General principles for design. International Organization for Standardization, 2023.
- Rockwell Automation 1756-RM094N-EN-P-2025: Logix 5000 Controllers Design Considerations. Rockwell Automation, 2025.
- NIST AI 100-1: Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, 2023.
Ready to Discuss a Material Handling System?
UTEC Industrial designs, engineers, machines, fabricates, and installs custom material handling systems for heavy industry, from the stress-relieved structure and drives to the Allen-Bradley PLC controls, tuning, and monitoring that run them, at its Spokane Valley, WA facility. Send UTEC the application, loads, and duty cycle to start a system review.
Questions? Call (509) 922-1832 or email sales@utec.co