Robot Risk Assessment with FMEA, HAZOP and STPA

Robot risk assessment identifies how people can be harmed across the application lifecycle and selects controls until residual risk is acceptable under the applicable process. It covers more than robot-arm faults: tools, process energy, workpieces, layout, software, people and operating modes share the boundary.

FMEA, HAZOP and STPA are complementary analysis techniques. FMEA begins with failure modes, HAZOP explores deviations from design intent and STPA examines unsafe control actions and inadequate constraints, including scenarios in which individual components have not failed.

This guide is educational, not a compliance determination. Use it with the Physical AI safety-layers guide and runtime safety-shield guide.

Set the complete system and lifecycle boundary

Define robot, tool, workpiece, fixtures, process energy, controls, networks, environment and interacting people. Cover transport, installation, setup, teaching, automatic operation, cleaning, recovery, maintenance and decommissioning.

Write intended use, reasonably foreseeable misuse, access routes, operator competence and environmental limits. Excluding a mode because production is not supposed to use it does not remove foreseeable exposure.

Industrial robotic welding cell enclosed by metal guards and welding curtains
A robot application includes guarding, process energy, fixtures, controls, access and maintenance modes; the photograph does not demonstrate completion of a risk assessment. Source: Wikimedia Commons contributor. License: CC BY-SA 4.0.

Define losses, hazards and acceptance rules first

Name unacceptable losses such as injury, exposure, dropped load or loss of control. Connect them to hazardous system states and credible scenarios before assigning scores.

Document the risk-estimation method and escalation rules. Numeric ranks help prioritize work but do not turn uncertainty into objective truth or permit serious hazards to disappear through multiplication.

MethodStarts fromFinds wellBlind spot if used alone
FMEAComponent or function failureFailure effects and detectionUnsafe normal interaction
HAZOPDeviation from design intentProcess and command deviationsComplex control structure
STPALosses and control constraintsUnsafe control interactionsDetailed component reliability
Incident reviewObserved eventReal operating evidenceUnknown unobserved cases
Task analysisHuman work sequenceExposure and misuseLatent technical failures

Use FMEA for failure effects and diagnostics

List functions, failure modes, local and system effects, causes, existing controls and detection. Follow failures through safety functions rather than stopping at a component symptom.

Treat common-cause and dependent failures separately. A diagnostic that shares power, timing or software with the failed function may not provide independent coverage.

Use HAZOP for deviations from intent

Apply guide words to commands and process variables: no, more, less, reverse, early, late, other than or part of. For a robot, deviations may involve speed, force, position, identity, sequence, timing and authority.

Record cause, consequence, safeguard and action. HAZOP works best with a multidisciplinary team that understands both process physics and control implementation.

Use STPA for unsafe control scenarios

Model controllers, controlled process, feedback and control actions. Ask whether an action is not provided, provided when unsafe, provided too early or late, or applied too long or stopped too soon.

The MIT STPA Handbook explains the method. Apply it within a defined project and retain evidence; using the name STPA does not guarantee complete scenario coverage.

Five-stage robot risk assessment validation
A risk worksheet is incomplete until each material scenario has an implemented control, verification case, owner and reassessment trigger. Source: Physical AI Lab.

Cover AI distribution shift and authority conflicts

AI robot hazards include unfamiliar observations, brittle language grounding, delayed inference, stale world state and a policy requesting motion outside the verified envelope. These are not always reducible to a failed hardware part.

Analyze conflicts among autonomy, teleoperation, safety PLC, human intervention and recovery controllers. Define who has authority, how transitions occur and what state is safe when channels disagree.

Select risk reduction in a hierarchy

Prefer eliminating hazards or reducing them through inherently safer design, then safeguards and complementary protective measures, then information and training. A warning cannot compensate for an avoidable pinch point or exposed process energy.

Verify that a control works in every relevant mode and does not create a new hazard. Record residual risks that must be communicated and managed.

Turn findings into testable safety requirements

Each material scenario should produce an owned requirement with trigger, state, response time, final condition, fault behavior and acceptance evidence. Avoid statements such as system shall be safe or AI shall be reliable.

Link requirements to design elements and verification cases in both directions. An implemented safeguard without a traced hazard and a hazard without a verification case are both gaps.

Use current standards and guidance carefully

The ISO catalog states that ISO 12100:2010 remains the current published machinery risk-assessment edition while a replacement draft is under development. Project applicability and transition rules must be checked at decision time.

Use the actual normative texts, Type-C standards and jurisdictional rules with competent professionals. This article explains analysis structure and does not assign a required risk level.

Include operations, incidents and near misses

Field observations reveal shortcuts, workload, nuisance stops and recovery practices missing from design workshops. Feed incidents, near misses, interventions and maintenance findings back into scenarios and controls.

NIOSH maintains a robotics safety research program relevant to workplace evidence. Treat public guidance as supporting material, not a substitute for application assessment.

Reassess every material change

Trigger review after tools, payload, program, speed, layout, sensors, firmware, AI model, data, network, operator role or operating environment changes. Compare the change with hazard assumptions and safety-function validation.

Connect emerging failures to the failure-mining guide but keep protective requirements independent of learning updates.

Trace linkRequired artifactAudit questionFailure
Hazard to controlRisk decisionWhy this controlUnjustified design
Control to requirementSpecificationWhat must happenVague intent
Requirement to testCase and resultHow verifiedPaper control
Test to baselineConfigurationWhat was testedVersion gap
Change to reassessmentImpact recordWhat became invalidStale safety file

Release a living risk record

Preserve boundaries, assumptions, participants, methods, scenarios, estimates, controls, residual risks, requirements, tests, configuration baseline, owner and reassessment triggers. Keep disagreements and uncertain evidence visible.

Close review with the following checks.

  • Cover all lifecycle modes and foreseeable misuse.
  • Combine component, deviation and control-scenario views.
  • Treat AI shift and authority transitions explicitly.
  • Trace every material hazard to a verified control.
  • Maintain the assessment through incidents and changes.

Frequently asked questions

Is FMEA alone a complete robot risk assessment?

No. It is useful for failure modes but can miss unsafe interactions and normal-operation hazards.

Can a low RPN be ignored?

Not automatically; severity, uncertainty, legal requirements and method-specific rules still require judgment.

Does STPA replace FMEA?

No. They answer different questions and are often complementary.

Who should participate?

Include design, controls, process, safety, operations, maintenance and affected user expertise.

Does an AI model update trigger reassessment?

Yes when behavior, distribution assumptions, timing, interfaces or verified controls may change.

Hazard-to-Requirement Traceability Boundary

A robot risk assessment succeeds through traceability from hazard scenario to implemented requirement, verified evidence and controlled change—not through the name or score of one analysis method.