Welcome Back

Sign in to your PART66Online account

One click — no password needed

or use email
Forgot password?

Don't have an account? Register here

Module 9 — Human Factors

9.1 — General

Free Preview
On this page (3)

Human Factors is the study of how humans interact with their working environment — the equipment, procedures, other people, and the physical surroundings. In aviation maintenance, understanding human factors is not optional; it is a regulatory requirement and a matter of life and death. Research consistently shows that approximately 80% of aviation accidents involve human factors. This module equips maintenance engineers with the knowledge to recognise, understand, and mitigate the human contributions to error.

The Need to Take Human Factors into Account

Aircraft maintenance is performed by people — and people are fallible. No matter how skilled, experienced, or conscientious an engineer is, they will make errors. The goal of human factors training is not to eliminate human error (which is impossible) but to:

  • Understand why errors occur — the underlying psychological, physiological, and organisational causes.
  • Recognise the conditions that make errors more likely — fatigue, time pressure, poor communication, inadequate resources.
  • Design systems that are error-tolerant — procedures, checklists, inspections, and organisational structures that catch errors before they cause harm.
  • Promote a safety culture where reporting errors and hazards is encouraged, not punished.

What "Human Factors" Actually Means

The term is used loosely in the industry, and the loose use costs marks. Outside maintenance it is heard most often in connection with flight-deck design and Crew Resource Management, but that is one application of the subject rather than the subject itself. In the maintenance context the UK Civil Aviation Authority's CAP 715, the guidance document written to introduce this syllabus, defines human factors as the study of human capabilities and limitations in the workplace: the interaction between maintenance personnel, the equipment they use, the written and verbal procedures and rules they follow, and the environmental conditions of the system they work in, with the aim of optimising that relationship so as to improve safety, efficiency and well-being.

Three consequences follow from that definition, and each of them is a distinction the examination tests directly.

  • It concerns the person inside a working system, not the aircraft. The strength of a lap joint, the fatigue life of a spar cap and the composition of an alloy belong to structures and materials. They become human factors questions only at the point where a person has to inspect, interpret, judge or repair them — which, in maintenance, is almost immediately.
  • It is not flight-deck human factors. Analysing instrument layout and pilot response under abnormal conditions is a legitimate and closely related field, but it studies a different population doing a different job under different constraints. Maintenance human factors studies the engineer, the hangar, the manual, the shift and the handover.
  • It is a systems subject, not a character assessment. The question it asks about an occurrence is not "who was careless" but "what about this task, this workplace and this organisation made that error likely, and what would have caught it". An investigation that arrives at the name of a person and stops has produced a disciplinary outcome, not a safety improvement.

Part-145 carries its own definition, and it is worth reading because it is deliberately broad. It describes human factors as the principles that apply to aeronautical design, certification, training, operations and maintenance, and that seek a safe interface between the human being and the other components of the system through proper consideration of human performance. Human performance, in the same sense, means the capabilities and limitations of the human being that have a bearing on the safety and efficiency of aviation. Notice what that pairing does: human factors is the body of principles, human performance is the raw material those principles have to work with. You cannot change the second, so the principles are entirely about arranging the five activities the definition names — design, certification, training, operations and maintenance — around it.

Because it is defined by the problem rather than by a technique, the subject borrows from several disciplines at once: human physiology; psychology, including perception, cognition, memory, social interaction and error; workplace design; environmental conditions; the human-machine interface; and anthropometrics, the measurement of the human body. A single maintenance task can call on all of them. Reading a torque value from a manual under a wing at three in the morning involves visual acuity and lighting (physiology and environment), working memory while you carry the figure to the tool (psychology), reach and posture in the space available (anthropometrics and workplace design), and the legibility of the page itself (human-machine interface). None of those is an aircraft systems question, and every one of them can produce a wrong torque.

Why the Human Share of Occurrences Keeps Growing

Two things have been happening in parallel for the whole of the jet age, and they move in opposite directions.

The first is that the machinery has become dramatically more reliable. Damage-tolerant structural design, redundancy in every critical system, better materials and processes, built-in test equipment that reports a fault before anyone has to look for it, and engine reliability good enough to make long-range twin-engine operations routine — each of these removes a class of event whose cause was simply that a component broke. That improvement is real, it is enormous, and it is a large part of the reason the overall accident rate has fallen.

The second is that human reliability has not improved to anything like the same degree, and the reason is structural rather than cultural. As CAP 715 puts it, aircraft become more and more reliable through design and manufacture, but the human being cannot be re-designed; we have to accept that the human being is intrinsically unreliable, and can only work around that unreliability through good training, procedures, tools and duplicate inspections. Vigilance still decays with time on task. Memory is still displaced by interruption. Performance at four in the morning is still worse than performance at four in the afternoon. No engineering programme changes any of that.

Put those two trends together and the consequence is arithmetic rather than opinion: as the number of purely mechanical events falls and the number of events with a human contribution does not, the human contribution becomes a larger fraction of a smaller total. This is the single most misunderstood statistic in the module, so it is worth working through with numbers.

How a safer fleet produces a higher human percentage

Take a fleet that suffers 100 airworthiness events in a year. Suppose 40 of them are a component failing on its own with nobody at fault, and the other 60 have a human contribution somewhere in the chain. The human share is 60 out of 100, which is 60%.

Now improve the hardware. Better materials, redundancy and built-in test halve the pure component failures, from 40 down to 20. Nothing whatever has changed about the people, so their 60 events are still 60. The fleet now has 80 events in the year instead of 100 — a genuine improvement of a fifth — but the human share is 60 out of 80, which is 75%.

The human percentage rose by fifteen points in a year in which the people got neither better nor worse and the fleet got measurably safer. That is the mechanism behind every statement that human factors are a growing proportion of the problem. It also shows why the percentage on its own is a poor measure of an organisation's performance: it can be pushed up by good engineering and pushed down by bad.

The historical record supports the point. CAP 715 records that as early as 1940 it was being calculated that roughly 70% of aircraft accidents were attributable to human performance, and that when the International Air Transport Association reviewed the position some thirty-five years later it found no reduction in the human-error component of the accident statistics. Over that period the aircraft changed out of all recognition. The human share did not fall.

Within that human share, maintenance is not a footnote. A study of 93 aircraft accidents carried out in the United States in 1986 by Sears, reported in CAP 715, ranked the significant causes and major contributory factors. Deviation by the pilot from basic operational procedures appeared in 33% of them and inadequate cross-checking by the second crew member in 26%, but maintenance and inspection deficiencies appeared in 12% — essentially level with design faults at 13%, and ahead of several factors that receive far more attention. A comparable exercise by the UK CAA in 1998, examining the causes of 621 global fatal accidents between 1980 and 1996 and published as CAP 681, again found maintenance or repair oversight, error or inadequacy among the top ten primary causal factors.

The practical conclusion for an engineer sitting this examination is blunt: accidents and engineering faults traceable to maintenance and inspection are not a declining or a stable problem. Measured as a share of what still goes wrong, they are significant and increasing, and that trend is the reason human factors became a mandatory element of the licence syllabus and of maintenance organisation training rather than an optional extra for those who found it interesting.

Reading Maintenance Error Statistics

Three quite different figures circulate in this subject and they are constantly confused with one another, in the hangar and in examinations. They are not in conflict; they count different things. Get the population right and each of them is useful.

FigureThe population it is drawn fromWhat it does not say
The figure quoted at the head of this note: around 80% of aviation accidents involve human factorsEvery human contribution anywhere in the system taken together — flight crew, air traffic control, design, manufacture and maintenanceThat maintenance engineers cause 80% of accidents. Maintenance is one contributor inside that total, alongside several larger ones.
Maintenance and inspection deficiencies were a factor in 12% of 93 accidents (Sears, 1986, reported in CAP 715)Accidents in which a maintenance or inspection shortfall was one of the significant causes. An accident could carry several such factors at once, so the column does not total 100%.That maintenance error is a minor issue. The same study put design faults at 13%, so the two are of comparable weight, and 12% of hull losses is a great many lives.
Omissions 56%, incorrect installation 30%, wrong parts 8%, other 6% (Reason's analysis of 122 maintenance incidents in one airline over three years, reported in CAP 715)The internal breakdown of maintenance errors that are already known to have occurred. These four categories are exhaustive, so they total 100%.That 56% of accidents are omissions. These are shares of maintenance errors, not of accidents, and the great majority of them never became an accident.

The first figure is also met in a narrower form, applied to maintenance rather than to aviation as a whole: that around 80% of maintenance-related incidents and accidents have human error as a contributing factor. In that form it is far less contentious, because a maintenance-related occurrence is very nearly by definition one in which something a person did, or did not do, is part of the story. Keep the two readings apart, because examination questions use both. The broad figure is about the whole system, in which maintenance is one slice; the narrow figure is about that slice on its own.

The third row deserves a moment on its own, because it tells an engineer where to concentrate. The largest single category of maintenance error is not doing the wrong thing; it is not doing something at all. Omission dominates because reassembly is intrinsically more error-prone than disassembly. There is normally one way to take an assembly apart and a very large number of ways to put it back together, of which only one is right and many differ from the right one only by something that has been left out — a seal, a lockwire, a clip, a cover, a torque. Several of the events summarised in the next section turned on something that was simply not done at all: screws not replaced, locking devices not fitted, covers not refitted.

Two ways to misuse these numbers. The first is to quote the human-factors share as though it measured maintenance alone: "engineers cause eighty per cent of accidents" is not what any of these studies found, and it invites the defensive reaction that kills a reporting culture. The second is subtler and more damaging. It is to treat "human error" as the conclusion of an investigation rather than its starting point. Naming the error tells you what happened; it tells you nothing about why a competent person made it on that night, why nothing downstream caught it, and whether the same conditions are still sitting in the hangar tonight. Those are the questions this module exists to answer.

Who the Requirement Reaches

The obligation itself is short. Point 145.A.30(e) of Part-145 requires the maintenance organisation to establish and control the competence of its personnel to a standard agreed with the competent authority, and states that competence must include, in addition to the expertise needed for the job function, an understanding of the application of human factors and human performance issues appropriate to that person's function in the organisation. The rule text stops there. The structure of the training, the list of who must receive it and the length of the continuation cycle quoted in the next section all sit in the Acceptable Means of Compliance rather than in the rule.

That distinction matters in practice, and it is not a loophole. An Acceptable Means of Compliance is the Agency's published way of satisfying the rule; an organisation may propose an alternative means to its competent authority and have it accepted, but it cannot simply ignore the published one. In practice almost every organisation adopts the AMC, writes it into its maintenance organisation exposition, and is then bound by its own exposition — which is audited.

The AMC's list of who should receive initial and continuation human factors training is much wider than the people who touch the aircraft. It covers, as a minimum:

  • post-holders, managers and supervisors;
  • certifying staff, support staff and mechanics;
  • technical support personnel such as planners, engineers and technical records staff;
  • quality control and quality assurance staff;
  • specialised services staff;
  • human factors staff and human factors trainers;
  • stores department staff and purchasing department staff;
  • ground equipment operators.

The breadth of that list is the answer to a question the bank asks in several forms: whom does human factors matter to? It is equally important to technicians, engineers, planners and managers, and the reason is visible in the causal chains of every case in the next section. A planner who compresses a task into a window that does not fit it has created time pressure that a technician will absorb. A purchasing decision that leaves the correct part out of stock creates the moment at which somebody has to decide whether a near-equivalent will do. A stores absence removes an independent pair of eyes from the issue and check of equipment, without anyone deciding to remove a defence. None of those three people ever picks up a tool, and each of them can set an error in motion. The person holding the spanner is usually the last link in the chain and almost never the only one.

Training is proportionate rather than identical. The AMC allows the syllabus to be adjusted to the nature of the organisation and to the nature of each function within it, and gives examples: planners may cover the scheduling and planning objectives in more depth and the shift-working material in less; a small organisation that does not work shifts may treat teamwork and communication more briefly. Adjusted, however, is not the same as omitted — the topics remain the topics, and the syllabus for initial training is published as Guidance Material to the same point of Part-145.

On timing, the AMC expects all personnel, including people recruited from another organisation, to have received initial human factors training that meets the organisation's own training standard before they start their actual job function, unless a competence assessment justifies otherwise. There is one narrow relaxation: newly employed personnel who are working under direct supervision may receive the training within six months of joining. That relaxation is worth understanding rather than memorising: it is granted only because direct supervision is itself a defence, so it lapses at the moment the supervision does.

The SHELL Model

The SHELL model (developed by Edwards, 1972, modified by Hawkins, 1975) illustrates the relationships between humans and their working environment. The model places the individual (Liveware) at the centre, surrounded by four elements they interact with. The edges between the centre and each surrounding block are irregular — representing the need to match and adapt the interfaces.

Layer 1 The SHELL Model S-L H-L E-L L-L Mismatches at any interface can lead to human error Software Procedures, manuals, checklists, rules Environment Temperature, noise, lighting, workspace Hardware Tools, equipment, aircraft, machines Liveware (The Person) Liveware Other people — colleagues, crew, ATC

SHEL, SHELL and the Second Liveware

Two spellings of the name are in circulation, SHEL and SHELL, and they refer to the same model. Edwards originated the concept in 1972, the name being taken from the initial letters of its components — Software, Hardware, Environment, Liveware — and Hawkins developed the modified diagram in 1975. The diagram above is that Hawkins version, and it is the one ICAO printed in Human Factors Digest No 1, published as Circular 216: the figure there is captioned "The SHEL model as modified by Hawkins", it places a person at the centre with a second Liveware block below for the other people that individual deals with, and the text alongside it gives the Liveware–Liveware interface a numbered paragraph of its own. So the spelling carries no information about the diagram — ICAO writes SHEL and its figure already has the second L, and the CAA's CAP 715 prints the same five-block figure under the same four-letter name. This note uses SHELL throughout, and the interfaces are named from the central person outwards: L–S, L–H, L–E and L–L.

The second L is not a fifth category of thing in the world. It is the recognition that other people are an interface in their own right and cannot be filed under Environment. Treating colleagues as part of the surroundings hides the failure modes that matter most in maintenance, because those failures are transactions rather than conditions: a handover that transfers a task but not its state, a briefing given verbally and remembered differently by two people, an instruction issued to somebody who is not qualified to question it, a junior engineer who sees something wrong and does not say so. None of those is a property of the hangar. Each is a property of an exchange between two people, and each has appeared as a causal factor in the accidents in the next section.

Why the Edges Are Ragged

The most important feature of the diagram is the one easiest to overlook: the blocks do not have straight edges, and they are drawn that way on purpose. The model's claim is that failures occur at the joints, not inside the blocks. Every block can be individually sound and the system can still fail, because two sound blocks do not fit each other.

Consider three cases where nothing is defective. A maintenance manual procedure that is entirely correct, but written for a different variant of the same type, is a good piece of software meeting a person who has no way of knowing it does not apply to the aircraft in front of him. A torque wrench that is in calibration and functioning perfectly, but whose scale cannot be resolved in the light available under a cowl, is good hardware meeting an environment that has disabled it. An accurate, complete verbal briefing given by one competent engineer to another, with no written record, is two competent people meeting an interface that has no memory. In each case an inspection of the parts finds nothing wrong. The mismatch is the fault.

There is a second-order effect that students almost always miss, and it is where real occurrences come from: the interfaces are not independent of one another. Each one silently sets the conditions under which the others can function. Degrade the lighting and you have not merely made the hangar unpleasant — you have disabled every defence that consists of somebody seeing something, without laying a finger on the hardware that provides those defences and without producing any record that they are now absent. Take a person's assistant away and you have done the same to every defence that consists of somebody else noticing. This is why an environmental or a staffing problem is never merely a comfort or an efficiency problem: it removes protection that the paperwork still says is in place.

One further asymmetry decides what to do about a mismatch. Four of the five blocks can be changed at will; the central Liveware cannot. Selection and training do adapt the person to the job, and they are genuine levers — a trained engineer performs a task quite differently from an untrained one. But the capacities that actually fail in these accidents are biological: the acuity of the eye in poor light, the persistence of working memory across an interruption, the level of alertness available at four in the morning. Instruction does not change them, and neither does willingness. So where a mismatch is found, the reliable fix is to change the software, the hardware, the environment or the way people are organised; asking the person to try harder is the weakest available response and the one that decays fastest once the briefing is over. CAP 715 states the aim as fitting the man to the job and the job to the man. Both halves are real. Only the second one still works on a bad night.

SHELL Interfaces Explained

InterfaceDescriptionExample of Mismatch
L–H (Liveware–Hardware)How the person interacts with tools, equipment, controls, and displaysControls placed out of reach; displays too small to read; tools that don't fit the user's hand size
L–S (Liveware–Software)How the person interacts with procedures, manuals, checklists, and regulationsProcedures that are ambiguous, too complex, or contradict each other; poorly written manuals
L–E (Liveware–Environment)How the person is affected by the physical work environmentExcessive noise preventing communication; poor lighting during inspections; extreme temperatures
L–L (Liveware–Liveware)How people interact with each other — communication, teamwork, supervisionPoor shift handover; language barriers; lack of assertiveness; personality conflicts

Working the Model Through One Job

The 1990 British Airways windscreen accident is summarised in the table at the start of the next section; it is used here only to show the model doing work. Reduce the whole event to a single sub-task — a shift maintenance manager selecting replacement bolts in a stores area during a night shift — and every interface in the model can be seen failing quietly, none of them fatally on its own. The account below follows CAP 715's description of the accident.

InterfaceWhat was actually the case on that jobWhy it mattered
L–SThe manual procedure for the job called for neither a pressure check nor a duplicate inspection afterwards.One person's judgement was the only barrier between the selection and the flight. There was no stage at which a second person or a test could disagree.
L–HThe correct bolt and the bolt actually fitted differed by a small amount of diameter and by one character of part number. Both went in and both held on the ground.Nothing in the hardware resisted the substitution. A design that cannot be assembled wrongly needs no vigilance; this one needed a great deal.
L–EThe selection was made at night, in a poorly lit stores area, by matching sight and touch against a sample bolt carried from the aircraft, and without spectacles.Discriminating a small difference in diameter by eye and hand is at the edge of human capability in good light. In that light it was not a realistic task at all.
L–LShort-handed, the manager elected to carry out the job himself and signed it off himself. The storeman advised that a different bolt was required, and the advice did not change the outcome.The supervisory layer that would normally review the work was occupied doing it. A correct challenge was made by the one person present who was not under time pressure, and it was absorbed.

Not one of those four rows is a failure of technical knowledge, and the engineer concerned was not judged incompetent by anybody. Each is a mismatch at an interface and each, on its own, was survivable. The accident required all four to line up on the same task on the same night.

The clearest illustration of interfaces acting on each other is also in this case. Once the bolts were in position the countersink sat lower than it should have done. That is a physical cue: the hardware was reporting, in the only language it has, that the fit was wrong. It was not noticed. A cue is only a defence when the conditions allow somebody to see it and the working arrangement gives somebody the job of looking — which is the environment interface and the person-to-person interface jointly deciding whether the hardware interface gets to do its job. Read that way, the model stops being a diagram to memorise and becomes a checklist for planning work: for this task, on this night, which of the four is carrying the load, and what happens if it does not hold?

Examination points from this section: human factors is the study of how human capabilities, limitations, behaviour and performance interact with the work environment and the systems around it to affect safety — it is not the study of structural design and materials, and it is not the analysis of cockpit instruments and pilot response. The primary reason for studying it is to reduce human error, which is a contributing factor in the majority of aviation accidents and incidents. It is equally important to technicians, engineers, planners and managers, because a planning or management decision can set up an error just as readily as a hands-on slip. And the proportion of accidents attributable to maintenance and inspection causes is not declining: it is significant and rising.

Incidents Attributable to Human Factors/Human Error

The aviation industry has learned many of its human factors lessons through tragedy. Key maintenance-related accidents that shaped human factors awareness include:

EventHuman Factors IssuesOutcome / Lesson
Aloha Airlines 243 (1988)Inadequate inspection of ageing structure; complacency; organisational failures18 feet of fuselage roof separated. Led to ageing aircraft safety programmes.
British Airways BAC 1-11 (1990)Wrong bolts used in windscreen replacement (84 of 90 too small); normalisation of deviance; time pressure; poor stores proceduresWindscreen blew out at 17,300 ft. Captain partially sucked out. Survived.
Continental Express 2574 (1991)Shift handover failure — 47 screws left out of horizontal stabilizer leading edge after de-icing boot removal. Incomplete task not communicated.In-flight breakup. 14 fatalities. Led to improved handover procedures.
Airbus A320 G-KMAM (1993)Spoilers left in maintenance mode after a flap change; collars and flags not fitted; spoiler operation not fully understood; engineers more familiar with another type; inadequate shift handover briefingUndemanded roll to the right after take-off from Gatwick. Returned and landed safely. A locked control surface passed the standard pilot functional checks.
Boeing 737-400 G-OBMM (1995)HP rotor drive covers not refitted after overnight borescope inspections of both engines; verbal-only task handover; approved data not used; work by torchlight; required ground idle runs omittedAlmost all the oil lost from both engines in flight. Diverted to Luton and landed safely. One error repeated on both engines removed the redundancy.
China Airlines 611 (2002)Inadequate repair of tail strike damage 22 years earlier; repair did not meet SRM requirements; fatigue crack grew undetectedIn-flight breakup. 225 fatalities. Highlighted importance of repair quality.
Chalk's Ocean Airways 101 (2005)Corrosion and fatigue in wing spar not detected during inspectionsWing separated in flight. 20 fatalities. Emphasised inspection quality.

Case Study: Both Engines, One Error

The double engine oil loss deserves a section of its own. It is one of the small number of maintenance cases that examinations use repeatedly, and it repays study in detail because nothing exotic happened in it: every single element of it was an ordinary night in an ordinary hangar. The account here follows the UK Air Accidents Investigation Branch report on the incident, published as Aircraft Incident Report 3/96.

In February 1995 a Boeing 737-400, registration G-OBMM, was due to receive routine 750-hour borescope inspections on both engines during the night before a charter flight from East Midlands Airport to Lanzarote. The task began normally. The line engineer in charge of the line night shift picked up the paperwork, went to the aircraft in a hangar some distance from the main line maintenance area, and started preparing the left engine. Preparation includes removing the high pressure rotor drive cover from the accessory gearbox, because that is where the core has to be turned from during the inspection. He removed it, and it hung on its lanyard.

He then needed the borescope equipment, which was in a tool store in another building, and a second person, because the HP spool has to be turned by somebody while the inspector looks through the probe. He walked across to base maintenance to ask for both. What happened next is the pivot of the whole event, and it was not a mistake by anybody at the time.

The base maintenance controller on that night shift had returned from a week's leave the previous evening. His shift should have had a foreman, four licensed inspectors and three licensed technicians; the foreman and three of the four inspectors were absent, leaving himself and one inspector to supervise the work. He was also aware that line maintenance had four people out of a nominal six and eight aircraft to deal with. And he had a personal interest in the task: company rules required him to perform at least two 750-hour borescope inspections in a twelve-month period to keep his company authorisation, opportunities were scarce, and he had recognised that he was in danger of letting it lapse. So he offered to take over the borescope inspections personally if the line engineer would take over moving another aircraft that was in the way. The offer was accepted and the transfer of the task was logged at 2200 hrs.

The handover that followed was verbal. The line engineer briefed the controller on where he had reached: the aircraft was dead, with no electrical or hydraulic power and nobody else working on it, the left engine was fully prepared, and the paperwork pack was with the aircraft. There was no written statement of the stage the job had reached and no annotated work stage sheet — and, as the report records, there was no suitable place on the line maintenance paperwork to note it. Procedurally, additional paperwork should have been raised. It was not, and the receiving engineer was content with what he had been told.

From there the conditions accumulated. The controller selected a fitter to assist him who was not licensed, though he had three years in the company and some previous experience of assisting on this inspection. The storeman was also absent, so the controller used his own key to draw the borescope equipment and he and the fitter checked it themselves — a check that would normally have been done by the storeman, and therefore by somebody who was not about to use the kit. The paperwork was line maintenance paperwork, unfamiliar to a base engineer and carrying no manufacturer's task cards; the controller considered it unnecessary to draw any additional reference material and worked from his own personal, annotated copy of the borescope training manual. The hangar was partly occupied by workshops and was not big enough to take the aircraft, so it was pulled in far enough for the doors to be partially shut against the fuselage sides just ahead of the tailplane. It had no sockets compatible with the engineers' lead lights, which reduced them to working by torchlight; the report notes that it was particularly dark under the engine cowls in the area of the accessory gearbox, and that against the lighting guidance it quoted, the hangar's level was graded as suitable only for infrequently used areas.

The controller pointed out the removed HP rotor drive cover on the left engine to the fitter and told him to prepare the right engine in the same way. Both inspections were carried out. Neither cover was refitted. The post-inspection ground idle engine runs required by the aircraft maintenance manual were not carried out either. The job was signed off in the aircraft technical log as complete.

The following morning the dispatching engineer reviewed the technical log before his pre-departure checks and saw two borescope inspections signed off as satisfactory. There was no outstanding item to query and nothing visible from outside a closed cowl. The one anomaly the crew did find was that the hydraulic power circuit breakers had been left open; the first officer asked the dispatching engineer about them, the engineer could see no reason for them to be open, and the first officer closed them. An unexplained departure from the expected state after an overnight maintenance visit is information about that visit. Here it was tidied away rather than accounted for, which is what usually happens to a small anomaly when the paperwork says everything is complete. The aircraft began its take-off roll at 1157 hrs. Climbing through about flight level 140 the commander noticed both oil quantity gauges reading around 15% and falling; the oil pressures then began to fall as well. The crew declared a Mayday, diverted, and landed at Luton, where both engines were shut down during the landing roll. Almost all the oil had been lost from both engines. Nobody was injured and the aircraft was undamaged.

Two details from the investigation are worth holding on to. The first is what the engine manufacturer's subsequent test showed: with the cover missing, motoring the engine over on the starter with the ignition off produced only a very fine oil mist from the open port in the gearbox casing, but starting the engine and running it at ground idle — exactly what the maintenance manual required after the inspection — produced a significant outflow of oil. The omitted check was not a formality. It was the action on that task that would have made the fault announce itself without anybody having to go looking for it, and it is specified after a borescope inspection for exactly that reason.

The second is how far up the report's causal factors reach. They record the aircraft being presented for service with the inspections signed off although the covers had not been refitted; failure to comply with the maintenance manual in a number of areas, most importantly the covers and the ground runs; that the operator's quality assurance department had not identified the non-procedural conduct of borescope inspections that had been prevalent among company engineers over a significant period; and that the Civil Aviation Authority, reviewing the company's procedures for its JAR-145 approval — the framework that preceded today's Part-145 — had detected limitations in aspects of the quality assurance system, including procedural monitoring, but had been satisfied they were being addressed and had not withheld approval. Fifteen safety recommendations were made. Two people got one job wrong on one night; the causal factors reach past both of them to the company's routine practice, to its quality department and to its regulator.

Why "coincidence" is not an available answer

Two engines on a twin are deliberately independent. They have their own oil systems, their own accessory gearboxes, their own indications, and the whole case for despatching a twin over water rests on the assumption that whatever fails on one of them has no bearing on the other.

Suppose an oil system failure of any kind has some small probability, call it p, on a given flight. If the two engines really are independent, the probability of both failing on the same flight is p multiplied by p. For an event as rare as total oil loss, p is already very small; p multiplied by itself is smaller again by that same factor. The result is a number so small that you would not expect the double event even once in the service life of a whole fleet.

So when both engines do lose their oil on the same flight, the observation is not evidence of very bad luck. It is evidence that the independence assumption was false — that something acted on both engines. The only thing that had acted on both engines was the maintenance. That is the reasoning behind the examination answer: the event is a direct result of human error, not a statistically expected coincidence, not something within accepted probability limits, and not loose bolts or a design weakness. The moment one person performs the same task on two supposedly independent systems, the redundancy that justified the design has quietly been cancelled.

The Rule That Came Out of Cases Like This

Point 145.A.48 of Part-145, on the performance of maintenance, is where the industry's answer to that reasoning now sits, and reading it against the case above shows how directly accident findings turn into rule text.

The procedures an organisation writes must be aimed at minimising multiple errors and preventing omissions. The Acceptable Means of Compliance spells out what that means: every maintenance task is signed off only after it has been completed; where tasks are grouped for the purpose of sign-off, the grouping must still allow critical steps to be identified; and work performed by personnel under supervision, such as temporary staff and trainees, is checked and signed off by an authorised person.

Then comes the provision that addresses the double oil loss directly. Procedures must minimise the possibility of an error being repeated in identical tasks and so compromising more than one system or function. In practice they should ensure that no one person is required to perform a maintenance task involving the removal and installation, or the assembly and disassembly, of several components of the same type fitted to more than one system on the same aircraft during a particular maintenance check, where a failure of those components could affect safety. Where unforeseen circumstances leave only one person available, the organisation may use a reinspection instead. Read that sentence with two engines and one borescope inspector in mind and its origin is obvious.

Alongside it sits the concept of the critical maintenance task and the error-capturing method. An error-capturing method is defined as an action the organisation has decided upon to detect maintenance errors made while performing maintenance, and the AMC is explicit that it must be adequate for the work and for how much the system was disturbed — a combination of actions such as a visual inspection, an operational check, a functional test and a rigging check may be needed for one task. The ground idle run in the case above was precisely such a method, and it was the last opportunity that job had to detect the omission. Dropping a check of that kind is a different and much larger decision than skipping an ordinary step, because it does not merely leave one thing undone: it removes the mechanism by which everything else would have been verified.

The AMC also lists where an organisation should look to decide which of its tasks are critical: information from the design approval holder, accident reports, the investigation and follow-up of incidents, occurrence reporting, flight data analysis, the results of audits and independent inspections, monitoring schemes, and feedback from training. That list is the formal mechanism by which the events in the table above stop being history and start being procedure.

What the Cases Have in Common

CAP 715 examines three of the UK events together — the 1990 windscreen accident, the 1993 spoiler incident and the 1995 oil loss — and draws out the features they share. Before the list, one finding that frames all of it.

These were not bad engineers. In all three of those UK cases, the engineers involved were considered by their own companies to be well qualified, competent and reliable employees, and in every one of them the hands-on work was being done by a supervisor. In the 1988 Aloha inspection failure the two inspectors who found no cracks had 22 and 33 years of experience respectively, and post-accident analysis established that the aircraft carried over 240 cracks in its skin at the time they looked at it. If your working model of a maintenance accident is that somebody incompetent was involved, you have no model at all, and you will not see the next one coming.

The features common to all three UK cases were these:

  • there were staff shortages;
  • time pressures existed;
  • all the errors occurred at night;
  • shift or task handovers were involved;
  • they all involved supervisors doing long hands-on tasks;
  • there was an element of a "can-do" attitude;
  • interruptions occurred;
  • there was some failure to use approved data or company procedures;
  • manuals were confusing;
  • there was inadequate pre-planning, equipment or spares.

The most useful thing about that list is what kind of thing it is. None of these ten items is a cause in the sense of being the thing that made a particular event happen; each is a precondition, a state of affairs that raises the probability of error across every task carried out in it. That difference is what makes the list operationally valuable. A cause can only be identified after the event. A precondition is visible beforehand — on a roster, in a workload plan, on a walk round the hangar at two in the morning — and can be counted before anything has gone wrong. The ten group naturally into four families.

Resourcing. Staff shortages, time pressure, and inadequate pre-planning, equipment or spares are the same problem seen from three angles: the work to be done exceeds the capacity available to do it properly. The important point is where the shortfall is absorbed. It is not absorbed by the aircraft leaving late; it is absorbed by the individual, in the form of shortcuts, deferred checks and tasks accepted without the right data. Every one of the three UK cases has a resourcing shortfall standing immediately behind the error, and in the oil loss case the shortfall existed on both shifts at once.

Time of day. All the errors occurred at night, and that is not an artefact of when maintenance is scheduled. Alertness, vigilance and the ability to sustain concentration all fall through the small hours, and they fall furthest at the end of a night shift — which is exactly when a task started at the beginning of the shift reaches its final reassembly and inspection. The night shift concentrates the most error-prone part of the work into the least capable part of the human day. The physiology behind this is developed later in the module; here it is enough to note that it is present in every one of the three cases.

Information transfer. Handovers, interruptions, confusing manuals and failure to use approved data are all failures of the same commodity: knowing the true state of the job. A handover is the point at which the state of a task has to be moved from one person's head into another's, and it is the only routine moment in maintenance when that transfer is guaranteed to be necessary. An interruption does the same thing to one person across time — the state has to survive the gap, and human working memory is poor at that. A manual that is confusing, or approved data that is not in front of you, means the correct state of the job is not available even in principle. In the oil loss case a handover, repeated interruptions and the absence of the manufacturer's task cards were all present on the same task.

Role. Supervisors doing long hands-on tasks, and the "can-do" attitude, are the subtlest pair. A supervisor absorbed in a two-hour task is not supervising during those two hours, but the organisation's chart still shows a supervisor on shift, so the layer is believed to be there when it is not. Worse, when the senior person does the job, the check that a supervisor normally provides is being applied by the individual who most needs it. That is exactly what happened in the 1990 windscreen accident and in the 1995 oil loss, in both of which the senior engineer present carried out the work personally; in the windscreen case he also signed it off himself. The "can-do" attitude is not a vice — it is what keeps a fleet flying and it is why the industry works at all. It becomes dangerous at one specific moment: when it is the reason a task is accepted that the available resources, data or conditions do not actually support.

From Error to Accident: the Chain and the Iceberg

Two models explain why these events are both preventable and rare, and they answer different questions.

The first is the error chain, illustrated in CAP 715 with a Boeing diagram in which links contributed by management, by maintenance and by the crew run in sequence to an accident. The claim is simple and it is the most encouraging thing in the subject: break any one link and the accident does not happen. It does not matter which one. Applied to the oil loss, the chain had at least four links that could have been broken independently by different people on different days — raising the additional paperwork at the handover so the state of the job was recorded; drawing the manufacturer's task cards rather than working from a personal copy; carrying out the ground idle run the manual specified; or a quality audit at any point over the preceding period noticing that borescope inspections were routinely not being done to procedure. Any single one of those four, on its own, ends the story before the aircraft flies.

The second is the iceberg. CAP 715 draws accidents as the visible tip, with serious incidents, incidents, errors and minor events in descending layers below the waterline. Most errors made by maintenance engineers never produce an accident, and that is the source of the most dangerous inference in the industry: because nothing happened, nothing was wrong. The correct inference is different. On that occasion, either a defence caught the error or the circumstances happened to be forgiving — and neither of those is guaranteed to repeat. The same error in different circumstances is the accident. The submerged part of the iceberg is not a record of harmless events; it is a preview.

That is the argument for occurrence reporting, and it is a data argument rather than a moral one. Accidents are too rare to learn from quickly; incidents and errors are common enough to reveal a trend while there is still time to act on it, and they are the only large-volume evidence anyone has about what is happening below the waterline. CAP 715 records that all maintenance incidents in the UK had to be reported to the CAA's Mandatory Occurrence Reporting Scheme so that the data could be used to disclose trends and drive action. Across the EU the equivalent obligation now sits in Regulation (EU) 376/2014 on the reporting, analysis and follow-up of occurrences, which also requires that reports are handled in a way that does not penalise the reporter for honest error. That last provision is not sentiment. A reporting scheme that people are afraid of collects nothing, and an organisation with no reports is not an organisation with no errors — it is an organisation that has lost sight of its own iceberg.

Aviation Context: EASA made Human Factors training mandatory under Part-145. Initial and recurrent training (every 2 years) is required for certifying and support staff, and for other personnel in proportion to their role. This reflects the understanding that technical knowledge alone is insufficient — engineers must also understand the human system they are part of.

What that training must contain is not left to taste. The syllabus for initial human factors training is published as Guidance Material to the same point of Part-145, and its opening topic is the need to address human factors — the subject of this note. Continuation training is meant to be evidence-driven rather than repetitive: the Acceptable Means of Compliance ties its duration and content to the organisation's own quality audit findings and to internal and external information about human error in maintenance, and expects feedback collected during the training to be passed formally to the quality department so that action can be initiated. That closes the loop opened by the accidents above. Occurrences produce reports, reports identify which tasks are critical and which conditions keep recurring, and both feed the training that the next shift receives. An organisation whose human factors training says the same thing every cycle regardless of what its own occurrence reports say is meeting the letter of the requirement and losing all of its value.

Murphy's Law

"Anything that can go wrong, will go wrong."

Attributed to Captain Edward A. Murphy Jr., a US Air Force engineer, in 1949 during rocket-sled deceleration tests at Muroc Army Air Field (now Edwards Air Force Base). A technician had wired all 16 strain gauges backwards. Murphy's observation was about designing systems to prevent human error — not about pessimism.

What the Law Claims, and What It Does Not

Taken as a statement about a single event, Murphy's Law is obviously false: most tasks go right. It is not a statement about a single event. It is a statement about exposure — about what happens to a small per-task probability when the task is repeated across a fleet, a network and a service life. The word doing the work in it is not "wrong" but "will", and the honest expansion is: given enough opportunities, every available failure path is eventually taken.

What "eventually" is worth in numbers

Suppose a particular reassembly step can be got wrong, and suppose a competent, rested engineer gets it wrong once in every thousand attempts — a rate most people would call excellent. Across a fleet the task comes round 5,000 times a year.

The expected number of wrong reassemblies is 5,000 multiplied by 0.001, which is 5 per year. The chance of a whole year passing with none at all is under one per cent. In other words, at this rate the error is not a risk. It is a schedule.

Now suppose the organisation responds with a briefing, and suppose the briefing is unusually effective and halves the error rate to one in two thousand. The expected number falls to 2.5 a year. That is a real improvement, and it is nowhere near enough, because the failure path is still open and the fleet still flies 5,000 of these tasks.

Now suppose instead that the component is modified so that it physically cannot be assembled the wrong way. The expected number is zero, and it stays zero on night shifts, in the rain, at the end of a twelve-hour duty, and for the engineer who joined after the briefing.

The numbers are an illustration rather than data from any particular fleet, but the shape of the result is completely general, and it is the whole argument for designing errors out rather than exhorting people not to make them.

Two misreadings are worth naming because both appear as distractors. The first is fatalism: if everything that can go wrong will, why bother? That inverts the law's meaning. Used fatalistically it is an excuse offered after an event; used as engineering it is an instruction issued before one, and the instruction is specific — find the failure paths and close them. The second is blame. Murphy's Law says nothing about whose fault anything is. It is not the proposition that errors are always due to neglect, and it is not the proposition that blame rests with the last person to touch the aircraft. Both of those are statements about attribution; the law is a statement about probability and design.

Complacency: the Attitude That Keeps Proving It

CAP 715 introduces the law by way of an observation about people rather than about probability: there is a tendency among human beings towards complacency, and the belief that an accident will never happen to me or to my company is a major obstacle to getting anyone to take human factors seriously, recognise risks and act on them rather than paying lip service. That is why the examination answer to what perpetuates Murphy's Law is complacency, and it is worth understanding why complacency specifically, rather than the two answers that sound equally plausible.

Inattention is a momentary state. Attention wanders, it comes back, and the person often knows afterwards that it wandered. It is exactly the failure mode that ordinary defences are built for: a checklist that has to be signed line by line, an independent inspection, a functional check. Inattention produces errors, and the system is designed on the assumption that it will.

Carelessness is not a mechanism at all. It is a judgement about how much somebody cared, made after the fact by somebody who knows the outcome. As an explanation it is unfalsifiable and it predicts nothing: it cannot tell you which task is at risk tomorrow, and it points at no defence you could install. Investigators avoid it for that reason.

Complacency is different in kind from both. It is a settled expectation built by success. Every previous occasion on which the task went well strengthens the belief that this occasion will too, and that belief acts not on the hands but on the defences. A complacent engineer does not merely make a slip; he stops looking for the slip. The inspection becomes a formality, the check becomes a signature, the manual stays in the office because "I have done this a hundred times". Inattention defeats the person and leaves the barriers standing. Complacency dismantles the barriers first.

It also has a mechanism that makes it self-reinforcing, and the direction of that mechanism is worth stating carefully. As a task is repeated without incident, the perceived risk falls. The actual risk does not change at all, because nothing about the hazard has been touched. The gap between the two therefore widens with every uneventful repetition, and it widens fastest for exactly the people with the most experience of the task. That is not a paradox; it is what happens when the only evidence available is the absence of a bad outcome, and most errors do not produce bad outcomes.

The double oil loss is the worked case. The report found that the non-procedural conduct of borescope inspections had been prevalent among the company's engineers over a significant period. Every one of those earlier inspections was completed and every aircraft came back. Read as evidence, that record said the shortcut was safe; in reality nothing about it had ever been safe, and the accumulated absence of consequences was simply a run of luck being mistaken for a result. That is also how a shortcut becomes a norm and eventually becomes the way the job is done — the process the accident table above calls normalisation of deviance in the 1990 windscreen case. Complacency is one of the twelve error precursors treated elsewhere in this module; what matters here is its specific relationship to Murphy's Law. The law says that an open failure path will eventually be taken. Complacency is the state of mind in which nobody closes it, because nothing has gone wrong yet.

One corollary follows immediately, and CAP 715 states it plainly: it is not true that accidents only happen to people who are irresponsible or sloppy, or to organisations that were already known to be poor. Errors are made by experienced, well-respected individuals, and accidents occur in organisations previously thought to be safe. An examination option claiming that errors only happen to untrained or uncertified personnel is not a near miss; it is the exact belief the law exists to attack.

Application in Aviation Maintenance

  • Design for error prevention — if a connector can be plugged in backwards, someone eventually will plug it in backwards. Use keyed connectors, poka-yoke (mistake-proofing), and asymmetric designs.
  • Assume errors will happen — design systems with multiple layers of defence (checklists, independent inspections, BITE).
  • Plan for the worst case — if a failure mode exists, it will eventually occur. Identify failure modes and mitigate them proactively.
  • Don't rely on human perfection — procedures, training, and good intentions alone do not prevent errors. System design must account for human fallibility.

How Error-Proofing Actually Works

"Design the error out" is easy to say and covers three quite different mechanisms, which are not equally reliable. The useful way to rank them is by how many human acts each one depends on, because every human act in a defence is a place where the defence can fail on a bad night. Fewer acts, stronger defence.

First, and strongest: make the wrong action impossible. Different thread forms and pitches so a line cannot be cross-connected. Connector keyways, differing shell sizes and different pin counts so a harness will not mate with the wrong socket. Asymmetric bolt patterns and mounting lugs of different sizes so a component only goes on one way round. The wrong assembly does not go together, so fatigue, time pressure, poor light, an interruption and inexperience have nothing to act on. This is the only category whose effectiveness does not depend on the state of any person, which is why it is worth a great deal of design effort and why it is the first thing a maintainability review should look for. Its limitation is that it has to be decided at design time; an engineer meeting the aircraft twenty years later cannot add it.

Second: make the wrong action announce itself under test. Functional checks, operational checks, rigging checks, leak checks and ground runs. These do not prevent the error; they cause the system to reveal it before the aircraft flies. They depend on two human acts — that the check is actually carried out, and that the check chosen is adequate for how much the system was disturbed, in the sense the AMC uses earlier in this section. With both right, the result no longer depends on anybody's vigilance, because the aircraft does the detecting. With the second wrong, a check can pass a fault it never exercised: the G-KMAM spoilers in the table above were left in maintenance mode and still passed the standard pilot functional checks. That is a much narrower dependency than it sounds, and it is why this category outranks the next one.

Third: make the wrong action visible. Lockwire, torque witness marks, streamered blanks and covers, warning tags, a removed part left hanging on a lanyard. These raise the probability that somebody notices, and they cost almost nothing, so they are everywhere and they are worth having. But they depend on two human acts rather than one: somebody must look, and the conditions must permit looking. Degrade the light, the access or the time available and a passive cue quietly stops being a defence while remaining fully in place. A duplicate or independent inspection sits at this level too, even though it adds a person rather than a cue. It is a second person looking, so it depends on the same two acts: the inspection has to be done, and that person has to see the defect in the light, the access and the time available. What it adds over a passive cue is real — the second person did not do the work, so they do not arrive carrying the first person's expectation of what it should look like — but what it does not add is any independence from human vigilance, which is why a check the aircraft itself answers is the stronger barrier wherever the task admits one.

Below all three: exhortation. Briefings, posters, "be careful", a stern word after the last event. These depend on the person remembering the message and acting on it while under exactly the load that caused the problem, and their effect decays from the moment the briefing ends. They are not worthless — awareness is a precondition for everything else — but treating them as a control is the error the whole subject exists to correct.

In plain sight, and still missed. The third category fails more often than engineers expect, and the double oil loss described in the previous section is the proof. The removed cover was not hidden in a toolbox. It was hanging from the accessory gearbox on its lanyard, on both engines, and it was explicitly pointed out and discussed during the job. It was still not refitted — because these people were working under a cowl at night with a torch in one hand, in a hangar whose lighting was rated for infrequent use, and a cue in those conditions is a cue nobody is going to look at. The defence that would have caught it sat in the second category and needed nobody to look at anything at all. The practical rule: when the environment is working against you, do not add another cue. Move the defence from noticing to testing.

There is a converse worth stating, because it is where the third category earns its place. A cue can be strengthened by making it impossible to ignore rather than merely possible to see: a blank with a long red streamer that hangs into the walkway, a cover that fouls the cowl so the cowl will not latch, a tool tag that stays in the toolbox slot until the tool comes back. Each of these takes a passive cue and gives it an active consequence, so that missing it costs the engineer something before the aircraft is released rather than after. That is the cheapest available upgrade in maintenance human factors and it needs no design authority approval, only somebody who has thought about how the job actually goes.

Murphy's Law in Engineering Terms:

If there are n ways to perform a task, and one of them leads to a catastrophic outcome, then someone, at some point, will perform it that way. The probability is not if but when. The correct response is to make the catastrophic way impossible — not to tell people to be more careful.

Murphy's Law on the Hangar Floor

The examination has a favourite scenario for this topic, and it is worth working through because the reasoning generalises. A spanner is placed on a wing surface while the engineer works. It is knocked or kicked off, falls into an open engine cowl, and breaks a sensor connector on the way down. What kind of event is that?

It is Murphy's Law. An available failure path was left open, the task was repeated enough times, and on this occasion the path was taken. It is not a team error, which would be a failure in the way two or more people coordinated their work — here nobody miscommunicated anything. It is not a lapse either: in the error taxonomies used elsewhere in this module a lapse is a failure of memory, and nobody forgot a step. Nothing was forgotten and nobody was misinformed. What happened is that the job was set up in such a way that a bad outcome was reachable, and reachable outcomes are eventually reached.

The apparent malice of it — the tool finding the one opening that mattered — has a mundane explanation, and understanding it converts the joke into a control. Where a dropped object ends up is not evenly distributed over the ramp. Aircraft surfaces are curved and sloped, so an object that starts moving keeps moving in a predictable direction; and the apertures that are open during maintenance are open precisely because people are working at them, so they are close to where objects get put down. The distribution of landing places is therefore heavily biased towards the openings, which is exactly the set of places where an object does damage. Murphy's Law, in this reading, is the working assumption that you will get the biased outcome rather than the average one — which, given the bias, is simply good estimation.

The controls follow directly, and they are all about the task set-up rather than about care:

  • Nothing rests on an aircraft surface. Not a tool, not a torch, not a cup, not a part. A trolley or a mat costs nothing and removes the failure path entirely, which puts it in the strongest category described above.
  • Openings are blanked as soon as they are opened. A cowl that is open for two hours is a target for two hours. Blanks and covers with streamers do two jobs at once: they close the path, and their absence is visible from a distance.
  • Tools are accounted for before anything closes. Shadow boards, foam cut-outs and tool tallies exist so that a missing item is discovered while the panels are still open and while the search area is still small. Discovering it afterwards means either an unserviceable aircraft or an unrecorded object loose in a structure.
  • The account is against the task, not the shift. Checking the box at the end of a twelve-hour shift tells you a tool is missing but not which of six jobs it is in. Checking at the end of each task tells you where to look.

CAP 715 offers a smaller example in the same spirit, and after two cases in which lighting was a factor it is worth taking seriously. A torch is very useful to an engineer, but Murphy's Law dictates that the batteries will run down at the moment the engineer is on the far side of the airfield from the stores. The response is not to be more careful with torches; it is to carry a spare set of batteries rather than gamble on attempting a job without enough light. The pattern is identical to the one in the accidents: identify the failure path, decide in advance what happens when it is taken, and remove the need for anybody to be lucky.

Examination points on Murphy's Law: state it as the notion that if something can go wrong, it will. It is perpetuated mainly by complacency — the settled belief that it will not happen to me or to my company — rather than by momentary inattention or by carelessness. Its constructive application in maintenance is to design systems, procedures and checks that assume errors will occur and to build in defences that catch them; it is not a licence to accept errors and merely document them without preventive action, and it is not answered by hiring only experienced staff and trusting their expertise. And where a question describes a chance chain such as a spanner falling from a wing into an open cowl and damaging a connector, the answer being tested is Murphy's Law.

Printing is not available

Please view study notes online at part66online.com

We use essential cookies to keep you signed in, plus anonymous analytics to understand how the site is used. Cookie-based analytics is set only with your consent. See our Privacy & Cookie Policy.