All of the interesting systems (e.g. transportation, healthcare, power generation) are inherently and unavoidably hazardous by the own nature.
所有值得关注的系统(如交通、医疗、发电)就其本性而言,都内在地、不可避免地危险。
The frequency of hazard exposure can sometimes be changed but the processes involved in the system are themselves intrinsically and irreducibly hazardous.
暴露于危险的频率有时可以改变,但系统所涉及的过程本身,是内在地、不可化约地危险的。
It is the presence of these hazards that drives the creation of defenses against hazard that characterize these systems.
正是这些危险的存在,驱动了针对危险的层层防御的建立——而这些防御恰是此类系统的标志性特征。
The high consequences of failure lead over time to the construction of multiple layers of defense against failure.
失败的高昂代价,随时间推移催生出针对失败的多层防御。
These defenses include obvious technical components (e.g. backup systems, 'safety' features of equipment) and human components (e.g. training, knowledge) but also a variety of organizational, institutional, and regulatory defenses (e.g. policies and procedures, certification, work rules, team training).
这些防御既包括显而易见的技术组件(如备份系统、设备的"安全"特性)和人的组件(如培训、知识),也包括形形色色的组织、制度与监管防御(如政策与流程、资格认证、作业规程、团队训练)。
The effect of these measures is to provide a series of shields that normally divert operations away from accidents.
这些措施的效果,是提供一系列盾牌,在通常情况下把运行引离事故。
The array of defenses works. System operations are generally successful.
防御阵列是起作用的。系统运行大体上是成功的。
Overt catastrophic failure occurs when small, apparently innocuous failures join to create opportunity for a systemic accident.
显性的灾难性失败,发生在多个看似无害的小失败汇合、为系统性事故创造出机会之时。
Each of these small failures is necessary to cause catastrophe but only the combination is sufficient to permit failure.
每个小失败对酿成灾难都是必要的,但只有它们的组合才足以放行失败。
Put another way, there are many more failure opportunities than overt system accidents.
换句话说,失败的机会远比显性的系统事故多。
Most initial failure trajectories are blocked by designed system safety components. Trajectories that reach the operational level are mostly blocked, usually by practitioners.
大多数初始失败轨迹被设计好的系统安全组件拦下;抵达运行层面的轨迹也大多被拦下——通常是被一线从业者拦下的。
The complexity of these systems makes it impossible for them to run without multiple flaws being present.
这些系统的复杂性决定了:它们不可能在不带多处缺陷的情况下运行。
Because these are individually insufficient to cause failure they are regarded as minor factors during operations.
因为这些缺陷单个都不足以引发失败,运行中它们被当作次要因素。
Eradication of all latent failures is limited primarily by economic cost but also because it is difficult before the fact to see how such failures might contribute to an accident.
根除全部潜伏失败之所以做不到,首先受限于经济成本,但也因为在事前很难看出这些失败会怎样凑成一场事故。
The failures change constantly because of changing technology, work organization, and efforts to eradicate failures.
而由于技术在变、工作组织在变、根除失败的努力本身也在变,这些潜伏失败始终处于变动之中。
A corollary to the preceding point is that complex systems run as broken systems.
上一条的推论是:复杂系统是作为"带病系统"在运行的。
The system continues to function because it contains so many redundancies and because people can make it function, despite the presence of many flaws.
系统之所以还能运转,是因为它内含大量冗余,也因为人有本事让它运转——尽管缺陷遍布。
After accident reviews nearly always note that the system has a history of prior 'proto-accidents' that nearly generated catastrophe.
事故后的复盘几乎总会指出:系统早有一串"准事故"(proto-accidents)前科,差一点就酿成灾难。
Arguments that these degraded conditions should have been recognized before the overt accident are usually predicated on naïve notions of system performance.
"这些降级状况本应在显性事故之前就被识别出来"——这类论调通常建立在对系统运行的天真想象之上。
System operations are dynamic, with components (organizational, human, technical) failing and being replaced continuously.
系统运行是动态的:组件(组织的、人的、技术的)持续地失效,又持续地被替换。
Complex systems possess potential for catastrophic failure.
复杂系统蕴藏着灾难性失败的势能。
Human practitioners are nearly always in close physical and temporal proximity to these potential failures – disaster can occur at any time and in nearly any place.
一线从业者几乎总是在空间与时间上紧贴着这些潜在失败——灾难随时可能发生,几乎无处不可发生。
The potential for catastrophic outcome is a hallmark of complex systems.
灾难性后果的可能性,是复杂系统的胎记。
It is impossible to eliminate the potential for such catastrophic failure; the potential for such failure is always present by the system's own nature.
消除这种灾难势能是不可能的;依系统本性,这种失败的可能永远在场。
Because overt failure requires multiple faults, there is no isolated 'cause' of an accident.
既然显性失败需要多重故障,事故就不存在孤立的"原因"。
There are multiple contributors to accidents. Each of these is necessarily insufficient in itself to create an accident. Only jointly are these causes sufficient to create an accident.
事故有多个贡献因素,每一个单独都必然不足以造成事故,只有合在一起才足够。
Indeed, it is the linking of these causes together that creates the circumstances required for the accident.
事实上,正是这些原因的相互勾连,才造出了事故所需的情境。
Thus, no isolation of the 'root cause' of an accident is possible.
因此,把事故的"根因"(root cause)单独剥离出来是不可能的。
The evaluations based on such reasoning as 'root cause' do not reflect a technical understanding of the nature of failure but rather the social, cultural need to blame specific, localized forces or events for outcomes.
基于"根因"式推理的评判,反映的不是对失败本质的技术理解,而是一种社会与文化需要:要为结果找到具体的、局部化的力量或事件来背锅。
Knowledge of the outcome makes it seem that events leading to the outcome should have appeared more salient to practitioners at the time than was actually the case.
知道了结局,就会觉得通向结局的那些事件当时对从业者来说"本应"更显眼——而实际远非如此。
This means that ex post facto accident analysis of human performance is inaccurate.
这意味着,事后对人的表现所做的事故分析是失准的。
The outcome knowledge poisons the ability of after-accident observers to recreate the view of practitioners before the accident of those same factors.
结局知识毒化了事后观察者的还原能力——他们再也无法重建从业者在事故之前看待同样这些因素时的视角。
It seems that practitioners "should have known" that the factors would "inevitably" lead to an accident.
于是看起来,从业者"早该知道"那些因素会"必然"导致事故。
Hindsight bias remains the primary obstacle to accident investigation, especially when expert human performance is involved.
后见之明偏差(hindsight bias)至今仍是事故调查的头号障碍,尤其当涉及专家级人类表现时。
The system practitioners operate the system in order to produce its desired product and also work to forestall accidents.
系统的从业者操作系统,既为产出想要的产品,也在努力预先阻止事故。
This dynamic quality of system operation, the balancing of demands for production against the possibility of incipient failure is unavoidable.
系统运行的这种动态性质——在生产压力与萌芽中的失败可能之间走钢丝——是不可避免的。
Outsiders rarely acknowledge the duality of this role.
局外人很少承认这一角色的双重性。
In non-accident filled times, the production role is emphasized. After accidents, the defense against failure role is emphasized.
太平时节,被强调的是生产者角色;事故之后,被强调的是防御者角色。
At either time, the outsider's view misapprehends the operator's constant, simultaneous engagement with both roles.
无论何时,局外人的视角都误解了一个事实:操作者始终同时扮演着两个角色。
After accidents, the overt failure often appears to have been inevitable and the practitioner's actions as blunders or deliberate willful disregard of certain impending failure.
事故之后,显性失败常常显得早已注定,从业者的行动则显得是昏招,或是对迫在眉睫的失败的蓄意无视。
But all practitioner actions are actually gambles, that is, acts that take place in the face of uncertain outcomes.
但从业者的一切行动实际上都是赌博——即在结果不确定的情况下做出的行为。
The degree of uncertainty may change from moment to moment.
不确定的程度可能时刻在变。
That practitioner actions are gambles appears clear after accidents; in general, post hoc analysis regards these gambles as poor ones.
"从业者的行动是赌博"这一点,在事故后看得很清楚;而事后分析通常把这些赌注评判为下得糟糕。
But the converse: that successful outcomes are also the result of gambles; is not widely appreciated.
但反过来那一面——成功的结果同样是赌博赌来的——却少有人领会。
Organizations are ambiguous, often intentionally, about the relationship between production targets, efficient use of resources, economy and costs of operations, and acceptable risks of low and high consequence accidents.
组织在生产指标、资源效率、运营经济性与成本、以及高低后果事故的可接受风险之间的关系上,总是含糊其辞——而且常常是故意含糊。
All ambiguity is resolved by actions of practitioners at the sharp end of the system.
所有这些模糊,最终都由系统"刀尖"(sharp end)上从业者的行动来了断。
After an accident, practitioner actions may be regarded as 'errors' or 'violations' but these evaluations are heavily biased by hindsight and ignore the other driving forces, especially production pressure.
事故之后,从业者的行动可能被判定为"差错"或"违规",但这些评判深受后见之明的偏倚,且无视了其他驱动力——尤其是生产压力。
Practitioners and first line management actively adapt the system to maximize production and minimize accidents.
从业者和一线管理者主动地调适系统,以求产出最大、事故最少。
These adaptations often occur on a moment by moment basis.
这些调适往往是逐时逐刻进行的。
Some of these adaptations include: (1) Restructuring the system in order to reduce exposure of vulnerable parts to failure. (2) Concentrating critical resources in areas of expected high demand. (3) Providing pathways for retreat or recovery from expected and unexpected faults. (4) Establishing means for early detection of changed system performance in order to allow graceful cutbacks in production or other means of increasing resiliency.
其中包括:(1)重组系统,减少脆弱部位对失败的暴露;(2)把关键资源集中到预期高需求的区域;(3)为预期内外的故障准备退路与恢复通道;(4)建立系统性能变化的早期侦测手段,以便从容减产或以其他方式增强韧性。
Complex systems require substantial human expertise in their operation and management.
复杂系统的运行和管理需要大量人类专长。
This expertise changes in character as technology changes but it also changes because of the need to replace experts who leave.
这种专长随技术变迁而改变性质,也因需要顶替离去的专家而不断更替。
In every case, training and refinement of skill and expertise is one part of the function of the system itself.
无论哪种情形,技能与专长的培养和精进,本身就是系统机能的一部分。
At any moment, therefore, a given complex system will contain practitioners and trainees with varying degrees of expertise.
因此在任一时刻,一个复杂系统里总是混杂着专长程度参差的从业者与受训者。
Critical issues related to expertise arise from (1) the need to use scarce expertise as a resource for the most difficult or demanding production needs and (2) the need to develop expertise for future use.
与专长相关的要害问题源自两个需要:(1)把稀缺专长用作资源,投向最困难、最苛刻的生产需求;(2)为将来储备和培养专长。
The low rate of overt accidents in reliable systems may encourage changes, especially the use of new technology, to decrease the number of low consequence but high frequency failures.
可靠系统中显性事故率低,这会鼓励变更——尤其是采用新技术——去削减那些低后果、高频率的失败。
These changes maybe actually create opportunities for new, low frequency but high consequence failures.
而这些变更实际上可能为新的、低频率但高后果的失败创造机会。
When new technologies are used to eliminate well understood system failures or to gain high precision performance they often introduce new pathways to large scale, catastrophic failures.
当新技术被用来消灭那些已被充分理解的系统失败、或去追求高精度性能时,它们常常开辟出通往大规模灾难性失败的新通道。
Not uncommonly, these new, rare catastrophes have even greater impact than those eliminated by the new technology.
屡见不鲜的是:这些新的、罕见的灾难,冲击力甚至大于新技术所消灭的那些旧失败。
These new forms of failure are difficult to see before the fact; attention is paid mostly to the putative beneficial characteristics of the changes.
这些新形态的失败在事前很难看见;注意力大多放在变更那些据称有益的特性上。
Because these new, high consequence accidents occur at a low rate, multiple system changes may occur before an accident, making it hard to see the contribution of technology to the failure.
又因为这类高后果事故发生率低,一场事故之前系统可能已历经多次变更,技术对失败的贡献也就更难辨认。
Post-accident remedies for "human error" are usually predicated on obstructing activities that can "cause" accidents.
针对"人为差错"的事后补救,通常建立在阻断那些会"引发"事故的活动之上。
These end-of-the-chain measures do little to reduce the likelihood of further accidents.
这些"链条末端"的措施,对降低后续事故的可能性几乎无益。
In fact that likelihood of an identical accident is already extraordinarily low because the pattern of latent failures changes constantly.
事实上,一模一样的事故重演的可能性本来就极低——因为潜伏失败的组合格局在不停变化。
Instead of increasing safety, post-accident remedies usually increase the coupling and complexity of the system.
事后补救非但没有提升安全,通常反而增加了系统的耦合度与复杂度。
This increases the potential number of latent failures and also makes the detection and blocking of accident trajectories more difficult.
这既增加了潜伏失败的潜在数量,也让事故轨迹的侦测与拦截变得更难。
Safety is an emergent property of systems; it does not reside in a person, device or department of an organization or system.
安全是系统的涌现属性(emergent property);它不寄居于某个人、某台设备,或组织与系统的某个部门之中。
Safety cannot be purchased or manufactured; it is not a feature that is separate from the other components of the system.
安全买不来,也造不出;它不是一个可以从系统其余组件中剥离出来的特性。
This means that safety cannot be manipulated like a feedstock or raw material.
这意味着安全没法像原料或库存那样被摆弄。
The state of safety in any system is always dynamic; continuous systemic change insures that hazard and its management are constantly changing.
任何系统的安全状态永远是动态的;系统的持续变化保证了危险及其治理也在不停变化。
Failure free operations are the result of activities of people who work to keep the system within the boundaries of tolerable performance.
无失败的运行,是人们努力把系统保持在可容忍性能边界之内的活动的成果。
These activities are, for the most part, part of normal operations and superficially straightforward.
这些活动绝大部分属于日常操作,表面看平平无奇。
But because system operations are never trouble free, human practitioner adaptations to changing conditions actually create safety from moment to moment.
但因为系统运行从来不会无事,从业者对变化情境的适应,实际上是在一刻接一刻地创造安全。
These adaptations often amount to just the selection of a well-rehearsed routine from a store of available responses; sometimes, however, the adaptations are novel combinations or de novo creations of new approaches.
这些适应常常只是从既有应对库里挑出一套演练纯熟的例行动作;但有时,它们是新颖的组合,甚至是全新方法的从零创造。
Recognizing hazard and successfully manipulating system operations to remain inside the tolerable performance boundaries requires intimate contact with failure.
识别危险、并成功地驾驭系统运行使之留在可容忍性能边界之内,需要与失败有亲密接触。
More robust system performance is likely to arise in systems where operators can discern the "edge of the envelope".
在操作者能辨认出"性能包线边缘"(edge of the envelope)的系统里,更强健的系统表现更可能出现。
This is where system performance begins to deteriorate, becomes difficult to predict, or cannot be readily recovered.
所谓边缘,就是系统性能开始劣化、变得难以预测、或难以轻易挽回的地方。
In intrinsically hazardous systems, operators are expected to encounter and appreciate hazards in ways that lead to overall performance that is desirable.
在天生危险的系统里,操作者被期望以某种方式遭遇并领会危险,从而带来整体上令人满意的表现。
Improved safety depends on providing operators with calibrated views of the hazards.
安全的改进,依赖于给操作者提供对危险的经过校准的认知。
It also depends on providing calibration about how their actions move system performance towards or away from the edge of the envelope.
也依赖于让他们校准另一件事:自己的行动正把系统性能推向边缘,还是拉离边缘。