Tec Nikan
فارسی
Talk to us
All posts

Why Operators Ignore Your Alarms

An alarm system that produces more signals than a person can act on has not made the plant safer; it has trained the control room to stop reading. What goes wrong, and the specific fixes that reverse it.

alarm managementcontrol roomsSCADAoperationsprocess safety

Walk into a control room where the alarm banner has three hundred active entries and ask the operator what they are. They will tell you which two matter and that the rest have been there for months. That is not a failure of discipline. It is the predictable outcome of a system that produces more alarms than a human can act on, and it is one of the most common defects in industrial software — common enough that there are standards written specifically about it, and rare enough for those standards to be followed that the problem persists.

The mechanism is simple. An alarm is a request for a person to do something. If the request arrives when there is nothing to do, or arrives faster than actions can be taken, or arrives about something the person cannot influence, then it is not an alarm but a message. Mixed in with real alarms, messages make the real ones harder to see, and the system's overall value drops as the number of alarms rises. Past a certain rate the value goes negative: the operator stops reading, and an alarm system nobody reads is worse than none at all, because the plant is designed on the assumption somebody is watching.

The rates that industry guidance considers workable are lower than most systems achieve. Long-standing practice puts a manageable steady-state load at roughly one alarm every ten minutes per operator, with short peaks tolerated during upsets. Systems in the field routinely run at ten or more per minute during an upset — which is exactly when the operator has least attention to spare and most need of the ones that matter.

Four specific failures produce most of the noise.

The first is alarms that are not actionable. Every alarm should have an answer to the question "what does the operator do about this, right now, that they would not otherwise do?" If the honest answer is nothing, or the same thing they were already doing, it is not an alarm. Configuring one on every point because the tool made it easy is how a system ends up with thousands of them.

The second is priority that means nothing. When ninety percent of alarms are configured high priority, priority has stopped carrying information, and the operator falls back on their own judgement about which ones matter — which is a private, undocumented, un-handover-able version of the rationalisation nobody did. A useful distribution is heavily weighted toward low: something like five percent high, fifteen percent medium, eighty percent low is the shape guidance suggests, and most systems are nowhere near it.

The third is standing alarms — the ones that have been active for weeks because a transmitter failed, a valve is in manual, or a limit was set for a mode of operation the plant has not been in since spring. Every standing alarm is permanent occupancy in the banner and a small, continuous lesson that the banner does not need reading. They accumulate because clearing one requires either fixing something or changing a setpoint, and both need somebody with authority to decide.

The fourth is chattering: an alarm on a measurement that sits near its threshold and toggles, producing dozens of entries an hour. Deadband and on-delay solve this and take five minutes to configure, and the reason it goes unfixed is almost always that nobody owns the alarm system as a system.

Which is the underlying issue. Alarms are configured project by project, by whoever built each subsystem, against no shared philosophy. The fix that actually works is neither clever nor quick. Write an alarm philosophy: what an alarm is for on this plant, what the priorities mean in terms of consequence and time to respond, and who may add or change one. Then rationalise the existing set against it — go through them, decide which are alarms, which become log entries and which are deleted. It is dull, it takes weeks, and it is the only intervention that reliably works.

Measure afterwards, and measure the boring things: alarms per operator per hour in steady state and during upsets, the top ten most frequent alarms as a fraction of the total, the count of standing alarms, and how the priority distribution actually looks. The top-ten figure is usually the most revealing number in the exercise, because a handful of chattering points routinely account for half of everything the system produces, and fixing those buys back the operator's attention faster than anything else.

None of this is exotic. It is the difference between a system that asks for attention when it needs it and one that asks constantly and is therefore never listened to.

Want to work with us?

Tell us what you're building and we'll help you scope the first deployment.