Asset Performance Reporting and KPIs for Rail Signalling
A wayside monitoring programme is approved on an expected outcome — fewer failures in service, fewer site attendances, less delay — and it is reviewed on whether that outcome can be shown to have happened. That review depends entirely on how the underlying measures are defined and how consistently they are applied. This guide covers the small set of performance indicators worth reporting for signalling assets, how to define a failure so the same event is counted the same way twice, how to combine condition data with maintenance records, the report set that actually changes a decision, and the data quality problems that quietly make a trend line meaningless.
What asset performance reporting is for
Reporting on signalling asset performance serves three separate audiences, and confusing them is the most common reason a reporting pack grows without becoming more useful. A maintenance supervisor needs to know which assets to attend to next week. An asset engineer needs to know which asset class or which site is degrading, and whether an intervention worked. A head of asset management needs to know whether the estate is getting better or worse, and where renewal money returns the most.
All three questions can be answered from the same underlying records, but not by the same chart. A dashboard that tries to serve all three at once usually ends up as a wall of gauges that nobody acts on. The practical structure is a small set of consistently defined measures at the top, and the ability to drill from any of them down to the individual events and the individual asset that produced the number.
The measures themselves are not novel. Reliability, availability and maintainability are defined for railway applications in EN 50126, maintenance terminology in EN 13306, and maintenance performance indicators in EN 15341. What differs between organisations is not the formula but the counting rules underneath it, and that is where reporting programmes succeed or fail.
A small set of measures, defined once
Most signalling asset teams are well served by six to eight indicators. The table below groups them by the question they answer, with the definition to pin down before the first report is issued.
| Measure | What it answers | Definition to agree first |
|---|---|---|
| Failures per asset per year | How often this asset class fails, normalised for fleet size | What counts as a failure, and whether the denominator is asset count or operating hours |
| Mean time between failures (MTBF) | Reliability of a repairable population | Whether the population is homogeneous enough for a mean to mean anything |
| Mean time to detect | How long a fault existed before anyone knew | Whether detection means the monitoring event, the alarm acknowledgement or the report from the control room |
| Mean time to repair (MTTR) | How long active repair takes once work starts | Clock start — arrival on site, or work order start, or first tool on the asset |
| Mean time to restore service | How long the operator was affected | Whether restoration means normal working or a degraded mode that let trains run |
| Availability | Proportion of required time the asset was able to perform | Inherent (repair time only) or operational (all down time), stated explicitly |
| Service-affecting failures | How many failures reached the timetable | The threshold at which a failure is deemed to have affected service |
| Delay consequence | Where failures hurt most, rather than where they are most numerous | Which attribution process the figure comes from, and its known limitations |
| Planned versus reactive ratio | Whether work is moving from unplanned response to scheduled intervention | How condition-based work is classified — as preventive, not as a fault |
| Repeat-failure rate | Where a defect is being reset rather than resolved | The window within which a second failure counts as a repeat |
| No-fault-found rate | Whether attendances are being generated without a diagnosable cause | Whether an intermittent fault confirmed later is retrospectively reclassified |
Reliability measures
Mean time between failures applies to items that are repaired and returned to service, such as a point machine or a level crossing controller. Mean time to failure applies to items that are replaced rather than repaired — a lamp, a plug-in card — and measures life to first failure. The two are frequently used interchangeably in reporting packs, which makes any comparison between asset classes unsafe.
For a wayside estate, failures per asset per year is often more useful than either. It is easier to explain, it does not imply a statistical distribution the data cannot support, and it can be broken down by asset class, region and manufacturer without becoming misleading. Reserve mean-time figures for populations large and homogeneous enough that an average describes something real.
Maintainability and response measures
Mean time to repair measures active repair once work has begun. It is a property of the design and the maintenance arrangement — access, spares, tooling, competence — rather than of the failure itself, and on a remote wayside site it is usually a small fraction of the outage the operator experienced.
The interval that matters operationally is the whole chain, and it is worth reporting the chain in segments because different interventions act on different segments:
- Detection. Fault occurs until it is known. This is the segment condition monitoring compresses most directly, and often the largest one before a monitoring programme exists.
- Notification and dispatch. Known until someone is assigned. Governed by alarm routing and escalation rather than by anything at the trackside.
- Travel and access. Assigned until on site and able to work. Governed by geography, protection arrangements and possession availability.
- Diagnosis. On site until the cause is understood. The segment that shrinks when the technician arrives already knowing what the data showed.
- Repair and test. Understood until the asset is proven and handed back.
Reported as a single number, an improvement in detection can be entirely masked by a bad month for possession availability. Reported in segments, the two are visibly separate, and the argument for the monitoring programme is made on the segment it actually influences.
Availability
Inherent availability is MTBF divided by the sum of MTBF and MTTR. It is a design comparison, and it flatters an operational estate because it counts only active repair. Operational availability substitutes mean down time — everything from detection to hand-back — and is the figure that corresponds to what the operator experienced. Both are legitimate; a report that does not say which one it used is not.
Tip: Write the definitions down in a short document that sits alongside the report itself, one paragraph per measure, and version it. When a figure moves sharply, the first question is always whether the assets changed or the counting changed — and a versioned definitions note answers it in seconds rather than in a week of reconciliation.
Counting a failure consistently
Every difficult question in asset performance reporting is a counting question. The formula for availability is not in dispute; what goes into it is.
The usual starting boundary is functional failure: the asset no longer performs its required function, or performs it outside specified limits. Around that boundary sit half a dozen decisions that have to be made deliberately, because each of them can move a headline figure by more than any real change in asset condition.
| Boundary case | The decision to record |
|---|---|
| Self-recovering events | A crossing that fails and resets itself without intervention. Counted, excluded, or counted in a separate transient-event measure — the last is usually most informative, because these are the events work orders never capture. |
| Degraded but functioning | A signal head with one lamp out of a redundant pair, or a point machine drawing high current but still throwing. Typically not a functional failure, but should appear in a degradation watchlist rather than vanishing. |
| Repeat within a window | Whether a second failure within, say, seven days is a new failure or a continuation of an unresolved one. Both conventions are defensible; only one can be in force. |
| External cause | Road-vehicle strikes, vandalism, cable theft, flooding, third-party power loss. Usually excluded from asset reliability and reported separately, because including them measures the environment rather than the asset. |
| Failures during maintenance | Faults introduced or revealed during planned work, or during commissioning of a change. Normally excluded from in-service reliability and tracked as a work-quality measure instead. |
| No fault found | An attendance that closed without a diagnosis. It is real cost and real evidence of an intermittent condition, so it belongs in the record rather than being written off. |
| Right-side failures | A failure that placed the system in its safe state — the intended behaviour, and still an availability event. Report as an availability consequence, not as a safety event. |
A structured failure classification makes these decisions durable rather than personal. Where an organisation already uses one, reuse it. ISO 14224, written for the petroleum and gas industries, is the most widely cited reference for the underlying distinction between failure mode (what was observed), failure mechanism (the physical process) and failure cause (the underlying reason). Applying that distinction consistently is what makes pooled failure data worth analysing later; recording only a free-text description is what makes it worthless.
Reporting from monitoring data and from maintenance records
Condition monitoring and the maintenance management system record different things, and a defensible report needs both.
| Question | Best source | Why |
|---|---|---|
| When did the fault begin? | Monitoring | Source-timestamped to the event rather than to when someone noticed and typed it in |
| How long did it last, and did it recover? | Monitoring | Captures duration and self-recovery, including events that never reached a work order |
| How often has it happened before? | Monitoring | Intermittent events are visible in the event history even when each one cleared |
| What was actually wrong? | Maintenance record | Only the technician on site can confirm the cause and the component |
| What did it cost, and what was consumed? | Maintenance record | Labour, spares and access costs live in the maintenance system |
| Was the detection correct? | Both, joined | Requires the detection record and the confirmed outcome against the same asset |
Joining them requires a stable shared identifier for the asset, which is the same problem that arises when integrating monitoring data with an EAM or CMMS. Without an explicit mapping between the monitored asset and the record in the asset register, the two data sets can be reported side by side but never reconciled, and any measure that spans them — false alarm rate, detection lead time, confirmed detections — cannot be produced at all.
Timing accuracy matters more than it appears to. A manually entered restoration time rounded to the nearest quarter hour introduces error comparable to the improvement most programmes are trying to demonstrate. Where sub-second ordering matters — establishing whether the supply failure preceded the asset failure or followed it — the record needs to come from sequence-of-events data with a disciplined clock, not from a reconstruction after the fact.
Attributing service impact
Failure counts describe asset health. They do not describe where that health matters. Two point machines with identical failure records, one on a lightly used siding and one on a junction that every service passes through, produce very different consequences for the timetable, and a renewal programme ranked on counts alone will pick the wrong one.
Most infrastructure managers already run a delay attribution process that assigns each delay to a prime cause and then to a responsible party and asset. Condition data does not replace that process, and reporting a competing set of delay figures is a reliable way to have the whole pack dismissed. What condition data does well is supply evidence into it: an accurate time at which the fault began, the asset it began on, and whether a related fault had occurred earlier on the same equipment. Attribution most often goes wrong on exactly those points.
Report consequence alongside frequency rather than instead of it, and make the ranking explicit: assets ordered by delay consequence, assets ordered by failure count, and the assets that appear near the top of both. That third list is usually short, and it is usually where the next intervention belongs.
The report set that changes a decision
A useful reporting pack is small and each item exists because someone does something differently as a result of it.
| Report | Audience | Decision it supports |
|---|---|---|
| Fleet health summary — failures, availability, service-affecting count, period on period | Asset management | Whether the estate is improving, and where to look next |
| Repeat offenders — assets with three or more failures in twelve months, ranked | Asset engineering | Which specific sites need investigation rather than another reset |
| Degradation watchlist — assets trending toward a threshold but not yet failed | Maintenance planning | What to schedule into the next planned window before it fails in service |
| Detection quality — detections confirmed on site, missed failures with no prior detection | Asset engineering | Whether thresholds should be tightened or relaxed, on evidence |
| Response segments — detect, dispatch, travel and access, diagnose, repair | Operations and maintenance | Which part of the response chain to attack next |
| Alarm system performance — alarm rate per operator, standing and chattering alarms, top ten by count | Control room | Which alarms need rationalising before they are ignored |
| Coverage and data quality — assets monitored, channels reporting, gaps and stale points | All | Whether the rest of the pack can be believed this period |
The alarm performance report deserves its place even though it measures the monitoring system rather than the assets. The benchmarks in EEMUA 191 and the alarm performance metrics in ISA-18.2 and IEC 62682 give it an external reference point, and an alarm system that has drifted out of those bounds will quietly undermine every other measure in the pack, because operators stop acting on the alarms that generate the events being counted. The guides to alarm flood reduction and alarm management cover that ground in detail.
Data quality determines whether the numbers are defensible
The fastest way to lose credibility with an audience of engineers is to present a trend that moved because the data changed. Several failure modes are common enough to design against directly.
- Coverage changes. Adding forty monitored crossings mid-period increases event counts with no change in asset condition. Normalise by monitored asset count, and show the denominator on the chart.
- Gaps in the record. A site offline for three weeks reports no failures. Unless outage periods are excluded from the exposure time, an unmonitored asset looks like a perfectly reliable one.
- Stuck and stale channels. A frozen input reports a plausible constant value indefinitely. Quality flags and staleness detection need to reach the report, not stop at the acquisition layer.
- Clock error. Durations computed across devices with undisciplined clocks are unreliable, and negative durations are the visible symptom of a much wider problem.
- Mapping gaps. Monitored assets with no asset-register mapping silently drop out of every report that groups by asset class or region.
- Retrospective correction. Failure classifications and delay attributions change after the fact. Reports should be reproducible — either stamped as at a date, or regenerated in full — so that two copies of the same month do not disagree without explanation.
Publishing the coverage and data quality report alongside the performance pack, rather than keeping it internal, is worth the small loss of polish. It sets the expectation that the figures have known limits, and it makes the difference between a pack that is challenged and one that is dismissed.
Common ways the numbers mislead
Beyond data quality, a handful of analytical traps recur in asset performance reporting.
- Small denominators. A class of six assets produces a failure rate that swings wildly on a single event. Report the count as well as the rate, and resist trend lines through very few points.
- Mixed populations. An MTBF computed across three generations of point machine describes none of them. Segment by type, age and duty before averaging.
- Rolling window artefacts. A twelve-month rolling figure changes when a bad month drops out of the window, which looks like an improvement and is not.
- Targets that change behaviour. A no-fault-found target encourages speculative component replacement so that something is recorded as found. Where a measure is used to judge people, expect the recording to adapt.
- Definition drift. The single largest source of unexplained step changes. Any change to a counting rule should be dated, noted on the chart, and ideally applied retrospectively so the series remains comparable.
- Averages hiding the tail. Restoration time is usually strongly skewed. A median with a ninetieth percentile describes the operator's experience far better than a mean.
Tip: Before publishing a new measure, run it against the last two years of history and look at what it would have said. If it would not have identified an event the team already remembers, or if it flags twenty assets a week, it is not ready. Backtesting a measure costs an afternoon and prevents a reporting pack that nobody trusts a quarter later.
Building the report set in stages
Reporting programmes typically fail by attempting complete coverage of every asset class and every measure before anything is issued. A narrower start produces something usable sooner and exposes the definitional arguments while they are still cheap to settle.
- Agree the failure definition and the counting rules for one asset class, and write them down. Point machines and level crossings are the usual starting points, because their failure modes are well understood and their consequences are visible.
- Establish the asset mapping between the monitoring platform and the asset register, so monitoring events and maintenance records can be joined at all.
- Publish coverage and data quality first. Issuing it before any performance figures establishes what the data can and cannot support, and usually surfaces several problems worth fixing before anyone sees a trend.
- Add three measures, not eleven — failures per asset per year, mean time to restore service reported in segments, and repeat offenders. These three carry most of the early decisions.
- Add detection quality once outcomes are flowing back from the maintenance system, and begin tuning thresholds against confirmed findings rather than opinion.
- Extend to further asset classes, reusing the definitions and the mapping rather than renegotiating them per class.
- Review the pack annually and remove anything that has not changed a decision in twelve months. A reporting set that only ever grows is one where nothing in it is trusted.
Sequenced this way, the first asset class carries the cost of settling the definitions, and each class after it is incremental. It also means the first numbers the organisation sees are ones the team can defend in detail, which is generally what determines whether the pack survives its first serious challenge.
Frequently asked questions
Which KPIs should a signalling asset team report?
A small set, defined once and left alone. Most teams are well served by six to eight measures: failures per asset per year (or mean time between failures) for reliability, mean time to repair and mean time to restore service for maintainability, availability derived from the two, service-affecting failures and their delay consequence for operational impact, the proportion of work that was planned rather than reactive, the repeat-failure rate, and the no-fault-found rate. Everything else is diagnostic detail that belongs in a drill-down rather than on the front page. A reporting set that grows every quarter is usually a sign that no measure is trusted enough to act on.
What is the difference between MTBF, MTTF and MTTR?
Mean time between failures (MTBF) applies to repairable items and measures the average operating time between one failure and the next — a point machine that is repaired and returned to service. Mean time to failure (MTTF) applies to items that are replaced rather than repaired, such as a lamp or a plug-in module, and measures average life to first failure. Mean time to repair (MTTR) measures the active repair time once work starts, and is a property of the design and the maintenance arrangement rather than of the failure. Because MTTR excludes waiting time, it is usually far shorter than the outage the operator experienced, so most rail reporting also needs a mean time to restore service that runs from detection to service resumption.
How is availability calculated for a signalling asset?
In its simplest form, inherent availability is MTBF divided by the sum of MTBF and MTTR. That figure is useful for comparing designs but flattering as an operational measure, because it counts only active repair time. Operational availability replaces MTTR with mean down time, which includes detection, notification, travel, access or possession waiting, diagnosis, repair and testing. The two can differ by an order of magnitude on a remote site, and the difference is precisely the part a monitoring programme is intended to reduce. Whichever is used, state which one it is on the report, because a number labelled only as availability invites comparison with numbers computed a different way.
What counts as a failure for reporting purposes?
Whatever the written definition says, applied identically every time. The usual boundary is loss of a required function: an asset that no longer performs its intended function, or performs it outside specified limits. Around that boundary sit decisions that must be made explicitly — whether a self-recovering event counts, whether a degraded but functioning asset counts, whether a repeat within a defined window is a new failure or a continuation of the last one, and whether third-party causes such as vandalism, road-vehicle strikes or extreme weather are counted, excluded or reported separately. None of these choices is uniquely correct. The damaging outcome is leaving them unwritten, because the definition then drifts and the trend line measures the drift rather than the assets.
Should performance reporting be based on monitoring data or on work orders?
On both, because they record different things. Monitoring data supplies accurate timing — when the condition began, how long it lasted, whether it recovered on its own — and it captures events that never generated a work order at all, such as intermittent faults and transient conditions. The maintenance record supplies what was actually found, what was replaced and what it cost, which monitoring cannot know. Reporting from monitoring alone tends to overstate event counts and cannot attribute cause; reporting from work orders alone loses every fault that cleared before anyone attended, understates event frequency and inherits the imprecision of manually entered times. Joining the two on a shared asset identifier is what produces a defensible picture.
How should delay minutes be attributed to a signalling asset?
Through the infrastructure manager's existing attribution process, with the monitoring record used as evidence rather than as a competing source of truth. Most networks attribute each delay to a prime cause and then to a responsible party and asset, and the value of accurate condition data is that it can establish when a fault actually began and which asset it began on, which is the part attribution most often gets wrong. Report delay consequence alongside failure counts rather than instead of them, because the two answer different questions: failure counts describe asset health, delay minutes describe where the same amount of asset health hurts the timetable most. Ranking renewal candidates on consequence rather than count is usually what changes an investment decision.
Why do repeat failures and no-fault-found rates matter more than headline reliability figures?
Because they identify work that can be acted on, where an aggregate cannot. A fleet mean time between failures that moves by a few percent tells nobody what to do on Monday. A ranked list of assets that failed three or more times in the last twelve months names a small number of sites where a defect is being reset rather than resolved, and those sites typically account for a disproportionate share of both attendances and delay. A high no-fault-found rate points somewhere else entirely — at intermittent faults that clear before the technician arrives, at alarm thresholds set too tightly, or at a diagnosis that cannot be made without data captured at the moment of failure. Both measures translate directly into an intervention; a headline average rarely does.
Does asset performance reporting from a monitoring platform form part of safety assurance?
It is management information drawn from a non-vital monitoring overlay, and it does not substitute for the reliability, availability, maintainability and safety demonstration the vital signalling carries under its own EN 50126 lifecycle. The two are related in one useful direction: operational failure data gathered in service is a legitimate input when reviewing whether in-service performance matches the reliability assumptions made at design time, and it is generally far better evidence than recollection. Treat the reporting as evidence that informs assurance activity, not as a replacement for it, and keep the provenance of every figure traceable so that anyone reviewing it can see which records it was derived from.
Can your monitoring data answer the questions your reporting pack asks?
RailNet Operations is being shaped around the view that a monitoring platform should produce evidence, not just alarms — event records accurate enough to compute a restoration time, asset identity stable enough to join to a maintenance record, and coverage figures honest enough to publish alongside the numbers they qualify. If you produce signalling asset performance reporting today, we would like to hear which measure caused the most argument and how you settled it. Interested in helping explore what a platform like this could do?
Start a conversation