Cloud vs On-Premise Deployment for Rail Monitoring
Where the monitoring platform runs is one of the first questions asked in a wayside monitoring procurement and one of the least well specified. It is usually framed as cloud or on-premise, as though those were the only two options and as though the choice were mainly about cost. In practice there are four deployment models, the parts of the system that matter most stay trackside in all of them, and the decision turns on network placement, data residency, connectivity dependence and exit cost far more than on the price of a server. This guide sets out the models, the criteria that separate them, and the requirements worth holding regardless of which one is chosen.
What is actually being decided
A wayside monitoring system is not one thing in one place. It is a chain: sensors and inputs in a location case, a site data logger or RTU that acquires and time-stamps them, a communications link, a central tier that stores history and evaluates alarms, and a presentation tier where operators and maintainers look at the result. Only the last two are genuinely in scope for the hosting decision. The acquisition tier stays at the trackside in every model, because that is where the equipment is.
Framing the question that way removes a good deal of the heat from it. The choice is not whether the railway is "in the cloud". It is where the storage, analysis and presentation tier of a non-vital overlay is operated, and by whom. That is a narrower and more answerable question, and it is the one procurement should be asking.
The four deployment models
Cloud and on-premise are the endpoints of a range, not a binary. Four models cover most of what is offered in this market, and they differ along two axes: who owns the infrastructure, and who operates the software running on it.
| Model | Infrastructure | Operated by | Suited to |
|---|---|---|---|
| On-premise | Operator-owned servers in an operator data centre or control centre | The operator's own IT/OT team | Strict residency constraints, existing data centre capability, large stable estate |
| Private cloud | Operator's own virtualised or co-located infrastructure | The operator, with pooled rather than dedicated hardware | Operators consolidating many systems onto shared internal infrastructure |
| Hosted (single-tenant managed) | Supplier-provided infrastructure, dedicated instance per customer | The supplier, under contract | Operators wanting supplier-run infrastructure without shared tenancy |
| SaaS (multi-tenant) | Shared supplier platform serving many customers | The supplier, on a common release cadence | Smaller estates, fast start, minimal internal platform capability |
The distinctions that matter in practice are not the labels but four consequences of them: who controls upgrade timing, whether the instance is isolated from other customers, which jurisdiction the data sits in, and what happens at the end of the contract. A hosted single-tenant deployment and a multi-tenant SaaS offering are both "cloud" in a sales conversation and materially different on all four.
What stays at the wayside regardless
Three functions belong at the trackside in every model, and a design that moves them is worse rather than more modern.
- Acquisition and time-stamping. Every reading and state change should carry an accurate time applied at the point of capture, not on arrival at the central tier. Ordering that survives a variable network is the whole basis of sequence-of-events analysis.
- Local evaluation. Fault detection that must be prompt, and any processing that reduces data volume before transmission, belongs on the site device. Processing at the edge is what keeps a hosted deployment viable on a constrained link.
- Store-and-forward buffering. Each site holds its own timestamped buffer so an outage of the link or the central tier delays the record rather than losing it. It back-fills in order on recovery.
With those three in place, the hosting decision stops being a question about data loss and becomes a question about live visibility. That is a real cost, and one worth quantifying, but it is a different and smaller risk than a gap in the historical record.
Network placement and the boundary out
The security architecture is where a hosted deployment is either sound or quietly broken, and the relevant frameworks are well established. IEC 62443 groups assets with similar security requirements into zones connected by defined conduits; CLC/TS 50701 derives its security models and risk-assessment process from that series for railway applications. The layered reference most practitioners picture — field devices at the bottom, site operations in the middle, enterprise systems at the top, with a demilitarised zone between the operational and enterprise layers — is the Purdue model, and it is still the common language for describing where something sits.
In that language, a hosted monitoring tier is simply another zone, at a lower trust level than the operational network, and the link to it is a conduit that has to be specified rather than assumed. A sound design has three properties:
- A single, named egress path. Data leaves the operational network through one controlled boundary — typically a broker or gateway in the demilitarised zone — not through whatever route each device happens to open.
- Outbound-only initiation. The site or the boundary opens the connection outward; the hosted tier never initiates a session inward. There is no inbound path from the hosted zone to the operational one.
- Enforcement proportionate to consequence. A filtering conduit — an authenticated, encrypted, firewalled path — covers most cases. Where the consequence of any inbound reachability is unacceptable, a unidirectional gateway or data diode makes the one-way property physical rather than configured, as discussed in the signalling cybersecurity guide.
The failure mode to watch for is the opposite pattern: an edge device with a mobile network connection that dials a supplier's cloud service directly, bypassing the site and enterprise layers entirely. It is quick to deploy and it puts an undocumented route out of the operational network into service, outside the zone model and usually outside anyone's asset register. If telemetry is leaving, it should leave through a conduit somebody has drawn.
Tip: Write the data flow down before comparing suppliers. One diagram showing every point at which monitoring data crosses a trust boundary — site to site network, site network to operational core, operational core to enterprise, enterprise to hosted tier — with the direction of session initiation marked on each arrow, will settle more architecture arguments than any amount of discussion about cloud in the abstract. It also makes the difference between two suppliers' offerings immediately visible.
Data residency, sovereignty and regulation
Railway infrastructure managers are generally designated as critical or essential infrastructure. Under the EU NIS2 Directive, rail infrastructure managers and operators fall in scope as essential entities, which carries the strictest tier of risk management, supply-chain and incident-reporting obligations — and those obligations extend to the cloud services and third-party suppliers that hold operational data. Comparable regimes apply in other jurisdictions, and national rules frequently constrain where infrastructure data may be held and who may access it.
Two terms get used interchangeably and should not be. Residency is the physical question: which country the data sits in. Sovereignty is the broader one: whose law applies to it, and which administrators — including the supplier's support staff, in whichever country they are based — can technically reach it. A region-pinned deployment satisfies residency and does not by itself satisfy sovereignty.
The controls that answer these are contractual as much as technical, and they belong in the procurement, not in a later architecture review:
- Named region or country of storage, and of any backup or disaster-recovery copy.
- Customer-managed encryption keys, where the operator can revoke access unilaterally.
- Stated limits on supplier support access, with logging of every access to operator data.
- A current list of sub-processors, and notice before it changes.
- Defined data handling at termination — export format, timescale, and proof of deletion.
Under a shared responsibility model, the infrastructure provider secures the platform and the operator remains accountable for the data on it. That accountability does not transfer with the hosting, which is why these questions are the operator's to ask rather than the supplier's to volunteer.
Connectivity, availability and latency
A hosted deployment adds a dependency: the operator's connection out. It is worth being precise about what that dependency actually costs, because it is often overstated in one direction and ignored in the other.
With store-and-forward at every site, an outage does not lose data. What it loses is live visibility and notification for its duration — nobody sees a developing fault, and no alarm reaches an on-call maintainer. For a control room that acts on wayside alarms in real time, that is a significant operational exposure and the egress path deserves the same redundancy analysis as any other single point of failure: diverse links, ideally diverse bearers, and monitoring of the path itself. The high availability guide covers how that analysis is structured.
Latency matters less than it is usually said to. A monitoring overlay observes; it does not close a control loop, so tens or hundreds of milliseconds of additional transport delay changes nothing about the quality of the record. Anything that genuinely must respond quickly — threshold detection on a point machine drive-current signature, for example — should be executing on the site device in any deployment model, precisely so that it is not sensitive to where the central tier lives.
| Concern | On-premise | Hosted or SaaS |
|---|---|---|
| Data loss during a network outage | Site buffering still required for the site-to-centre link | Site buffering still required; equivalent exposure |
| Live visibility during an outage | Depends on the internal network only | Also depends on the egress path; needs diverse links |
| Latency to the central tier | Lower, rarely materially so for an observing system | Higher, immaterial where prompt logic runs at the edge |
| Scaling the central tier | Capacity planned and purchased ahead of need | Elastic; capacity is a commercial rather than lead-time question |
| Patching and platform upgrade | Operator's effort, operator's schedule | Supplier's effort, supplier's schedule — confirm the notice period |
| Access from the field | Requires remote access into the operational network | Available to a maintainer with a browser and credentials |
Cost and lifecycle over a long-lived estate
The cost comparison is routinely made over the wrong period and with the wrong items included. A hosted model is almost always cheaper to start: no capital purchase, no data centre footprint, no platform administration to staff. On-premise front-loads capital and then carries a largely fixed recurring cost in refresh and staff.
An honest comparison holds both sides to the same period and counts what is usually left out. On the on-premise side: hardware refresh at five to seven years, operating-system and database licensing, patching effort, backup and disaster recovery, and the specialist staff time to run it. On the hosted side: subscription escalation across the term, data egress and integration charges, the effort of any integration the supplier does not provide, and the cost of leaving. Rail assets are specified for decades and monitoring systems are expected to last a long time alongside them, so a three-year total cost view flatters the subscription model and a ten-year view is closer to the operator's actual exposure.
Lifecycle mismatch is the underrated risk. A trackside asset may be in service for twenty-five years or more; a hosted software service is contracted in three-year increments and its supplier may change ownership, pricing or direction well inside the life of the equipment it monitors. That is not an argument against hosting. It is an argument for making the deployment decision reversible, which is a design and contractual property rather than a hosting one — the same property that keeps a supplier relationship from becoming a dependency.
A decision framework
Rather than starting from a preference, work through the criteria that actually discriminate between the models and see where the weight falls. Most operators find some criteria are decisive and the rest are noise.
| Criterion | Points towards on-premise or private cloud | Points towards hosted or SaaS |
|---|---|---|
| Regulatory and residency constraints | Data must remain in-country, under operator-only access | Compliant region and access controls demonstrably available |
| Internal platform capability | Established data centre and OT platform team | Small team, no appetite to run infrastructure |
| Estate size and growth | Large, stable, well-understood estate | Growing or uncertain scale; pilot before commitment |
| Egress connectivity | Constrained or unreliable connectivity out | Diverse, reliable links already in place |
| Integration surface | Deep integration with on-site systems and existing historians | Integration mainly with systems already hosted |
| Upgrade control | Change windows must align with the operator's own process | Continuous improvement is welcome; notice period acceptable |
| Time to first value | Long lead time acceptable | Needs to be running in weeks, not a procurement cycle |
| Capital vs operating budget | Capital funding available and preferred | Operating expenditure preferred, capital constrained |
A hybrid answer is common and legitimate: acquisition and prompt logic at the wayside, a minimal on-premise tier holding the operational record and serving the control room, and a hosted tier for long-term history, cross-estate analytics and access by field maintainers. It keeps the time-critical path inside the operational network while putting the parts that benefit from elasticity and easy access where they are cheapest to run. The cost of the hybrid is a second environment to manage and a clear rule about which tier is authoritative for what.
Requirements worth holding in every case
The strongest position an operator can take into this decision is to make it reversible. If the platform can be moved later without a rewrite and without losing the history, then the choice made now is a preference rather than a commitment, and the pressure to get it right first time largely disappears. Five requirements do most of that work:
- Open, documented protocols for acquisition and integration — OPC UA, MQTT, Modbus and similar — so the data path does not depend on one supplier's proprietary transport. The open protocols guide covers the trade-offs.
- A documented API and a bulk history export in an open format, exercised once during acceptance so it is known to work rather than assumed to.
- Deployment portability — the same platform able to run on-premise, in a private cloud or hosted, without a different product or a migration project.
- A written data statement covering residency, access, retention and termination, referenced by the contract rather than by a datasheet.
- Local buffering at every site, so no deployment model makes the wayside dependent on the central tier being reachable.
A platform that meets these can change hosting model as the operator's circumstances change. One that does not has made the decision permanent on the day of purchase, whichever way it went — and that, rather than cloud or on-premise, is the outcome worth avoiding.
Frequently asked questions
Can a rail monitoring platform run in the public cloud?
Yes, provided the data path is one-directional and the wayside layer keeps working without it. A monitoring overlay collects and presents data; it does not command signalling equipment, which makes it a candidate for hosting in a way that a control system is not. The conditions are that trackside collection, local logic and store-and-forward buffering stay on site, that the connection out is outbound-only through a controlled boundary, and that residency, contractual and exit questions are answered before signing.
What is the difference between on-premise, private cloud, hosted and SaaS deployment?
They differ in who owns the infrastructure and who operates the software. On-premise means the operator owns the servers and runs the platform. Private cloud means the operator's own virtualised infrastructure, pooled rather than dedicated. Hosted means supplier-provided infrastructure running a dedicated instance for one customer. SaaS means a shared multi-tenant service. What matters in practice is control over upgrade timing, isolation from other tenants, where the data sits, and what happens at end of contract.
Does cloud monitoring break IEC 62443 zone and conduit segmentation?
Not inherently, but it has to be designed into the zone model rather than bolted on. IEC 62443, and CLC/TS 50701 which derives its security concepts from it, group assets into zones connected by defined conduits. A hosted tier is another zone at a lower trust level, and the link to it is a conduit that must be filtered, authenticated, encrypted and — where consequence demands — unidirectional. What breaks the model is an edge device opening its own path out, bypassing the site and enterprise layers entirely.
What happens to monitoring when the link to the cloud goes down?
In a well-designed system the record is delayed, not broken. Each site data logger holds a timestamped store-and-forward buffer, so capture continues locally and back-fills in order once connectivity returns. What is lost during the outage is live visibility and notification — which is why the design question is what the operator cannot see while the link is down, and for how long. That egress path deserves the same redundancy thinking as any other single point of failure.
Where do data residency and sovereignty come into the decision?
Rail infrastructure managers are typically designated critical or essential infrastructure — under the EU NIS2 Directive they fall in scope as essential entities — and national rules often constrain where operational data may be stored and who may access it. Residency is which jurisdiction the data physically sits in; sovereignty is whose law and whose administrators can reach it. Region pinning, customer-managed keys, limits on support access and a sub-processor list are the practical controls, and they are procurement questions rather than technical afterthoughts.
Does putting the monitoring overlay in the cloud affect the signalling safety case?
It should not. The overlay is non-vital and advisory, and the vital signalling retains its own EN 5012x (RAMS) assurance wherever the monitoring tier runs. The condition attached is architectural: the deployment must not create an inbound path from the hosted tier into the vital layer. Where the data path out is read-only and the boundary is enforced, the safety argument is unchanged by the hosting decision.
Is cloud cheaper than on-premise for rail monitoring?
Usually cheaper to start, not necessarily cheaper to own. Hosting removes the up-front purchase, the data centre footprint and the administration effort. On-premise front-loads capital and then carries refresh and staff. The comparison only means something over the intended service life with the hidden items counted on both sides — refresh, patching, backup, disaster recovery and staffing against subscription escalation, egress, integration and the cost of exit. Rail lifecycles are long, so a ten-year view is more honest than a three-year one.
What should you require from a monitoring platform regardless of where it runs?
The properties that keep the decision reversible: open, documented protocols so the data is not trapped in one supplier's transport; a documented API and a bulk history export in an open format, exercised at least once; deployment portability across on-premise, private cloud and hosted without a rewrite; a written statement of residency, access and termination handling; and local buffering at every site. A platform that satisfies these can be moved later. One that does not has made the hosting choice permanent.
Where would you want a monitoring platform to run?
RailNet Operations is being shaped around ideas like these — the same platform deployable on-premise, in a private cloud or hosted, with acquisition, prompt logic and store-and-forward buffering always at the wayside, a single outbound-only path through a controlled boundary, open protocols in and a documented API and bulk export out. If you run a wayside estate, the constraints that decide this for you are the ones worth hearing. Interested in helping explore what a platform like this could do?
Start a conversation