Case Study
Before Building the Factory of the Future, Find the Data Problem
Industry 4.0 · Data Architecture · Enterprise Technology Strategy
When a manufacturing organization begins an Industry 4.0 transformation, the obvious questions are usually about machines, IoT, AI, dashboards and automation.
I found a more fundamental question: what happens to the data once everything starts producing it?
In a manufacturing transformation I worked on, the organization already had multiple sources of operational and business data. The challenge wasn't simply connecting them. It was making sense of information arriving from different places, in different formats, at different times.
The real problem isn't collecting data
Consider a simple manufacturing event.
A machine stops at 10:15 AM. At 10:23 AM, an operator records the reason. At 11:05 AM, maintenance closes the incident.
Three updates. Potentially three systems. Three different timestamps. But the business doesn't want three records. It wants to know:
- When did the machine actually stop?
- Why did it stop?
- How long was it down?
- What was the impact on production?
Putting all three records into a data lake doesn't answer those questions. The architecture has to understand that they are different pieces of the same operational event.
That, to me, is one of the less visible challenges of Industry 4.0.

Two architectural decisions
Make data the foundation
Rather than building individual applications around individual systems, I designed the transformation around a data-led architecture.
The first foundation was a data lake capable of ingesting different types of information from operational and enterprise sources. But ingestion was only the beginning.
The architecture needed to progressively turn raw information into something the business could actually use: ingest, classify, normalize, correlate, reconcile, derive, then present.
A production event, a quality event, a maintenance event and an enterprise transaction may each describe a different part of the same business reality. They need to be related, contextualized and reconciled before they can reliably support a KPI, dashboard, application or AI system.
This is where the data fabric becomes important.
Don't overbuild the first phase
If the long-term objective is a sophisticated digital manufacturing environment, should the organization build the entire target platform from day one? I didn't think so.
The first phase was deliberately designed to be on-premise, lightweight and open-source based. The reason wasn't simply cost.
A new digital platform also creates an operational responsibility. Someone has to deploy it, monitor it, secure it, back it up, troubleshoot it, upgrade it and eventually scale it. The IT team needs to develop the capability to operate that environment.
The proposed progression was: start manageable → build operational capability → learn from real workloads → increase scale → introduce more sophisticated or proprietary components when justified.
The architecture could therefore evolve alongside the organization.
From data to meaning
The proposed architecture separated this into distinct stages. The idea was simple: don't make every application interpret the data differently. Put that intelligence into the data foundation so that different consumers can work from the same understanding of the underlying business events.
- Ingestion
Bring information in from machines, shop-floor systems and enterprise applications.
- Processing
Clean, validate, transform and enrich incoming information.
- Semantic layer
Establish common definitions, relationships, master data and business context.
- Consumption
Make trusted information available to dashboards, analytics, applications and AI.
The architecture was designed to evolve
The initial platform was not intended to be the final destination. It was the foundation.
As data volumes increased, use cases expanded and the organization gained experience operating the platform, the infrastructure could progressively become more sophisticated.
That approach also reduces the risk of making a large technology commitment before there is enough operational experience to understand where that investment will actually create value.
Industry 4.0 is more than connecting machines
Machines, sensors, IoT gateways, enterprise systems and AI are all important pieces of an Industry 4.0 environment. But connecting them is only the beginning.
The harder problem is turning fragmented events into a shared understanding of what is actually happening in the business.
That's why I would approach an Industry 4.0 transformation from the data outward.
- Where does the data originate?
- How does it move?
- How is it interpreted?
- How are events related?
- How is business context applied?
- And only then — what should we build on top of it?
A blueprint, not a finished factory
In this engagement, that thinking resulted in a strategic and technical blueprint covering the current state, data architecture, integration architecture, technology choices, security considerations and phased transformation roadmap.
The proposal is currently under discussion, so the operational improvements defined in the plan remain targets rather than measured outcomes.
For me, the interesting part of the exercise wasn't designing a future factory. It was figuring out what foundation would allow that future factory to make sense of itself.
Start with the data problem
If an Industry 4.0, IoT or data platform program is forming, the first work is to understand how information will be interpreted—not only how it will be collected.