Smart Manufacturing

Your Factory Has Data. That Does Not Mean It Is Ready For AI

Photo by Andres Siimon (@johnmcclane) on Unsplash

A production manager can usually point to years of machine logs, quality reports, maintenance records and sensor readings stored somewhere across the organisation. On paper, the factory appears to possess exactly what an industrial AI project needs: large quantities of operational data accumulated through daily production.

The first pilot often reveals a less convenient reality. Machine names differ between systems, timestamps do not align, maintenance notes are stored as free text, and the quality database records a defect without preserving the process conditions that preceded it. Engineers understand what the individual values mean because they know the line, the product and the history behind the equipment, but an AI system sees disconnected measurements whose significance has never been made explicit.

Industrial companies rarely suffer from a complete absence of data. They struggle because the information they possess was collected for maintenance, control, compliance or reporting rather than for machine learning, which leaves much of it fragmented, poorly labelled and detached from the operational context required to support a reliable decision.

A Measurement Is Not Yet Industrial Knowledge

A temperature of 82 degrees, a vibration peak or a sudden increase in energy consumption may appear meaningful, although none of these values can be interpreted properly in isolation. The same temperature could indicate normal operation during one production stage and an emerging fault during another, while a vibration pattern that looks unusual for one machine may be entirely expected after a tool change or when a different material enters the process.

For an AI model to recognise those distinctions, the measurement must be connected to the machine that produced it, the product being manufactured, the recipe and operating mode, the tool installed at the time and the events that followed. It may also need information about the shift, ambient conditions, supplier batch, maintenance history and whether an engineer had temporarily adjusted a setting.

This surrounding information is what turns a signal into usable industrial context. Without it, a model may discover statistical relationships that appear convincing while confusing normal changes in production with signs of failure or poor quality.

A factory can therefore collect billions of data points and still lack the information required to answer a relatively simple question: under which precise conditions does this process begin to produce defects?

Factory Data Was Built Around Different Purposes

The difficulty begins with the architecture of industrial operations. Programmable logic controllers were designed to control machinery, historians to preserve time-series signals, manufacturing execution systems to manage production and enterprise platforms to handle orders, materials and finance. Quality, maintenance and engineering teams often introduced additional applications according to their own needs, creating a landscape in which each system performs a legitimate function but describes the same factory differently.

The identifier used for a machine in the maintenance system may not match the tag used in the historian. A product can have one name in engineering, another in production and a third in the commercial system, while changes to a component may be recorded in a document without being reflected consistently in the data collected from the line.

This fragmentation does not always prevent employees from working because experienced teams learn how the systems correspond. They know that “Line 4”, “L04” and an internal asset number refer to the same equipment, and they remember that a sensor was replaced during a particular shutdown even when the change was never documented in a structured form.

AI cannot rely on that informal knowledge unless the organisation captures it. The model needs consistent identities, relationships and definitions that connect data across systems, otherwise much of the factory’s meaning remains locked inside the experience of individual employees.

More Sensors Will Not Correct A Weak Data Foundation

When an industrial AI project lacks useful information, the instinctive response is often to install additional sensors. In some cases that is necessary, particularly when an important physical condition has never been measured, but the factory may already collect enough signals and simply be unable to interpret them together.

Adding more devices to a fragmented environment can increase volume without improving understanding. The company receives additional vibration, pressure and temperature streams, yet still cannot determine which production order was running when the anomaly occurred or whether the resulting component later failed inspection.

Before expanding its sensor network, a manufacturer should establish which decision the AI application is expected to support and what evidence would be needed to make that decision reliably. Predicting tool failure requires different data from optimising energy use, while identifying the cause of a surface defect may demand a combination of process parameters, material genealogy, machine condition and image data.

Beginning with the operational question prevents the project from turning into an indiscriminate collection exercise. The objective is not to create the largest possible data lake but to preserve the relationships that explain what happened in production.

Historical Records Often Contain A Distorted Picture

Industrial AI systems learn from the examples available to them, but factory history is rarely balanced or complete. A well-maintained machine produces years of normal operating data and relatively few examples of failure. Serious faults may occur only occasionally, while smaller problems are corrected by experienced technicians before they become formal incidents.

The resulting dataset contains an abundance of ordinary production and only a small number of the events the model is expected to recognise. Even those examples may be unreliable because failure codes were entered inconsistently, maintenance notes lack detail or different technicians used different descriptions for the same problem.

Operational practices also change. A machine may receive new tooling, software or components, and the material supplied today may behave differently from the material used when the historical data was collected. A model trained on the past can lose relevance even though its technical accuracy once appeared strong.

This makes industrial data preparation more demanding than removing duplicate rows from a database. Engineers and operators must determine whether records still represent the current process, which events were labelled correctly and where the absence of a recorded failure merely reflects the fact that someone intervened in time.

The factory’s history is not a neutral account of production. It is a partial record shaped by equipment changes, human judgement and the reporting habits of different teams.

Unstructured Information Often Contains The Missing Explanation

Some of the most valuable industrial knowledge does not sit in a sensor database. It appears in maintenance comments, shift books, inspection photographs, engineering drawings, supplier reports and conversations between experienced operators.

A numerical record may show that vibration increased before a machine stopped, while the maintenance note explains that a replacement bearing from a new supplier had been installed two weeks earlier. A quality system records a dimensional deviation, but a shift report mentions that the material was unusually difficult to process that morning.

These connections are precisely what AI needs, yet the information is frequently stored in PDFs, handwritten notes or incompatible applications. Industrial language introduces another complication because abbreviations, local terminology and machine nicknames may be obvious to the workforce while remaining incomprehensible outside a particular site.

Large language models can help extract and organise some of this material, although they still require a controlled vocabulary and links to authoritative operational data. Turning a technician’s note into a structured event is useful only when the company can associate it with the correct asset, component, time and production order.

The task is not simply digitising documents. It is preserving the meaning they contain and connecting it to the physical process they describe.

Contextualisation Is The Work Most Pilots Underestimate

Industrial AI projects are often presented as model-development initiatives, so teams spend considerable time comparing algorithms while treating data preparation as a preliminary technical task. In practice, the demanding work lies in connecting signals to assets, products, events and process conditions in a form that can be used consistently.

This process is known as contextualisation. It gives data a position within the factory rather than leaving it as an isolated value, so a pressure reading becomes the pressure at a particular stage of a specific production cycle, on a named machine, processing a defined material batch under a recorded recipe.

Once these relationships exist, the same foundation can support several applications. Engineers can investigate quality deviations, maintenance teams can compare equipment behaviour and energy managers can distinguish productive consumption from waste. Digital twins and AI agents can also draw on a shared operational structure rather than building a separate interpretation of the factory for every project.

The alternative is a series of isolated pilots in which each team cleans and reconnects the same sources for one narrow use case. The demonstration may succeed, but the work cannot be scaled because the underlying data remains fragmented and every new application begins from the same difficult starting point.

Recent manufacturing research continues to identify industrial big data, heterogeneous sensing and effective data management among the critical barriers to reliable AI deployment, particularly where systems must operate safely and explainably in high-stakes environments.

Data Quality Has To Be Defined Operationally

Corporate discussions of data quality often focus on completeness, accuracy and consistency. These dimensions matter in manufacturing, but industrial data also has to be judged against the physical process.

A sensor can produce technically valid values while being poorly calibrated. A timestamp can be recorded correctly in one system but remain useless when another application uses a different clock. A machine state labelled “running” may include setup, testing and productive operation, even though those conditions should not be analysed together.

The people closest to production are therefore indispensable. Data engineers can identify gaps and inconsistencies, but operators, maintenance specialists and process engineers understand whether a signal is physically plausible and whether its interpretation changes under specific conditions.

This collaboration also prevents a common mistake in which a statistically clean dataset represents the process incorrectly. Removing an apparent outlier may eliminate the earliest sign of a fault, while averaging measurements over a convenient period can conceal the brief event that caused a defect.

AI readiness cannot be certified by the IT department alone. It depends on whether technical data and operational reality still describe the same thing.

Governance Determines Whether The Data Can Be Trusted Later

An AI model may begin with a carefully prepared dataset and gradually become unreliable because the factory changes around it. Sensors are replaced, recipes are adjusted, products evolve and systems receive software updates, often without the changes being communicated to the team responsible for the model.

Manufacturers therefore need clear ownership of important data sources. Someone must know what a variable means, how it is generated and when its definition changes. Calibration history, missing periods and manual corrections should remain visible rather than disappearing during preparation.

Lineage matters as well. When an AI system recommends a maintenance intervention or identifies a quality risk, engineers need to understand which data contributed to the result and whether those records came from trusted sources. This becomes particularly important when the recommendation affects safety, compliance or a costly production decision.

A factory data foundation should not create a polished version of reality from which uncertainty has been removed. It should show where information is incomplete, estimated or affected by a change in equipment, allowing the model and its users to treat the result with appropriate caution.

Brownfield Plants Need A Selective Approach

Few manufacturers can replace their entire technology landscape before beginning with AI, and mature factories often contain machines from several generations, proprietary interfaces and systems that were never designed to exchange information.

Waiting for complete modernisation would postpone useful applications indefinitely. Attempting to connect everything at once can be equally unproductive because the programme becomes an expensive infrastructure exercise without a clear operational return.

A more credible approach begins with one production problem whose economic value is understood. The company maps the data required to investigate it, identifies the most important gaps and builds enough shared context to support that use case properly. The work should still follow standards and architecture that can be extended later, but it does not need to solve every historical inconsistency in the plant.

For example, a manufacturer trying to reduce defects on one line may initially connect machine settings, quality results, material batches and selected maintenance events. Once those relationships are established, the same structure can be expanded to adjacent lines or reused for predictive maintenance.

This creates a data foundation through operational progress rather than a prolonged preparation programme whose value remains abstract.

A Data Lake Is Not The Same As Accessible Data

Many companies have already consolidated large volumes of industrial information in cloud platforms or central repositories. That is an important step, yet physical co-location does not automatically make the data understandable or usable.

A data lake can contain signals from every site while preserving the same inconsistent naming, missing relationships and uncertain ownership that existed in the original systems. Teams may gain a central place to search but still spend weeks determining which fields are relevant and whether different plants measure the same process in comparable ways.

AI requires more than storage. It needs data that is findable, accessible under the correct permissions, interoperable across systems and reusable with enough description to preserve its meaning. Semantic models, knowledge graphs and industrial data fabrics are increasingly used to represent these relationships, although the terminology matters less than the function they perform.

The useful layer is the one that allows a model to understand that a particular sensor belongs to a specific machine, that the machine performed an operation on a defined product and that the resulting component later received a particular quality result.

Siemens, for example, describes industrial data fabrics as a means of connecting fragmented sources and adding the context required by AI, while its wider industrial AI architecture links shop-floor collection with contextualisation and cross-domain use.

AI Agents Raise The Standard Further

A predictive model that produces an inaccurate maintenance warning may waste an engineer’s time. An AI agent that acts on unreliable data can create a much wider operational problem.

Agents are designed to retrieve information, reason across systems and initiate actions with less direct supervision. In manufacturing, they may prepare work orders, adjust schedules, recommend process changes or coordinate a response to an equipment event. Their usefulness depends on having access to information that is not merely available but correctly related and sufficiently current.

An agent that sees an alarm without knowing that the machine is in maintenance mode may initiate an unnecessary response. One that reads an obsolete work instruction can recommend a procedure that no longer applies, while inconsistent asset identifiers may cause it to retrieve the history of the wrong machine.

As industrial AI becomes more autonomous, the cost of weak context rises. A person can often notice that information looks wrong because it conflicts with experience; an agent may continue confidently through several steps unless the data foundation and operating rules make the inconsistency visible.

This is why industrial data readiness is becoming a question of trust rather than only technical access. Agentic systems require reliable product and operational context before companies can permit them to act with confidence.

Start With The Decision, Not The Dataset

A practical readiness assessment begins by choosing the decision the company wants to improve. The team should specify who makes it today, which information they use, what constitutes a correct result and what happens when the decision is wrong.

Only then should the manufacturer examine whether the necessary data exists, whether it can be connected and whether historical examples are representative. This reveals the difference between a genuine data gap and a contextual gap, which often requires documentation and integration rather than another sensor.

The company also needs a baseline. If the purpose is to reduce unplanned downtime, it should know current downtime, maintenance effort and false-alarm rates before introducing the model. Without that comparison, an AI pilot may look technically impressive while producing no measurable improvement in operations.

The first deployment should remain close enough to the production team that errors can be investigated quickly. Engineers and operators need to see how the model reached its conclusion, challenge incorrect assumptions and identify changes in the process that the data pipeline has not captured.

When the use case proves valuable, the company can extend the same data relationships rather than starting again. This is how isolated industrial information becomes a reusable asset.

Data Readiness Is An Operating Capability

Factories are often advised to clean their data before adopting AI, as though readiness were a project that could be completed and closed. Industrial environments do not remain still long enough for that interpretation to work.

New machines arrive, suppliers change, products are redesigned and operators discover better ways to run the process. Every alteration changes the meaning or relevance of some data, which means that contextualisation, governance and quality control must continue after the model enters production.

The companies best positioned to scale industrial AI will not necessarily be those with the largest historical databases. They will be those that know where their operational information comes from, how it relates to the physical process and who is responsible when those relationships change.

A factory becomes ready for AI when its data can support a decision that engineers trust, not when a dashboard shows that millions of records have been collected. Until the signals are connected to machines, materials, events and outcomes, the organisation has data in abundance but knowledge in short supply.