Tacit knowledge is the missing dataset

Industrial AI has plenty of sensor data, but some of the most important context still lives in people's experience. The challenge is capturing that knowledge before it disappears.

Share

Why better models and more documents still leave industrial AI without some of the context it needs.

When I first started working with industrial AI, I thought the data problem was mostly about sensors. Collect enough vibration data, improve its quality, label the important events, and the model should get better over time. That turned out to be only part of the problem.

The harder issue often appeared after the data arrived.

An abnormal vibration pattern might look significant until an engineer explained that the machine behaves differently during a particular production cycle. A technician might remember seeing the same pattern after a coupling replacement. At another site, the sensor might have been installed in a slightly different position because the recommended location was physically inaccessible.

These details could completely change how the signal should be interpreted, yet none of them necessarily appeared in the sensor data.

Sometimes they were not in the maintenance record either.

The missing information is often between systems

Industrial environments already generate plenty of information. Sensors capture physical signals, maintenance systems record work orders and repairs, manuals describe expected behavior, and production systems provide operating history.

But knowing what happened is different from knowing why someone interpreted it the way they did.

Suppose a monitoring system raises an alert and an engineer decides not to take immediate action. The database may eventually contain a timestamp, an acknowledgement, and perhaps a status such as closed. The decision itself, however, may have depended on a physical inspection, the operating load at that moment, knowledge of a similar incident six months earlier, or the fact that the equipment could not be stopped until the next scheduled shutdown.

The event survives. Much of the reasoning does not.

This problem has been discussed for much longer than AI has existed. Michael Polanyi's work on tacit knowledge begins with a deceptively simple observation:

“We can know more than we can tell.” [1]

Polanyi was not writing about predictive maintenance, but the idea translates surprisingly well to industrial work. An experienced technician may recognize a machine problem from its sound. An engineer may notice that a spectrum looks unusual before any individual feature exceeds a threshold. Operators gradually learn which behaviors are normal during startup, seasonal changes, or specific production conditions.

Much of this is not mysterious intuition. It is experience compressed through repeated exposure. The difficulty is turning that experience into something a system can reuse.

Research in industrial maintenance suggests that the gap is real. A 2021 study of maintenance technicians compared their perceived personal knowledge with knowledge recorded by their companies. Depending on the activity, the amount of recorded knowledge relative to technicians' own knowledge ranged from 43.9% to 68.78%. [2]

Maintenance areaRecorded knowledge relative to technicians' own knowledge
Operations43.90%
Energy efficiency49.61%
Reliability & breakdowns51.25%
Preventive, predictive & corrective maintenance68.78%

Source: Cárcel-Carrasco & Cárcel-Carrasco (2021). These values reflect technicians' perceptions in the study, not a universal measurement of industrial knowledge.

The exact percentages should not be generalized to every factory, but the pattern is more interesting than the numbers themselves: a meaningful part of operational knowledge remained with the people doing the work rather than in the organization's recorded information. The authors also point to the risk of losing that knowledge when experienced technicians leave.

More retrieval does not fix missing context

This becomes particularly relevant as companies adopt LLMs.

One of the obvious ways to improve an enterprise AI system is to give it access to more information: manuals, maintenance reports, sensor history, databases, emails, troubleshooting guides, and previous incidents. Better retrieval clearly helps when useful information already exists somewhere.

But retrieval cannot retrieve something that was never captured.

A model might have the complete maintenance manual and every work order for a motor, yet still not know that a recurring vibration peak comes from resonance in a connected structure unless somebody documented it.

I find it useful to separate three problems that are often grouped together as a single “data problem”:

  • Missing data: something important was never measured.
  • Fragmented data: the information exists, but is scattered across systems.
  • Missing context: someone understood what happened, but the reasoning was never recorded.

The first can often be improved with sensing. The second is increasingly addressable through integration, search, and retrieval. The third is different because the information may disappear as soon as the person moves on to the next task.

Knowledge-management research has dealt with this problem for decades. Nonaka's work on organizational knowledge creation focused on the relationship between tacit knowledge held by individuals and knowledge that can be shared within an organization. [3] Industrial-maintenance researchers have also explored systems for capturing and reusing expertise rather than leaving it isolated with individual experts. [4]

AI makes this old problem much more visible. A model can process an extraordinary amount of explicit information, but giving it more processing capacity does not create context that the organization never preserved.

The product can create the dataset

This changed the way I think about feedback loops in industrial software.

Consider two alerts that engineers both mark as false positives. From a machine-learning perspective, they have the same label. But perhaps the first occurred because production load temporarily changed, while the second came from a sensor that had partially detached from the machine.

The label is identical. The explanation is not.

Capturing every explanation through detailed forms is not a realistic solution. Maintenance engineers and operators have machines to operate, not datasets to annotate. If documenting context becomes additional administrative work, much of it simply will not happen.

The more interesting design problem is therefore how to make operational work leave useful traces naturally. That could include:

  • capturing a short reason when an alert is dismissed;
  • linking an alert automatically to the resulting work order;
  • attaching an inspection photo or short voice note;
  • recording the operating condition at the moment a decision is made;
  • surfacing previous similar incidents while the engineer is investigating.

None of these features is particularly sophisticated on its own. Their value comes from what accumulates over time.

A sensor dataset tells us how a machine behaved. A maintenance history tells us what people did to it. If we can also preserve why those decisions were made and what happened afterward, the system begins to contain something much closer to operational experience.

That is a different kind of dataset.

What I would build for

This is why I no longer think the most important question at the beginning of an industrial AI project is simply, “How much historical data do we have?”

There is another question worth asking:

What new data will this product create as people use it?

Models will improve. Sensor hardware will improve. Retrieval across enterprise systems will improve as well. Those technologies make existing information easier to use, but they do not automatically preserve the knowledge being created while people diagnose problems and make decisions.

For physical systems, that may be one of the more durable opportunities for AI products: not merely analyzing machine data, but gradually capturing the relationship between signals, human judgment, actions, and outcomes.

The missing dataset may already exist in people's experience. The challenge is designing the product so that, little by little, it stops disappearing.


References

[1] Polanyi, M. (1966). The Tacit Dimension. Polanyi's central argument begins from the observation that human knowledge contains a tacit component that cannot always be fully articulated.

[2] Cárcel-Carrasco, J., & Cárcel-Carrasco, J.-A. (2021). “Analysis for the Knowledge Management Application in Maintenance Engineering: Perception from Maintenance Technicians.” Applied Sciences, 11(2), 703. DOI: 10.3390/app11020703.

[3] Nonaka, I. (1994). “A Dynamic Theory of Organizational Knowledge Creation.” Organization Science, 5(1), 14–37.

[4] Potes Ruiz, P. A., Kamsu-Foguem, B., & Noyes, D. (2013). “Knowledge reuse integrating the collaboration from experts in industrial maintenance management.” Knowledge-Based Systems, 50, 171–186. DOI: 10.1016/j.knosys.2013.06.005.