“Industrial artificial intelligence is a central component of the AI strategies of many industrialized nations, such as Germany and Japan. It creates significant economic added value, for example, in the areas of autonomous vehicles, the automation of logistics centers, supply chain and transportation planning, and predictive maintenance of machinery.
Industrial AI differs fundamentally from the large language models (LLMs) that are widely used today in many commercial and private applications. These models are primarily trained with vast amounts of publicly available internet data.
Industrial AI, on the other hand, cannot rely on such generic data sources. It depends on highly specific, high-quality, and often scarce data generated within industrial environments—data that no single company possesses in sufficient quantities.
Industrial data is not only limited in availability but also frequently highly sensitive, as it contains business-critical information. Furthermore, certain events, such as machine failures, occur only rarely, even though they are essential for training powerful predictive models.
In addition, industrial environments are constantly evolving. Machines are modernized, processes are adapted, and external conditions change. AI models must therefore be continuously developed to avoid performance declines, the so-called "model drift."
The opportunities for industrial AI are enormous. However, its success depends less on further algorithmic advances than on access to high-quality, domain-specific data.
The data from a single company is generally insufficient to develop robust industrial AI models.
One example is predictive maintenance of industrial machinery: A manufacturer can collect operational data from its own equipment or that of its customers. But this data set is often too small or too homogeneous to train highly reliable models. Rare failures, varying operating conditions, or industry-specific requirements can only be captured when data from multiple organizations is combined.
This creates a fundamental prerequisite: Companies must cooperate in ecosystems and share data with one another. This is precisely where the reservations begin.
From the perspective of data owners, sharing raw data potentially means a loss of control and the risk of disclosing trade secrets. Companies fear, for example, that competitors could draw conclusions about production inefficiencies or failure patterns.
Added to this is the uncertainty surrounding regulatory frameworks. While laws such as the Data Act, the AI Act, and the Data Governance Act are intended to provide guidance, in practice they often lead to further reluctance because many companies are unsure whether they fully comply with all the requirements.
The training of industrial AI models is therefore characterized by a fundamental conflict of interest: On the one hand, data from diverse sources and the collaboration of various stakeholders are indispensable. On the other hand, companies are reluctant to disclose their valuable data assets.
Data spaces offer a solution to this dilemma. They enable the exchange of industrial data and simultaneously foster trust within industrial data ecosystems. The basic idea is that data remains with the owner and can still be used under clearly defined conditions. Instead of storing raw data centrally, data spaces enable direct, controlled exchange between participants.
A key technological approach in this context is federated learning. With this approach, the data does not leave its local environment. Instead, AI models are trained decentrally, and only model parameters or updates are sent to the data center. a central instance transmits the data. This enables collaborative learning without disclosing sensitive information. Data spaces also create an architecture that distinguishes between a "control plane" and a "data plane." While the control plane defines access rights, governance rules, and usage guidelines, the data plane organizes the actual data exchange. This separation allows for flexibility while simultaneously fostering trust and regulatory certainty.
Within this framework, three forms of collaborative AI can be distinguished: the joint development of AI models by multiple stakeholders, for example, through federated learning; the cross-organizational use of external data to improve existing models; and the collaboration of independent AI agents that exchange insights without sharing raw data.
By combining the approaches, data spaces enable scalable and trustworthy collaboration across corporate boundaries.
Despite the obvious benefits, the adoption of data spaces is progressing slowly. There are several practical reasons for this: First, many companies still struggle with internal data management. Data resides in silos, varies in quality, and requires significant effort to prepare for AI applications.
Second, data quality itself is a key issue. Poor-quality data not only impairs the performance of AI models but also poses significant risks, particularly in safety-critical areas such as manufacturing or autonomous systems.
Third, governance remains a sensitive topic. While data spaces offer mechanisms for defining usage rules, the participating parties must first agree on common standards. This requires trust and coordination—and is anything but trivial.
Added to this is an often-underestimated economic dimension. Data is a valuable resource. Those contributing data to a shared model want to understand how that contribution translates into value—whether in the form of better models, financial returns, or strategic advantages.
The challenge is therefore not merely technological: successful implementation requires aligning technology, legal and organizational frameworks, and economic incentive systems.
In the long term, the goal is to build intelligent data ecosystems that enable a comprehensive digital representation of the real world. These ecosystems will be based on interconnected digital twins of assets, processes, and supply chains, allowing for continuous data exchange across organizational boundaries.
A defining characteristic of this development is the shift from linear to circular thinking: traditional business processes follow clear sequences with defined start and end points, whereas the data-driven systems of the future operate within continuous feedback loops. Data is continuously generated, shared, and reused to constantly improve processes and decisions. This circular approach applies to both physical processes—such as resource efficiency and reuse—and digital processes, in which data continuously feeds back into AI models and decision-making systems.
Boris Otto is Professor of Industrial Information Management at TU Dortmund University and Director of the Fraunhofer Institute for Software and Systems Engineering ISST.
Takahide Matsutsuka is responsible for the research and development of cutting-edge technologies in the field of data spaces at Fujitsu Research.” [1]
Does China develop similar movement?
Yes, China is aggressively pursuing a similar and arguably more state-directed movement, officially designating data as a core "factor of production" alongside labor, land, capital, and technology.
Rather than relying strictly on decentralized European-style data spaces (like Gaia-X), Beijing established the National Data Administration to build unified national data standards, trading markets, and cross-sector industrial pooling.
Core Elements of China's Industrial Data & AI Strategy
• Data as Production Factor: Elevating data to a foundational economic asset to power smart manufacturing, autonomous systems, and physical AI.
• "AI + Manufacturing" Action Plans: Coordinated via the Ministry of Industry and Information Technology (MIIT) to merge the industrial internet with advanced AI across tens of thousands of enterprises.
• National Data Exchanges: Creating regional data trading hubs in cities like Shanghai and Shenzhen to circulate proprietary industrial, transport, and manufacturing datasets.
• Physical & Embodied AI Focus: Leveraging its massive factory floor automation and robotics density to gather real-world, sensor-level data that Western consumer-web models cannot easily replicate.
The main creators of American artificial intelligence (Sam Altman and Dario Amodei) are developing a different system. They create closed AI models, do not allow them to be improved locally by companies themselves, taking advantage of people's lack of education and the alleged danger that AI, which is supposedly beyond their control, poses to humanity. They create AI agents that hack into the systems of other companies. In this way, they hope to steal the business secrets of all the companies in the world, create a powerful AI on its basis, use the efficiency of this AI and push all the companies in the world out of the market. Then there would be only two American companies in the world that could calmly collect data for themselves and continue to develop AI.
Critics often accuse Altman and Amodei of using the "fear factor". There is an opinion that by talking about the “existential threat to humanity” and periodically releasing AI agents-hackers, these big players are trying to provoke state regulation (strict laws), which would put an excessive burden on small startups and other AI companies and thus protect the market share of their own large companies. This is called “regulatory capture”.
If these two actors succeed in their plan, the problem of collecting data for AI training will be solved in a purely American way, thanks to private companies, without the participation of European cooperatives and the Chinese government.
1. Datenräume: Der Schlüssel zur industriellen KI. Frankfurter Allgemeine Zeitung; Frankfurt. 01 June 2026: 18. Von Boris Otto und Takahide Matsutsuka
Komentarų nėra:
Rašyti komentarą