Scenario and context

When discussing Artificial Intelligence applied to the Industrial Internet of Things, the first mistake is to assume that one AI model can handle everything.

It cannot.

An intelligent industrial plant does not need a single large, centralised brain. It needs a distributed chain in which each layer uses the right model for the right task.

The sensor must detect. The microcontroller must react. The gateway must aggregate. The AI model must interpret. The local agent must turn data into operational decisions.

This is the real difference between simply “putting AI in a factory” and building an industrial architecture that is genuinely useful.

Why industrial AI cannot depend only on the cloud

In the consumer world, we are used to thinking of AI as something that lives in the cloud: a request is sent, a remote server processes it and a response comes back.

The factory context is different.

Industrial data can be highly sensitive. It may describe production efficiency, machine cycles, energy consumption, anomalies, quality, maintenance, production recipes or strategic company information.

Sending all of this to the cloud is not always acceptable. An industrial system must also continue working when connectivity is unstable, minimise latency, remain under company control and integrate with PLCs, SCADA, local databases, sensors, gateways and existing systems.

This is where local AI becomes relevant: not as a replacement for certified industrial systems, but as a new layer for analysis, interpretation and decision support.

First layer: TinyML close to the machine

The first layer is the one closest to the physical world: microcontrollers, smart sensors, ESP32, STM32, industrial Arduino boards, embedded devices and small dedicated CPUs.

Large language models do not belong at this level. No one should expect to run a complete LLM directly inside a microcontroller.

This is the domain of TinyML.

TinyML runs small machine-learning models directly on devices with extremely limited resources. Runtimes such as LiteRT for Microcontrollers are designed for microcontrollers with only a few kilobytes of memory, without an operating system, standard C/C++ libraries or dynamic memory allocation.

These models do not need to “reason” in natural language. They need to detect.

They can recognise abnormal vibration, an out-of-curve temperature, an unusual sound, irregular power consumption or a small recurring change in machine behaviour.

Their job is simple, fast and local. They do not explain what is happening. They signal that something deserves attention. That simplicity is precisely what makes them valuable.

Operational implications

Second layer: edge gateways and compact models

The second layer is the edge gateway. Here we can use more capable devices: Raspberry Pi systems, fanless industrial mini-PCs, local edge servers, DIN-rail gateways or small internal servers.

This is where data from sensors is aggregated, normalised and converted into readable information. It is also where compact AI models begin to make sense.

Examples include Llama 3.2 1B/3B, Qwen2.5 1.5B/3B or Phi-3.5 Mini. These smaller models are intended for local, mobile or edge scenarios and can be deployed through runtimes such as Ollama.

They must not replace a PLC, control safety interlocks, command emergency stops or make millisecond real-time decisions. Those responsibilities remain with PLCs, SCADA, industrial controllers and certified systems.

A compact model on an edge gateway has a different task: read logs, classify events, interpret anomalies, produce structured responses and convert technical data into operational guidance.

For example: “During the last 12 cycles, the press showed a progressive variation in its pressure curve compared with historical data. The behaviour remains within the threshold, but the trend suggests a possible loss of hydraulic efficiency. A preventive inspection is recommended before the next production shift.”

That is the point: not another dashboard full of charts that someone must interpret, but a system capable of explaining what is happening.

Which model should run on a Raspberry Pi or edge device?

For a Raspberry Pi or a small local gateway, model selection must be realistic. The goal is not to chase the largest model, but to choose one light enough to run well, stable enough to behave predictably and disciplined enough to produce useful output.

For simple classification, log reading and short messages, Llama 3.2 1B can be a practical choice. On a Raspberry Pi 5 with adequate RAM, Llama 3.2 3B may offer a useful compromise between quality and size.

When the system must produce clean structured output, such as JSON that will be sent back to management software or a dashboard, Qwen2.5 1.5B or 3B can be particularly interesting.

Phi-3.5 Mini is another valid option when more sequential reasoning is required, but it should be used cautiously on a Raspberry Pi. If the same device must collect data, handle MQTT, write to a database, feed a dashboard and run continuously, a smaller model is often the safer choice.

The correct logic is: TinyML on microcontrollers, a compact LLM on the gateway, and Python as the orchestrator.

What changes for the company

PLCs and SCADA remain responsible for safety and deterministic real-time control.

Third layer: the local AI agent

At the third layer, the model no longer works alone. It is integrated with Python, local databases, MQTT, Modbus, internal APIs, dashboards, company rules and historical data.

This is where the system stops being a simple collection of signals and becomes a real operational assistant.

The agent can read events, compare them with historical records, check thresholds, query a database, consult company-defined rules and generate a diagnosis. It can issue a technical alert, prepare a shift report, suggest a preventive check, classify an anomaly or explain why a parameter looks suspicious.

All of this can happen locally, without sending sensitive data to the cloud, depending on an internet connection or exposing strategic industrial information outside the company.

This is a substantial difference. We are not describing a general AI system that answers general questions, but a local agent configured to understand the operating context of a specific plant.

The right chain for local industrial AI

A credible IIoT architecture can be organised as follows:

sensors → TinyML → edge gateway → compact model → local agent → operational diagnosis

Each layer has a precise role. Sensors collect data. TinyML detects elementary patterns near the machine. The gateway aggregates and normalises. The compact model interprets logs, events and anomalies. The local agent connects everything and produces understandable guidance. The PLC continues to manage real-time control.

This separation is essential because a credible industrial architecture must never confuse AI with safety automation.

Method and next steps

AI must not replace what already works and must remain deterministic. It should add a higher layer: one capable of reading plant behaviour over time, identifying weak signals, clarifying data and helping people make better decisions.

Why this approach is more concrete than generic cloud AI

The promise of industrial AI is not a chatbot on the factory floor. The real promise is a system capable of understanding operational context.

A sensor that detects a vibration is not enough. A dashboard with twenty charts is not enough. An alarm that triggers only after a threshold has already been exceeded often arrives too late.

Value emerges when the system combines several signals: historical trends, machine cycles, temperatures, power draw, pressure, product quality, previous maintenance, recurring anomalies and operating conditions.

At that point, local AI can turn a fragmented set of data into a comprehensible interpretation. It does not replace the technician. It helps the technician see earlier what would normally become visible too late.

Conclusion

Industrial IoT becomes genuinely interesting when intelligence is not distant, but close to the machine: inside the plant, inside the company and under the control of the people who generate those data every day.

The future of industrial AI will not be made only of enormous cloud models. It will also include micro-models on sensors, compact models on gateways, Python-orchestrated local agents and distributed architectures that can work offline.

Not enormous models everywhere. Not compulsory cloud processing for every data point. Not automation described as magic.

Instead: TinyML to detect, edge AI to interpret, local agents to explain, and PLCs or industrial systems to control.

This is one of the most solid directions for artificial intelligence applied to industry.