Scenario and context

AI no longer lives only in the cloud

For a long time, advanced artificial intelligence was associated with one precise architecture: large models hosted on remote servers, access through APIs, usage-based costs and data sent to external platforms.

That was, and still is, the dominant model.

ChatGPT, Gemini, Claude and other cloud systems have brought AI into companies, professional firms, marketing departments, technical teams and content production. Their advantages are clear: immediate computing power, continuous updates, simple access and no need to manage complex local infrastructure.

But this model also has limits.

When a company works with confidential data, internal documents, operational procedures, contracts, legal records, medical files, invoices, financial reports or proprietary code, the question is no longer only how powerful the model is. It also matters where data is processed, who controls it, how much each request costs, how far the system can be customised and how deeply it can be integrated into daily processes.

This is where local AI enters the picture: not as a total replacement for cloud services, but as a second possible architecture.

A dedicated machine installed in a company or professional office, with a capable CPU, suitable GPU, sufficient RAM and fast storage, can become a private AI environment. It can run local models, build RAG systems, index documents, host specialised agents and support custom vertical applications.

In this context, Gemma 4 from Google DeepMind is relevant because it represents another step towards more efficient models that can run not only on servers, but also on workstations, advanced laptops and local devices.

What is Gemma 4?

Gemma 4 is a family of open-weight models developed by Google DeepMind. It should not be confused with Gemini, Google’s proprietary and cloud-oriented model family.

The distinction matters. Gemini is primarily delivered through Google services, APIs and cloud platforms. Gemma is designed as a family of models whose weights can be downloaded and used by developers in local or customised environments.

This allows a company, developer or consultant to run a Gemma model on its own hardware, integrate it into an application, connect it to local documents, use it to create agents or place it inside an existing software workflow.

Gemma 4 introduces several relevant capabilities:

  • multimodal support for text and images, with audio available on some models;
  • very long context windows, up to 128K or 256K tokens depending on the version;
  • improved reasoning capabilities;
  • coding support;
  • features designed for agents;
  • native system-prompt support;
  • function-calling capabilities;
  • several model sizes for different hardware and use cases.

The point is not that Gemma 4 is automatically “better” than the most powerful cloud models. That would be an oversimplification. The point is that models of this kind make a different architecture more realistic: local, private and customisable AI integrated into company processes.

Why local AI really matters to companies

Discussions about local AI often fall into two extremes. Some assume that installing a model on a PC immediately creates a private equivalent of the best cloud systems. It does not. Local models, particularly small or heavily quantised ones, can have important limits in reasoning, accuracy, speed and reliability.

Others still regard local AI as a hobbyist or developer experiment. That view is now equally reductive.

Operational implications

Local AI is interesting not because it does everything better than the cloud, but because it solves different problems.

The first is confidentiality. When a model runs on an internal machine, documents can remain inside the company environment. This is central for law firms, notarial offices, tax advisers, healthcare organisations, administrative departments and businesses that manage proprietary know-how.

The second is control. A local system lets the organisation choose the model, configuration, indexed documents, applied rules, retained logs and integrated functions.

The third is operational continuity. A local system can continue to work without constant dependence on external APIs, third-party policies, price changes or usage limits.

The fourth is specialisation. A local model connected to a document archive, database, internal procedure or company files can be far more useful than a generic chatbot. Not because it “knows everything”, but because it is placed in the correct context.

This is the real distinction: useful business AI is not only the model. It is the architecture around the model.

Gemma 4 models and the resources they require

The Gemma 4 family includes several sizes, each with a different balance of capability, speed, memory use and intended scenario.

The figures below are indicative inference requirements and depend on the precision used. BF16 consumes more memory but preserves greater precision; 8-bit reduces consumption; 4-bit quantisation is lighter and more practical for local systems, although it can introduce qualitative compromises.

ModelBF16 memory8-bit memory4-bit memorySuggested use
Gemma 4 E2Babout 11.4 GBabout 5.7 GBabout 2.9 GBTesting, edge, lightweight applications, less powerful devices
Gemma 4 E4Babout 17.9 GBabout 8.9 GBabout 4.5 GBLocal assistants, automation, small agents, light text analysis
Gemma 4 12Babout 26.7 GBabout 13.4 GBabout 6.7 GBAdvanced laptops, consumer workstations, more capable multimodal agents
Gemma 4 26B A4Babout 57.7 GBabout 28.8 GBabout 14.4 GBHigh-end workstations, local servers, more complex agentic workflows
Gemma 4 31Babout 69.9 GBabout 34.9 GBabout 17.5 GBWorkstations or servers, advanced reasoning, demanding applications

These numbers must be interpreted carefully. They do not mean that having exactly that amount of memory will always be sufficient. Requirements can grow with context length, generated tokens, runtime, GPU management, operating system and inference configuration.

For example, a 6.7 GB 4-bit model may appear suitable for an 8 GB GPU, but a long context or an inefficient runtime can exhaust the available memory. Part of the workload may then spill into system RAM, reducing speed.

E2B and E4B: the lightweight models

Gemma 4 E2B and E4B are the lightest models in the family. The “E” refers to effective parameters. They are designed to extract as much capability as possible from compact architectures, particularly for on-device or edge scenarios.

Gemma 4 E2B is suitable for quick tests, small assistants, prototypes and applications where lightness matters more than deep reasoning. It can make sense on less powerful devices or where a model must respond quickly to simple tasks.

Gemma 4 E4B is probably the first size that becomes interesting for many practical uses. In 4-bit quantisation it requires a manageable amount of memory and can support local assistants, text analysis, classification, light automation, information extraction and small agents.

I would not choose it for complex strategic reasoning, very long analyses or advanced code generation. But it can be a useful foundation for controlled local applications.

Gemma 4 12B: the most interesting balance

Gemma 4 12B may be the most interesting model for organisations that want to build a serious local system without immediately entering the world of extreme workstations.

It is substantially more capable than the small models while remaining compatible with advanced consumer hardware, especially when quantised. It can make sense on powerful laptops, PCs with 12 or 16 GB GPUs, compact workstations and local development environments.

Its role is to bridge the gap between lightweight models and heavier server-class systems. For a consultant, developer or small company, Gemma 4 12B can be a good foundation for practical work.

What changes for the company

  • document analysis;
  • internal assistants;
  • RAG over company procedures;
  • technical-writing support;
  • SEO analysis;
  • coding support;
  • document classification;
  • AI-agent prototypes;
  • local tools for professional offices.

If the hardware allows it, this is probably the model size from which I would start a concrete project.

Gemma 4 26B A4B: the Mixture-of-Experts model

Gemma 4 26B A4B uses a Mixture-of-Experts architecture. It has many total parameters, but activates only part of them for each token during inference. In theory, this increases overall capacity without using the full model densely at every step.

There is an important caveat: even if only part of the parameters is active during generation, the system still has to load the complete structure required for the model to operate. It is therefore wrong to assume that “only four billion parameters are active, so it weighs the same as a 4B model”. It does not.

Gemma 4 26B A4B requires substantial resources. At 4-bit it may be manageable on a workstation with a large GPU or a mixed GPU/RAM configuration, but it should not be regarded as lightweight. It is more appropriate for agentic workflows, more complex reasoning, tool use, structured coding assistants and ambitious local applications.

Gemma 4 31B: more capability, more hardware

Gemma 4 31B is the heaviest model in the main family. It may deliver greater capability, but it needs suitable hardware. Even at 4-bit, its approximate footprint exceeds what many consumer GPUs can handle comfortably, particularly with long contexts and acceptable performance.

It belongs in a serious workstation or local server environment, not as the first model installed on an ordinary office PC. For many companies, starting directly with a 31B model would be a mistake, not because the model lacks value, but because it increases complexity: more memory, more technical management, higher cost and greater configuration effort.

In many cases, it is better to begin with E4B or 12B, validate the use case and only then decide whether a larger model is justified.

How to run Gemma 4 locally

There are several ways to use models such as Gemma 4 locally. The simplest testing route is through tools such as LM Studio or Ollama, which can download models, start them locally and expose them through a graphical interface or local API.

LM Studio is useful for experimenting with different models without immediately entering deep technical complexity. A model can be loaded, a local OpenAI-compatible server can be started and then connected to a Python or Streamlit application.

Ollama is another practical route, particularly for running local models through the command line or an HTTP API.

For greater control, developers can use llama.cpp, llama-cpp-python, Hugging Face Transformers, vLLM and other inference runtimes. These are more technical, but make it possible to build more integrated applications with less dependency on external software.

The choice depends on the objective. LM Studio is convenient for testing models and prompts. Ollama is practical for rapid prototyping. For a professional installable application, a more integrated solution such as llama-cpp-python or an application-controlled local server may be more appropriate.

The real value: RAG and specialised agents

The most important point is not merely running Gemma 4 locally. Real value appears when the model is connected to company data.

A local model on its own is still a general model. It can answer, write, summarise, reason and produce code, but it does not automatically know a firm’s internal procedures, a client’s documents, archived invoices, contracts, legal records, company policies or sales reports.

To make it useful, an architecture must be built around it. This is where Retrieval-Augmented Generation, or RAG, becomes important.

Company documents are indexed in a local system. When the user asks a question, the system retrieves the most relevant passages and provides them to the model as context. The model should not invent the answer; it should work from the retrieved text. This reduces hallucinations, improves relevance and transforms AI from a generic chatbot into an operational tool.

A notarial office, for example, could use a local AI system to consult deeds, drafts, internal procedures and document templates. A private clinic could query administrative documents, management-control reports and procedures. A technical company could connect the model to manuals, product sheets, quotations, standards and internal documentation.

In all these cases, the model is not the final product. The final product is the agent built above the model.

Cloud and local AI are not enemies

One of the most misleading interpretations is that local AI must completely replace cloud AI. It does not. In many cases, the best architecture will be hybrid.

Cloud services remain useful when very powerful, continuously updated and highly multimodal models are required, or when the task demands advanced reasoning and complex external integrations. Local systems become valuable when privacy, control, customisation, predictable costs and integration with internal data matter.

An intelligent hybrid architecture might use:

  • local AI for confidential documents, internal procedures and initial analysis;
  • cloud AI for more complex tasks, advanced generation, validation or processing that requires stronger models;
  • local RAG to retain control of data;
  • internal databases and logs to trace activity;
  • simple interfaces for end users.

Method and next steps

The objective is not to choose a technological flag. The objective is to design the architecture correctly.

Why Gemma 4 matters to consultants, developers and SMEs

For digital consultants, developers, companies and professional firms, Gemma 4 makes a local-AI proposition more credible. Until recently, suggesting a computer dedicated to AI sounded like a laboratory project. It can now become a concrete business proposal.

A correctly configured workstation can host local models, vector databases, indexing tools, Streamlit interfaces, login systems, administration panels, usage logs and vertical agents.

This opens several scenarios: a local SEO-analysis agent; an assistant for reading and querying PDF documents; a system that analyses invoices and generates reports; a management-control agent; an assistant for legal or notarial offices; an internal copilot for technical departments; a company knowledge-management system; or a private AI application for an administrative or sales department.

In these cases, value is not only in the model’s power. It is in the ability to build a complete solution: hardware, model, data, interface, rules, workflow and maintenance.

Limits to consider

Local AI is not a magic solution. The first constraint is hardware. More capable models require GPUs with substantial memory. An 8 GB GPU may be enough for small or medium quantised models, but not for everything.

The second constraint is speed. A local model may be slower than a cloud model optimised on very large infrastructure.

The third is quality. For the same task, a frontier cloud model may still be superior, particularly in complex reasoning, general knowledge, advanced coding and demanding multimodality.

The fourth is maintenance. Models, runtimes, libraries and dependencies change. Someone must be able to update, test and maintain the system.

The fifth is application security. Running a model locally does not automatically make the system secure. Access, permissions, backups, logs, encryption, data isolation and source control still have to be managed.

Local AI therefore has to be designed. Installing a model is not enough.

Conclusion: the question is not whether to use AI, but which architecture to use

Gemma 4 is interesting because it confirms a direction that is becoming increasingly clear: artificial intelligence is becoming more distributed. It will not live only in large data centres. It will also enter company computers, professional workstations, local servers, edge devices and custom vertical systems.

For companies, this means that asking “which model should we use?” is no longer sufficient.

The correct questions are broader: Which data do we want to use? Where must they remain? Which activities should be automated? Which users will operate the system? Do we need cloud, local or hybrid architecture? Which risks must be controlled? Which operational value should be produced?

Gemma 4 is not simply another model to add to a list. It signals that local AI is becoming technically and economically more accessible. The future of business AI will not be made only of chatbots. It will be made of architectures.

Those who can design these architectures by connecting models, data, applications and real processes will have a far more durable advantage than those who merely “use artificial intelligence”.