Enterprise AI Startups

Why More AI Companies Are Letting Clients Run Models Themselves

Photo by Steve A Johnson (@steve_j) on Unsplash

For most companies, using generative AI still means sending a request to somebody else’s system. The model sits in an external cloud environment, the provider controls its technical development, and access continues for as long as the commercial relationship does. This has made artificial intelligence remarkably easy to adopt. It has also left many organisations with limited control over one of the technologies they increasingly expect to embed in sensitive business processes.

A different proposition is now becoming more prominent. Some AI companies are allowing clients to download models, adapt them and operate them within infrastructure selected by the client. Thinking Machines Lab, founded by former OpenAI technology chief Mira Murati, has taken this route with its first model, Inkling. Rather than competing only for the highest benchmark score, the company is presenting adaptability, control and dependable operation as priorities for corporate users. Clients can download the model and run it themselves, while an accompanying platform is intended to make customisation easier.

The development reflects a broader change in enterprise AI. The market is no longer divided simply between companies that use AI and companies that do not. Technology leaders must now decide where models should run, who should operate them, how deeply they should be modified and how easily the organisation could move to another provider later.

Those decisions require more precision than the language surrounding open AI usually provides.

Open does not always mean open source

Terms such as open model, open weights and open source are often used as though they describe the same product. They do not.

A conventional hosted model is accessed through an application programming interface, or API. The client sends data to the provider and receives a response. The model itself remains under the provider’s control. This is the simplest arrangement for many companies because there is no need to acquire specialist computing infrastructure, maintain the model or manage the underlying software stack.

An open-weight model gives users access to the numerical parameters created during training. Those weights can usually be downloaded and operated on the client’s own infrastructure or through a cloud provider of its choice. The organisation gains considerably more control over deployment and adaptation, although the licence may still restrict certain commercial uses, modifications or forms of redistribution.

A genuinely open-source system goes further. It provides the source code and licensing rights required to inspect, modify and redistribute the software. Even here, however, the complete training data and training process may remain undisclosed. A model can therefore be open in important technical respects without offering full transparency into how it was created.

Self-hosting describes something else again. It concerns where the system is operated, not whether the technology is open. A company may run an open-weight model in its own data centre, deploy it through a Swiss cloud provider or use a dedicated environment managed by an external technology partner. In each case, the client has greater control over the operating environment than it would have with a standard public API. That does not automatically make the model open source, secure or economical.

The distinction matters because each option gives the client control over a different layer. Access to model weights provides technical flexibility. An open-source licence provides legal permissions. Self-hosting provides operational control. None of these alone guarantees data sovereignty or commercial independence.

Control is becoming part of the product

The largest AI providers built their early advantage around performance and ease of access. A company could connect to a sophisticated model without buying servers, hiring machine-learning engineers or understanding how the system had been trained.

That remains attractive for experiments, general office applications and services that do not handle particularly sensitive information. It becomes less comfortable when the model is expected to read internal contracts, interact with operational systems, support investment decisions or work with confidential client data.

For businesses in Switzerland, the concern is not merely whether information is physically stored within the country. They also need to understand who can access it, whether prompts are retained, how subprocessors are involved, what happens when the provider changes its terms and how the system could be transferred if the commercial relationship ends.

Running a model within a controlled environment can answer some of those concerns. Data may remain within infrastructure selected by the company, access controls can be aligned with existing security policies, and connections to internal systems can be governed more closely. Model updates can also be tested before they are introduced into production rather than arriving according to the provider’s timetable.

This is particularly relevant in banking, insurance, healthcare, pharmaceuticals and advanced manufacturing, where valuable information may include far more than personal data. Technical specifications, production records, pricing logic, investment research and internal decision processes can all become part of the context supplied to an AI system.

Self-controlled deployment allows an organisation to treat the model as part of its own technology architecture rather than as a remote digital service.

The highest benchmark score may not be the best corporate choice

Inkling is notable partly because its developers do not claim that it leads the market in standard performance rankings. The model sits behind several open alternatives in comparative tests. Thinking Machines Lab nevertheless appears to be targeting companies that place greater value on adaptability, reliability and control than on a marginal advantage in a benchmark.

That trade-off is familiar in enterprise technology. Companies rarely select core accounting software, databases or industrial systems solely because they perform best in a laboratory test. They also examine integration, documentation, support, security, compatibility and the expected life of the product.

AI procurement is gradually moving in the same direction.

A slightly less capable general model may perform better inside a particular company once it has been adapted to the organisation’s terminology, documents and processes. It may also be easier to test, cheaper to operate at high volumes and more predictable after deployment.

Benchmark performance still matters, especially for complex reasoning and technically demanding tasks. Yet the difference between the leading model and a credible alternative may have little commercial relevance if the system is primarily classifying internal documents, retrieving information from a controlled knowledge base or assisting employees with defined procedures.

A company needs the model that performs reliably within its own workflow, not necessarily the one that wins the broadest public comparison.

Self-hosting changes costs rather than eliminating them

The economics of running a model independently can appear attractive. An organisation is no longer paying a provider for every input and output, and high-volume applications may become less expensive once the required infrastructure is in place.

The calculation is more complicated than replacing an API bill with a server.

Models require computing capacity, storage, monitoring and technical maintenance. Security patches and model updates must be managed. Performance can deteriorate when the system encounters new types of data or when the organisation changes its processes. Someone must remain responsible for testing, access management and incident response.

Large models may also require expensive graphics processors, although smaller models can often handle narrowly defined corporate tasks with far lower infrastructure demands. The financially sensible choice therefore depends on the application. A public API may be cheaper for an occasional assistant used by a few employees. A self-hosted model may be more attractive for a system processing thousands of repetitive requests or handling information that cannot comfortably leave a controlled environment.

Hybrid architectures are likely to become common. A company might run a smaller model locally for confidential or routine tasks while using an external frontier model for occasional work requiring greater reasoning capacity. Requests can be routed according to their sensitivity, complexity and cost.

The result is not complete independence from external providers. It is a more deliberate allocation of dependence.

Customisation can create a new form of lock-in

The ability to adapt a model is one of the strongest arguments for self-controlled deployment. A company can connect it to internal knowledge, adjust its behaviour, refine it using specialist examples or optimise it for a particular task.

Every layer of customisation, however, can make replacement more difficult.

A model that has been fine-tuned extensively, connected to numerous internal systems and embedded in employee workflows cannot necessarily be exchanged without significant redevelopment. The organisation may avoid dependence on a hosted API only to become dependent on a particular model architecture, implementation partner or infrastructure provider.

Portability should therefore be considered before deployment begins. Companies need to know whether their prompts, evaluation data, retrieval systems and fine-tuning datasets can be reused with another model. Interfaces should be designed so that the model can be replaced without rebuilding the entire application. Performance tests should also compare several models against the company’s own tasks rather than assuming that one vendor will remain the best option.

A self-hosted model gives a company the possibility of greater independence. Technical architecture determines whether that possibility becomes real.

More choice does not remove the need for governance

The expansion of open models gives European companies a broader range of suppliers from the United States, China and Europe. It also weakens the assumption that the enterprise market will inevitably be dominated by a small group of closed American providers.

The underlying models are becoming internationally interconnected. Inkling reportedly draws on an architecture associated with DeepSeek V3 and training data produced by another Chinese model. The example shows how difficult it is to classify an AI system neatly by the location of its corporate headquarters.

For corporate buyers, the provider’s country of origin is only the beginning of the assessment. They also need to examine the model’s lineage, licence, training methodology, software dependencies and operating environment. A model offered by an American or European company may incorporate research, code or synthetic data originating elsewhere. Conversely, a model developed in China may be operated entirely within European infrastructure.

Self-hosting does not resolve questions about intellectual property, model provenance or embedded vulnerabilities. It transfers more responsibility for answering them to the client.

The enterprise AI market is becoming more modular

The significance of downloadable models extends beyond data protection. It indicates that AI is beginning to resemble an enterprise technology stack rather than a single service purchased from a single provider.

Companies may obtain the model from one developer, operate it through another infrastructure provider, use a separate platform for adaptation and appoint an implementation partner to connect it to internal applications. They can change individual components without necessarily abandoning the entire system.

This modular market is less convenient than opening an account with a hosted AI service. It also gives companies more negotiating power and greater freedom to design systems around their actual requirements.

The relevant procurement question is therefore no longer simply which model performs best. Companies must decide which parts of the AI system they are prepared to outsource and which they need to control themselves.

For many Swiss businesses, especially those handling regulated or commercially sensitive information, the answer will not be complete self-hosting or complete reliance on external APIs. It will be an architecture that keeps sensitive processes close, uses external services selectively and preserves the ability to change providers.

AI companies are allowing clients to run models themselves because corporate buyers have begun to ask for more than access. They want a credible route to control.

  Why More AI Companies Are Letting Clients Run Models Themselves