
For most of the last three years, building with AI meant choosing a frontier model, calling an API, paying per token, and letting someone else's infrastructure do the work.
That was the right default in 2023. In 2026, a new stack is emerging.
Open-weight models have reached the point where many enterprises no longer need to trade control for capability. The important change is that they are now good enough to run production workloads. The hard questions are no longer about model performance, they are about where models run, how they access data, and whether their answers can be trusted.
This shift did not come from a single release. It came from steady improvements across the open ecosystem.
A year ago, self-hosting usually meant accepting a quality penalty. That tradeoff has largely disappeared. Model families such as GLM, DeepSeek, Qwen, Kimi, MiniMax, Mistral, Nemotron, and Gemma now compete with leading closed models across reasoning, coding, long-context analysis, and agentic workflows. On many benchmarks, the gap between the strongest open and closed models is small enough that it is no longer the deciding factor for enterprise adoption.
Two developments made this possible.
The first is licensing. Many of the models enterprises want to build on now ship under permissive licenses such as Apache 2.0 or MIT. Companies can deploy, fine-tune, and commercialize them without worrying about restrictive terms, user caps, or unexpected licensing changes. For organizations building AI into core products, that predictability matters more than a small difference on a benchmark.
The second is the surrounding ecosystem. Production inference runtimes, quantization, optimized serving infrastructure, and practical fine-tuning pipelines have turned self-hosting into standard engineering work. Running a frontier model inside a private environment is no longer a research project.
As model quality converges, deployment becomes a business decision. Three factors increasingly drive that decision: control, compliance, and cost.
Organizations handling financial, healthcare, or other sensitive data are understandably reluctant to send that information through third-party APIs. Even with strong contractual protections, critical data still passes through infrastructure they do not control. Running models inside a private environment keeps governance, auditing, and security within the organization's own boundary.
Compliance follows naturally. Data residency, sovereignty, and industry regulations are easier to satisfy when data never leaves the environment in which it already resides. Rather than depending on vendor policies, compliance becomes part of the system architecture.
The economics also change. API pricing is attractive during experimentation but scales directly with usage. As AI becomes embedded across more workflows, infrastructure costs become easier to forecast and often less expensive than paying for inference one request at a time.
None of this requires building a hyperscale data center. Modern open models run effectively in private clouds, on-premises environments, or fully air-gapped deployments. Many organizations still use external APIs for lower-risk tasks while keeping sensitive workloads inside their own infrastructure.
Deploying a frontier model locally is no longer the difficult part.The difficult part is making them answers trustworthy.
A language model knows nothing about your business.
It does not know what net revenue means at your company. It does not know how your general ledger maps across reporting entities, which operational system is the source of truth for a customer, or why a restatement changed how a metric should be calculated.
Point a model directly at a data warehouse and it will often produce answers that sound plausible but are wrong. The problem is not reasoning ability. The model simply lacks the business context needed to interpret the data correctly.
This is where many enterprise AI projects slow down. The model is capable of reasoning, but it is not a source of truth about the business.
What closes that gap is governed context. Metrics, entities, relationships, lineage, and access policies give the model the information it needs to reason correctly. Without that layer, the model is matching patterns across raw tables. With it, the model can answer questions using the same business definitions that finance and operations already trust.
Preql is an agentic semantic layer built for enterprise finance. It sits between governed enterprise data and the models running inside a company's own infrastructure.
Rather than asking a model to infer what net revenue means from thousands of database columns, Preql provides the governed definition, the relationships between entities, the business logic behind the metric, and the lineage required to explain every answer.
The result is not simply better responses. It is AI that produces numbers finance teams can validate, audit, and defend.
Because Preql runs alongside enterprise data and enterprise models, organizations keep the same security, sovereignty, and compliance benefits that motivated private deployment in the first place.
We have believed this from the beginning: governed, semantically defined data is not an enhancement for enterprise AI, it’s a prerequisite.
First, frontier-capable open-weight models under permissive licenses.
Second, deployment inside private cloud, on-premises, or air-gapped infrastructure so data remains within the organization's security boundary.
Third, a governed context layer that gives those models an accurate understanding of the business they are reasoning about.
Open models have made high-quality reasoning widely available. The remaining challenge is grounding that reasoning in enterprise context. For organizations deploying AI at scale, that’s the only way to get past the proof of concept stage and start delivering.

