What is model governance?

Quick answer

Model governance is the set of policies, processes, and decision rights that determine which AI models an organization is allowed to use, exactly how a model gets approved before it ever reaches production, how its actual behavior is tracked once it’s live and running, and who holds the actual authority to change, restrict, or fully retire it, functioning as the connective structure that turns individually sound practices around evaluation, guardrails, and access control into one coherent, organization-wide discipline, rather than a loose, scattered collection of disconnected, uncoordinated efforts each individual team happens to apply inconsistently and entirely on its own, with no shared standard tying any of it together.

Summary slides
Model governance
Why model governance is a distinct concern from any single system's…
How to structure decision authority so governance has genuine teeth
How model governance needs its own explicit connection to incident…
Common mistakes teams make around model governance

Why model governance is a distinct concern from any single system’s reliability practice

A team can build excellent evaluation, well-calibrated guardrails, and access control specifically for the one system it owns and still quite easily, leave the broader organization with unmanaged risk if nothing coordinates these efforts across every system relying on AI models throughout the organization, since individually sound engineering at the level of one system doesn’t automatically produce coherent oversight at the level of the organization as a whole. A second team, building a different system on a different model with no visibility into what the first team already learned, can easily repeat the same mistakes the first team already worked through, can adopt a model whose known limitations the first team had already discovered and carefully documented but never had any way to share, or can deploy a model version the organization has already decided, somewhere else entirely, carries a risk that hasn’t fully been resolved yet.

Model governance exists specifically and deliberately to close exactly this kind of coordination gap, providing the organization-wide visibility, the shared standards, and the centralized decision authority that no single team’s good practice can provide on its own, however well that team happens to be operating in isolation. This doesn’t mean centralizing every decision about how a system uses a model, since the engineering choices within a system still belong with the team that owns it, but it does mean establishing organization-wide answers to a smaller set of cross-cutting questions: which models are approved for use at all, what evaluation a model has to pass before any team can deploy it, what monitoring has to continue once it’s live, and who has the authority to restrict or withdraw approval when new information reveals a problem nobody previously knew about.

How model approval processes work in practice

A model approval process establishes a concrete checkpoint well before a model can ever be used in production at all, verifying that it meets an organization’s baseline standards for reliability, safety, and compliance before any individual team is free to build on top of it, rather than leaving each team to independently evaluate every model it’s considering entirely on its own, with predictably inconsistent thoroughness depending on how much time any team happened to have available for that evaluation.

This particular checkpoint needs concrete evaluation criteria specifically, deliberately defined well in advance, rather than an essentially ad hoc judgment call made somewhat differently each and every time a new model comes up for review, since a consistent process requires the same baseline questions asked of every model under consideration: has it been tested against the organization’s representative tasks, what’s known about its failure modes and its limitations, what data was it trained or fine-tuned on and does that data raise any concern, and what’s the actual vendor or provider’s track record for transparency about the model’s behavior and its known issues. A model that successfully passes this checkpoint deserves to be carefully recorded in an accessible, organization-wide registry, giving every team a single, authoritative place to check whether a model is approved, under what conditions, and with what known limitations already documented, rather than each team having to independently rediscover this same information on its own.

The approval process also needs a concrete path specifically for exceptions, since a rigid process that only ever accepts a small, fixed list of pre-approved models will eventually block some team from a legitimate use case a general approval process never specifically anticipated, and a governance structure with no exception path tends to get quietly bypassed under deadline pressure rather than respected, which defeats the entire purpose of having the process in the first place. A well-designed exception path lets a team request approval for a narrower use case even when the general model itself hasn’t cleared full, organization-wide approval, provided that narrower use goes through its appropriately scoped review rather than being treated as an unreviewed workaround that simply avoids the broader process altogether.

How ongoing monitoring differs from the initial approval gate

A model approved just once and then never monitored again quietly, silently drifts out of the very same standards that originally justified its initial approval in the first place, since the model itself can change through vendor-side updates, the actual tasks it’s being used for within the organization can shift well beyond whatever was originally evaluated, and the broader landscape of known risks around a model can evolve as new research or new incidents surface information that simply wasn’t available at the time the original approval decision was made.

Ongoing monitoring needs to track a model’s actual, real-world performance across every system that’s using it, rather than merely relying passively on the vendor’s separate claims about that same model’s behavior, since a model that performs well on a vendor’s published benchmarks can still perform considerably worse on an organization’s tasks, and only monitoring tied to actual production use will ever catch that particular gap. This monitoring also needs to specifically track incidents and near-misses across every team using a model, aggregated at the organization level rather than staying siloed within whatever individual team happened to encounter a problem, since a pattern that looks like an isolated, one-off issue within a single team’s experience can reveal itself as a systemic problem with the model itself once the same pattern is compared across several teams independently reporting something similar.

Version tracking deserves particular, dedicated attention within this ongoing monitoring, since a vendor’s model update can shift behavior in ways that meaningfully invalidate an earlier approval decision entirely without the organization ever explicitly re-approving anything, and a governance structure with no mechanism for tracking which model version each team is running, and no trigger for re-evaluation when a vendor pushes a significant update, will eventually let an unapproved, unreviewed model version quietly operate in production with nobody having consciously decided that version was acceptable.

How to structure decision authority so governance has teeth

A governance policy that carefully identifies standards but assigns nobody the actual authority to enforce them is, in practice, really just a recommendation rather than a policy at all, and model governance specifically needs a clear, unambiguous answer to a deceptively simple question: who can really say no, and to precisely what, when a team wants to deploy a model or continue using one that a governance review has flagged as a concern worth taking seriously.

This particular authority needs to be concrete rather than vaguely, uselessly diffuse, since a governance structure that nominally sits with an entire committee, with no individual member of that committee accountable for a decision, tends in practice to produce exactly the kind of consensus paralysis where nobody feels personally responsible for blocking a launch, and a decision eventually gets made by default, through inaction, rather than through any deliberate judgment anyone exercised. Assigning clear, individual accountability for each category of decision, exactly who approves a new model for organization-wide use, who can grant a properly scoped exception when one’s warranted, and who can order an existing, live deployment restricted or fully withdrawn once new evidence justifies it, gives governance the same kind of traceable accountability chain discussed as essential to responsible AI more broadly, rather than leaving these decisions to an informal, collective process that dissolves the moment anyone tries to trace who specifically made a call.

This same authority also urgently needs well-defined escalation paths built directly in, specifically for the cases where a team honestly disagrees with a particular governance decision, since a governance structure with no appeal mechanism tends to accumulate quiet resentment and quiet workarounds from teams who feel a decision was wrong but have no legitimate channel to contest it, whereas a structure with a respected escalation path lets disagreement surface and get resolved through the process itself rather than through teams simply working around governance entirely once they’ve concluded, rightly or wrongly, that the process itself isn’t going to listen to their concern.

How model governance needs to scale with organizational size and model diversity

A small, comparatively modest team building purely on top of a single model doesn’t need the exact same, heavyweight governance apparatus that a considerably larger organization running dozens of different models across many different teams truly requires, and treating model governance as a single, fixed structure that every organization needs to adopt identically, regardless of its actual scale and its actual diversity of AI use, produces either crushing, unnecessary overhead for a small team that never really needed it, or seriously inadequate oversight for a large organization that critically, urgently did need it.

A small team can often satisfy the underlying goals of model governance, shared standards, centralized visibility, and clear decision authority, through comparatively quite lightweight, informal practice: a shared document tracking which models are in use, a regular review meeting where the team collectively discusses any concerns that have come up, and a single person or a small group with clear, explicit authority to make the relevant calls. A larger organization running many different models across many separate teams needs considerably more formal, structured infrastructure specifically to achieve those exact same underlying goals, only now at that considerably larger, more complex scale: a properly maintained model registry with well-structured and consistently accurate metadata attached, a truly dedicated governance function staffed with specialized expertise, and formal, carefully documented review processes that hold up under scrutiny that can scale to match the volume of models and the volume of teams a large organization has to coordinate across day to day.

The particular structure chosen matters considerably less than whether the underlying goals are being met in practice at all, and a team evaluating its governance maturity should ask directly whether it has organization-wide visibility into which models are in use, whether it has shared standards those models are evaluated against, and whether decision authority is clear and respected, rather than asking whether it has adopted some formal governance framework that might be considerably heavier, or considerably lighter, than what its actual scale requires.

How model governance connects to vendor and third-party model risk specifically

Most organizations don’t train their models entirely from scratch, they instead build directly on models provided by external, third-party vendors, which introduces a distinct category of governance concern beyond what an organization’s internal engineering practice alone can address, since a vendor’s decisions, about training data, about model updates, about the vendor’s security practices, directly shape risk the organization is ultimately exposed to but has comparatively limited direct visibility into or direct control over.

Vendor evaluation, specifically and carefully assessing a provider’s track record for transparency, for responsible handling of known issues, and for responsiveness whenever a customer reports a problem, deserves to be built into the model approval process discussed earlier as its explicit, dedicated component, since the same underlying model can carry meaningfully different risk depending on whether its provider is forthcoming about known limitations or is inclined to minimize and downplay problems a customer directly raises with them. Contractual protections, specifically and deliberately negotiating directly for advance notice of any significant model update, for transparency about training data and known limitations, and for a defined process when a vendor’s model is found to have a problem, give an organization contractual leverage beyond merely hoping a vendor happens to behave responsibly on its own initiative.

Concentration risk, specifically the risk of an organization depending far too heavily on a single vendor’s models across far too much of its actual AI use overall, deserves its explicit, dedicated governance attention in its own right, since a significant problem discovered in one vendor’s models, whether a reliability issue, a security vulnerability, or simply the vendor’s business circumstances changing in a way that affects service availability, has outsized impact on an organization that’s concentrated most of its actual AI use around that single provider, relative to an organization that’s deliberately maintained some meaningful diversity across its model dependencies specifically and intentionally to avoid exactly this kind of concentrated, correlated exposure in the first place.

How to build governance that gets used rather than quietly bypassed

A governance process that teams experience as pure, unmitigated friction, adding delay and overhead without ever providing any perceived value back to the teams subject to it, tends to get worked around in practice regardless of how formally, explicitly it’s documented as mandatory, since a determined team under deadline pressure will generally find a way to route around a process that feels like pure obstruction rather than worthwhile protection.

Building governance that teams want to use, rather than governance they merely, grudgingly tolerate because they have no other practical choice, means the process itself needs to provide value directly back to the teams going through it: a model registry that helps a team quickly find a model already approved and already suited to their need, an evaluation process that surfaces useful information about a model’s limitations the team would otherwise have had to painstakingly discover entirely on their own, and a review that catches a problem before it becomes the team’s considerably more painful, production incident later on. Governance that’s fast specifically where speed is safe, approving well-understood, already thoroughly-vetted models quite quickly while reserving deliberate, careful scrutiny specifically for novel or higher-risk cases instead, respects a team’s legitimate need to ship work, rather than treating every single model use as equally deserving of the same slow, heavyweight review regardless of how well-understood or how low-risk that use already is.

Measuring whether a governance structure is truly working in practice, carefully tracking compliance rates, tracking how often teams route through the official exception process versus how often a workaround is later discovered operating quietly outside it, and tracking teams’ honest, direct feedback about whether the process is providing perceived value, gives an organization empirical evidence about whether its governance structure is functioning as intended, rather than relying entirely on the comfortable but unverified assumption that a policy documented on paper is automatically the policy being followed in the messier, more complicated reality of how work gets done day to day.

How model governance needs its explicit connection to incident response

When something does go wrong with a model already running in production, a security issue, a reliability failure, a fairness or privacy problem, model governance needs an explicit connection to incident response specifically, since diagnosing whether an incident is isolated to one team’s particular use or reflects a broader, more systemic problem with the model itself requires exactly the organization-wide visibility that model governance, done well, is specifically designed to provide.

An incident response process specifically for model-related issues needs a clear, well-defined path for escalating a single-system problem directly up to the broader, organization-wide governance function whenever there’s reason to suspect the underlying issue might extend meaningfully beyond that one system, since a team debugging its single incident in isolation has no direct visibility into whether other teams using the exact same model have quietly encountered something similar, and only a centralized, organization-wide view can connect those separate dots into a coherent, complete picture of what’s going on across the organization as a whole. This connection also needs to run in the other direction just as importantly, with governance actively, proactively pushing incident information back out to every team using a model once a problem has been confirmed and understood, rather than leaving each individual team to somehow discover on its own, entirely independently, that a model it’s currently, actively relying on has a known, documented issue other teams already identified and already reported through the exact same governance channel weeks or months earlier.

How model governance needs to account for fine-tuning and derived models specifically

Governance built only for the models an organization approves directly from a vendor misses an important, distinct category of risk introduced quite specifically by fine-tuning, adaptation, and other derived models built on top of those approved foundations, since a team that takes an already-approved base model and fine-tunes it on their own internal data has, in a meaningful sense, created a new model whose behavior may no longer resemble whatever the original approval decision evaluated.

A fine-tuned model deserves its entirely separate review, distinct from whatever approval the underlying base model itself already, originally received on its own terms, since fine-tuning can meaningfully shift a model’s behavior in ways the original evaluation never anticipated, introducing new failure modes, new fairness concerns, or new privacy risks specifically tied to whatever internal data the fine-tuning process used, none of which the original base-model approval could have possibly accounted for since that data simply didn’t exist as part of the model at the time the original decision was made. Treating a fine-tuned model as automatically, implicitly covered by the base model’s existing approval, with no distinct review of its own, is exactly the kind of gap that lets new risk enter production without ever passing through the governance process specifically designed to catch it.

This exact same underlying principle extends to any system that meaningfully, substantively wraps or adapts a base model in some way, a retrieval-augmented system built on top of an approved model, an agentic system layering tool use on top of an approved model, each of which introduces new behavior and new risk surface beyond whatever the base model’s original approval covered, which means model governance needs an explicit answer to the question of when a system built on an approved model needs its additional, dedicated review, rather than assuming every system built on an already-approved foundation automatically inherits that same approval without any further scrutiny at all.

How to build governance documentation that stays current

Governance documentation that’s accurate on the day it’s written, and then never revisited afterward, becomes actively misleading over real time, since teams relying on outdated documentation to understand which models are approved, under what conditions, and with what known limitations, will make decisions based on information that’s quietly, silently stopped reflecting the organization’s actual, current state of affairs.

Treating the model registry and its associated documentation as a living, actively maintained artifact that gets updated as a natural, ongoing part of the normal, everyday workflow, rather than as a static reference occasionally revisited during a periodic audit, is what keeps it trustworthy over the long run, since documentation updated only sporadically, disconnected from the actual, ongoing decisions it’s meant to reflect, drifts out of sync exactly as quickly as any other system property that’s monitored infrequently rather than continuously. Building the update itself directly into the actual approval workflow, so that a model’s registry entry gets created or updated as a natural, entirely unavoidable part of the approval process itself rather than as a separate, easily-forgotten task someone has to remember to do afterward, removes the reliance on anyone’s individual diligence to keep documentation honest and current.

Periodic audits still deserve a concrete place within this broader practice, specifically to catch the actual drift that inevitably, quietly accumulates even within an otherwise well-designed workflow, comparing what the registry claims against what’s running in production and reconciling any discrepancy that’s found, since even a workflow carefully designed to keep documentation current will eventually miss some edge case, some exception handled informally outside the normal process, or some model deployed under time pressure with the registry update simply deferred and then, in the pressure of whatever came next, never properly completed at all.

Common mistakes teams make around model governance

A first mistake, and the foundational one nearly every other mistake on this list traces back to in some form, is building excellent reliability practice within individual systems while leaving absolutely no organization-wide coordination across them, letting different teams independently repeat the same mistakes or adopt models whose known limitations were already discovered and carefully documented, but never shared, elsewhere within the very same organization.

A second mistake is running model approval as an essentially ad hoc judgment call made somewhat differently each and every time, rather than against consistently applied evaluation criteria that were defined explicitly well in advance and applied identically to every single model that comes up for consideration.

A third mistake is building an approval process with no exception path built into it at all, leaving that process either rigid enough to actively block legitimate use cases, or vulnerable to being quietly bypassed under deadline pressure once teams eventually conclude the process itself has no room whatsoever for their legitimate need.

A fourth mistake is treating model approval as a purely one-time gate with no ongoing monitoring at all afterward, letting a model quietly drift out of the very standards that originally justified its approval as the model itself, its actual usage, and the broader, evolving landscape of known risk surrounding it all continue to change over time.

A fifth mistake is tracking incidents and near-misses only within whatever individual team happened to encounter them, entirely missing a systemic pattern that only becomes visible once similar reports are compared directly across several separate teams at the full organization level.

A sixth mistake is having no mechanism at all for tracking exactly which model version each individual team is currently running, letting a significant, entirely unreviewed vendor update quietly operate in production with literally nobody ever having consciously, deliberately re-approved that new version.

A seventh mistake is assigning governance authority diffusely across an entire committee with no individual member personally accountable for any decision, producing exactly the kind of consensus paralysis where a consequential decision effectively ends up getting made by default, through simple, unremarked inaction.

An eighth mistake is building a governance structure with no escalation path or appeal mechanism whatsoever, leaving disagreement to quietly accumulate as quiet resentment and quiet workarounds rather than openly surfacing and getting resolved through the governance process itself as it was originally intended to work.

A ninth mistake is adopting a governance structure sized for an organization considerably larger, or alternatively considerably smaller, than the organization adopting it, producing either crushing, unnecessary overhead nobody ever needed in the first place, or inadequate oversight relative to the actual scale of AI use that organization has.

A tenth mistake is relying entirely and exclusively on a vendor’s claims about a model’s actual behavior, with no independent evaluation ever run directly against the organization’s actual representative tasks, missing the meaningful gap that often exists between published benchmark performance and real-world performance on that organization’s particular work.

An eleventh mistake is concentrating an organization’s entire AI use tightly around a single vendor with no deliberate diversification whatsoever, leaving that organization with outsized, concentrated exposure to whatever problem that single vendor eventually, and quite inevitably, ends up encountering at some point down the line.

A twelfth mistake is building a governance process that teams experience purely as friction, with essentially no perceived value ever returned to the teams subject to it, virtually ensuring it eventually, quietly gets worked around regardless of how formally, officially it’s documented as mandatory somewhere on paper.

A thirteenth mistake is treating a fine-tuned or otherwise derived model as automatically, implicitly covered by its base model’s existing approval, with no distinct review of its own required at all, letting new risk introduced by the fine-tuning data itself, or by whatever wrapping system sits around it, quietly enter production without ever passing through governance in any meaningful sense.

A fourteenth mistake is treating governance documentation as though it were merely a static reference, revisited only occasionally, during periodic audits, rather than as a living, actively maintained artifact updated as a natural, unavoidable part of the approval workflow itself, letting the registry quietly, gradually drift out of sync with whatever’s running in production at any moment.

A fifteenth and truly final mistake is skipping periodic audits entirely, on the comfortable assumption that a well-designed workflow alone is sufficient on its own, missing the edge cases, the informal, undocumented exceptions, and the deferred updates that inevitably, quietly accumulate even within a well-designed governance process given enough sustained time.

What connects all fifteen of these mistakes is treating model governance as a bureaucratic formality layered on top of engineering work, rather than as the coordination structure that turns individually sound practice at the level of a single system into coherent, organization-wide oversight of every model an organization depends on. Organizations that build evaluation standards, ongoing monitoring, decision authority, and value back to the teams going through the process tend to catch model-related problems before they become full-scale, organization-wide incidents, while organizations that treat governance as paperwork tend to discover, usually during an actual incident that governance was specifically supposed to prevent, that the coordination they assumed was happening across their organization was never happening at all.

The deeper principle underlying all of this is that coordination, at organizational scale, doesn’t happen automatically simply because individual teams are each acting responsibly within their particular scope, since responsible local behavior has no inherent mechanism for aggregating into responsible organizational behavior unless something explicit is built to connect the two together. A dozen teams each independently evaluating the same model, each independently discovering the same limitation, each independently working around the same known issue, represents a considerable waste of effort that good local practice alone can never solve, precisely because good local practice, by its very nature, stays local unless a deliberate structure exists specifically to lift what any one team learns up to the level where every other team can benefit from it too.

This is exactly why model governance deserves to be understood as an extension of the same underlying discipline that runs through every other part of responsible, reliable AI system design, treating coordination itself as a first-class concern with its explicit structure, rather than as something that will simply emerge naturally once enough individual teams happen to be doing their individual jobs well. Organizations that build this coordination deliberately, with registries, evaluation standards, decision authority, and feedback loops connecting incidents back to the teams who need to know about them, tend to get considerably more value out of their AI investment than the simple sum of what each individual team could have achieved entirely on its own, precisely because problems get caught once, at the organizational level, rather than being independently rediscovered, team by team, at whatever painful cost each individual rediscovery happens to carry.