A bank just Open-Sourced AI Governance. What can enterprises learn from Santander?
Santander open-sourced 13 AI governance repos. Here is the three-layer enterprise governance stack that release fits into, and what your team should build.

AI governance at large enterprises used to live inside compliance binders. It is moving into the open. Santander quietly published 13 AI governance repositories to GitHub under a new organization called SantanderAI, licensed mostly under Apache 2.0. The repos are real engineering work. The signal is bigger than the code.
If you are running IT, innovation, or a Centre of Excellence at a regulated enterprise, this release is a calibration point. The bar for what "we have in place for AI governance" has moved, in public.
This post is the strategic frame around it: how the enterprise AI governance stack actually works, where Santander's release sits inside it, and what your team should be building for the layers Santander does not cover.
What Santander shipped
The SantanderAI organization is run by Santander AI Lab. The headlines repositories include:
- mech-gov-framework: mechanical governance for LLM decisions, with hard gates and entropy-based controls
- autoguardrails: a scaffold for LLM alignment research around policy surfaces
- mutatis-mutandis: situation testing for discrimination analysis using counterfactual comparators
- gen-fraud-graph: synthetic financial transaction graph generator for training GNN-based fraud and AML detection models
- sota-stressed-datasets: open benchmark datasets in stressed form to evaluate ML and LLM robustness
- auto-bayesian: config-driven Bayesian network training for relational tabular data
- causal-perception-implementation: causal-perception research code
Real engineering work. Not a marketing release. And already prompting other regulated institutions to ask the obvious question: what have we published, and what should we?
The enterprise AI governance stack has three layers
An enterprise governance strategy for AI does not sit on one axis. It has three, and each layer requires different tools, different owners, and different measurement.
Layer 1: Model governance. How the models your organization uses or trains behave, including alignment, fairness, hallucination controls, and hard gates on high-stakes decisions. This is where Santander's mech-gov-framework and autoguardrails sit. This is where your data-science and MLOps teams live.
Layer 2: App-level governance. How the individual AI-generated apps your teams ship actually behave, across security, compliance, reliability, maintainability, and commercial readiness. This is the layer that catches the gap between "the AI wrote the code" and "the code is safe to put in front of real users, real customers, and real regulators." This is where NEKOD's 360° review sits.
Layer 3: Portfolio governance. How the whole set of AI-generated apps across your organization is discovered, tracked, prioritized, and reported to the board. This is where a Centre of Excellence sits. It is the layer that answers "how many AI-built apps do we actually have in production, and what is our exposure?"
Most enterprises are heavily invested in Layer 1 and have very little at Layers 2 and 3. Santander's release makes Layer 1 more accessible, which is good. But it does not close the gap on Layers 2 or 3. That gap is where most of the actual risk lives.
Why the two layers matter the same
The Lovable breach, the Anthropic Mythos hold-back, and the Odido, Booking, and Basic-Fit consumer data exposures were all Layer 2 and Layer 3 problems, not Layer 1 problems. The models were fine. The apps and the portfolios were not.
For a regulated enterprise, the exam question at your next risk committee is not "which alignment framework are we using." It is:
- How many AI-generated apps are currently in production across our teams?
- Which of them handle personal, payment, or regulated data?
- What is our Launch Readiness Score across the portfolio?
- Who owns remediation when something is flagged?
- What does our audit trail look like?
These are Layer 2 and Layer 3 questions. Santander's release does not answer them for you.
What enterprise IT and innovation leaders should build
Three concrete moves for enterprise governance teams.
Adopt or adapt at Layer 1. Pick two of the SantanderAI repos most relevant to your stack (mech-gov-framework and autoguardrails are the strongest starting points for most organizations). Assign an engineer to read the code. Document what you would use as-is, what you would adapt, and what does not apply. Treat this as a calibration exercise for your model-governance program.
Stand up a Layer 2 assessment for every AI-generated app going to production. Five areas per app: security, compliance, reliability, maintainability, commercial. Context-driven, so the assessment adapts to what each app actually does. A landing page and a customer-facing fintech app produce different reviews because their stakes are different. This is what NEKOD's 360° review is built for.
Build the Layer 3 portfolio view as a Centre of Excellence. One team responsible for discovering AI-generated apps across the organization, running Layer 2 assessments consistently, reporting risk exposure to the board, and enabling internal builders instead of blocking them. This is the layer that turns shadow AI from a governance problem into a governed capability.
How we work with enterprise governance teams
NEKOD is built for Layers 2 and 3. We assess AI-generated apps across the five areas above, regardless of which platform they were built on (Lovable, Replit, Cursor, V0, Claude, or internal builds using open frameworks like the ones Santander published). We produce a single Launch Readiness Score per app and portfolio-level rollups for governance teams. Our dev team can pick up remediation, advisory or hands-on, using the same audit context.
Where enterprise teams have already built Layer 1 tooling (their own, adapted from Santander, or licensed from vendors), we sit alongside it. Model governance and app governance are two different problems. Both matter. Neither replaces the other.
We wrote up how this pattern applies to platform-level security in Replit's Security vs NEKOD. Same shape, different layer. Platform security and app readiness are two different problems. Model governance and app governance are also two different problems. The five most common code-layer issues that show up in AI-generated apps across the enterprise portfolio are in The 5 Security Gaps Hiding in Every Vibe-Coded App.
Key takeaways
- Santander's SantanderAI open source release is a signal that enterprise AI governance is moving from internal binders into public frameworks
- Enterprise governance has three layers: Model, App, and Portfolio. Santander's release strengthens Layer 1
- Layers 2 and 3 are where most of the actual risk for AI-generated apps sits, and they are the layers your risk committee will ask about
- A Centre of Excellence approach ties all three layers together and turns shadow AI into a governed capability
- NEKOD is built for the Layer 2 and Layer 3 gap: platform-agnostic, across five assessment areas, with a dev team to close the loop on findings
Our take
If you are building or running an enterprise AI governance program and want to see how NEKOD's Layer 2 and Layer 3 tooling maps to your existing Layer 1 investments, a consultation walks through the 360° review at portfolio scale and how a Centre of Excellence uses it in practice.


