Two numbers sit on the same page of a health plan services organization’s published program record, and between them they describe the state of AI in regulated industry better than any conference talk this year.
The first is a build figure. A CMS Enhanced Direct Enrollment platform for brokers and agents was delivered on what the record calls a six-week regulated build pathway, at “approximately 650–900 delivered hours against an 8,500–10,500-hour conventional estimate”.97
The second is further down the same page, under evidence: “1 review candidate, 11 pending captures, 1 source pending, and 0 export-ready — of 13 FIT cases”.97
The build compressed by roughly a factor of ten. The auditor’s package is at zero.
Here is the thesis: in a regulated business the binding constraint on AI-native delivery was never capability, and it is not capability now. It is evidence — the one part of the work that agents did not make cheaper.
Any team can move fast with agents. Very few can answer, on demand and without a person reconstructing it afterward, which model saw the work, whether that model was permitted to see it, what the run actually did, and who accepted the result. Those four questions are what a security review exists to ask, and an organization that cannot answer them does not get to keep its speed. It is why “become AI-native” travels so badly into healthcare from the industries the advice was written in, where nobody has to answer for the output.
What follows is drawn from that record, published by Wipro Health Plan Services across a portfolio register, a delivery standard, a traceability posture and a set of product dossiers. It is first-party and self-reported, which is a limit on it and also why it is usable: every assertion below links to the page it came from, and where the record marks its own figures as unvalidated, that marking is repeated rather than dropped.
Decide where the model runs before deciding what it does
Almost every transformation plan opens with a use case. This one opened with hardware. The first phase of the record, dated 2024, is a sentence about physical infrastructure: “Three on-premises NVIDIA DGX systems established a protected place to run AI.”100 Before a member-facing use case existed, the question a security review opens with already had an answer.
The ordering is the transferable part. Most published advice puts the data-boundary decision after the pilot, as a hardening step. In a health plan it cannot go there, because the boundary decides which use cases are available at all. The definition doing that work is not an architecture preference. A business associate, under the HIPAA rules, is a person who on a covered entity’s behalf “creates, receives, maintains, or transmits protected health information for a function or activity regulated by this subchapter”.106 That is a description of an inference endpoint. Which models may see what is a contracting question before it is an engineering one, and the pilot that answers it late answers it in front of counsel.
The architecture that followed reads as a sequence of boundary decisions rather than model decisions. Identity and an allowed context scope the data before anything is retrieved. Retrieval returns source-backed context with citations. A model gateway holds routing policy and treats the models themselves as replaceable. Inference runs inside a protected local runtime.102 The model is the swappable component in that list, which is the correct place for it and the opposite of how most programs are scoped.
Skipping this stage has a specific failure signature. Every pilot re-argues the same question in front of the same reviewers, on its own terms, and each one loses on a schedule nobody planned for. A model whose data boundary is undecided is not a governed system. It is an open question with a budget attached.
Settling it is also not sufficient, and the record says so about itself. Its target-state phase lists what remains to be evidenced for model governance: an approved model inventory, runtime configuration, training or fine-tuning standing, dataset lineage, and evaluation baselines.100 Deciding where inference happens answers one question. It does not produce a model register.
The obligation to hold one is also moving, and in a direction worth watching. Certified health IT today must supply nine categories of source attribute for any predictive decision support intervention it carries — intended use, the inclusion and exclusion criteria that shaped the training set, the external validation process, quantitative measures of performance in local and external data, and the schedule on which the thing is revalidated — plus a risk analysis covering validity, reliability, robustness, fairness, intelligibility, safety, security and privacy.107 In December 2025 ASTP/ONC proposed to “fully remove” those requirements: every source attribute, and the risk-management provision beside them.108 That proposal is not final. If it lands, the model card does not stop being asked for; it stops being produced by the certification body and starts being owed by whoever deployed the model.
Prove it where the control model already exists
The first production workload was member service. The record reports five call types on the AI contact-center path and, for the late-2025 production stack, 3.6 million member interactions at 99.99% uptime — figures it labels “context, not independently validated”.102 That label is the sentence most vendor pages would not write about their own headline number.
Volume is not what makes member service the right place to start. It is that member service already had a control model, and the AI was installed inside it rather than beside it. Identity and consent scope the work before retrieval happens. The system proposes a response. A service representative retains the decision to resolve, edit or escalate before anything is written back. Incomplete identity, a source mismatch, low confidence or a privacy policy violation returns the work to manual service.102
None of that was designed for AI. It is the service control model that already existed, with a new proposer inside it, which is precisely why it could be reviewed at all. Where a determination touches a member’s coverage, the human decision is not a courtesy laid over the automation. CMS told Medicare Advantage organizations in February 2024 that “for inpatient admissions, algorithms or artificial intelligence alone cannot be used as the basis to deny admission or downgrade to an observation stay”.109
Where that constraint comes from is the instructive part. CMS proposed binding AI guardrails for Medicare Advantage in December 2024 and then declined to finalize them, saying it would continue to consider whether future rulemaking is appropriate.110 The February guidance is sub-regulatory, and what actually stops the algorithm is the ordinary medical-necessity rule applied to a new participant. An organization waiting for AI-specific regulation to tell it what is permitted has misread which rule binds it, and will keep waiting while the existing one is already being enforced against its output.
The failure mode at this stage is quiet, because the pilot succeeds. A use case chosen because it presents well, in a domain with no existing control model, produces a result that generalizes to nothing: the controls were invented for it, and they leave with it. What an organization needs out of its first workload is not a win. It is a control model it can use again.
Make the delivery method the change, not the purchase
The methodological core of this record is a delivery standard rather than a tool. It sets seven canonical lifecycle phases, thirteen security and compliance gates, an accountable human role named at each phase, and six agent roles that exchange structured artifacts under role contracts.98 Its statement of intent is one of the more useful sentences published on the subject: “Intent enters once. An independently verified product, evidence package, and named human release decision leave together.”100
Read that as an output contract. A change does not leave the system as code with a description attached. It leaves as three artifacts produced in the same motion, one of which is a person’s decision. An organization that cannot state its own version of that sentence has not defined what finished means, and no amount of agent throughput will define it for them. NIST’s AI Risk Management Framework makes the same structural claim from a standards body’s side: governance is designed as a cross-cutting function, infused throughout the other three rather than standing beside them as a stage.105
The common alternative is a purchase. Licenses arrive, engineers get faster, and nothing about the acceptance path changes. A year later the organization has more code, a better throughput number for the board, and exactly the same inability to say who accepted what.
Verify at the point of generation, not after it
The strongest line in the record is a criticism of what most enterprises have already bought: “A log that records a violation has already permitted it.”99
The doctrine sorts verification by timing rather than by subject. Preventive means the agent cannot — policy as code, tool allowlists, scoped credentials, protected paths, data boundaries. Inline means checked while working rather than after. Gate means nothing ships unverified, closing at a named human authority. Continuous is described as the last line rather than the control, because it only observes.99 Twelve named verifiers sit across those classes, covering approved-context packages, credential scope, PHI boundary scans, protected-lane guards, intent against implementation, evidence completeness, and release-authority sign-off.98
The rule that makes any of it load-bearing is the independence one: “No agent verifies its own output. A verifier must be independent of the generator — a different model lineage, or a deterministic tool.”99 The record’s own name for this is segregation of duties applied to a new kind of actor, which is the reason an auditor will recognize the control without being taught a new vocabulary.
It is also not a novel demand. NIST separates, as a stated best practice, the actors building and using a model from those verifying and validating it.105 The doctrine is extending an existing rule to a participant that did not exist when the rule was written.
What is actually being drawn here is the line between observability and control, and most AI governance programs have not drawn it. The EU AI Act requires high-risk systems to “technically allow for the automatic recording of events (logs) over the lifetime of the system”, and requires those logs to be kept for at least six months.111 That is a genuine obligation, and it is satisfied entirely by the continuous class. Nothing in a record-keeping obligation prevents anything.
Instrument traceability before the first commit
This is where the two opening numbers reconcile, and where the money is.
The traceability row in this program is not a link between a ticket and a commit. It carries the CMS source — a requirement, toolkit, control, FIT row or change request — alongside the agile plan that owns it, the controlled delivery it became, and the proof and release state: test kit, case ID, security gate, readiness, open gap, disposition.101 Raw evidence stays in a controlled repository under a named evidence owner; the matrix carries readiness metadata and reviewer state rather than the artifacts themselves.
Then the sentence worth taking whole: “Traceability is complete; auditor evidence readiness is a separate status.”101 Twenty-one of twenty-one rows are mapped, and no packet is export-ready against thirteen test cases.97 Those two facts are not in tension. They answer different questions, and most organizations have a word only for the first.
Retrofitting the second is where an agentic program hands the savings back. A capture that was not taken during the run has to be reconstructed later by a person who was not there, out of a system that was never asked to keep it. If the build costs 900 hours and the package costs a specialist the better part of a year, the program’s economics are the old economics wearing a new shape. Evidence produced as a byproduct of the run is the only version of this that compounds, and it has to be decided before the first commit, because that is the last moment it is free. NIST’s phrasing for why is that measurement “provides a traceable basis to inform management decisions”.105 A basis assembled afterward out of recollection is not traceable, whatever the column heading says.
Keep the acceptance human, and keep it named
The record names a release authority as a role and states that human authority is non-delegable to agents at the release gate.98 Its model release contract runs five steps: register, evaluate, independently verify, human approve, then observe or revoke.104 The last step is the one usually missing, and it is the one that makes the others mean something — an approval that can be withdrawn is an approval that was real.
The advisory firms have arrived at the same shape from the other direction. BCG Platinion describes delivery where as few as three engineers run the factory and every stage gate still has a human accountable for approval.58 Forrester keeps the human accountable while AI does more of the execution.60 PwC’s agent-governance guidance reads like an auditor’s checklist: a verified identity, a defined role, task-specific permissions, auditable activity records, and clear limits on autonomous action.59 All three describe fewer people producing, with no reduction in what has to be proved.
In most estates today the agent has a name and a token and none of the rest. An acceptance either has a person against it or it did not happen, and it is the first record anyone from outside will ask for.
Publish a baseline, including what has not moved
The financial claim in this record is stated as a model and labeled as one. The legacy estate runs at $55.0M a year, $1.28 per member per month, against a $26.5M and $0.60 target and a 52% run-rate reduction goal.100 The caveat arrives in the same breath: a January 2026 modeled scenario rather than realized savings, with value recognized only after operating change and verified retirement.104 Mainframe exit is described as sequenced on evidence rather than on calendar.100
The portfolio register does the same thing with status. It lists twenty-two initiatives — two live or operational, eleven in active delivery, two decision-dependent, eight queued — with a note that the categories overlap and do not sum to the total.103 The record’s phrase for that arrangement is that mixed status is visible rather than averaged away.100
A baseline stated that way can be held against you later, which is exactly what makes it a baseline. A vendor’s percentage cannot, which is what makes it marketing. METR’s randomized trial is the standing warning about the substitute: experienced developers working in their own repositories were 19% slower with AI tools while believing they had been 20% faster.61 METR has since redesigned that experiment, reported a later estimate pointing the other way, and cautioned that selection effects leave it “only very weak evidence”.112 Read the two together and the finding is not that AI slows people down. It is that a team’s belief about its own speed is not a measurement in either direction, and it is the thing most programs are currently reporting.
Where this genuinely costs speed
Thirteen gates and twelve verifiers on every change is a tax, and a piece that will not say so is an advertisement. The doctrine’s own concession is modest enough: verification “slows a single generation step and speeds up everything after it”.99
The larger concession is sitting on the same page as the headline number. A build that compressed by roughly an order of magnitude has produced no export-ready evidence packets, inside a program that had already written seven phases and thirteen gates before that build started. Read without sympathy — and a reviewer will read it without sympathy — that says the governance apparatus did not keep pace with the delivery it was governing. The honest response is not that the framework is wrong. It is that on this record the evidence stage has not been shown to get cheaper, and nobody should say that it has until a package closes and an independent reviewer dispositions it.
There is a second cost that is structural rather than temporary. A control that can block is a control that can block the wrong thing, at an hour when the person who understands the policy is asleep. The doctrine’s answer is that a verifier is “a named, independently-owned check with a defined trigger, a defined scope, the authority to block, and a recorded disposition”99 — which is right, and which means a person who has to be reachable. That is headcount, not software. A program that budgets for the verifiers without budgeting for their owners has bought a system that fails closed with nobody to open it.
The strongest form of the objection is that plenty of organizations should not do any of this. Where the work carries no evidence obligation, gates are pure overhead and the team that skips them wins on every dimension that matters to it. Nothing above is general advice about building software. It is the price of operating somewhere a regulator can ask, and the industries where “become AI-native” works as a slogan are largely the ones where the question never comes.
What the record has to hold
We built SprintLoop around the record rather than the run. A model register holds which models an organization approved, where each one runs — hardware the organization controls, a tenant-isolated deployment under contract, or a public API — and whether work carrying regulated or member-identifying content may reach it, which is its own decision and starts at no. An approval with no stated basis fails the write, and withdrawing one records the withdrawal instead of removing the row, so a run that happened under it stays explainable.
A run register records what a run was asked to do, which approved model it ran under, and its outcome at each stage and in each parallel lane, down to the commit it landed at. Policy verdicts are sealed with a digest over the bytes they covered, in a table with no update path: a correction is a new run, and a policy with no verdict against a run reads as not evaluated rather than as passed. SprintLoop does not run the agents, hold a model key, or accept anything on anyone’s behalf, and it holds none of the security certifications a reviewer will ask about — the security page names them.
That program’s record shows a regulated build that got roughly ten times cheaper in hours, beside an evidence package that has not moved. Closing the second gap is what becoming AI-native in a regulated business actually consists of, and it is the half a regulator will actually look at.