Finance resilience | September 23, 2026

Count failure domains, not AI vendors

A second model logo does not create resilience when both services depend on the same cloud, identity provider, data pipeline, vector store, safety layer, network route, or two people who know how the workflow runs. Finance should map the critical business service end to end, set a measurable disruption tolerance, and prove a fallback or exit path under time pressure. Untested portability is an assumption, not evidence of operational continuity.

Service dependency graph DORA-style exit evidence Timed disruption tests Official sources checked Sep 23

One-click AI pack

Export the AI concentration and exit test workflow

Paste this into ChatGPT, Claude, Gemini, or an approved enterprise AI tool. It builds the dependency and evidence pack, but it cannot determine regulatory scope, accept risk, approve a provider, or declare an untested exit viable.

AI concentration is an operating problem before it becomes a systemic one

EIOPA's September 16 analysis says AI is moving from experimentation toward established use across financial services. Its 2025 survey covered 347 insurers in 25 EEA countries: 65% reported already using generative AI and another 23% planned to. The operational warning is not simply that many firms use AI. It is that they may depend on the same small group of model, cloud, and infrastructure providers while deploying AI into claims, fraud, customer service, data analysis, and other important services.

EIOPA's recommendation is concrete: make material dependencies visible, test provider or model failures, verify that exit plans work, retain human expertise and fallback capacity, and preserve provider diversity. It also makes an important policy judgment: another broad layer of AI regulation is not the immediate answer. Rigorous implementation of existing DORA and AI Act controls, coordinated monitoring, and targeted supervisory action matter more.

The same concern appears outside EU insurance. The Bank of Canada's 2026 Financial System Survey found almost all respondents were using AI, usually at limited or moderate levels, and planned further use in investment work, operations, back-office processes, financial-crime prevention, and customer service. Respondents pointed to concentration among AI/cloud providers, weak backup plans, talent constraints, and the ability to maintain critical functions during outages or disruption.

DORA supplies a demanding benchmark for in-scope EU financial entities. For ICT services supporting critical or important functions, exit strategies must address provider failure, degraded service, disruption, material risk, and contract termination. Plans must be comprehensive, documented, sufficiently tested, periodically reviewed, and capable of preserving business activity, regulatory compliance, and service continuity.

This page uses that tested-exit discipline as an operating model. It is not legal advice, and not every AI tool or organization falls into the same regulatory scope. Legal and compliance owners must determine which service and entity obligations apply.

The unit of resilience is not the vendor contract. It is the critical business service and every dependency required to keep that service within tolerance.

Start with the critical business service, not the model inventory

An inventory that says “OpenAI, Anthropic, Microsoft, Google” is too shallow. Finance needs to know which business service stops, which obligations are missed, how quickly harm accumulates, and which minimum service must continue. Define the service before scoring providers.

Service factQuestionEvidence
OutcomeWhat must customers, staff, or regulators still receive?Service map, process owner, product terms, obligations.
Disruption toleranceHow long can the outcome be degraded or unavailable?Impact analysis, approved tolerance, peak calendar.
Minimum viable serviceWhich steps can pause, degrade, or become manual?Runbook, manual capacity test, exception policy.
Data stateWhich records, prompts, indexes, and work queues must survive?Data map, RPO, exports, retention, reconciliation.
Control stateWhich approvals, checks, logs, and evidence cannot disappear?Control register, audit trail, model-risk requirements.
Financial impactWhat revenue, cost, liquidity, claims, or penalty exposure grows over time?Scenario model and Finance-approved assumptions.

Only then map AI use cases. One service may contain several: document extraction, retrieval, summarization, fraud triage, draft communications, recommendation, decision support, and autonomous action. They do not share the same criticality. A draft-summary assistant may be removable; a model that routes claims or freezes payments may require a tested fallback before production.

Freeze the population from evidence, not self-report alone. Reconcile procurement records, accounts-payable vendors, cloud spend, SSO applications, API keys, model-gateway logs, repositories, data-platform connections, and business-owner attestations. The focused community scan found finance practitioners asking whether AI is actually removing work or moving it around. That is relevant because “manual fallback” often means an invisible queue of review and correction work. Measure the real capacity.

Map the hidden common failure domains

critical business service
  -> AI-enabled workflow / SaaS application
  -> model endpoint and version
  -> hosting cloud, region, network, gateway
  -> identity, keys, secrets, policy enforcement
  -> source data, retrieval index, vector store, work state
  -> safety, observability, ticketing, billing
  -> fourth-party infrastructure and package registries
  -> internal operators, reviewers, and scarce specialists
  -> fallback / degraded / manual / replacement path

Two model vendors may both run in the same cloud region. Two SaaS products may call the same underlying model. A self-hosted model may still depend on the same identity provider, object store, container registry, GPU fleet, and network route as the primary. A manual fallback may depend on the same document repository that the outage has made unavailable.

Map identity separately. If the fallback uses the same SSO tenant, secret store, API gateway, or enterprise key-management service, an identity or control-plane incident can remove both paths. Do the same for data: exporting prompts is not enough if embeddings, fine-tunes, evaluation fixtures, conversation state, workflow queues, or audit logs cannot move.

Map fourth parties from current evidence. Provider subprocessor lists, architecture material, assurance reports, contract schedules, status pages, data-location records, and support responses can reveal common clouds or services. Mark unknowns explicitly. Do not treat “vendor confidential” as evidence of diversity.

Finally, map people. The Bank of Canada found talent and AI literacy constraints to be significant. An exit plan that requires a specialist who left six months ago is not a plan. Record who can operate the fallback, when they last rehearsed it, and how long they can sustain manual throughput.

Score substitutability from observed evidence

A single composite score can hide the reason a service is fragile. Keep the dimensions visible and link every rating to a dated artifact or test.

DimensionLow risk evidenceHigh risk evidence
CriticalityAdvisory use; process continues without it.Critical outcome or legal/customer obligation stops.
Provider commonalityIndependent model, cloud, region, identity, and network.Primary and fallback share material control planes.
Data portabilityCurrent, tested export reconstructs open work.Prompts export but indexes, fine-tunes, logs, or state do not.
Control equivalenceFallback preserves approvals, logging, privacy, and quality checks.Fallback removes guardrails or audit evidence.
CapacityRepresentative peak workload passed.Demo traffic passed; peak throughput unknown.
Contract supportTransition, data return, deletion, audit, and termination are evidenced.Rights absent, ambiguous, or commercially prohibitive.
Human capabilityNamed trained owners rehearse the path.Knowledge is vendor-held or concentrated in one person.
Test freshnessScenario tested within policy and defects closed.Document-only plan or expired rehearsal.

Use a traffic-light or 1-5 scale only after defining anchors. For example, portability is not green because a contract mentions export. It is green when a representative export was imported into the fallback, open cases resumed, records reconciled, and required evidence remained readable.

Aggregate to the business-service level carefully. A severe failure in identity, data state, control equivalence, or capacity should remain visible even if other dimensions score well. Do not let ten low-impact diversified assistants dilute one concentrated, critical decision workflow.

Test five failures and one actual exit

ScenarioWhat the test provesSuccess evidence
Model endpoint unavailableRouting, alternate capacity, compatibility, and backlog handling.Measured switch time, output reconciliation, no control bypass.
Cloud or region unavailableThe fallback is not in the same physical/control-plane failure domain.Independent region/cloud path, sustained representative workload.
Identity or key service unavailableBreak-glass access works without uncontrolled shared credentials.Scoped emergency identity, logged use, rotation and revocation.
Quality degrades without outageMonitoring catches harmful drift and can move work to review/degraded mode.Trigger fires, queue routes correctly, reviewers preserve tolerance.
Quota or price shockFinance limits consumption without silently stopping critical work.Budget trigger, priority rules, approved degraded mode, cost model.
Contractual/regulatory exitData, state, identities, configuration, work queues, and evidence can move.Former provider removed; replacement continues within tolerance.

Define the clock before the test. Measure detection time, escalation time, decision time, switch time, backlog peak, recovery time, output variance, control exceptions, data loss, customer impact, and incremental cost. A successful login or HTTP 200 proves availability, not business-service recovery.

Test under representative load and timing. A month-end close, market open, renewal window, claims catastrophe, payroll cut-off, or regulatory reporting date may expose capacity and review bottlenecks absent from a weekday demo. Use synthetic or approved masked data, but keep the process shape and volume realistic.

Separate fallback from exit. Failover may route new requests to another endpoint while old work, logs, embeddings, or evidence remain trapped. Exit proves the former provider can be removed, required data is returned or deleted, identities and secrets are rotated, open work is reconciled, and the service can continue without hidden calls back to the old dependency.

Worked case: claims intake has two vendors and one failure domain

An insurer uses Vendor A to extract fields from claim documents and Vendor B to summarize the file for an adjuster. Procurement records show two contracts. The initial inventory labels the service diversified.

The dependency map changes the conclusion. Vendor B calls Vendor A's model for image processing. Both applications run in the same cloud region, use the same enterprise identity provider, retrieve documents from the same object store, and write to one vector index. The manual runbook requires exports from that index, but the export has never been imported. Only one analyst knows how the rule-based extraction fallback works.

Control claimObserved evidenceDisposition
Two AI providersTwo contracts, but Vendor B depends on Vendor A for vision.Common model failure domain; no model diversity.
Multi-provider cloudBoth production paths use one region and identity tenant.Regional/identity concentration remains.
Data is portableJSON export exists; embeddings and open-work links not tested.Portability unproved; exit test required.
Manual fallbackOne analyst; 18% of peak daily capacity.Only a short degraded mode, not full continuity.
Controls are equivalentFallback omits fraud-score evidence and reviewer reason codes.Control gap blocks release for high-risk claims.

The test disables the primary model route, uses an independent region and rule-based extractor for a representative batch, queues ambiguous files for human review, and measures throughput. The service stays within its four-hour tolerance for priority claims but exceeds tolerance for the full population after six hours. Finance approves temporary capacity funding; IT builds an independent object-store replication path; the business owner narrows the minimum viable service; procurement negotiates transition assistance; and the team schedules a full exit rehearsal.

The result is not “red vendor.” It is a precise gap register: shared vision model, shared region, unproved index portability, insufficient manual capacity, and missing fraud-control evidence. Each gap has an owner, cost, interim control, deadline, and retest.

Keep an evidence ledger that distinguishes facts from assumptions

RecordRequired evidenceOwnerExpiry or trigger
CriticalityApproved impact analysis, tolerance, RTO/RPO, minimum service.Business + FinanceService or volume change.
DependencyArchitecture, contract/subprocessor, logs, provider confirmation.IT + procurementProvider/model/region change.
PortabilityExport schema, sample, import result, reconciliation.Data ownerFormat or state change.
ContractTransition, audit, incident, data return/deletion, termination rights.Procurement + legalRenewal or material change.
TestScenario, timestamps, versions, workload, results, defects, reviewers.Resilience leadPolicy interval or incident.
Residual riskGap, impact, interim control, cost, owner, acceptance authority.Management bodyDeadline or exposure change.

Label every field as observed, provider-asserted, contract-stated, tested, estimated, assumed, or unknown. An AI assistant can reconcile dates, surface contradictions, and build the gap queue. It must not promote a provider claim to tested fact or infer a cloud dependency from branding.

Link tests to immutable configuration and output evidence. Record model version, endpoint, region, gateway policy, data snapshot, prompt/instruction version, fallback configuration, workload, reviewers, and timestamps. Preserve enough detail to replay the test without exposing secrets.

Finance should connect operational gaps to money: outage loss by hour, backlog-clearance cost, alternative capacity, exit fees, duplicate-run costs, transition staffing, and control remediation. This complements the AI seat and usage-billing review, which reconciles ongoing consumption rather than resilience.

Failure modes that make concentration registers look safer than reality

FailureWhy it happensControl
Vendor-count diversificationProcurement counts logos, not shared infrastructure.Service-level graph with common model/cloud/identity/data domains.
Endpoint-only failoverNew requests route, but open work and evidence stay trapped.Export/import and open-case reconciliation in the exit test.
Paper capacityFallback limits are quoted, not load-tested.Representative peak workload and sustained-duration test.
Control sheddingDegraded mode removes review, logging, privacy, or fraud controls.Minimum-control profile and explicit exception approval.
Shared identity blind spotBoth paths use one SSO, key store, or gateway.Independent scoped break-glass identity with monitored use.
Unknown fourth partiesSubprocessors are stale or too generic.Dated evidence, contract-change notification, unknown-as-risk rule.
Manual fallback fictionNo throughput, skill, or duration test exists.Named trained staff, measured capacity, fatigue and backlog limits.
Annual checkbox testA demo never exercises the actual critical period.Risk-based scenarios, peak timing, incidents and change-triggered retests.
AI approves its own evidenceModel summary becomes the risk conclusion.AI prepares; named owners verify, decide, and accept residual risk.

Run a 30-day pilot on one critical service

Days 1-5: choose one service with a clear business owner and meaningful AI dependency. Legal/compliance records scope. Finance and the owner approve minimum viable service, disruption tolerance, RTO/RPO, peak period, financial impact assumptions, and manual-capacity limits.

Days 6-10: freeze the AI use-case population. Reconcile procurement, spend, SSO, gateways, repositories, data connections, logs, and owner declarations. Build the dependency graph through direct providers, common control planes, data/state, fourth parties, and people. Mark unknowns.

Days 11-15: collect contract and assurance evidence. Map incident notice, service levels, audit/access, subcontractors, continuity, transition, data return/deletion, termination, pricing, and exit assistance. Procurement and legal record gaps without asking the AI tool for legal conclusions.

Days 16-22: run endpoint failure, quality degradation, identity failure, and peak-capacity scenarios. Then execute one exit slice: export representative state, import it into the fallback, rotate identities, route work, reconcile outputs, and prove the old provider is no longer called.

Days 23-27: quantify defects and cost. Classify architecture, capacity, data, control, contract, skills, and governance gaps. Assign interim controls, budget, owners, dates, and retests. Update the AI vendor contract workflow for rights that blocked testing.

Days 28-30: hold a human release review. The business owner confirms continuity, Finance validates cost and impact, IT resilience/security validates the test, procurement/legal review contracts and scope, model risk/privacy confirm controls, and authorized management accepts any residual exposure. No test evidence means NOT READY, not “low risk.”

FAQ

Is using two AI model vendors enough to reduce concentration risk?

No. They may share cloud, region, identity, data, gateway, network, model family, safety service, or internal operators. Demonstrate independent failure domains and capacity at the business-service level.

Does DORA require an AI-specific exit plan?

DORA governs ICT third-party risk for in-scope financial entities and requires tested exit strategies for ICT services supporting critical or important functions. Whether a specific AI arrangement falls within that scope requires qualified legal/compliance review. The workflow also works as a control pattern outside DORA.

Can AI perform the concentration assessment?

AI can reconcile inventories, contracts, diagrams, logs, and tests; identify missing fields; and prepare an evidence matrix. It cannot invent dependencies, decide legal scope, accept risk, or treat a provider assurance statement as a successful test.

What proves an AI exit plan works?

A timed rehearsal that moves representative work and required state to an approved path, preserves controls and evidence, stays within tolerance, reconciles results, removes calls to the old provider, and receives named human review.

Sources and further reading

Sources were checked on September 23, 2026. Exact practitioner discussion of EIOPA's September analysis was thin, so the guide is grounded primarily in official financial-sector and regulatory sources. It provides an operational control pattern, not legal, regulatory, accounting, or investment advice.