Identity resilience is your ability to keep users signing in, or to fail in a safe and predictable way, when your identity provider (IdP) or one of its dependencies stops working. For a SaaS product it comes down to three questions: which parts of your login flow break first during an outage, how long the sessions you already issued keep working, and what you can do to shorten the gap or route around it.
You do not need a perfect answer to all three, but you do need to have asked them before the incident. Teams that decide in advance how long their tokens live, how admins get in, and what they tell customers can handle an IdP outage as a routine incident. Teams that have not decided usually work those answers out during the outage, with customers waiting.
This post is the planning view, and it applies to any identity provider, hosted or self-run.
A note on the term “identity resilience”
If you search for this phrase you will mostly find backup and recovery products aimed at Active Directory, Entra ID and Okta, which focus on restoring a directory after an attack. That is a real and separate problem. This post is about something narrower and closer to a product team’s day job: keeping the login path of your own application available when the provider behind it has trouble.
What breaks first when an IdP goes down?
Not everything fails at once, and the order matters because it tells you where to put your effort. In a typical OpenID Connect (OIDC) setup, the pieces degrade roughly like this:
- New sign-ins fail first. A user who is not already signed in cannot get an authorization code, so they cannot start a session. This is the most visible symptom.
- MFA fails with sign-in, and can also fail on its own. If the second factor depends on a separate service such as an SMS gateway or push service, users with the right password still cannot complete sign-in when that service is down.
- Token refresh fails once tokens expire. Your app keeps working for as long as access tokens are valid, and then the refresh call fails and users are signed out.
- Provisioning and sync stop. SCIM pushes from your customers’ directories and any scheduled sync jobs will fail or queue, which usually shows up later as stale access rather than as an error.
- Admin consoles may be unreachable. If admin access is federated to the same provider, the people who need to change configuration are locked out too. Our post on break-glass accounts covers this case.
What keeps working is anything that validates a token locally. An access token that is a signed JWT can be checked by your API against the provider’s cached public keys without calling the provider at all, so requests carrying a valid token continue until the token expires. Our explainer on JWKS shows how that validation works and why caching the key set matters.
How long do existing sessions keep working?
The answer depends on three lifetimes that you control in most providers: the access token lifetime, the refresh token or session lifetime, and any idle timeout. During an outage, users keep their access until the shortest relevant one runs out, and then they are signed out.
That creates a real trade-off. Short access tokens limit the damage of a stolen token and let you revoke quickly, which is good security. Longer lifetimes give you more runway during an outage. Neither extreme is right for every product, so decide on purpose. Our guides on session management in Keycloak and offline tokens cover the settings and what each one exposes.
Two cautions apply. First, do not stretch token lifetimes in the middle of an incident, because the people most likely to benefit from a longer token are also an attacker holding one. Second, remember that sessions only help people who were already signed in, so they do nothing for new customers on the day of the outage.
What design choices make login more resilient?
Most of the useful levers are architectural, and you can adopt them in stages.
- Validate tokens locally and cache the keys. Make sure your APIs fetch the provider’s keys once and keep them, rather than calling the provider on every request. This keeps already-authenticated traffic flowing.
- Set token and session lifetimes deliberately. Pick values that balance revocation speed against outage runway, and write down the reasoning.
- Run the IdP without a single point of failure. Inside one region this means redundant nodes, a replicated database and tested failover. Across sites it means a second cluster. Keycloak’s supported multi-site setup is the production option today, and the 26.7 multi-cluster v2 preview (for metro distances, with under 10 ms between clusters) is covered in our multi-region stateless architecture post. It is still a preview, so test it before relying on it.
- Put a broker between your app and the upstream providers. If your customers sign in through their own corporate IdPs, an identity hub means a problem at one upstream provider affects only that customer instead of your whole login. The broker itself then becomes the shared dependency, so it needs the redundancy described above. See federated SSO versus a single IdP and how enterprise SSO survives IdP changes.
- Test restore, not just backup. A backup is only useful once you have proved you can restore it, and our Keycloak backup and restore guide walks through a restore drill.
- Have a degraded mode. Decide whether read-only access, a limited feature set, or a clear maintenance message is acceptable when sign-in is down, and build it before you need it.
If you run Keycloak yourself, the production readiness checklist covers the baseline settings.
What do uptime numbers actually cover?
A service level agreement (SLA) is a contractual promise with a remedy, usually service credits, and it says nothing about how well you will handle the hours it is down. Most SLAs are measured monthly, and a 99.9% target allows about 43 minutes a month. Over a year that is a little under nine hours of downtime, and a 99.99% target allows a little under an hour, so a single long incident can use a year’s allowance at once. Our post on how many nines your Keycloak SLA needs lays out the downtime math and which architectures deliver each tier.
Read the definition carefully before comparing numbers. Ask what is counted as downtime, whether planned maintenance is excluded, whether MFA and admin access are in scope or only the sign-in endpoint, and how the provider measures it. You can see how we describe our own commitments on the SLA page, and the trust page lists the related documents.
How should you communicate during an IdP outage?
A short plan, agreed in advance, covers most of it:
- Name an owner for communication separate from the person fixing the problem.
- Post on your status page early, say what is affected (new sign-ins, MFA, or all access), and say when you will update next.
- Tell enterprise customers’ admins directly if their SSO is affected, since their helpdesks will get the calls first.
- Link to the upstream provider’s status page when the cause is theirs, and still own the customer experience.
- Write a short review afterwards covering what happened, what you changed, and how you will detect it sooner next time.
Ten questions to ask your IdP vendor about resilience
Enterprise buyers may ask these in security reviews, and you can ask the same of your own provider:
- What is the uptime commitment, and what exactly counts as downtime?
- Does the commitment cover MFA and admin access, or only the sign-in endpoint?
- How is the service deployed across zones and regions?
- What is the recovery time and recovery point objective, and when was a restore last tested?
- How are customers notified of incidents, and how quickly?
- Is there a status page with history, not only a current state?
- What happens to issued tokens and sessions during an outage?
- How do I reach an administrator if normal admin sign-in is unavailable?
- Can I export my configuration and users so I can move if I need to? Our exit plan guide explains why that matters.
- Where can I read the post-incident reviews for past outages?
What about managed Keycloak?
A managed service moves the operations work, such as patching, failover and backups, to the provider, and the architecture questions above still apply to it. For Skycloak, only the Enterprise plan carries a contractual uptime SLA with service credits, and support hours depend on the plan, with 24/7 support on Enterprise only. The SLA page and the managed Keycloak solution page have the current terms.
Frequently asked questions
What is identity resilience?
It is the ability of your sign-in and access system to keep working, or to degrade in a controlled way, when the identity provider or one of its dependencies fails. For a SaaS team it covers token and session lifetimes, redundancy in the IdP itself, admin recovery, and communication.
What happens to logged-in users during an IdP outage?
Users who are already signed in generally keep working until their access token expires, provided your APIs validate tokens locally. After that, refresh requests fail and they are signed out. New sign-ins fail from the start, and MFA prompts fail with them.
Does an SLA protect me from login downtime?
Only financially, and usually in small amounts. An SLA sets a target and a credit, and does not prevent an outage. You still need your own plan for token lifetimes, a degraded mode and customer communication.
Should I stretch token lifetimes to survive outages?
Consider it as a design choice made in advance, weighed against the security cost of longer-lived tokens. Avoid changing it mid-incident, because a longer lifetime also helps anyone who has stolen a token.
How does an identity broker help with resilience?
A broker sits between your app and the upstream providers, so a failure at one customer’s corporate IdP affects that customer’s sign-ins only, and you can keep serving everyone else.