Mail Credential Vault deprecated
Motivation
Some features reach a user's mailbox when that user is not logged in: a snoozed mail returns to its folder, a scheduled mail is sent, a data export gathers mail, permanent push keeps an IMAP connection open, and a personal access token lets an MCP client read mail on the user's behalf. Each of those used to keep its own copy of whatever the mailbox needs - the session password or OAuth tokens - in its own table, in its own format.
The mail credential vault replaces those copies with one place. A feature asks the vault to keep what the session provides and gets a handle back; the row it writes carries that handle instead of the credential. When the time comes, the vault turns the handle into a session that reaches mail, and the feature never sees the credential itself.
What is stored
The table mail_credential holds one row per kept credential:
| Kind | Content |
|---|---|
master | Nothing. Master authentication was in effect, so the mail stack needs nothing from the session |
session | The session password, the session's OAuth tokens, or both - whatever the configured authentication type needs |
exchanged | Tokens obtained through an OAuth token exchange, logged out at the identity provider along with the user's last exchanged credential for their back-end |
The content is encrypted with com.openexchange.sessiond.encryptionKey, the same key that protects session passwords, through the session obfuscator. A caller may instead hand in a secret of its own; the row is then sealed with that secret and unreadable without it. A change of the user's password is written into the rows the vault can read; a row it cannot read right then - a key not configured yet, the obfuscator being reloaded - is left as it is, not dropped. Personal access tokens (feature access-token) do that with the token's secret, which only the token's holder presents: such a row does not depend on com.openexchange.sessiond.encryptionKey, so a change of that key leaves it readable; it is not repaired after a password change, the maintenance job does not keep its tokens alive, and its tokens are never obtained through an exchange, since the vault could not log them out later. Where its password or tokens stop working, the user mints a new token.
Configuration
What a row holds follows from the mail configuration that already applies to the user: com.openexchange.mail.passwordSource, com.openexchange.mail.authType and com.openexchange.mail.transport.authType. Permanent push is the exception: it keeps a credential only where com.openexchange.push.credstorage.enabled says it may, and the same switch decides whether its old rows are migrated. com.openexchange.push.credstorage.rdb no longer decides where the credential is kept - the vault keeps it in the database whatever it says; it only decides where rows of the former storage still live: with false, in cluster memory, where they are never migrated and go with the user's next login.
Tokens from an exchange
Where the deployment says so, the tokens a row holds are not the session's own but ones obtained for it through an RFC 8693 token exchange - typically longer-lived, which is what a credential kept for months wants. That is configured once, at the vault, and applies to every feature it keeps credentials for:
com.openexchange.mail.credential.oauth.tokenExchange=true
com.openexchange.mail.credential.oauth.tokenExchange.backendPath=
com.openexchange.mail.credential.oauth.tokenExchange.scope=
com.openexchange.mail.credential.oauth.tokenExchange.additionalParameters=
com.openexchange.mail.credential.oauth.tokenExchange.requestedTokenType=
Each of the five also exists with a feature identifier appended - snoozed-mail, scheduled-mail, gdpr-export, push - and that value wins for that feature. So a scope that covers sending goes to …tokenExchange.scheduled-mail.scope, the provider's own marker to …tokenExchange.snoozed-mail.additionalParameters=usecase=mail-snooze, and a feature left out of it all gets …tokenExchange.push=false. Each setting is resolved on its own: a feature that says nothing about its scope gets the shared one, even where its switch comes from a setting of its own. All five are config-cascade aware and reloadable; which level a value came from is logged on DEBUG by com.openexchange.mail.credential.impl.TokenExchangeConfig.
The shared switch reaches every feature, permanent push and the data export included - two that had no such setting before. Where only snoozed and scheduled mail are meant, they are taken out with …tokenExchange.push=false and …tokenExchange.gdpr-export=false.
Switched off by default, so nothing changes for an existing deployment. Snoozed and scheduled mail configured this themselves before it moved here (com.openexchange.mail.snoozed.oauth.tokenExchange and com.openexchange.mail.scheduled.oauth.tokenExchange, with the same three settings behind them). Those names are still read: they beat the shared value, but not the feature's own new name, so a deployment that set them keeps exactly what it had and has nothing to migrate; they will go one release after this one.
What the exchange has to deliver is a token that outlives the user's session, or at least a refresh token that does - otherwise it buys nothing over the session's own tokens. By default an access token is asked for (requestedTokenType=access_token), and an identity provider that hands out nothing else leaves a row that expires with that token: snoozed and scheduled mail then depend on a live session of the user. requestedTokenType=refresh_token asks for a refresh token as well, so that the row is refreshed and kept alive like any other; the full URN of either type is taken too.
Keycloak's standard token exchange hands out a refresh token only for refresh_token and only where the client allows it (Allow refresh token in Standard Token Exchange: Same session). That refresh token belongs to the user's SSO session and dies with the logout; Keycloak refuses one for a session that is itself an offline one, and an exchange never yields an offline token. So with Keycloak, leave the exchange off and add offline_access to com.openexchange.oidc.scope instead: the session's own refresh token is then an offline one, survives the logout, and a kept row refreshes with it for as long as the offline session lives.
A row that keeps the session's own tokens holds the same refresh token as that session, and a copy of a credential - the one a push registration keeps from a personal access token - holds the same refresh token as its original. Each of them refreshes on its own. Where the identity provider rotates refresh tokens, invalidating the old one with every refresh, the first refresh of one strands the others: their next refresh is rejected, and an identity provider with reuse detection may even end the user's session over it. Kept tokens therefore need an identity provider that keeps a refresh token valid across refreshes, as Keycloak does unless Revoke Refresh Token is switched on - or a token exchange, whose tokens are the row's alone.
An exchange applies only where the token exchange handler takes the session - an OpenID Connect session of a registered back-end. A session it does not take, one from a SAML login say, is captured as it is, with its own tokens, and the feature is still offered on the terms the switch promises; where that is not wanted, leave the switch off for such deployments.
Where an exchange applies, the row's kind is exchanged, which is what gets those tokens logged out at the identity provider. Exchanges made from the same SSO session share one client session there - with Keycloak, revoking one token ends that client session and every token exchanged in it. So exchanged tokens are logged out only along with the user's last exchanged credential for their back-end: when that one is released or expires, or the user or context is deleted. Tokens a newer credential of the same feature replaces, and those a repair replaces with tokens of the same back-end, are left to expire, as is a released credential's while others remain. Without a back-end of its own for the exchange (com.openexchange.mail.credential.oauth.tokenExchange.backendPath), the exchange runs through the client the user logged in with, and its tokens share the client session of the user's App Suite login: revoking them would end that session at its next token refresh. So the tokens are also left to expire while the user has a live session of the same back-end, or while the sessions cannot be looked up. That logout is an end-session request and takes the ID token; exchanged tokens that came without one are revoked instead at the back-end's revocation end-point (com.openexchange.oidc.opRevocationEndpoint, per back-end, e.g. …oidc.background.opRevocationEndpoint). Left empty, it is the revocation_endpoint of the issuer's discovery document (<opIssuer>/.well-known/openid-configuration), fetched on first need and kept for an hour; Keycloak names one there. Where the document names none or cannot be fetched, the tokens are left to expire on their own, as are those of a logout that finds the document being fetched for another one for more than half a second. A released credential's logout runs off the request that released it, in the queue the deletion logouts use (20,000 credentials per node); beyond that, the tokens are left to expire, with a warning. A password change is not written into the tokens of such a row; a password the row holds beside them - mail access by password, transport by exchanged tokens - is rewritten like any other. A live session repairs such a row the way it was captured, by exchanging again - the session's own tokens are no replacement for what an identity provider handed out. Rows a migration moves are taken over as they are, no exchange is attempted for them.
Refreshing kept tokens
A row's access token is refreshed when it is used and expires within a minute; one node at a time does that, under a lease on the row, so a refresh token is never presented twice. Where the identity provider cannot be reached for now, the access token is used while it lasts; one that has expired is not presented to the mail server, which would take it for a bad credential: the use fails with MCRED-0010 and is tried again later, so a snoozed or scheduled mail waits and push keeps its credential.
Where the configuration does not allow a refresh - the identity provider refuses the client (invalid_client, unauthorized_client, e.g. after the client secret was rotated) or the OpenID Connect back-end the tokens name is not configured any more - no later run helps. The row is kept as it is, neither rejected nor released, and refreshes again once the configuration is fixed. Its access token is used while it lasts; once it has expired, a use fails with MCRED-0013: a snoozed or scheduled mail is then returned or sent through a live session of the user, or the user is told once that it waits for a login, as for a refused credential. Push keeps its credential; a data export retries as for any other failure for now, a limited number of times. The cause is logged on ERROR once an hour per back-end; every such refresh is counted (appsuite_mailcredential_refresh_misconfigured_total).
An identity provider may leave expires_in out of a refresh (RFC 6749 only recommends it). The new access token's expiry is then taken from its exp claim where it is a JWT - read, not verified, since it only schedules the next refresh - and otherwise from the lifetime of the token it replaces, counted from now. Where neither is known, or that lifetime is within a minute, the expiry stays unknown and the token is not refreshed on use, only by the keep-alive below. A row that is not used is not refreshed that way - and an identity provider may end the session behind an idle refresh token: Keycloak ends an offline session after 30 days without a refresh (Offline Session Idle). A snoozed mail that returns after two months, or a push registration whose device stays silent, would then find its refresh token gone.
The maintenance job therefore refreshes kept tokens that went without a refresh for a while, whatever their own expiry:
com.openexchange.mail.credential.oauth.keepAliveInterval=3D
A time span (3D, 1W, 36h); 0 switches the keep-alive off, and anything shorter than an hour is raised to an hour. While it is off, rows are still scheduled by the default interval, so that switching it on again takes up the rows kept meanwhile; until then the rows_due gauge stays at zero and the administration lists them as not_kept_alive. Rows that nodes of an earlier version keep unscheduled meanwhile, during a rolling upgrade, are scheduled by the maintenance job, 1000 per schema and hour. Keep it well below the identity provider's idle limit: the job runs hourly, and a row comes due up to a fifth of the interval early, so that rows kept at the same time do not come due at the same time ever after. Server-wide and reloadable; an invalid value is logged once and the default used. Where a refresh token that is kept, exchanged or refreshed names its own expiry (exp, as Keycloak's do) and that comes before the row is due, the keep-alive will come too late for it: that is counted (appsuite_mailcredential_keepalive_too_late_total) and logged on WARN once a day per back-end; the row is scheduled as ever. A normal Keycloak refresh token lives as long as the SSO session may idle (SSO Session Idle, 30 minutes by default), far shorter than any interval; an offline token names no expiry unless Offline Session Max is on.
Rows with a refresh token of an OpenID Connect back-end are kept alive; an ID token is not needed, so tokens from a token exchange, which come without one, are refreshed and kept alive as well. The job only takes rows the server key opens. The credential of a personal access token is sealed with the token's secret, which only its holder presents: it is kept alive whenever the token is presented, whatever the request is about, at most once an hour per token and without holding the request up. It is not refreshed while the session it was created in holds the same refresh token, nor while a copy of it exists, e.g. the one of a push registration: the job keeps the copy alive, and with it the identity provider's session behind both. A token that is not used at all keeps its credential alive only through such a copy. A row used regularly is refreshed by its use and never comes due; the job only refreshes the idle ones, at most 1000 per schema and hour, the longest due first. A row whose session is alive and holds the very same refresh token is left to that session. Where the user's sessions cannot be looked up, e.g. while the session storage is unavailable, the row is not refreshed and stays due, for the next run or presentation to look again. Where the identity provider's signing keys cannot be fetched to check the new ID token, the refreshed tokens are kept with the former ID token and the next refresh checks again; an ID token that fails the check is a rejection. A refresh that fails for now puts the row off by an hour, and a run gives up on a schema after 20 failures in a row; while no refresher is registered - the OpenID Connect bundle restarting, say - a row is put off by an interval. A row whose refresh the configuration does not allow (see above) is put off by an interval as well, counts as misconfigured and does not count towards the 20 failures, so that such rows do not hold up the others; an administrator's refresh of it answers failed.
The keep-alive outlasts the identity provider's idle limit on purpose: a push registration keeps its credential without expiry, so its offline session stays alive as long as the registration does, also for a user who does not come back. Where that limit is meant to cut off inactive users, switch the keep-alive off.
Where the identity provider rejects a refresh token for good - invalid_grant, or new tokens the refresher cannot use, such as an ID token that does not pass - the row notes that (refresh_due is -1) and no use asks the identity provider again. A live session of the user repairs the row: the access that ran into the rejection tries that right away, a login with OAuth tokens does as soon as the session exists, and the maintenance job does, 500 rows per schema and hour, for every user that has a session at the time it gets to the row. The credential of a personal access token, which only the token's secret seals, is repaired the next time the token is presented, from a live session of its user; until then it reaches no mail. Where a session cannot repair it - e.g. one the identity provider ended that is still stored - the next one is tried, up to three, those whose tokens were refreshed last first; an access that could not repair the row does not try again for a minute. A repair that takes over a session's own tokens refreshes them at the identity provider first, except right after a login, so that a session the identity provider ended does not leave the row rejected again; the provider's refusal ends that session, and the next one is tried. A repair that exchanges or refreshes tokens takes the row's lease first, so concurrent repairs ask the identity provider once. Where the user's sessions cannot be looked up, nothing is repaired for now; an administrator's repair answers failed then. Until then the row's access token is used as it is, and the mail server refuses it; the feature behaves as it would without a credential.
A refresh whose new access token expires within the refresh threshold of the OpenID Connect back-end (com.openexchange.oidc.oauthRefreshTime) is no rejection: the refresh token is fine, and where the identity provider rotates refresh tokens the old one is spent by then. The new tokens are kept, and a node uses them while they last instead of refreshing on every use; each such refresh is counted (appsuite_mailcredential_shortlived_total), and logged on WARN at most once an hour, since the cause is the token lifetimes of the identity provider, not the row.
A repair takes its credentials from a session of the user's own: not from a guest, a restricted or an OAuth client's session, not from one of a personal access token or of a kept credential, and not from one established on behalf of the user. Among the user's sessions it takes one that holds what mail access needs - tokens where mail is reached by OAuth, a password where it is reached by password.
Changing the encryption key
A change of com.openexchange.sessiond.encryptionKey makes the rows sealed with the former key unreadable - unless the former key is configured alongside. In a cluster that is updated node by node, change it in two steps, so that every node reads what any other one writes at any time:
- On every node, configure the new key as the one that only reads; nothing is sealed with it yet:
com.openexchange.sessiond.encryptionKey=<former key>
com.openexchange.sessiond.previousEncryptionKey=<new key>
- Once every node runs with that, swap them:
com.openexchange.sessiond.encryptionKey=<new key>
com.openexchange.sessiond.previousEncryptionKey=<former key>
Both are reloadable. During the first step, the job reports the "former" key as no longer needed once its walk is through; that is about the new key and means nothing yet.
Sessiond reads the former key as well, so sessions stored before the change stay readable instead of being lost; keep it at least as long as the longest session lives. The vault takes com.openexchange.mail.credential.previousEncryptionKey where that is set, and the one of sessiond otherwise.
The vault is the only one that seals its data again. Other data obfuscated with the key and kept in the database - e.g. the passwords of secondary mail accounts provisioned through the admin interface, or what a snoozed mail keeps besides its credential - is read with the former key as long as that is configured, but stays sealed with it. Removing the former key makes that data unreadable, whatever the vault reports.
A row the former key opens is sealed again with the new one as soon as it is read or written, and the maintenance job walks the table for the others, 500 rows per schema and hour. A walk that sealed something again is followed by one more, which checks that nothing was left behind its cursor; a table of N rows thus takes up to 2·N/500 hours. Once a walk from beginning to end found nothing left to seal again, the job logs that the former key is no longer needed in that schema, and the progress row of the rotation in mail_credential_migration carries the time in finished. Its name is mail_credential# followed by a digest of the key pair, so that every rotation walks afresh:
SELECT table_name, FROM_UNIXTIME(finished/1000) FROM mail_credential_migration WHERE table_name LIKE 'mail_credential#%';
Read it only once every node runs with the new key: a node that still seals with the former one writes rows behind the cursor. When every schema says so, remove the former key; the job then drops the progress row. Rows that neither key opens do not keep the former key needed - it cannot help them either - but the job names up to ten of them per run. A row that cannot be looked at for the moment, e.g. while the session obfuscator is reloaded, the database fails a read or no key derivation slot becomes free in time, keeps the walk from finishing and is not counted as unreadable; the next walk looks at it again. The former key protects what the rows hold just as the current one does: keep it out of reach like that one.
Without the former key, the rows are lost, and they are not repaired either: whether such a row is a copy can no longer be told, and a copy that forgot its original would outlive that original's revocation. Snoozed and scheduled mail fall back to a live session of the user where one exists; where none exists, the return or the send fails and is retried on every run, past the credential's expiry as well, until a live session of the user returns or sends it. The owner of a scheduled mail is told once, by mail, that the mail waits for a login; the owner of a snoozed mail is told once through the notification queue (com.openexchange.notification.queue), where the deployment runs it, and the notification goes once the job returns the mail. App Suite UI does not show these notifications yet. A data export and permanent push drop their row - for push that means the listener falls back to the old credential storage, but only as long as that row still exists: the migration empties that table. Once it has, a change of the key without the former one costs permanent push its listeners until each device registers again. It used to be independent of this key, because it kept its own copy under com.openexchange.push.credstorage.passcrypt. That setting stays required as long as com.openexchange.push.credstorage.enabled is set: it protects the rows of the former storage until the migration has emptied it, and the push bundle refuses to start without it.
Life cycle
A credential lives as long as the row that references it:
- Snoozed and scheduled mail release it once the mail is back or sent, and keep it for seven days past the due date so a retry after a mailbox outage still works. A mail scheduled from a personal access token is sent with a copy of the token's credential, and is refused where the token expires before the mail is due plus those seven days, so that the copy is not cut short in silence.
- A data export releases it when the task completes or is aborted, at the latest thirty days after it started.
- Permanent push keeps one credential per user with no expiry and releases it when the registration goes - but only where
com.openexchange.push.credstorage.enabledis set, which is what lets it keep credentials at all. Where a push subscription is made with a personal access token, as apps of the Mobile API do, the credential is a copy of the one the token opens, sealed with the server key: it expires with the token and is released when that token is revoked - the copy remembers which credential it came from and goes with it, on its next use at the latest; a device keeps a new one when it next registers. A listener on another node goes on with the connection it has open until it reconnects, which it can no longer do. An app signed in with the operator's identity provider gets a credential only through a token exchange (…tokenExchange.push=trueand…tokenExchange.push.backendPath), never with its own access token. - Deleting a user or a context drops their rows. Exchanged tokens are read before the delete and logged out at the identity provider on a thread pool, so a slow provider does not hold the deletion up; a node that goes down in between leaves that logout undone. At most 10,000 rows of one deletion and 20,000 rows of all deletions on a node wait for their logout, four pages of 1,000 at a time; beyond that, as in a mass deletion against a slow identity provider, the tokens are left to expire there, with one warning until the node has caught up.
- An hourly clean-up job drops rows whose expiry has passed.
- An hourly maintenance job, on one node at a time, keeps idle tokens alive, repairs rows whose refresh token was rejected and, while a former key is configured, seals rows again with the current one. Until the update task has added the
refresh_duecolumn to a schema, it leaves that schema alone. The column defaults to due, for the rows kept before as for those a node of the former release inserts during the upgrade, so the first runs after an upgrade look at each of them once - 1000 per schema and hour, refreshing those that hold refreshable tokens.
Every statement of the vault reads the refresh_due column. Run the update tasks of a schema before nodes of this release serve it, as the update job of the deployment does; a schema they have not reached yet fails every access to a kept credential until they ran.
Rows written before the vault existed keep working through the old code, and clean-up jobs move them over - one per table, so up to 1500 rows per schema and round. Each of the two mail jobs also gives up once it has come across 1000 rows it could do nothing with - neither move nor even read - so a run against something that turns every row down does not work through a whole table for nothing. Each runs every six hours and starts an hour after a node comes up. That hour is a node-local timer, not a cluster-wide interlock: during a rolling upgrade a node that still runs the old code cannot serve a row that was already migrated, and leaves it for the next run. A row that is being returned or sent at the time is left for a later run as well, unless its lock is older than a day and thus abandoned; and a mail that is returned or sent reads its row again once it holds the lock, so it goes on with - and releases - the credential the row points at by then.
While a rolling upgrade lasts, a node with the old code that picks up a snoozed or scheduled mail whose row was migrated logs an error for it and leaves it locked until the lock expires; a node with the new code returns or sends it then. A re-snooze on an old node writes the old format back over the handle, and the vault row it referenced goes on until it expires. Neither loses a mail.
Rollback. The tables can stay, the old code ignores them. What the old code cannot serve: snoozed and scheduled mails whose rows were migrated or created by the new code hold a handle instead of a credential - without master authentication they are never returned or sent again, only logged; a data export started by the new code has no generator the old code knows; and permanent push registrations whose credential the migration moved out of the former storage have no credential until the user signs in again. Once the first migration run is through - an hour after the first node with the new code came up - a rollback has to be weighed against that.
Throughput. A migration run moves at most 500 rows per table and schema every six hours, so a table of a million rows takes about a year to empty; nothing waits for it, since the features read both formats.
How far the two mail migrations have come is kept in mail_credential_migration, one row per legacy table (rows named mail_credential#… belong to the walk for a former key, see above): resume_at is where the next run carries on, finished says when a run last walked the table from its beginning to its end without finding anything of the old format, and walked when a walk from beginning to end last rewrote nothing - such a walk may take several runs, e.g. when each stops after twenty rows it cannot read, and walked is negative while it is under way. finished is the signal an operator needs before the release that drops the old readers:
SELECT table_name, FROM_UNIXTIME(finished/1000), FROM_UNIXTIME(walked/1000) FROM mail_credential_migration WHERE table_name NOT LIKE 'mail_credential#%';
A zero in finished means the schema is not through, and so does a missing row. Read it only once every node runs the new code: a node still running the old one keeps writing rows of the old format, and a row written while a run was under way is not part of what that run saw. A table that is through, or that holds nothing but rows that refuse to move, is therefore walked again once a week; a run that moved something stops a few pages behind its last rewrite and carries on from there next time, so a large table is not walked to its end on every run.
Each row a run moves is rewritten on a connection of its own and committed at once; only the progress row waits for the end of the run. Where a rewrite fails without a clear outcome - e.g. the connection drops before the commit is confirmed - the credential it sealed is kept as long as the row still holds what was read; the next run looks at the row again. Only a rewrite the database rolled back for sure (lock wait timeout, deadlock) or a row that now holds something else drops it.
The third migration - the one that empties the credentials table of permanent push - keeps no such marker; it is done when that table holds no rows of its own any more.
If a table does not come through, the nodes say why.
Could not move the credentials of N row(s)- the rows are still served by the old code, but something about them cannot be kept: a user that no longer exists, say. As long as such a row is in the table,finishedstays at zero. The message names up to ten of the rows by theiruuid; all of them are logged forcom.openexchange.mail.credential.migrationatDEBUG. The same message also appears on its own when two nodes went for the same rows and one of them lost - once, and gone on the next round. To get such a row out of the way, either return or send the mail from a session of its user, or delete the row: it cannot be served without a credential anyway.Could not read the meta of N row(s)- the rows can be read but not decrypted, usually becausecom.openexchange.sessiond.encryptionKeyhas changed. They are lost to the old code as well; the migration leaves them alone and keepsfinishedat zero. Twenty of them end a run withStopped the migration of table ... row(s) whose meta could not be read, since the cause is then the key, not the row.Stopped the migration of table ... early- a page could not be read. The run says nothing about the table and writes no progress at all, so the next one starts where this one began.
To make a schema start over - after a finished that turned out to be wrong, or to pick up rows written behind the cursor - run this in every context schema, and while no migration is under way, since a run in flight writes its own position at the end:
UPDATE mail_credential_migration SET resume_at = NULL, finished = 0, walked = 0 WHERE table_name = 'snoozedMail';
A migrated credential lives a year past the date its row works towards, where the old format had no expiry at all. That is generous on purpose - a snoozed mail whose return keeps failing still has its credential, however far out its date is - but it is not unlimited, so a run that is interrupted between keeping the credential and pointing the row at it does not leave a password behind for good.
Metrics
The vault counts what the keep-alive, the repairs and a change of the key do, on the usual /metrics endpoint. Every tag takes one of the values listed; nothing names a user, a context or a schema.
| Meter | Tags | Counts |
|---|---|---|
appsuite_mailcredential_keepalive_total | trigger: job, presentation, admin; outcome: refreshed, not_refreshable, rejected, deferred, failed, misconfigured | keep-alives |
appsuite_mailcredential_keepalive_capped_total | - | runs per schema that found more rows due than one run keeps alive - the job falls behind |
appsuite_mailcredential_keepalive_too_late_total | - | refresh tokens kept, exchanged or refreshed that expire before their row is due for a keep-alive |
appsuite_mailcredential_rejected_total | - | refresh tokens the identity provider rejected for good |
appsuite_mailcredential_repaired_total | trigger: access, login, job, feature, presentation, admin | credentials repaired from a live session |
appsuite_mailcredential_resealed_total | trigger: read, walk | credentials sealed again with the current key |
appsuite_mailcredential_unreadable_total | - | credentials the walk found that neither key opens |
appsuite_mailcredential_shortlived_total | - | refreshes whose new access token expires within the refresh threshold |
appsuite_mailcredential_refresh_misconfigured_total | refresh: use, keepalive | refreshes the configuration does not allow, e.g. invalid_client or a missing OpenID Connect back-end |
Two gauges tell the state rather than events. The maintenance job counts both per schema at the end of each run, from the index, and shares the count in a cluster map, one entry per schema that the next count replaces wherever the job runs. Every node reports the sum over all schemas, so take the value of any one node (or the maximum), not the sum over the nodes; a schema the job moved to another node counts once. A schema no node counted for three hours drops out. Without the cluster map a node reports only the schemas it ran the job for in the last 70 minutes; then sum over the nodes, and a schema the job moved meanwhile may count twice for up to that long.
| Gauge | Value |
|---|---|
appsuite_mailcredential_rows_rejected | credentials whose refresh token the identity provider rejected, waiting for a repair |
appsuite_mailcredential_rows_due | credentials sealed with the server key that are due for a keep-alive and that the run left due |
A rows_rejected that does not go down means users who do not come back to repair their credentials; rows_due above zero after several runs means the same as a rising keepalive_capped. A rising keepalive_capped means the keep-alive cannot keep up at the configured interval: raise the interval or look into why so many tokens go unused. A rising failed outcome points at the identity provider, a rising rejected count at its idle or lifetime limits. A rising refresh_misconfigured means no kept token of the affected back-end can be refreshed: check its client secret and OpenID Connect back-end configuration; the ERROR log names the back-end and the cause. A rising keepalive_too_late means kept credentials that go unused lose their refresh token before the keep-alive reaches them, and with it their mail access until a repair: lower com.openexchange.mail.credential.oauth.keepAliveInterval below the refresh token's lifetime, or have the identity provider issue longer-lived ones - offline tokens, or exchanged tokens with a longer idle timeout. The WARN log names the back-end and both durations.
Administration
Operators inspect and release a user's kept credentials through the provisioning API, as the master administrator or the administrator of the user's context. GET /prov/v1/contexts/{context_id}/users/{user_id}/mail-credentials lists them, oldest first, with feature, kind (session, exchanged, master), callerKey (only the caller's secret opens it, as for personal access tokens), created, expires and state: kept_alive (refreshed next at refreshDue), due, not_kept_alive, rejected or expired. A row with callerKey is only refreshed when its token is presented, so due there just means the token was not used lately; and since only the token's secret seals it again, rejected there is repaired the next time the token is presented while its user has a live session. Nothing returned reveals what a credential holds or the login it was kept for. DELETE …/mail-credentials/{id} releases one, DELETE …/mail-credentials?feature=push all of a feature; a release logs out exchanged tokens along with the user's last exchanged credential for their back-end, and whatever relied on the credential falls back to a live session of the user.
Two operations act on one credential right away instead of waiting for the maintenance job; both answer with an outcome and the credential as it stands afterwards (absent if the user has no such credential). POST …/mail-credentials/{id}:repair repairs a rejected one from a live session of the user: repaired, no_live_session, failed (e.g. the identity provider no longer exchanges the live session's tokens, or the user's sessions cannot be looked up), deferred (another node holds the credential and works on it), not_rejected or not_found. POST …/mail-credentials/{id}:refresh refreshes its tokens at the identity provider whatever their schedule: refreshed, not_refreshable (a password, or tokens without a refresh token), rejected (repair it instead), deferred (another node or a live session of the user refreshes it, or the vault is still starting), failed (e.g. the identity provider was not reachable) or not_found. A credential with callerKey answers caller_key to both: only a presentation of its token opens it. Over HTTP the service com.openexchange.grpc.provisioning.MailCredentialAdminService has to be added to the gateway's services; over gRPC and RMI (OXMailCredential) it is always available to administrators.
Administrative mailbox access
Managing a mailbox - its ACLs, its METADATA, its subscriptions - is a different question from reaching it as the user, and it has its own service. It manages IMAP mailboxes only, through DoveAdm where that is configured and otherwise over IMAP with the master password, and reports MBADM-0002 where the mailbox is not an IMAP one or neither is available. Deputy permissions for mail, cross-context mail shares and Dovecot push registrations all go through it.