Stuck service operations recovery - #5367
Draft
kathap wants to merge 4 commits into
Draft
Conversation
When CCDB is briefly unavailable during a broker polling cycle, the CC polling job can fail permanently (max_attempts=1) while the broker is still processing. This leaves an update stuck: last_operation.state stays 'in progress' with no delayed job working on it, requiring operator intervention. Add ServiceOperationsUpdateInProgressCleanup, a periodic job that detects stuck 'update' operations whose polling job has permanently failed and marks the operation and its pollable job as 'failed', giving clients a definitive final state. A FOR UPDATE SKIP LOCKED guard prevents double processing across concurrent CC instances. Unlike the create cleanup, no orphan mitigation is triggered: an update targets a resource that already exists, so it must not be deprovisioned.
…tions_update_stuck_in_progress_failed
Add ServiceOperationsDeleteStuckInProgressRetry, a periodic clock job that detects service-instance delete operations stuck in 'in progress' (broker still working, CC polling job permanently failed after a transient DB error) and re-enqueues the original delete polling job instead of marking it failed. The failed delayed_job's serialized handler is reused, preserving the original user_audit_info and start_time so the ReoccurringJob max-duration expiry still marks the operation failed once the original polling window elapses. No orphan mitigation, since delete targets a resource that should be removed.
kathap
marked this pull request as draft
August 14, 2026 09:39
3 tasks
Detect credential-binding and service-key delete operations stuck in 'in progress' with a permanently-failed polling job, and internally retry by re-enqueuing the original DeleteBindingJob. Mirrors the existing service-instance delete-retry job. Route bindings are skipped.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When CCDB is briefly unavailable during a broker polling cycle, the CC polling job can fail permanently (
max_attempts=1) while the broker is still processing. This leaves service operations stuck:last_operation.statestays'in progress'with no delayed job working on it, requiring operator intervention (more frequent on CCEE landscapes where CCDB restarts more often than on public IaaS).This PR adds three periodic clock jobs that bring stuck operations to a definitive state automatically. Each uses
FOR UPDATE SKIP LOCKEDto prevent double processing across concurrent CC instances.ServiceOperationsUpdateStuckInProgressFailed— detects stuck'update'operations whose polling job has permanently failed and marks the operation and its pollable job as'failed'. No orphan mitigation is triggered: an update targets a resource that already exists, so it must not be deprovisioned.ServiceOperationsDeleteStuckInProgressRetry— detects stuck service-instance'delete'operations and internally retries by re-enqueuing the originalDeleteServiceInstanceJob(reset pollable toPOLLING). Delete should ultimately succeed, so retry is the correct recovery. The original@start_timeis preserved, soReoccurringJob's max-duration expiry still bounds retries and a permanently-broken broker still ends in a terminalfailedstate.ServiceOperationsBindingDeleteStuckInProgressRetry— same retry mechanism for credential bindings and service keys (re-enqueues the originalDeleteBindingJob). Route bindings are skipped, matching the create-cleanup precedent.All three are wired into the clock scheduler with a configurable
frequency_in_seconds(default 1h), with matching properties in capi-release.An explanation of the use cases your change solves
A service instance/binding create, update, or delete gets stuck
'in progress'after a transient CCDB outage kills the polling job. Previously this required manual operator intervention (which the CFP requirement forbids). With this change:update: the client sees a definitive
failedstate instead of a hang, and the existing resource is left untouched (no unwanted deprovision).instance/binding delete: the delete is automatically resumed and completes, or reaches a terminal
failedstate once the max polling window elapses.Links to any other associated PRs
Stuck service operations recovery capi-release#681
I have reviewed the contributing guide
I have viewed, signed, and submitted the Contributor License Agreement
I have made this pull request to the
mainbranchI have run all the unit tests using
bundle exec rakeI have run CF Acceptance Tests