We would like to share details about an incident that affected Phrase Orchestrator on July 27–28, 2026. During this period, workflow executions in the Next-Gen Workflow Engine were unable to progress and remained stuck in an "executing" state. No data was lost during the incident. This post-mortem explains what happened, when it was resolved, and the steps we have taken to prevent a recurrence.
27 July 2026 at 18:55 CEST – The workflow engine began producing errors as database query performance degraded. Workflow executions stalled and stopped progressing.
27 July 2026 at 20:54 CEST – The first customer report of executions stuck in "executing" was received.
27 July 2026 at 22:39 CEST – The incident was formally declared.
28 July 2026 at 02:18 CEST – A service restart provided temporary relief; workflow executions resumed.
28 July 2026 at 05:15 CEST – The issue recurred as the underlying database performance problem persisted.
28 July 2026 at 09:56 CEST – The root cause was identified and addressed. No executions were lost; however, due to partial service restarts, some actions within executions were retried, which may have caused a small number of executions to fail that otherwise would have succeeded.
28 July 2026 at 13:46 CEST – The full backlog of stalled executions was confirmed as cleared. The system was declared stable.
28 July 2026 at 13:48 CEST – Incident resolved.
The incident was caused by progressive bloat in database indexes used by the workflow job scheduling system. The performance of these particular indexes gradually degraded over time as they accumulated dead index entries from prior writes and updates.
The job scheduling engine acquires database-level coordination locks while querying these indexes to determine which jobs to dispatch. As the index lookups grew slower, they began exceeding the database's configured statement timeout. When a lookup was canceled by the timeout, the scheduling process responsible for that work crashed and restarted. With no schedulers running, no workflow steps could be dispatched and all in-progress workflow executions became stuck.
The database server itself remained healthy throughout the incident, with normal CPU and connection levels. The problem was exclusively lock and latency contention within the scheduling layer. A service restart cleared the crashed processes and temporarily restored execution. However, because the index bloat was still present, the same degradation recurred once query load resumed. A manual index rebuild fully restored performance and resolved the issue.