A preview that fails is tried once more, a stack whose lock is held is busy, and the job log shows the tool's last lines
Decision record 0117
Amends 0012 (a preview failure is a preview that failed twice, and the
strictinput does not count a busy stack), 0022 (one reason is picked from the tool’s words: a held state lock, which has no exit code of its own), 0009 (the row marker gains the optional keybusy), 0029 and 0114 (a Busy section right under Preview failed), 0040 (the counts line gainsN busy, only when there is one), 0037 (the summary lists busy stacks apart). Built as slice 5.54, for onboarding log hurdle 30.
On 2026-09-30 the first real user’s dashboard turned red for one stack of 54. The scan had previewed that stack three minutes after a deploy of it, the preview failed in 11.8 seconds, the other 53 stacks were in sync, and the row read preview failed: the tool exited with an error (exit code 1). It stayed that way until the next scan, a day later. The owner’s words: users do not want to see this.
What the job log of that scan shows:
- The cause was a registry that answered 502 once. The tool’s diagnostic reads
failed to authorize: failed to fetch oauth token: unexpected status from POST request to https://ghcr.io/token: 502 Bad Gateway. Nothing was wrong with the stack, and no lock was held. A second preview a few seconds later would have worked. - The tool’s error was in the log, and nobody found it. While a preview runs, the scan prints the tool’s stderr behind the stack id (0022, slice 5.9). Pulumi with
--jsonwrites its error as a diagnostic in the document on stdout, which holds values and is never printed live (0021). So the three lines printed for the stack were SDK warnings, the line after them saidpreview failed, the tool exited with an error (exit code 1), and the diagnostic came 175 lines later, inside the closed group of the stack, one of 54 groups. The warning on the run and the error of thestrictinput name the reason from the fixed list and not where the words are.
Decision
- The job log shows what the tool wrote right under the line that says a preview did not work. At most the last 20 lines, each behind the stack id in square brackets, as the live lines are, under
What the tool wrote for <stack id>:, orThe last 20 of the 57 lines the tool wrote for <stack id>, which the group of the stack holds in full:. A tool that wrote nothing getsThe tool wrote nothing for <stack id>., so the log always says something. The lines are the tool log the group already holds: stderr and the diagnostics, never the rest of stdout (0021). They go to the job log and nowhere else (0022): the row, the summary, the warning on the run and the result file keep their reasons from the fixed list, and the failure line of 0113 is untouched. - A preview that failed is tried once more in the same scan. After the pool is done, every stack whose preview failed for a reason a second run can change is previewed again, through the same pool, after one pause of 10 seconds for all of them. Only a preview that failed twice is a preview failure. A second try that works gives the row of its preview, with no warning, and
strictstays green. - Which reasons get a second try is decided from the reason, a fact of Sluiceway’s own, never from the tool’s words: the tool exited with an error, the tool could not authenticate, a resource operation failed in the tool, the tool gave up on a time limit of its own, and a busy stack. Not tried again: what the same commit gives again (the stack does not exist in the backend, the configuration is invalid, output that cannot be read or held, a step Sluiceway does not know, an env file that does not load), the preview’s own time limit, which a second try would spend a second time, and a bug of Sluiceway’s own.
- When every preview of the round failed, none is tried again. That is the broken environment of 0012, and a second round would take as long again to say the same. The job log says so.
- The job log says each step:
2 previews did not work and are tried once more after a pause of 10 s: a:prod, b:prod., thenPreviewed a:prod again in 9.1 s: in sync. The group of a stack that was previewed twice keeps the first try: its reason and the tool’s words of that try. The warning of a stack that failed both times readsThe preview of a:prod failed twice: <reason>. - The time of a stack is both tries added up, in the result file and in the summary. The preview that counts for the late read (0004) is the second one.
- A stack whose lock another update holds is busy, not failed. The fixed list of 0022 gains the reason
another update holds the stack's lock. It gets a second try like any other. A stack that is busy both times gets a row that says so:- **network:dev** · busy: another update holds the stack's lock, the next scan previews it · [run](…). No warning on the run, thestrictinput does not count it, and “every preview failed” does not either. The job log saysnetwork:dev is busy: another update holds the stack's lock. It was tried twice. Its row says busy, and the next scan previews it. - The busy row has the marker state
preview-failedand the keybusy="true". The state is the one that fits, as it did for the row of 0113: no box, no diff, and every scan previews the stack, a narrowed one too (0010), so “the next scan previews it” is true. A reader of the published shape (0096) that does not know the key draws a preview failure, which is a row that could not be previewed and so still true. The dashboard itself tells the two apart by the key: a busy row does not make the header failing, the counts line counts it asN busywith the white dot, shown only when there is one, and it is listed under## Busy, right under Preview failed wherever the layout puts that (0114), and never turned off. The alt text of the picture ends with, 1 stack is busy. The summary counts and lists busy stacks apart in the same way. The key is a display cache likefailed: nothing is decided from it. - Busy is recognised in the OpenTofu family only, because that is where a preview takes a lock. Measured and recorded:
tofu planandterraform plantake the state lock. While a deploy of the stack runs, the plan exits with 1, the code of any failed plan, and its JSON log holds one error diagnostic with the summaryError acquiring the state lock. Recorded as the scenariostate-lockedwith tofu 1.11.0 and 1.12.6 and terraform 1.14.0 and 1.16.3: the recorder holds a real deploy of a resource that takes a minute behind the plan. A Terragrunt unit and a cdktf stack go through the same code, and no scenario of theirs records it.pulumi previewtakes no stack lock on a file backend. The scenariopreview-lockedrecords a preview under a held lock with v3.229.0 and v3.263.0, and it runs. The drift check takes none either (scenariodrift-locked). Pulumi Cloud’s page on update conflicts names a preview-only run as the one that does not collide with a deploy; that is read from the docs and not recorded, because the recorder has no account. So no Pulumi preview is known to be busy, and the adapter looks for nothing.helm diff upgradedoes not fail on a release with an install or upgrade in progress: run against a release inpending-installwith helm v3.18.0 and the diff plugin v3.15.11, it exits with 0 and prints the diff. “Another operation is in progress” is whathelm upgradesays, which is the deploy.
- How the adapter knows a held lock: the tool gives it no exit code, so this is the one reason picked from the tool’s words. The adapter reads the plan’s JSON log for a diagnostic of severity error whose summary is that phrase, compared whole. A phrase in a detail, in a warning or in a line that is no diagnostic is not read, so a program that quotes it stays a tool error. Nothing of the words is shown: the reason is a constant. A plan that ran out of its time limit is still a plan that timed out.
Why reading the tool’s words here keeps the promise of 0022
Record 0022 is about what leaves the job log: no text the tool wrote reaches the issue, a comment, a deployment record or a summary. That holds. What changes is how one constant is picked. The exit code was the only thing read until now because it is the most stable fact a tool gives, and for a missing stack the tool documents one. For a held lock no tool does. The cost of being wrong is small in both directions: a tool that changes its phrase gives the tool error of before, tried twice, and a false match would call a failed preview busy for one scan, which the next scan previews again. Neither can lead to a deploy: a busy row has no box, no hash and no fingerprint, and every deploy is held to its own fresh preview (0008).
Rejected
- Trying again inside the pool slot, right after the failure. Each failed stack would hold its slot through a pause of its own, so a scan with many failures would wait many pauses. One pause after the pool costs 10 seconds a scan, and the first failures get the longest wait, which is what a lock or an outage needs.
- More than one second try, or a longer pause. A scan is on the clock of whoever waits for the dashboard. One try more catches the blip, and the next scan is the third try.
- Keeping the row the stack had and adding a note. It is the calmest dashboard: a busy stack in sync would stay in the fold. But a row block is never patched inside (0004), the scan would show a diff it did not see under a scan line that says it scanned, and a narrowed scan would carry the row with nothing to make it preview the stack again.
- A new row state
busy. Honest in the marker, and a new state in the published shape (0096) for a row that lives until the next scan. The same choice as 0113. - Trying the preparation again too. An init that fails for a registry blip is the same kind of failure. It runs alone and before the pool (0053), and a second try there is a change of its own. It is on
docs/later.md. - Calling a deploy that meets a held lock busy. A deploy that did not start because another one runs is a failure line on the row today, with the exit code. It is the case where Pulumi and Helm do say “in progress”. On
docs/later.md: it needs the deploy reasons of 0022 looked at as a whole. - A short excerpt of the tool’s error on the row. Rejected in 0022, and still: the row is emailed and kept in edit history. The job log is one click from the row, and now shows the words where the eye lands.
Consequences
- A scan with a failed preview takes 10 seconds and one preview longer. A scan without one takes no longer.
- A transient failure no longer reaches the dashboard, so the rows that say preview failed are the ones a person has to look at.
src/core/preview-retry.tsholds the rule,src/adapters/opentofu/state-lock.tsthe one phrase. The recorder gains aholdstep, a command left running behind the steps after it, for scenarios about a lock.- The second try is the scan’s alone. The fresh preview of
apply, the previews after a merge (0071) and of a pull request (0101), andsluiceway previewrun once, as before.applythat meets a held state lock ends its record withthe preview before the deploy failed: another update holds the stack's lock. - The result file keeps the state
preview-failedfor a busy stack, with the reason from the list. Its dashboard counts do not count busy rows as preview failures, and neither does thepreview-failedoutput.