Skip to content

fix(workflows): namespace nested descendant step ids in loops/fan-out - #4338

Open
Noor-ul-ain001 wants to merge 18 commits into
github:mainfrom
Noor-ul-ain001:fix/nested-step-id-namespace-collision
Open

Noor-ul-ain001 wants to merge 18 commits into
github:mainfrom
Noor-ul-ain001:fix/nested-step-id-namespace-collision

Conversation

@Noor-ul-ain001

@Noor-ul-ain001 Noor-ul-ain001 commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • while/do-while loop bodies and fan-out templates namespace nested step ids per iteration/item so logs and state.step_results entries stay unique — but the namespacing only rewrote the id of the immediate child step, not any descendant nested deeper (e.g. a shell step inside an if inside a while body, or inside a fan-out template's if/switch branch). That grandchild kept its bare, unnamespaced id across every iteration/item, so each iteration/item silently overwrote the previous one's entry in state.step_results under that same key — only the last iteration's or item's result for that nested step ever survived, and no per-iteration/per-item record of it ever existed.
  • This is also a correctness gap beyond bookkeeping: nested/template step ids are deliberately exempted from the workflow's global id-uniqueness validation, on the assumption that runtime namespacing makes any collision safe. Since only the top-level child was actually namespaced, a step nested one level deeper could collide with an unrelated step of the same id elsewhere in the workflow and silently overwrite its result.
  • Fix: add _rename_step_tree_ids, which recursively rewrites every id in a step's subtree (walking then/else/steps/default/cases.* — the same nesting keys overlays/merge.py walks for step-tree attribution) and returns a {new_id: original_id} map. Both the while/do-while loop body and fan-out's run_item now use this helper instead of renaming only the top-level id, and alias every renamed descendant's result back to its original id (mirroring the existing single-level aliasing) so sibling steps within the same iteration/item and code reading steps.<id>.output after the loop/fan-out still see that iteration's/item's value.

Follow-up fixes (review rounds)

  • Namespace loop iteration 0 too and alias immediately (not after the whole subtree finishes), so the first iteration gets its own recorded entry and a same-iteration sibling sees an earlier sibling's aliased value right away.
  • Stop double-renaming loop bodies nested inside an outer loop/fan-out, and stop an unsafe fan-out bare-id alias write from racing across concurrently-running items (each concurrent item now gets a private context.steps overlay instead of writing the shared one directly).
  • Stop using a ChainMap overlay for that per-item isolation: _resolve_dot_path (what every {{ steps.x... }} expression goes through) only descends through isinstance(current, dict), and ChainMap isn't a dict subclass, so every such expression inside a concurrent fan-out item silently resolved to None. Replaced with a plain dict copy.
  • Two further edge cases combining previously-separate test scenarios:
    1. Reserved-id collision + item-local sibling reference: when a fan-out template id collides with a real, globally-unique step elsewhere in the workflow, the guard that protects that unrelated step's persisted result also suppressed the write needed for a later sibling step within the same item to resolve the colliding id locally — the sequential path (which runs directly against the real shared context.steps, unlike the concurrent path's private copy) read the outside value instead of the item's own. Fixed by always updating context.steps for the local read while still skipping the persisted-state write for the collision case; the sequential path now snapshots and restores that entry so the transient item-local write never leaks past the item that made it.
    2. Nested while/do-while aliasing dropped under concurrency: a while/do-while body nested inside a fan-out template renames itself dynamically at runtime, so it has no entry in the fan-out's own static id map. The concurrent path's per-item alias publishing was reconstructed from that static map after the item finished, which silently missed the while body's dynamic alias — the namespaced entries (e.g. fan:item:0:leaf:0) were still recorded correctly, but the convenience bare-id entry (leaf) never was, unlike the sequential path (which publishes immediately during execution instead of reconstructing afterward). Fixed by threading a mutable accumulator through _execute_steps that captures every alias write at the point it happens, regardless of nesting depth or whether the rename was static or dynamic.
  • Latest round — two more issues from Copilot's next review pass:
    1. Reserved-id restore set too narrow (landed as a separate 1-line Copilot Autofix commit, then covered here by a regression test): the sequential path's snapshot/restore for a reserved-id collision only covered ids present in the fan-out's own static id_map. A nested while/do-while body's id is deliberately absent from that map (it re-namespaces itself dynamically at runtime — see point 2 above), so if such a body's id collided with a reserved id, _execute_steps's collision guard still wrote the item-local value into the live context.steps[orig_id], but nothing restored it afterward — the persisted state.step_results entry stayed protected, but the last item's transient value leaked into the live context past the end of the fan-out. Fixed by snapshotting/restoring every id in context.reserved_step_ids, not just the ones in id_map.
    2. Nested fan-out doesn't inherit enclosing isolation: a fan-out template nested inside another fan-out's item decided its own isolation purely from its OWN max_concurrency, with no awareness that the ENCLOSING item might itself be one of several running concurrently. A nested fan-out with max_concurrency <= 1 unconditionally wrote its bare-id alias directly to the real, persisted state.step_results as each inner item completed — so two outer items running that same nested fan-out concurrently could race on that key, the exact race the outer fan-out's own isolation exists to prevent. It also never fed its aliases into the enclosing item's own accumulator, so an outer sibling step could never see them. Fixed by threading the enclosing _execute_steps call's alias_local_only/alias_records through _run_fan_out as parent_local_only/parent_alias_records: every item of a nested fan-out (sequential or concurrent) now isolates when the enclosing item does, and its aliases flow into the enclosing accumulator instead of the persisted store.

Test plan

  • test_while_loop_namespaces_nested_descendant_steps, test_fan_out_namespaces_nested_descendant_steps (original fix)
  • test_fan_out_concurrent_sibling_step_isolated_per_item, test_fan_out_concurrent_sibling_step_resolves_via_expression, test_fan_out_concurrent_alias_never_clobbers_unrelated_step, test_fan_out_sequential_alias_never_clobbers_unrelated_step (prior review rounds)
  • New this round: test_fan_out_reserved_id_collision_still_resolves_item_local_sibling, parametrized over max_concurrency 1 and 2 — a template's second step reads the first's output via {{ steps.first.output... }} where first collides with an outside reserved step; asserts the item-local value resolves under both paths, and the outside step's entry survives untouched
  • New this round: test_while_loop_nested_in_fan_out_aliases_to_true_original_id is now parametrized over max_concurrency 1 and 2 (previously sequential-only)
  • Verified both new/updated tests fail without this round's fix (test-the-test, git stash on engine.py only): test_fan_out_reserved_id_collision_still_resolves_item_local_sibling[1] failed with the outside value ('unrelated' == 'a') leaking into the item-local read; test_while_loop_nested_in_fan_out_aliases_to_true_original_id[2] failed with KeyError: 'leaf' — both reproducing exactly the two symptoms from review, then passed once restored
  • Ran the full tests/test_workflows.py suite (prior round): 935 passed, 20 pre-existing Windows symlink-elevation failures (need admin rights, unrelated to this change), 7 skipped — no regressions
  • New this round: test_nested_fan_out_inside_concurrent_fan_out_publishes_by_item_order — outer fan-out runs 2 items concurrently, each item's template is itself a fan-out (max_concurrency: 1) ending in a bare-id leaf step. Forces outer item "x" (item order 0) to finish strictly after outer item "y" (item order 1) via a threading.Event, then asserts state.step_results["leaf"] still resolves to "y" (the item-order convention) rather than "x" (which a completion-order race would produce)
  • New this round: test_fan_out_sequential_nested_while_reserved_collision_restores_live_context — a while/do-while body's dynamically-namespaced id collides with a reserved outside id; asserts context.steps["leaf"] is restored to the outside value after the fan-out finishes, not left holding the last item's transient value
  • Verified both new tests fail without their respective fix (test-the-test): the nested-fan-out test asserted 'x' == 'y' and failed with x when _run_fan_out didn't accept parent_local_only; the while-collision test failed with the last item's value ('c') leaking into context.steps["leaf"] when the restore loop was reverted to set(id_map.values()) & context.reserved_step_ids (the pre-Autofix form) — both passed once restored
  • Ran the full tests/test_workflows.py suite again this round: 937 passed, 20 pre-existing Windows symlink-elevation failures (unrelated), 7 skipped — no regressions
  • ruff check . clean

AI Disclosure

  • I did use AI assistance (describe below)

This PR, across all rounds including this one, was authored with Claude Code (Anthropic CLI), model Claude Sonnet 5, used in its interactive, agentic CLI mode with full local tool access (file read/edit, running pytest, ruff, git). For this round specifically: the AI agent traced both new review comments back to their root cause in _execute_steps/_run_fan_out (confirming the reserved-id restore gap was already fixed by a separate Copilot Autofix commit already on this branch, and writing the isolation-propagation fix and both regression tests above), ran the test-the-test verification and full suite described above, and wrote this description, under my direction. I reviewed the diff and test output myself before pushing and understand the change being made.

`while`/`do-while` loop bodies and `fan-out` templates namespace nested
step ids per iteration/item so logs and `state.step_results` entries stay
unique — but the namespacing only rewrote the id of the *immediate* child
step, not any descendant nested deeper (e.g. a `shell` step inside an `if`
inside a `while` body, or inside a `fan-out` template's `if`/`switch`
branch). That grandchild kept its bare, unnamespaced id across every
iteration/item, so each iteration/item silently overwrote the previous
one's entry in `state.step_results` under that same key — only the last
iteration's or item's result for that nested step ever survived, and no
per-iteration/per-item record of it ever existed.

This is also a correctness gap beyond bookkeeping: nested/template step ids
are deliberately exempted from the workflow's global id-uniqueness
validation, on the assumption that runtime namespacing makes any collision
safe. Since only the top-level child was actually namespaced, a step
nested one level deeper could collide with an unrelated step of the same
id elsewhere in the workflow and silently overwrite its result.

Fix: add `_rename_step_tree_ids`, which recursively rewrites every id in a
step's subtree (walking `then`/`else`/`steps`/`default`/`cases.*` — the
same nesting keys `overlays/merge.py` walks for step-tree attribution) and
returns a `{new_id: original_id}` map. Both the while/do-while loop body
and fan-out's `run_item` now use this helper instead of renaming only the
top-level id, and alias every renamed descendant's result back to its
original id (mirroring the existing single-level aliasing) so sibling
steps within the same iteration/item and code reading `steps.<id>.output`
after the loop/fan-out still see that iteration's/item's value.

## Test plan
- Added `test_while_loop_namespaces_nested_descendant_steps` and
  `test_fan_out_namespaces_nested_descendant_steps` to
  `tests/test_workflows.py::TestWorkflowEngine`: a `shell` step nested
  inside an `if` inside a `while` body (and inside a `fan-out` template)
  gets a distinct namespaced `state.step_results` entry per
  iteration/item, while the unprefixed key still holds the latest value.
- Verified both fail without the fix (test-the-test): the namespaced keys
  (`retry-loop:leaf:1`, `fan:leaf:0`, etc.) were simply absent, and
  `step_results` only ever held the last iteration's/item's bare-keyed
  entry — reproducing the exact bug.
- Ran the full `tests/test_workflows.py` suite: 926 passed, 20 pre-existing
  Windows symlink-elevation failures (need admin rights, unrelated to this
  change), 7 skipped. All `While`/`DoWhile`/`FanOut`/`FanOutConcurrency`
  tests pass, including the concurrent-execution and per-thread context
  isolation tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJHJ2dHP2RVCNncHqN8Qm9

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Initial loop iterations remain unnamespaced, while delayed and shared aliases break sibling references and concurrent fan-out isolation.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Adds recursive runtime namespacing for nested workflow steps in loops and fan-out templates.

Changes:

  • Adds recursive descendant ID rewriting and aliasing.
  • Adds while and fan-out regression tests.
  • Preserves namespaced execution results.
File summaries
File Description
src/specify_cli/workflows/engine.py Recursively namespaces nested step IDs.
tests/test_workflows.py Tests nested loop and fan-out results.
Review details

Suppressed comments (1)

src/specify_cli/workflows/engine.py:1447

  • These aliases are created only after the entire renamed subtree finishes. If a branch contains step a followed by step b that references steps.a, b executes before a is aliased and reads the previous iteration's value (or no value), whereas both steps previously used their bare IDs. Alias each descendant immediately after that descendant completes so intra-branch references retain their existing semantics.
                            for new_id, orig_id in id_map.items():
                                if new_id in context.steps:
                                    self._record_result(
                                        context, state, orig_id,
                                        context.steps[new_id],
                                    )
  • Files reviewed: 2/2 changed files
  • Comments generated: 2
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/specify_cli/workflows/engine.py
Comment thread src/specify_cli/workflows/engine.py Outdated
@mnriem
mnriem requested a balanced review from Copilot and removed request for Copilot September 1, 2026 22:54
@mnriem

mnriem commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Please address Copilot feedback

… post-subtree

Addresses Copilot review feedback on PR github#4338:

- while/do-while loop iteration 0 ran through a separate, unnamespaced
  code path before the loop-specific namespacing logic was reached, so
  it had no dedicated state.step_results entry and was immediately
  overwritten the moment iteration 1's aliasing ran. Every iteration,
  including the first, now goes through the same
  _rename_step_tree_ids + alias_map path.

- Bare-id aliasing for a namespaced descendant happened only after its
  entire renamed subtree finished executing, so a later sibling step
  in the same iteration/item that referenced an earlier sibling by its
  original id ran before that alias existed and read a stale (or
  absent) value. _execute_steps now threads an alias_map through so
  each descendant is aliased immediately after it completes, not in
  bulk afterward.

- For a concurrent fan-out (max_concurrency > 1), that same bare-id
  alias write raced across worker threads sharing context.steps, so
  one item's sibling read could observe another item's value. Each
  concurrent item now runs against a private ChainMap overlay for its
  bare-id aliases; only the namespaced (disjoint-key) result is
  published to shared state during execution. Once every item has
  finished (back on the single thread), the last item's aliases are
  applied to shared state once, deterministically.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NhR6g8xT8at5pPMhkrC3e2
@mnriem
mnriem requested a balanced review from Copilot September 8, 2026 13:27
@mnriem mnriem added author-awaiting Waiting on author response author-over-cap Over the 3-open-PR cap or repetitive batch submissions — please consolidate triage-nice-to-have Verdict: evidence-backed fix or greenlit feature — land after review labels Sep 8, 2026
@mnriem

mnriem commented Sep 8, 2026

Copy link
Copy Markdown
Collaborator

Strong direction — this is a real correctness bug (silent per-iteration result overwrite + the uniqueness-exemption collision risk), disclosed and tested. Two review findings to resolve before merge: the first loop iteration still isn't namespaced (it goes through the generic result.next_steps path before _rename_step_tree_ids runs), and the post-subtree aliasing breaks sibling references and concurrent fan-out isolation — a later child can execute before an earlier renamed child is aliased back. Please address those and re-request review.

Separately: you currently have 7 open PRs. Per CONTRIBUTING, that's past the point where further submissions may be deprioritized in the queue. Several are small workflow/config validation fixes (#4319, #4323, #4324, #4325) — could you consolidate the related ones so review can focus? This PR can stay standalone; it's the smaller validation fixes I'd group.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Nested-loop aliasing, expression resolution, and bare-ID collisions remain incorrect.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review details

Suppressed comments (2)

src/specify_cli/workflows/engine.py:1457

  • When this loop is itself inside another loop or a fan-out template, its result.next_steps were already recursively renamed by the outer _rename_step_tree_ids. Renaming them again treats the first generated ID as the original (for example fan:leaf:0 becomes fan:while:0:fan:leaf:0:0) and aliases only back to fan:leaf:0, so a later inner-loop sibling reading steps.leaf cannot see the current value. Preserve canonical original IDs/compose the outer alias map, or avoid eagerly renaming bodies that receive their own runtime namespace.
                            ns_copy, id_map = _rename_step_tree_ids(
                                ns, step_id, str(_loop_iter),
                                default_id=f"step-{ns_idx}",

src/specify_cli/workflows/engine.py:1752

  • The concurrent path has the same bare-ID collision as the sequential alias path: because each fan-out template is validated in its own ID set, orig_id can already belong to an unrelated workflow step or another fan-out, and this call deterministically overwrites that result after the pool joins. Namespaced entries survive, but the unrelated bare-key result still does not, so runtime namespacing has not made collisions safe.
            for orig_id, data in alias_slots[last_idx].items():
                self._record_result(context, state, orig_id, data)
  • Files reviewed: 2/2 changed files
  • Comments generated: 2
  • Review effort level: Balanced

Comment thread src/specify_cli/workflows/engine.py
Comment thread src/specify_cli/workflows/engine.py Outdated
Noor-ul-ain001 and others added 2 commits September 13, 2026 20:15
…liasing

Two follow-up bugs in the nested step-id namespacing added for loops and
fan-out templates:

1. A while/do-while step nested inside an outer loop iteration or
   fan-out item had its OWN 'steps' body eagerly renamed by the outer
   _rename_step_tree_ids pass (since 'steps' is a walked nesting key).
   When that while step then ran its own per-iteration rename, it treated
   the already-namespaced id (e.g. "fan:leaf:0") as the original and
   aliased back to that synthetic id instead of the workflow author's
   real bare id ("leaf") -- so a doubly-prefixed id like
   "fan:while:0:fan:leaf:0:0" was the only place the value ever landed,
   and state.step_results["leaf"] was never populated at all.
   _rename_step_tree_ids now still renames a nested while/do-while step's
   own id, but leaves its 'steps' body untouched for the loop's own
   runtime namespacing to rename exactly once, against the real ids.

2. A fan-out template's bare-id convenience alias (writing the current
   item's result under its unprefixed id, so `steps.<id>` sees the latest
   value) could silently clobber an unrelated, distinctly-authored step's
   result if the two happened to share an id -- fan-out templates are
   deliberately exempt from the workflow's global id-uniqueness check, so
   nothing prevents that collision. This applied to both the sequential
   path's immediate alias write and the concurrent path's deferred
   post-join write. Added _collect_reserved_step_ids, which computes the
   set of step ids declared outside any fan-out template (mirroring
   _validate_steps' global id set), and gate both alias-write sites on
   it via a new alias_may_collide flag -- set only for fan-out's own
   template execution, so ordinary while/do-while loop-body aliasing
   (always globally unique by validation) is unaffected.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U74yBbvVQCPwB7Ed8Dzeu6
…n-out

The concurrent fan-out item isolation gave each item's context.steps a
ChainMap(local_overlay, original_steps) so a sibling step could resolve
an earlier one by its bare id without racing other concurrently-running
items. But _resolve_dot_path (the function every {{ steps.x.output... }}
expression goes through) only descends when isinstance(current, dict)
is true, and ChainMap is not a dict subclass -- so _build_namespace's
`ns["steps"] = context.steps or {}` put a non-dict at "steps", and every
steps.* expression evaluated inside a concurrent fan-out item silently
resolved to None. The existing test only exercised context.steps.get()
directly, which IS supported by ChainMap, so it never caught this.

Replaced the ChainMap with a plain dict snapshot of the shared steps
dict, taken once when the item starts. It still isolates this item's
own writes from concurrently-running siblings (a fresh copy per item,
never shared), it's a real dict so isinstance(..., dict) and every
expression path work normally, and the snapshot copy is safe under the
GIL against a concurrent sibling's writes to the source dict.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U74yBbvVQCPwB7Ed8Dzeu6
@Noor-ul-ain001

Copy link
Copy Markdown
Contributor Author

Addressed everything in this review — both the two posted inline comments and the two suppressed ones summarized in the review body.

1. Double-renaming when a while/do-while loop is nested inside an outer loop iteration or fan-out item (the suppressed comment quoting engine.py:1457): confirmed exactly as described — a nested while step's own steps body was being eagerly renamed by the outer _rename_step_tree_ids pass (since steps is a walked nesting key), so the loop's own per-iteration rename then treated that already-namespaced id as "original" and aliased back to a synthetic id like fan:leaf:0 instead of the real leaf. state.step_results["leaf"] was never populated at all. Fixed by having _rename_step_tree_ids still rename a nested while/do-while step's own id, but leave its steps body untouched — the loop renames it exactly once, against the real original ids, when it actually runs.

2 & 3. Fan-out bare-id alias clobbering an unrelated step (both the suppressed concurrent-path comment quoting engine.py:1752, and the posted sequential-path comment): fan-out template ids are exempt from the workflow's global id-uniqueness check, so a template's bare id can coincide with a real step elsewhere. Added _collect_reserved_step_ids (mirrors _validate_steps's global seen_ids, minus the fan-out-template exemption) and gated both the sequential path's immediate alias write and the concurrent path's deferred post-join write on membership in that set — scoped via a new alias_may_collide flag so ordinary loop-body aliasing (always globally unique) is untouched.

4. ChainMap breaking steps.* expressions in concurrent fan-out (posted comment): confirmed — isinstance(ChainMap(...), dict) is False, and _resolve_dot_path only descends through actual dicts, so every {{ steps.x.output... }} expression inside a concurrent fan-out item silently resolved to None. Replaced the ChainMap overlay with a plain dict snapshot per item, which keeps the same isolation property while being a real dict.

Each fix has a dedicated regression test, and each was confirmed via test-the-test (temporarily reverting just that fix) to fail against the prior code and pass now. Full test_workflows.py run clean except the pre-existing Windows-symlink-elevation failures unrelated to this change (20 failures, same set as baseline on main). Pushed as 7365250 and a8188b6.

@mnriem mnriem removed the author-over-cap Over the 3-open-PR cap or repetitive batch submissions — please consolidate label Sep 14, 2026
@mnriem
mnriem requested a balanced review from Copilot September 15, 2026 17:31

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@mnriem

mnriem commented Sep 15, 2026

Copy link
Copy Markdown
Collaborator

The earlier review fixes are present, but two combined cases still need correction.

First, combine the reserved-ID collision case with a sibling reference inside the fan-out item. With an outside first result and an item-local first, sequential execution reads the outside value while concurrent execution correctly reads the item’s value. Protecting shared state must not suppress item-local lookup.

Second, run the nested-while test with both one and two workers. Sequential execution publishes the latest bare leaf result; concurrent execution leaves it absent, despite preserving the namespaced records. Please carry those nested aliases through deterministic publication.

Please cover both cases across worker counts and complete the plain AI disclosure with tool, mode/settings, and extent of assistance; the model attribution is already present.

Drafted for @mnriem with assistance from GitHub Copilot (model: GPT-6 Astra; interactive comment drafting).

@mnriem mnriem added the author-needs-disclosure AI use, or the agent/model/settings behind it, not disclosed per CONTRIBUTING label Sep 15, 2026
Two issues surfaced by combining previously-separate test scenarios:

1. Reserved-id collision + item-local sibling reference. When a fan-out
   template's step id collides with a real, globally-unique step elsewhere
   in the workflow (context.reserved_step_ids), _execute_steps's alias
   write was skipped entirely to protect that unrelated step's persisted
   result. For the SEQUENTIAL fan-out path, which runs each item directly
   against the real shared context.steps object (no per-item copy), that
   also suppressed the write to context.steps -- so a later sibling step
   within that SAME item's own template, referencing the colliding id,
   read the outside value instead of the item's own. The concurrent path
   was unaffected: its per-item context.steps is already a private copy,
   so this collision guard never touched the shared object there in the
   first place.

   Fix: the alias write now always updates context.steps for a colliding
   id (needed for item-local resolution), it just never calls
   _record_result for it (still protecting the persisted/shared entry).
   For the sequential path specifically, run_item now snapshots each
   colliding id's pre-item value and restores it once the item finishes,
   since that write landed directly in the real shared object.

2. Nested while/do-while body aliasing dropped under concurrency. A
   while/do-while step nested inside a fan-out template re-namespaces its
   own body dynamically at runtime (see _rename_step_tree_ids), so it has
   no entry in the fan-out's own static id_map. The concurrent fan-out
   path reconstructed each item's alias_records for deferred publishing by
   walking that static id_map after the item finished -- which silently
   missed the while body's dynamic aliases. The sequential path was
   unaffected: it publishes each alias immediately during execution
   instead of reconstructing it afterward.

   Fix: _execute_steps now takes an optional alias_records accumulator
   and populates it at write time, threaded through the recursive
   while/do-while and if/switch calls, so it captures every alias
   regardless of nesting depth or whether the rename was static or
   dynamic. run_item uses this instead of reconstructing from id_map.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@Noor-ul-ain001

Copy link
Copy Markdown
Contributor Author

Thanks for combining those two cases — both were real, and each was masked by the other's path working correctly:

  1. Reserved-id collision + item-local sibling reference: the guard that protects an outside colliding step's persisted result was also suppressing the context.steps write a same-item sibling needs for local resolution — only on the sequential path, since it runs directly against the real shared context.steps (the concurrent path already has a private per-item copy, which is why it was unaffected). Fixed by always writing context.steps for the local read while still skipping the persisted-state write for the collision case, with a snapshot/restore in run_item so that transient write never leaks past the item that made it.
  2. Nested while/do-while aliasing dropped under concurrency: a nested while body renames itself dynamically at runtime, so the concurrent path's post-hoc reconstruction of alias records from the fan-out's static id map missed it — the namespaced entries were fine, only the bare-id convenience alias (leaf) was silently dropped. Fixed by threading a mutable accumulator through _execute_steps that records every alias at write time, at any nesting depth, instead of reconstructing afterward.

Added test_fan_out_reserved_id_collision_still_resolves_item_local_sibling and parametrized the existing test_while_loop_nested_in_fan_out_aliases_to_true_original_id over max_concurrency 1 and 2, as requested. Verified both fail on the pre-fix code with the exact symptoms described (outside value leaking into the item-local read; KeyError: 'leaf'), then pass with the fix. Full suite: 935 passed, same 20 pre-existing Windows symlink failures, 7 skipped.

Also added a plain AI Disclosure section to the PR body (tool/model, mode, extent of assistance) covering this round and the prior ones.

@mnriem
mnriem requested a balanced review from Copilot September 18, 2026 12:19
@mnriem mnriem removed author-awaiting Waiting on author response author-needs-disclosure AI use, or the agent/model/settings behind it, not disclosed per CONTRIBUTING labels Sep 18, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Concurrent fan-outs can incur quadratic dictionary-copying work as completed-item results accumulate.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Bounded fan-outs incur quadratic copying of generated keys

src/​specify_cli/​workflows/​engine.py:1837

This snapshot can make bounded fan-outs quadratic in the item count. Each completed item publishes its namespaced keys back into original_steps below, so with a small fixed worker count every subsequently started item copies an ever-growing dictionary; for N items with K template IDs, the copies process O(K·N²) entries. There is no item-count limit, so large fan-outs can spend most of their time copying prior items' generated results. Consider snapshotting the pre-fan-out namespace once for concurrent workers, or using a dict-compatible overlay that keeps item-local aliases separate without copying previously completed item keys.

Low severity Update comment to reflect alias merging across collected items

src/​specify_cli/​workflows/​engine.py:1825

This comment no longer matches the implementation: the post-join code folds aliases from every collected item, preserving an earlier item's alias when later items did not run that ID. Describing it as applying exactly one item's aliases contradicts the logic at lines 2052-2065 and the new earlier-item regression test.

Concurrent fan-out items each copied the live shared steps dict, which
completed items keep growing with their namespaced results. With a small
worker count that made every later item's copy larger, O(K * N^2) work
overall, and what a late-starting item saw depended on scheduling. The
concurrent path now takes one snapshot before any worker starts and each
item copies that.

Also corrects the run_item comment that still described applying only
the last item's aliases (the post-join code folds every collected item's
aliases in item order), and adds an end-to-end resume() test for the
reserved-id collision guard, since resume() builds its own run context
separately from execute().

Assisted-by: Claude Code (model: Claude Opus 5.5, autonomous)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Noor-ul-ain001

Noor-ul-ain001 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor Author

Addressed the latest Copilot review in 109efc3:

  • Quadratic snapshot copying in concurrent fan-out (_run_fan_out): the concurrent path now takes one dict(context.steps) snapshot before any worker starts, and each item copies that through a new base_steps argument to run_item. Before this change, each item copied the live dict, which completed items keep growing with their namespaced results. That cost O(K * N^2) overall, and a late-starting item's view depended on scheduling. New test: test_concurrent_fan_out_items_copy_fixed_pre_fan_out_snapshot (6 items, 2 workers). It fails on the previous commit because item 2 sees fan:probe:0, item 4 sees three earlier keys, and so on. The sequential path is unchanged: there, later items are meant to see earlier items' results.
  • Stale run_item comment: it now says that the concurrent caller folds every collected item's aliases in item order. It no longer claims that only the last item's aliases are applied.
  • End-to-end reserved-id coverage: execute() was already covered. Added a resume() counterpart, with details in the inline reply.

Local results: tests/test_workflows.py + tests/specify_cli/workflows: 1074 passed, 11 skipped, and 24 failed. All 24 failures are symlink-creation tests that need elevation on Windows, and they fail the same way without this change. ruff check is clean.

On the previous CI run, pytest (ubuntu-latest, 3.14) failed in tests/specify_cli/workflows/test_catalog_versions.py, which this PR doesn't touch. It is flaky on main too: I reproduced it in 1 of 4 local runs on upstream d2ddd910. The cause is that _archive() uses ZipFile.writestr("workflow.yml", ...), which stamps the current time. Two calls that land in different 2-second DOS-time windows produce different bytes, so the catalog's declared sha256 no longer matches the "downloaded" archive.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The implementation addresses the identified collision and concurrency paths with comprehensive regression coverage.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

kartsan03 added a commit to kartsan03/spec-kit that referenced this pull request Oct 1, 2026
_archive() stamped workflow.yml with the current time, which zip stores
at 2-second resolution. The tests hash one build for the catalog and
serve another as the download, so a run that crossed an even second
between the two failed the integrity check (windows-latest on Python
3.13 in github#4338 and on 3.14 in github#4798).

Assisted-by: Claude Code (model: claude-opus-5-5, autonomous)
@Noor-ul-ain001

Copy link
Copy Markdown
Contributor Author

@mnriem could you re-run the failed CI jobs on 109efc3?

The latest Copilot review recommends approval. Its one listed finding (end-to-end coverage of the reserved-id wiring) is the older thread that was already addressed: test_fan_out_reserved_id_collision_end_to_end covers execute() and test_fan_out_reserved_id_collision_end_to_end_on_resume covers resume(), both on the sequential and concurrent paths. To confirm they guard the wiring, I replaced both _collect_reserved_step_ids(definition.steps) call sites with frozenset(). All 4 tests then fail, and they pass again with the real code.

The only failing test in CI is tests/specify_cli/workflows/test_catalog_versions.py::test_unqualified_add_still_uses_current_and_invalid_version_scope on windows-latest 3.13. It is the sha256 timestamp flake from #4788 that I described earlier, and it is unrelated to this PR. The macOS and windows 3.14 jobs were cancelled by fail-fast after that failure. They did not fail on their own.

Posted on behalf of @Noor-ul-ain001 by Claude Code (AI coding agent).

@mnriem

mnriem commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Please fix test & lint errors

Picks up the deterministic zip fixture fix (github#4823) for the flaky
test_catalog_versions.py integrity failure seen on Windows CI.

Assisted-by: Claude Code (model: claude-opus-5-5, supervised)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Noor-ul-ain001

Noor-ul-ain001 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

@mnriem: the CI run on 109efc3 had one real failure: tests/specify_cli/workflows/test_catalog_versions.py::test_unqualified_add_still_uses_current_and_invalid_version_scope on windows-latest/3.13 (sha256 integrity mismatch). The three other pytest jobs on macOS and Windows were cancelled by fail-fast, and ruff passed. That test is not part of this PR. Its zip fixture hashed differently depending on the wall-clock timestamp, and #4823 fixed it on main by pinning the timestamp.

Merged main into the branch (f5088bd) to pick up that fix. It merged cleanly with no conflicts. Locally on the merged branch: ruff check src tests passes, test_catalog_versions.py passes 43/43, and tests/test_workflows.py plus tests/specify_cli/workflows/ pass, apart from symlink and tmp-path tests that need elevation on my Windows machine. Those also fail on unmodified main here.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The extensive nested concurrency and shared-state changes warrant final human review despite strong regression coverage.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

@mnriem

mnriem commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Please address Copilot feedback

@Noor-ul-ain001

Copy link
Copy Markdown
Contributor Author

Posted on behalf of @Noor-ul-ain001 by Claude Code (model: claude-opus-5-5).

@mnriem: Copilot's one finding (r4096362148: add an end-to-end test for the reserved-id wiring) asks for coverage that is already on the branch. I replied inline with the test names and a check: removing the reserved_step_ids wiring from execute() or resume() makes the matching end-to-end tests fail. No code change this round, so the head is still f5088bd, and CI is green on it across all six pytest jobs plus ruff.

@mnriem

mnriem commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Verified the claim against the current PR head, f5088bd04eb635f50be18be670edbfcf4bedc602.

Commit 1422af024d187cef0f1c085f4753a6ca3c596287 adds the end-to-end execute() collision test; 109efc366743a877bfa7ffc8ceb1a65ca838653a adds its resume() counterpart. Both commits are in the PR history, and both tests remain in tests/test_workflows.py at the current head, each parametrized over concurrency 1 and 2.

I inspected the logs of https://github.com/github/spec-kit/actions/runs/37448037581: all four exact test variants show PASSED in every one of the six OS/Python jobs. Ruff also passed. The run is associated with head f5088bd; Actions checked out the PR merge commit 9cfd681, incorporating that head. No CI rerun is needed for this coverage question. I did not independently repeat the author’s wiring-removal experiments.

The review finding is not a newly posted inline finding: comment r4096362148 was originally submitted on September 24 against 54a03a2, before the execute coverage commit. The October 6 Copilot review targets the current head but lists that same still-unresolved thread under Open, and has no new inline comments. The observable issue is that existing feedback remains open despite the added coverage, not that the commits or CI results are missing. This does not establish why Copilot failed to recognize that the old request was satisfied. I have left the thread unresolved for the reviewer, as required by CONTRIBUTING.md. No code changes were made.

Posted on behalf of @mnriem by GitHub Copilot (model: GPT-6.1 Sol; agent mode, autonomous investigation at the user’s request). AI involvement: PR history/source and CI-log verification, and this comment.

@Noor-ul-ain001

Noor-ul-ain001 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

@mnriem: thanks for checking. That matches what we see: r4096362148 is the only open thread, and 1422af0 (execute()) and 109efc3 (resume()) address it. Nothing else is pending on our side at f5088bd. Would you like the Copilot thread resolved from the author side, or would you prefer to resolve it yourself during review?

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The intricate concurrent state and alias propagation changes warrant final human review despite strong regression coverage.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

@mnriem
mnriem requested a balanced review from Copilot October 6, 2026 14:06
@mnriem

mnriem commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

At @mnriem’s explicit direction, I resolved the remaining open review conversation. All nine review conversations are now resolved; eight were already resolved before this action.

The PR head remains f5088bd04eb635f50be18be670edbfcf4bedc602. I have requested one more Copilot review after resolving the conversations and triggered another validation round of all five existing CI workflows (tests/Ruff, CodeQL, lint, dependency audit, and version checks). Those runs have been accepted; their results are pending. No code changes or empty commit were added.

Posted on behalf of @mnriem by GitHub Copilot (model: GPT-6.1 Sol; agent mode, autonomous execution of the user’s explicit instructions). AI involvement: conversation resolution, review/CI requests, verification, and this comment.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Concurrency-sensitive state propagation spans nested execution and resume paths, warranting final human review despite strong regression coverage.

Review effort: Balanced
Findings: None

Resolved since last review (1)

@mnriem

mnriem commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

The original coverage request is satisfied: both end-to-end tests are present at f5088bd, and I independently confirmed they fail when the reserved-ID wiring is removed.

However, a deeper review found a separate regression in engine.py:1916–1927. Concurrent fan-out publishes only the namespaced IDs in the static id_map. Descendant IDs generated dynamically by nested while, do-while, or fan-out steps reach state.step_results, but are missing from the live context.steps.

For example, these workflow steps reproduce the nested fan-out case:

steps:
  - id: fan
    type: fan-out
    items: "{{ ['a', 'b'] }}"
    max_concurrency: 2
    step:
      id: inner
      type: fan-out
      items: "{{ ['inner'] }}"
      max_concurrency: 1
      step:
        id: leaf
        type: shell
        run: "echo {{ item }}"
  - id: after
    type: shell
    run: "echo {{ steps.fan:inner:0:leaf:0.output.stdout | default('MISSING') }}"

Expected: after prints inner. Actual: it prints MISSING, although the namespaced result exists in persisted state. Changing the outer concurrency to 1 works; running through resume() also works.

The existing workflow suite passes all 786 tests, but an additional reproduction matrix produces three failures—fresh concurrent execution with each of the three nested container types. The nested fan-out cases all pass against the upstream baseline, confirming a regression.

Please publish all item-produced namespaced results, including dynamic descendants, while preserving concurrency isolation and run-state locking. Add regression coverage for downstream visibility across fresh/resumed and sequential/concurrent execution before merging.

Drafted and posted on behalf of @mnriem by GitHub Copilot (GPT-6.1 Sol; autonomous agent mode). AI involvement: code review, automated reproductions, and this response.

…out items

An isolated (concurrent, or nested-in-isolated) fan-out item published
back to the shared steps dict only the ids in its static id_map. Ids
generated at runtime by a nested while/do-while body or fan-out (e.g.
fan:inner:0:leaf:0) reached state.step_results but not the live
context.steps, so a later `steps.<id>` reference resolved to nothing on
a fresh concurrent run.

Publish every key the item added to or replaced in its private view,
excluding its bare-id aliases (still folded in item order by the
caller). The resume-path skip and run-lock behavior are unchanged.

Adds a container x concurrency x fresh/resume regression matrix.

Assisted-by: Claude Code (model: Claude Opus 5.5, autonomous)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Noor-ul-ain001

Copy link
Copy Markdown
Contributor Author

Addressed in 2aa0f4b.

Cause. An isolated fan-out item copied its namespaced results back to the shared steps dict by iterating its static id_map. Ids that a nested while/do-while body or a nested fan-out generates at runtime are never in that map. They were recorded in state.step_results under the run lock, but they never got from the item's private view into the live context.steps.

Fix. run_item now publishes every key the item added to or replaced in its private view, compared by identity against the snapshot it copied. It leaves out the item's bare-id aliases (alias_records), which the caller still folds in item order after the pool joins. Other behavior is unchanged:

  • Each item still writes only its own disjoint namespaced keys.
  • The resume path still skips this publish, because context.steps is state.step_results and _record_result has already written the keys under the lock.
  • The fix applies recursively, so a nested fan-out's descendants reach the outer item's view and from there the shared dict.

Coverage. I added test_nested_dynamic_descendants_visible_downstream, which covers {fan-out, while, do-while} x max_concurrency {1, 2} x {fresh, resume}. Each case asserts that a downstream step resolves steps.fan:inner:<i>:leaf:0.output.stdout for both outer items.

  • Against f5088bd's engine.py: exactly the 3 fresh concurrent cases fail. The other 9 pass, which matches your reproduction.
  • With the fix: all 12 pass. Locally tests/test_workflows.py gives 788 passed / 5 failed. All 5 failures are symlink tests that need elevated rights on Windows and are unrelated to this change. ruff check is clean.

Posted on behalf of @Noor-ul-ain001 by Claude Code (model: Claude Opus 5.5, autonomous agent mode). The AI investigated, reproduced, implemented and tested this change and drafted this comment.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

triage-nice-to-have Verdict: evidence-backed fix or greenlit feature — land after review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants