What would you like?
Reject newly created durable operations whose parent context has already been
marked orphaned after an early-completing map or parallel operation.
Today, _mark_orphans() marks the descendants known when the parent context
completes. If an in-flight orphaned branch subsequently creates a new durable
operation, ExecutionState.create_checkpoint() adds the new operation to
_parent_to_children and checks only whether the new operation's own ID is in
_parent_done.
Because the new operation did not exist when _mark_orphans() took its
snapshot, its ID is not in _parent_done, so its checkpoint can be accepted
even though its parent_id is already orphaned.
The orphaned branch's own terminal result is normally rejected later, and the
parent BatchResult remains consistent across replay:
- Normal payloads replay the parent's serialized result.
ReplayChildren summaries preserve the branch as STARTED through
startedIndexes.
This is therefore an improvement to lifecycle enforcement rather than a
customer-visible replay-result correctness bug. It would still prevent
unnecessary durable work, side effects, and history entries from being created
inside a branch whose parent has already completed.
This issue is separate from #640, which covers the race where an already-known
branch passes orphan validation before parent completion and enqueues its own
terminal result afterward.
Possible Implementation
When admitting an operation checkpoint under _parent_done_lock, reject the
operation when either:
operation_update.operation_id in self._parent_done
or:
operation_update.parent_id in self._parent_done
If a parent can be indirectly orphaned without appearing directly in
_parent_done, propagate orphan state when registering the new child or walk
the known parent chain under the same lock.
The check should also be repeated in the atomic validation-and-enqueue section
proposed by #640 so parent completion cannot race queue insertion.
Suggested tests:
- A new step created under an already orphaned nested map iteration.
- A new invoke/callback/wait created under an orphaned parallel branch.
- The equivalent behavior for
NestingType.FLAT.
- The new operation never reaches the checkpoint service.
- Parent results remain identical on first execution and replay for normal and
ReplayChildren payloads.
Is this a breaking change?
No. The operation is already outside the lifetime of its completed parent
context and its enclosing branch cannot contribute a terminal result to the
completed batch.
Does this require an RFC?
No.
Additional Context
A deterministic state-level reproduction shows:
branch orphaned: True
late descendant rejected: NO (accepted)
late descendant orphaned: False
Relevant code:
packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py
in ExecutionState.create_checkpoint()
packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py
in _mark_orphans()
Observed against SDK version 1.8.0, repository HEAD
7ac7acc6a7dae231f2abbb8e37f9780cc9b89af0, using Python 3.14.
What would you like?
Reject newly created durable operations whose parent context has already been
marked orphaned after an early-completing
maporparalleloperation.Today,
_mark_orphans()marks the descendants known when the parent contextcompletes. If an in-flight orphaned branch subsequently creates a new durable
operation,
ExecutionState.create_checkpoint()adds the new operation to_parent_to_childrenand checks only whether the new operation's own ID is in_parent_done.Because the new operation did not exist when
_mark_orphans()took itssnapshot, its ID is not in
_parent_done, so its checkpoint can be acceptedeven though its
parent_idis already orphaned.The orphaned branch's own terminal result is normally rejected later, and the
parent
BatchResultremains consistent across replay:ReplayChildrensummaries preserve the branch asSTARTEDthroughstartedIndexes.This is therefore an improvement to lifecycle enforcement rather than a
customer-visible replay-result correctness bug. It would still prevent
unnecessary durable work, side effects, and history entries from being created
inside a branch whose parent has already completed.
This issue is separate from #640, which covers the race where an already-known
branch passes orphan validation before parent completion and enqueues its own
terminal result afterward.
Possible Implementation
When admitting an operation checkpoint under
_parent_done_lock, reject theoperation when either:
or:
If a parent can be indirectly orphaned without appearing directly in
_parent_done, propagate orphan state when registering the new child or walkthe known parent chain under the same lock.
The check should also be repeated in the atomic validation-and-enqueue section
proposed by #640 so parent completion cannot race queue insertion.
Suggested tests:
NestingType.FLAT.ReplayChildrenpayloads.Is this a breaking change?
No. The operation is already outside the lifetime of its completed parent
context and its enclosing branch cannot contribute a terminal result to the
completed batch.
Does this require an RFC?
No.
Additional Context
A deterministic state-level reproduction shows:
Relevant code:
packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.pyin
ExecutionState.create_checkpoint()packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.pyin
_mark_orphans()Observed against SDK version
1.8.0, repository HEAD7ac7acc6a7dae231f2abbb8e37f9780cc9b89af0, using Python 3.14.