Skip to content

[Feature]: Reject new durable operations under orphaned map/parallel branches #641

Description

@zhongkechen

What would you like?

Reject newly created durable operations whose parent context has already been
marked orphaned after an early-completing map or parallel operation.

Today, _mark_orphans() marks the descendants known when the parent context
completes. If an in-flight orphaned branch subsequently creates a new durable
operation, ExecutionState.create_checkpoint() adds the new operation to
_parent_to_children and checks only whether the new operation's own ID is in
_parent_done.

Because the new operation did not exist when _mark_orphans() took its
snapshot, its ID is not in _parent_done, so its checkpoint can be accepted
even though its parent_id is already orphaned.

The orphaned branch's own terminal result is normally rejected later, and the
parent BatchResult remains consistent across replay:

  • Normal payloads replay the parent's serialized result.
  • ReplayChildren summaries preserve the branch as STARTED through
    startedIndexes.

This is therefore an improvement to lifecycle enforcement rather than a
customer-visible replay-result correctness bug. It would still prevent
unnecessary durable work, side effects, and history entries from being created
inside a branch whose parent has already completed.

This issue is separate from #640, which covers the race where an already-known
branch passes orphan validation before parent completion and enqueues its own
terminal result afterward.

Possible Implementation

When admitting an operation checkpoint under _parent_done_lock, reject the
operation when either:

operation_update.operation_id in self._parent_done

or:

operation_update.parent_id in self._parent_done

If a parent can be indirectly orphaned without appearing directly in
_parent_done, propagate orphan state when registering the new child or walk
the known parent chain under the same lock.

The check should also be repeated in the atomic validation-and-enqueue section
proposed by #640 so parent completion cannot race queue insertion.

Suggested tests:

  • A new step created under an already orphaned nested map iteration.
  • A new invoke/callback/wait created under an orphaned parallel branch.
  • The equivalent behavior for NestingType.FLAT.
  • The new operation never reaches the checkpoint service.
  • Parent results remain identical on first execution and replay for normal and
    ReplayChildren payloads.

Is this a breaking change?

No. The operation is already outside the lifetime of its completed parent
context and its enclosing branch cannot contribute a terminal result to the
completed batch.

Does this require an RFC?

No.

Additional Context

A deterministic state-level reproduction shows:

branch orphaned: True
late descendant rejected: NO (accepted)
late descendant orphaned: False

Relevant code:

  • packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py
    in ExecutionState.create_checkpoint()
  • packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py
    in _mark_orphans()

Observed against SDK version 1.8.0, repository HEAD
7ac7acc6a7dae231f2abbb8e37f9780cc9b89af0, using Python 3.14.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions