Skip to content

Add Run.raise_for_status to raise a finished run's failure as a typed error - #1627

Merged
kumare3 merged 3 commits into
mainfrom
run-raise-for-status
Sep 29, 2026
Merged

kumare3 merged 3 commits into
mainfrom
run-raise-for-status

Conversation

@kumare3

@kumare3 kumare3 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Why

A task can call flyte.run to compose whole runs into a larger "run of runs". Until now it couldn't handle a child run's failure the way it handles a sub-action's failure. It had to inspect phases and error strings.

Change

  • Run.raise_for_status() (named after requests.Response.raise_for_status). It does nothing if the run succeeded. Otherwise it raises the same exception that awaiting the task as a sub-action would raise:
    • TIMED_OUT → the TaskTimeoutError subclass for the bound that fired (MaxQueuedTimeExceededError, MaxRuntimeExceededError, DeadlineExceededError), from Raise a cause-specific timeout error when a sub-action times out #1626
    • ABORTED → ActionAbortedError, including the abort reason
    • FAILED → the error converted from the run's error code, using the same convert_error_to_native path the controller uses
    • not yet terminal → RuntimeUserError("RunNotDoneError"); call wait() first
  • The error is read from the root action's own details (ActionDetails.get_details). GetRunDetails returns the root action without error_info or attempts; confirmed on playground. That looks like a server-side gap, which also affects run.details().action_details.error_info.
  • convert_error_to_native maps timeout class names back to their classes. A task that re-raises a child's MaxQueuedTimeExceededError records it under the code MaxQueuedTimeExceededError. Before this, the next level up got a generic RuntimeUserError. Now the type survives the whole chain, for runs and for sub-actions.
  • Example examples/advanced/run_of_runs.py. A driver runs a B300 child run, catches MaxQueuedTimeExceededError and falls back to a CPU child run.
    • Each child run is one @flyte.trace step that runs it, waits and calls raise_for_status(). The child run's name is an input to the trace: the driver's run name plus a step key, the same on every attempt.
    • No duplicate runs on retry. A finished step replays its recorded result or error. A step still waiting when the driver died re-runs under the same name, and CreateRun returns the existing run instead of creating a second one, so the driver waits on it again.

Tests

  • tests/flyte/remote/test_run_raise_for_status.py covers every terminal phase, each timeout code, a re-raised timeout class name (also behind a RetriesExhausted| prefix), and a missing error. It also checks that the error comes from the root action's details.
  • tests/flyte + tests/internal: 4485 passed. I deselected the two failures that also fail on main (the keyring test and a real Docker image build).
  • E2E on playground.canary, SDK wheel from this branch:
    • Example, driver run uj5m7z495xshwmdnzcwf: SUCCEEDED, output "trained 100 steps on CPU".
      launched uj5m7z495xshwmdnzcwf-b300
      no B300 capacity, falling back to CPU: Run uj5m7z495xshwmdnzcwf-b300 timed out: timed out waiting for resources: queued_timeout of 1m0s exceeded (attempt=0)
      launched uj5m7z495xshwmdnzcwf-cpu
      
    • Driver dies mid-wait (throwaway probe undh4z2bstd47c28zj7h, retries=1; os._exit(1) inside the traced step while its 30s child is running): attempt 2 re-created undh4z2bstd47c28zj7h-child and got the existing run back (phase=RUNNING), then waited on it. The child ran exactly once (one attempt, started during driver attempt 1). Output: run=undh4z2bstd47c28zj7h-child out=42.
    • Driver retried after the trace recorded (throwaway probe u4mt6bjnxz8b8dtjhflk): attempt 2 replayed the recorded step without launching again.
    • The first attempt, before switching to the action-details read, got no error info from GetRunDetails and raised a bare TaskTimeoutError. That run is what surfaced the server gap above.

🤖 Generated with Claude Code

kumare3 and others added 3 commits September 28, 2026 21:29
… error

A task that launches whole runs with flyte.run had no way to get the
typed exception a sub-action would raise; it had to inspect phases and
error strings. Run.raise_for_status() (named after requests) is a no-op
on success and otherwise raises what awaiting the same task as a
sub-action raises: the TaskTimeoutError subclass for the bound that
fired, ActionAbortedError, or the error converted from the run's code.

The error is read from the root action's own details because
GetRunDetails omits the root action's error_info.

convert_error_to_native now also maps timeout class names
(MaxQueuedTimeExceededError, ...) back to their classes, so a timeout
re-raised by an intermediate task keeps its type up the chain, for runs
and for sub-actions.

Adds examples/advanced/run_of_runs.py: a driver launches a B300 run
through @flyte.trace, catches MaxQueuedTimeExceededError, and falls back
to a CPU run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…run name

The traced step now only launches the child run, under a name the
driver derives from its own run name plus a step key, and returns that
name. Waiting and raise_for_status() run outside the trace, after
re-attaching with Run.get. A replayed launch returns the same name, and
a launch whose driver died before the trace was recorded re-creates
under the same name and gets the existing run back, so a retried driver
never launches a duplicate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the child run name passed into the trace, the split between a
traced launch and an untraced wait is unnecessary. A step that was
waiting when the driver died re-runs under the same name, gets the
existing run back from CreateRun, and waits on it again, so the
simpler single traced run-to-completion never duplicates a run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@kumare3
kumare3 merged commit 7d4deb1 into main Sep 29, 2026
65 checks passed
@kumare3
kumare3 deleted the run-raise-for-status branch September 29, 2026 18:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants