Add Run.raise_for_status to raise a finished run's failure as a typed error - #1627
Merged
Merged
Conversation
… error A task that launches whole runs with flyte.run had no way to get the typed exception a sub-action would raise; it had to inspect phases and error strings. Run.raise_for_status() (named after requests) is a no-op on success and otherwise raises what awaiting the same task as a sub-action raises: the TaskTimeoutError subclass for the bound that fired, ActionAbortedError, or the error converted from the run's code. The error is read from the root action's own details because GetRunDetails omits the root action's error_info. convert_error_to_native now also maps timeout class names (MaxQueuedTimeExceededError, ...) back to their classes, so a timeout re-raised by an intermediate task keeps its type up the chain, for runs and for sub-actions. Adds examples/advanced/run_of_runs.py: a driver launches a B300 run through @flyte.trace, catches MaxQueuedTimeExceededError, and falls back to a CPU run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…run name The traced step now only launches the child run, under a name the driver derives from its own run name plus a step key, and returns that name. Waiting and raise_for_status() run outside the trace, after re-attaching with Run.get. A replayed launch returns the same name, and a launch whose driver died before the trace was recorded re-creates under the same name and gets the existing run back, so a retried driver never launches a duplicate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the child run name passed into the trace, the split between a traced launch and an untraced wait is unnecessary. A step that was waiting when the driver died re-runs under the same name, gets the existing run back from CreateRun, and waits on it again, so the simpler single traced run-to-completion never duplicates a run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
cosmicbboy
approved these changes
Sep 29, 2026
EngHabu
approved these changes
Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
A task can call
flyte.runto compose whole runs into a larger "run of runs". Until now it couldn't handle a child run's failure the way it handles a sub-action's failure. It had to inspect phases and error strings.Change
Run.raise_for_status()(named afterrequests.Response.raise_for_status). It does nothing if the run succeeded. Otherwise it raises the same exception that awaiting the task as a sub-action would raise:TIMED_OUT→ theTaskTimeoutErrorsubclass for the bound that fired (MaxQueuedTimeExceededError,MaxRuntimeExceededError,DeadlineExceededError), from Raise a cause-specific timeout error when a sub-action times out #1626ABORTED→ActionAbortedError, including the abort reasonFAILED→ the error converted from the run's error code, using the sameconvert_error_to_nativepath the controller usesRuntimeUserError("RunNotDoneError"); callwait()firstActionDetails.get_details).GetRunDetailsreturns the root action withouterror_infoor attempts; confirmed on playground. That looks like a server-side gap, which also affectsrun.details().action_details.error_info.convert_error_to_nativemaps timeout class names back to their classes. A task that re-raises a child'sMaxQueuedTimeExceededErrorrecords it under the codeMaxQueuedTimeExceededError. Before this, the next level up got a genericRuntimeUserError. Now the type survives the whole chain, for runs and for sub-actions.examples/advanced/run_of_runs.py. A driver runs a B300 child run, catchesMaxQueuedTimeExceededErrorand falls back to a CPU child run.@flyte.tracestep that runs it, waits and callsraise_for_status(). The child run's name is an input to the trace: the driver's run name plus a step key, the same on every attempt.Tests
tests/flyte/remote/test_run_raise_for_status.pycovers every terminal phase, each timeout code, a re-raised timeout class name (also behind aRetriesExhausted|prefix), and a missing error. It also checks that the error comes from the root action's details.tests/flyte+tests/internal: 4485 passed. I deselected the two failures that also fail on main (the keyring test and a real Docker image build).uj5m7z495xshwmdnzcwf: SUCCEEDED, output"trained 100 steps on CPU".undh4z2bstd47c28zj7h,retries=1;os._exit(1)inside the traced step while its 30s child is running): attempt 2 re-createdundh4z2bstd47c28zj7h-childand got the existing run back (phase=RUNNING), then waited on it. The child ran exactly once (one attempt, started during driver attempt 1). Output:run=undh4z2bstd47c28zj7h-child out=42.u4mt6bjnxz8b8dtjhflk): attempt 2 replayed the recorded step without launching again.GetRunDetailsand raised a bareTaskTimeoutError. That run is what surfaced the server gap above.🤖 Generated with Claude Code