Skip to content

Zend async Scheduler ABI - #22561

Open
EdmondDantes wants to merge 34 commits into
php:masterfrom
true-async:async-core
Open

EdmondDantes wants to merge 34 commits into
php:masterfrom
true-async:async-core

Conversation

@EdmondDantes

@EdmondDantes EdmondDantes commented Jul 2, 2026 •

Copy link
Copy Markdown

A lightweight asynchronous core without complex logic.

The idea is:
https://github.com/true-async/php-async-core-rfc/blob/main/scheduler_rfc.md

A detailed explanation of the core integration can be found here:
https://github.com/true-async/php-async-core-rfc/blob/main/core-integration.md

At the moment, both the code and the RFC are still under development. I'd be happy to hear your ideas and feedback.

@EdmondDantes
EdmondDantes marked this pull request as draft July 2, 2026 17:48
@EdmondDantes
EdmondDantes force-pushed the async-core branch 4 times, most recently from 615d050 to ed9fe03 Compare July 2, 2026 18:11
@EdmondDantes
EdmondDantes force-pushed the async-core branch 10 times, most recently from f3b272d to 6e1a5ce Compare July 3, 2026 09:58
Comment thread Zend/zend_async_API.c Outdated
Comment thread Zend/zend_scheduler_hook.stub.php Outdated
Comment thread Zend/zend_gc.c Outdated
Comment thread Zend/zend_scheduler_hook.stub.php Outdated
# Conflicts:
#	Zend/zend_fibers.c
#	win32/build/config.w32
ZEND_MODULE_API_NO changes once per PHP version and does not tell two
revisions of this API apart; `size` only covers slots appended at the end.
ZEND_ASYNC_API_VERSION is the date of the last incompatible change, as
PDO_DRIVER_API is for PDO drivers. test_scheduler stays loaded but inert
when refused instead of failing startup without a message.
ZEND_COROUTINE_IS_STARTED read the status, which cannot tell a coroutine
waiting for its first run from one that yielded: both are QUEUED. The new
F_STARTED bit is set by the scheduler just before the body's first
instruction and never for a coroutine cancelled before it ran. A fiber
dropped while its coroutine is still queued now cancels it as well; one
whose enqueue failed is only released.
A coroutine may catch its cancellation and go on, so a cancelled fiber
may suspend again in finally, and test_scheduler lets a cancelled
coroutine park and delivers a later cancel once the earlier one was
thrown. Only a running coroutine or one with an error still pending
ignores a cancel, which keeps a fiber destroyed inside its own
force-close from queueing itself twice.
EG(error_handling) and EG(exception_class) were global across a switch:
a fiber suspended inside zend_replace_error_handling(EH_THROW) turned
every warning raised in main or in another fiber into an exception of
that window's class. The switch saves both per context and resets them
before the jump, so a context entered for the first time starts outside
any window.
The enqueue slot did not say what happens to a FINISHED coroutine, and
test_scheduler accepted it as a no-op. A finished coroutine cannot run
again, so the contract is false with an Error, as the core's callers
already expect from Fiber::resume() and Fiber::throw(). No core caller
reaches the case from PHP, so it has no test.
Under the scheduler, gc_collect_cycles() awaited the GC coroutine even in
a tick function, a signal handler or the scheduler's own work, where no
coroutine may switch. It now only starts the run, which collects when the
scheduler next picks the GC coroutine, and returns 0. A run that was
started but not awaited no longer feeds that 0 to the threshold
heuristic, which raised the threshold by a step each time.
GC_G(gc_coroutine) was cleared only at the end of the run's body, which a
bailout skips: a shutdown function's gc_collect_cycles() then awaited the
dead coroutine and returned 0 without collecting. A finish handler now
clears it however the run ends and drops the iterator microtask on a
bailout, and gc_reset() clears the async fields, so none of them reaches
the next request (FPM, ZTS); that half has no test in a CLI run.
After exit() in a fiber coroutine the core called the shutdown slot and
then took a reference to EG(exception) for the waiting caller. A scheduler
that consumes the unwind_exit there (the reference ext/async does) left
it NULL, and the fiber's caller crashed. A cleared exception now ends the
fiber without an error and wakes the caller. test_scheduler keeps the
exception, so there is no phpt; checked by hand with a shutdown slot
that clears it (segfault before, clean exit after).
ZEND_ASYNC_DEACTIVATE only switched the state off, so output handlers,
RSHUTDOWN and the next request's RINIT still saw the request's current
and main coroutine after the scheduler had torn them down; an internal
context lookup with no coroutine given read the freed one.
The await slot gains its third false case: the scheduler's own work cannot
wait, so the slot returns false without an exception and the caller goes
on; test_scheduler follows it and keeps the Error for userland await().
The execute_data slot speaks of a parked coroutine, which includes one
that yielded and is QUEUED. CREATED means allocated and not yet enqueued,
and bits 16-31 of the flags belong to the scheduler.
A destructor run by the scheduler (releasing a finished coroutine) could
call Fiber::start(), resume() or throw() and switch the scheduler out of
its own loop: the fiber never ran and the request ended with a stray
GracefulExit. The methods now throw FiberError there, as they do where
switching is blocked, before they change any state.
zend_fiber_switch_context() called zend_observer_fiber_switch_notify() on
every switch, and under a scheduler every await, suspend and resume is a
switch. The call now sits behind ZEND_OBSERVER_FIBER_SWITCH_ENABLED (fcall
observers or a registered switch handler), as in the TrueAsync fork. The
zend_test observer tests pass before and after and fail when the call is
dropped entirely.
A scheduler that allocates its fiber contexts apart from their stacks
frees the context in its cleanup; zend_fiber_destroy_context then read
context->stack from freed memory (heap-use-after-free under ASAN on the
first spawned coroutine of true_async). The stack is now read first, as
TrueAsync's core does.
Fiber::__construct() and ZEND_ASYNC_FCALL_DEFINE copied the callable's
fci without a reference on fci.object: the $this that [A::class, 'm']
resolves to inside a method of A. When A died before start(), the fiber
called the method on freed memory (reproduced on master and PHP 8.3
without a scheduler). Both now hold it, release it with the callable
and report it to the GC. ZEND_ASYNC_API_VERSION becomes 20261003: a
provider's get_gc reports the new reference.
php_call_shutdown_functions() catches the bailout of a fatal error or
exit() in a shutdown function and does not re-raise it, so the
scheduler never learned of it: coroutines queued before it ran after
it, and the destructors ran with a finished coroutine as the current
one. The catch now ends the scheduler as for a bailout in the script.
The globals lose a field; the Async API version 20261003, moved the
same day by the Fiber callable fix, covers it.
The call_on_main_stack and coroutine_from_object slots, the OBJ_REF object
model, ZEND_ASYNC_GET_EXCEPTION_CE (an alias of ZEND_ASYNC_GET_CE),
zend_async_is_enabled() and the empty globals destructor go.
The error stayed in EG(exception) across the switch and reached the coroutine entered: a destructor's await saw it and the coroutine it awaited finished unrun. The driving coroutine now finishes the pass itself when it resumes, as when no iterator could be created.
…uld not create

The enqueue error stayed in the GC coroutine, which then ended with an
exception no caller handles (a provider may end the request on it), and
the run gave up its rerun. Clear it, as a missing iterator leaves none:
the run is redone with the destructors called.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants