Skip to content

Zend: Compile a list assignment from an array literal without the array - #24180

Open
ArtUkrainskiy wants to merge 1 commit into
php:masterfrom
ArtUkrainskiy:destructure-array-literal
Open

ArtUkrainskiy wants to merge 1 commit into
php:masterfrom
ArtUkrainskiy:destructure-array-literal

Conversation

@ArtUkrainskiy

Copy link
Copy Markdown
Contributor

Closes GH-23048.

[$a, $b] = [$b, $a]; builds an array, fetches both elements back out of it and frees it. When the result is unused and the list is flat, unkeyed, without references or spread, with one value per target, this assigns the values directly:

T2 = QM_ASSIGN CV1($b)
T3 = QM_ASSIGN CV0($a)
T4 = COPY_TMP T2
ASSIGN CV0($a) T4
T6 = COPY_TMP T3
ASSIGN CV1($b) T6
FREE T2
FREE T3

That is the shape from the issue plus a copy per target. Without the copies the four opcodes are 467 instructions instead of 523, but the array held every value until after the last assignment, and dropping that is observable: a setter that discards the value, two targets that are references to each other, and a destructor that throws — on master [$a, $a, $b] = [new Boom, new stdClass, 42] still assigns $b, without the copies the exception cuts the statement short. I tried keeping the direct assignment for plain variables, but whether a variable still holds its value at the end of the statement is a runtime question (references, user code in a later target's expression), so every target gets a copy and the values are freed where the array was. With that I found no remaining difference, destruction order under exceptions included.

The expression lists of a for loop get the same treatment, since their results are unused too. Nested lists stay as they are: the inner list needs an owner whose destruction order matches the old fetches, and a COPY_TMP freed later can't be one — its live range is computed for the ?? pattern. Skipped or extra values stay too: the copy that would keep an untaken value alive is QM_ASSIGN + FREE, which the block pass turns into CHECK_VAR.

Release builds, perf stat instructions per statement:

statement instructions
[$a, $b] = [$b, $a] 922 → 523
[$a, $b, $c] = [$b, $c, $a] 1,119 → 629
ten Fibonacci steps [$a, $b] = [$b, ($a + $b) % M] 6,114 → 2,064
[$p[0], $p[1]] = [$p[1], $p[0]] 1,290 → 879
swap through a temporary variable, for reference 493

A loop of 300,000 swaps with opcache: 5.1 → 1.1 ms, with the tracing JIT 3.8 → 0.5 ms. In two large vendor trees (327,462 op_arrays) 56 statements change and nothing else does, so this is for the odd hot swap, not for applications.

Verified against master: 97 behaviour cases and a sweep of 3,376 generated programs with printing destructors and setters, identical output on release and debug builds, with opcache and both JITs. Test expectations are generated on master; the lifetime test also runs with opcache.optimization_level=-1. Details and reproduction: https://github.com/ArtUkrainskiy/php-src-bench/tree/main/reports/list-assign-from-array-literal

Possible follow-ups: keyed lists with matching constant keys, and an optimizer rule dropping COPY_TMP/FREE for values known not to be refcounted, which would give typed swaps the four-opcode shape.

[$a, $b] = [$b, $a]; built an array of the values, fetched every element
out of it again and freed it. When the result of the assignment is not
used, the right side is an array literal and the left side a flat list
with a target for every value, assign the values to the targets directly
instead.

The values are still evaluated in order before the first assignment, a
variable among them is copied first, and every value is released only
after the last assignment, as the array was: each target is assigned a
copy of its value, so a setter that drops the value, a target that is a
reference to another one, or a destructor that throws see the same
order as before, and so does an assignment that throws.

Assignments in the expression lists of a for loop, whose results are not
used either, are compiled the same way. Nested lists, keys on either
side, references, spread, a different number of values and targets and
a used result keep the usual compilation.

A swap goes from 922 to 523 instructions, a loop of swaps is 5 times
faster without the JIT and 7 times with it.
@ArtUkrainskiy
ArtUkrainskiy force-pushed the destructure-array-literal branch from eb63ed6 to 5bfa8e3 Compare October 7, 2026 20:39
@ArtUkrainskiy
ArtUkrainskiy marked this pull request as ready for review October 7, 2026 21:47
@ArtUkrainskiy
ArtUkrainskiy requested a review from dstogov as a code owner October 7, 2026 21:47
0001 CV1($b) = RECV 2
0002 T2 = QM_ASSIGN CV1($b)
0003 T3 = QM_ASSIGN CV0($a)
0004 T4 = COPY_TMP T2

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be great if these COPY_TMPs could be avoided (and thus the FREE as well); that would make the code of this swap function optimal

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The copies keep destructor timing as it is: the array held a reference to every value until the statement was done, so a value its target doesn't keep (a setter that drops it, [$a, $a, …], targets that are references to each other) was only released at the end. Without the copies it dies at that assignment, and a destructor is observable — if it throws, the rest of the statement doesn't run: [$a, $a, $b] = [new Boom, new stdClass, 42] leaves $b unassigned where master assigns it. I had the four-opcode form first and the lifetime test caught exactly this.

For values that can't be refcounted the copies are pure overhead, and that's the optimizer's job: with optimization_level=-1 the typed swap is already T2 = QM_ASSIGN $b; $a = COPY_TMP T2 — the only missing rule is that COPY_TMP of a scalar is QM_ASSIGN, after which the existing contraction gives the shape you have in mind. I'd send that as a follow-up if that works for you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Optimize destructuring of array literals to elide intermediate array

2 participants