Skip to content

perf(vm): make PrepareCall allocation-free for native fold loops - #794

Merged
nooga merged 1 commit into
nooga:mainfrom
nnunley:perf/prepare-call-alloc-free
Sep 6, 2026
Merged

nooga merged 1 commit into
nooga:mainfrom
nnunley:perf/prepare-call-alloc-free

Conversation

@nnunley

@nnunley nnunley commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator

Fixes the allocation regression from #727 and restores a green make bench-ratchet on main. Bisect and numbers in #791.

#727 routed reduce and some through PrepareCall, which heap-allocated a PreparedCall and an argument slice per invocation and resolved the target through resolveBytecodeCall, which packs a rest list for variadic fns before the preparation could reject them. The IR compiler performs thousands of short reductions per compile, so BenchmarkIRCompile gained ~1,900 allocs/op, over the ratchet's 2% deterministic bar.

  • PrepareCallInto(p, fn, arity) fills a caller-owned PreparedCall, so the hot loops keep it on their stack.
  • Argument slots live in a new Frame.prepArgs that survives pool reuse and does not alias argbuf, which installBytecodeCall overwrites on a same-frame tail call.
  • unwrapBytecodeFn, split out of resolveBytecodeCall, inspects the target without allocating, so variadic and wrong-arity fns are rejected before any packing. PrepareCall remains as the allocating form.
  • TestPrepareCallIntoDoesNotAllocate pins the zero-allocation contract; the existing PreparedCall tests (tail-call rebind, error reuse, arity mismatch) still pass.

On an Apple M3, plugged in, BenchmarkIRCompile [bytecode]:

allocs/op bytes/op ns/op
#780 baseline 114,255 4,963,229 15.05M
main (bdd8268) 116,149 5,080,290 17.01M
this PR 113,722 4,935,179 15.48M

make bench-ratchet on this tree: deterministic OK, anchor +0.7%, IRCompile [bytecode] +2.0% ok, InitFromLGB -2.4%, IRCompile [gogen_ir] -10.2%, 0 regressions. go test ./... -skip TestClojureTestSuite and go test ./test/e2e pass; make check-generated clean (manifest and sums refreshed).

@mparrett mparrett left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed head 1b4d8f0. No findings.

I checked the PreparedCall lifecycle and frame-pool ownership paths, including same-frame tail-call rebinding and rejection of unsupported targets. go test ./pkg/vm ./pkg/rt -count=1, the corresponding race run, 100 repetitions of the PreparedCall tests, make check-generated, and git diff --check all pass.

Independent base/head allocation checks reproduced the intended improvement. On the ordinary bytecode path, steady BenchmarkIRCompile samples dropped by 2,427 allocs/op and about 145 KB/op. With generated IR enabled, steady samples also dropped by about 1,486 allocs/op and 55 KB/op. I did not use wall-clock results as review evidence because this host was under heavy unrelated load.

The current merge conflict appears limited to generated.manifest / generated.sums; rebasing and regenerating those files should be the remaining integration step.

PreparedCall and an argument slice on every invocation, and resolved the
target through resolveBytecodeCall, which packs a rest list for variadic
fns before the preparation could reject them. The IR compiler performs
thousands of short reductions per compile, so BenchmarkIRCompile gained
~1,900 allocs/op (+2.4% bytes/op, over the ratchet's 2% deterministic bar)
and InitFromLGB slowed ~10%. Bisect and numbers in nooga#791.

PrepareCallInto(p, fn, arity) fills a caller-owned PreparedCall, so the
hot loops keep it on their stack; the argument slots live in a new
Frame.prepArgs that survives pool reuse and does not alias argbuf, which
installBytecodeCall overwrites on a same-frame tail call. The target is
inspected with unwrapBytecodeFn, split out of resolveBytecodeCall, so
variadic and wrong-arity fns are rejected before any packing. PrepareCall
remains as the allocating form.

On the M3, BenchmarkIRCompile goes from 116,149 to 113,722 allocs/op and
5,080,290 to 4,935,179 bytes/op, below the nooga#780 baseline (114,255 /
4,963,229) because variadic reducers no longer pay for a preparation they
cannot use. TestPrepareCallIntoDoesNotAllocate pins the zero-allocation
contract.

Refs nooga#791.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants