perf(vm): make PrepareCall allocation-free for native fold loops - #794
Conversation
mparrett
left a comment
There was a problem hiding this comment.
Reviewed head 1b4d8f0. No findings.
I checked the PreparedCall lifecycle and frame-pool ownership paths, including same-frame tail-call rebinding and rejection of unsupported targets. go test ./pkg/vm ./pkg/rt -count=1, the corresponding race run, 100 repetitions of the PreparedCall tests, make check-generated, and git diff --check all pass.
Independent base/head allocation checks reproduced the intended improvement. On the ordinary bytecode path, steady BenchmarkIRCompile samples dropped by 2,427 allocs/op and about 145 KB/op. With generated IR enabled, steady samples also dropped by about 1,486 allocs/op and 55 KB/op. I did not use wall-clock results as review evidence because this host was under heavy unrelated load.
The current merge conflict appears limited to generated.manifest / generated.sums; rebasing and regenerating those files should be the remaining integration step.
PreparedCall and an argument slice on every invocation, and resolved the target through resolveBytecodeCall, which packs a rest list for variadic fns before the preparation could reject them. The IR compiler performs thousands of short reductions per compile, so BenchmarkIRCompile gained ~1,900 allocs/op (+2.4% bytes/op, over the ratchet's 2% deterministic bar) and InitFromLGB slowed ~10%. Bisect and numbers in nooga#791. PrepareCallInto(p, fn, arity) fills a caller-owned PreparedCall, so the hot loops keep it on their stack; the argument slots live in a new Frame.prepArgs that survives pool reuse and does not alias argbuf, which installBytecodeCall overwrites on a same-frame tail call. The target is inspected with unwrapBytecodeFn, split out of resolveBytecodeCall, so variadic and wrong-arity fns are rejected before any packing. PrepareCall remains as the allocating form. On the M3, BenchmarkIRCompile goes from 116,149 to 113,722 allocs/op and 5,080,290 to 4,935,179 bytes/op, below the nooga#780 baseline (114,255 / 4,963,229) because variadic reducers no longer pay for a preparation they cannot use. TestPrepareCallIntoDoesNotAllocate pins the zero-allocation contract. Refs nooga#791.
1b4d8f0 to
5e26821
Compare
Fixes the allocation regression from #727 and restores a green
make bench-ratchetonmain. Bisect and numbers in #791.#727 routed
reduceandsomethroughPrepareCall, which heap-allocated aPreparedCalland an argument slice per invocation and resolved the target throughresolveBytecodeCall, which packs a rest list for variadic fns before the preparation could reject them. The IR compiler performs thousands of short reductions per compile, soBenchmarkIRCompilegained ~1,900 allocs/op, over the ratchet's 2% deterministic bar.PrepareCallInto(p, fn, arity)fills a caller-ownedPreparedCall, so the hot loops keep it on their stack.Frame.prepArgsthat survives pool reuse and does not aliasargbuf, whichinstallBytecodeCalloverwrites on a same-frame tail call.unwrapBytecodeFn, split out ofresolveBytecodeCall, inspects the target without allocating, so variadic and wrong-arity fns are rejected before any packing.PrepareCallremains as the allocating form.TestPrepareCallIntoDoesNotAllocatepins the zero-allocation contract; the existing PreparedCall tests (tail-call rebind, error reuse, arity mismatch) still pass.On an Apple M3, plugged in,
BenchmarkIRCompile [bytecode]:make bench-ratcheton this tree: deterministic OK, anchor +0.7%, IRCompile [bytecode] +2.0% ok, InitFromLGB -2.4%, IRCompile [gogen_ir] -10.2%, 0 regressions.go test ./... -skip TestClojureTestSuiteandgo test ./test/e2epass;make check-generatedclean (manifest and sums refreshed).