Skip to content

Commit 6d275be

Browse files
authored
Merge pull request #5988 from ab9rf/lua-performance
document potential levers for performance improvement
2 parents bee8d8f + 9f1b218 commit 6d275be

2 files changed

Lines changed: 292 additions & 0 deletions

File tree

‎docs/dev/index.rst‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -21,6 +21,7 @@ These are pages relevant to people developing for DFHack.
2121
/docs/dev/github-workflows
2222
/docs/dev/release-process
2323
/docs/dev/Memory-research
24+
/docs/dev/performance-plan
2425
/docs/dev/Binpatches
2526
/docs/dev/Remote
2627
/docs/NEWS-dev

‎docs/dev/performance-plan.rst‎

Lines changed: 291 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,291 @@
1+
=========================
2+
Performance plan
3+
=========================
4+
5+
Working plan for improving DFHack's per-tick CPU budget. Based on
6+
microarchitecture analysis (Intel VTune capture ``r120ue``, Jan 2026)
7+
and code review of the Lua integration. Not user-facing documentation.
8+
9+
.. contents::
10+
11+
Baseline findings
12+
=================
13+
14+
From ``r120ue`` (uarch exploration, whole-system sampling; ~24k
15+
``CPU_CLK_UNHALTED.THREAD`` samples):
16+
17+
- DFHack-side code is ~21% of sampled CPU: ``lua53.dll`` 3322,
18+
``dfhooks_dfhack.dll`` 1750, plugins ~110.
19+
- DFHack owns ~11% of ``STALLS_L3_MISS`` and ~11% of DTLB
20+
``WALK_ACTIVE`` samples.
21+
- Top lua53 stallers are data-layout problems, not dispatch:
22+
``luaH_getshortstr`` (hash-chain walk, 17 stall samples),
23+
``luaD_precall``, ``lua_geti``, ``luaH_next`` (``pairs`` iteration),
24+
``lua_getmetatable``, ``mainposition``, ``luaC_checkfinalizer``.
25+
- Top CPU consumers are the C<->Lua transition layer:
26+
``luaV_execute`` 759, ``luaD_precall`` 260, ``index2addr`` 115,
27+
``luaV_finishget`` (metamethod resolution) 90, ``luaG_traceexec``
28+
82 (the always-armed ``LUA_MASKCOUNT`` interrupt hook),
29+
``match``/``singlematch`` ~88 (Lua pattern matching per frame),
30+
plus ``lua_tolstring``/``luaS_hash`` churn.
31+
- C++ side: MSVC ``std::_Hash``/``std::_Tree`` internals,
32+
``std::string`` construct/destroy churn, ``matchFocusString``,
33+
``Units::isActive`` — chained-node STL containers have the same
34+
miss problem as Lua's tables.
35+
36+
Working hypothesis (operator model): DF's own memory traffic evicts
37+
essentially all DFHack data from L3 between simulation ticks, so every
38+
tick starts cold and each pointer-chase (hash chain, GC list,
39+
metatable walk) pays full miss latency. Goal: make the per-tick cold
40+
walk *sequential* (prefetcher-trackable) rather than pointer-chasing,
41+
and reduce the number of misses by shortening dependent-load chains.
42+
43+
Overlay processing is believed to dominate the per-tick budget; this
44+
is consistent with the data but not yet proven — see the attribution
45+
step below.
46+
47+
Constraints and rejected directions
48+
===================================
49+
50+
- **No PGO/LTO (deferred).** DFHack does questionable things to DF's
51+
vtables and binary; whole-program optimization changes codegen
52+
assumptions the interpose/binpatch machinery relies on. Would need
53+
a dedicated validation effort before it is safe to enable. Revisit
54+
only if cheaper items are exhausted.
55+
- **Hugepages: optional-only.** ``VirtualAlloc(MEM_LARGE_PAGES)``
56+
needs ``SeLockMemoryPrivilege`` and often fails for normal users;
57+
Linux needs THP/hugetlbfs. Steam Deck (Linux) is a significant
58+
user population and we have no Linux profiling. Any allocator work
59+
must function with ordinary pages; hugepages may be an opt-in
60+
bonus path (``madvise(MADV_HUGEPAGE)`` on Linux,
61+
``MEM_LARGE_PAGES`` on Windows) with mandatory fallback.
62+
- **No LLVM-JIT Lua (Ravi etc.).** The miss profile is in
63+
``ltable``/``lgc``/``ldo`` (data layout + transitions), not
64+
``lvm`` dispatch. A JIT keeps identical object layout and would
65+
also have to reproduce our ``lua_lock`` threading patch,
66+
``LUA_MASKCOUNT`` interrupt hooks, ``lstate.h`` field access,
67+
yieldable pcall, and debug-hook fidelity. Wrong tool for this
68+
problem.
69+
- **LuaJIT** has the best-in-class data layout but is Lua 5.1
70+
semantics; incompatible with our 5.3 corpus.
71+
72+
Action items
73+
============
74+
75+
Ordered by leverage-per-effort. All are architecture-independent
76+
unless noted.
77+
78+
Measurement
79+
-----------
80+
81+
1. Per-tick attribution. ``perf_counters.update_lua_ms`` /
82+
``update_plugin_ms`` (``Core.cpp:1668-1682``) already isolate Lua
83+
update time; ``script-manager.lua`` prints them. Baseline a real
84+
fort first.
85+
2. Callsite-stack attribution on VTune captures: ``dd_callsite``
86+
parent links in ``dicer.db`` allow rebuilding stacks and splitting
87+
Lua-side stalls between overlay widgets, timers, and plugin
88+
events. (Scratch query script: ``%TEMP%/vtune_agg.py``; accepts a
89+
``dicer.db`` path for other captures.)
90+
3. Lua-level: run ``profiler.lua`` scoped to the update window, or a
91+
count hook aggregating by Proto, to identify which script
92+
functions dominate ticks.
93+
4. Object-level (if needed): we own the Lua source — log
94+
``luaC_newobj``/``luaM_realloc_`` (addr, type, size, phase) during
95+
a session, dump live objects by walking ``allgc`` at a tick
96+
boundary, cross-reference VTune miss addresses.
97+
98+
Cheap knobs
99+
-----------
100+
101+
5. Raise the interrupt-hook interval. ``interrupt_init`` arms
102+
``LUA_MASKCOUNT, 256`` unconditionally (``LuaTools.cpp:508``) —
103+
``luaG_traceexec`` shows 82 samples. Raising the count trades
104+
runaway-script interrupt latency for less hook dispatch. Tune and
105+
verify interrupt still fires acceptably.
106+
6. GC tuning on the core state (``lua_gc`` pause/stepmul) — never
107+
configured today; only ``scripts/dwarf-op.lua`` calls
108+
``collectgarbage``. One-line experiments; watch memory growth.
109+
7. Audit per-frame Lua string work: ``match``/``singlematch``
110+
samples imply ``string.find``/patterns in a per-frame path;
111+
``luaS_hash``/``lua_tolstring`` churn implies string building per
112+
frame. Fix at the call site once attribution identifies it.
113+
8. Audit ``pairs()`` iteration over hash-part tables in hot loops
114+
(``luaH_next`` stalls): array-part iteration is contiguous and
115+
prefetchable; hash-part is not.
116+
117+
Structural: Lua heap layout
118+
---------------------------
119+
120+
9. Custom ``lua_Alloc`` with type/lifetime-segregated arenas.
121+
122+
- All GC objects funnel through ``luaC_newobj(L, tt, sz)`` — one
123+
place to tag by type; other allocs via ``luaM_realloc_`` by
124+
size class.
125+
- Permanent bump region for everything allocated during
126+
``Lua::Open`` + ``require dfhack`` + script load (protos,
127+
interned strings, metatables, fieldtables, module tables) —
128+
init order approximates tick access order; GC's ``allgc`` walk
129+
becomes quasi-sequential. Phase flag can live in
130+
``lua_extra_state`` (already used by ``dfhack_llimits.h``).
131+
- Separate churn region for post-init garbage.
132+
- Only 3 state-creation sites: ``Core.cpp:1394``,
133+
``LuaTools.cpp:1793``, ``LuaTools.test.cpp``.
134+
- Portable: works with plain VirtualAlloc/mmap; hugepages are a
135+
strictly optional overlay on top (see constraints).
136+
- Caveat: ``lua_Alloc`` must implement realloc semantics
137+
(alloc+copy+free is fine); Lua GC never moves objects
138+
(``push_adhoc_pointer`` relies on this already).
139+
10. Software prefetch at tick entry in ``Lua::Core::onUpdate``
140+
(``LuaTools.cpp:2073``): prefetch ``G(State)``, registry, and
141+
each due coroutine's ``stack``/``ci`` before ``lua_resume`` —
142+
portable via ``_mm_prefetch``/``__builtin_prefetch``. Cheap;
143+
converts serial cold misses into overlapped ones.
144+
145+
Binding layer (the userdata proxy question)
146+
-------------------------------------------
147+
148+
11. Kill the double lookup in the field path.
149+
``meta_struct_index`` → ``find_field`` → ``lookup_field``
150+
(``LuaTypes.cpp:422-454``) probes a Lua field-table *and* walks
151+
the enum metatable per access. Short strings are interned, so
152+
``TString*`` pointer equality = name equality: a small
153+
direct-mapped cache keyed ``(fieldtable, TString*)``, or a
154+
per-type dense index assigned to each field name at metatable
155+
build time, collapses this to one load. Highest-confidence fix —
156+
hits ``luaV_finishget``/``luaH_getshortstr``/``lua_getmetatable``
157+
simultaneously.
158+
12. Reduce per-access userdata churn. ``push_object_ref`` allocates
159+
a fresh userdata (+ ``object_ref_header``: tag_ptr, tag_identity,
160+
tag_attr, field_info) for every nested struct/container access —
161+
``unit.pos.x`` garbage-per-hop. Options, in increasing scope:
162+
163+
a) arena allocation makes the churn cheap and local (item 9);
164+
b) memoize refs via weak-valued cache keyed ``(ptr, identity)``
165+
— adds a hash probe, only worth it for expensive chains;
166+
c) bulk getters for hot access patterns (e.g. one call returning
167+
``x,y,z`` for ``pos``) — semantic addition, avoids N
168+
allocations and N metamethod hops per vector;
169+
d) pack/shrink ``object_ref_header`` — review which of its four
170+
pointers are needed on the common path; union-tag fields could
171+
live in a side table populated only for unions.
172+
13. Keep the proxy model; make the ref cheaper. A full move away
173+
from userdata proxies isn't practical (scripts rely on
174+
reference identity, ``__index`` polymorphism, ``_field``), but
175+
the identity-layer work in PR #5959 already removes one virtual
176+
dispatch from every primitive field read — that direction
177+
(more ``if constexpr`` static knowledge in the access path) is
178+
the right way to slim the proxy rather than replacing it.
179+
180+
C++-side containers
181+
-------------------
182+
183+
14. MSVC STL ``unordered_map``/``unordered_set``/``map`` on per-tick
184+
paths are chained-node structures with the same miss signature
185+
(``_Fnv1a_append_value``, ``_Hash`` loops appear in the stall
186+
list — e.g. inside ``Units::isActive``). Swap hot-path instances
187+
for flat maps / sorted vectors / precomputed indices.
188+
15. ``matchFocusString`` allocates ``std::string`` per call per
189+
frame — make it allocation-free (``string_view``, cached split).
190+
191+
Overlay / script layer
192+
----------------------
193+
194+
16. Once attribution lands (item 2/3): batch the overlay render
195+
path. ``dfhack.penarray`` bulk tile ops already exist
196+
(``LuaApi.cpp:997+``). Per-cell ``Pen`` writes through
197+
metamethods per frame should move to bulk APIs or C++ rendering.
198+
If overlay is indeed the budget owner, this is likely the
199+
largest single user-visible win.
200+
201+
Verification methodology
202+
========================
203+
204+
For each change: measure ``update_lua_ms`` / ``update_plugin_ms``
205+
before/after on the same save; VTune Hotspots (not uarch exploration
206+
— less noise) filtered to the sim thread; compare
207+
``MEMORY_ACTIVITY.STALLS_L3_MISS`` and ``MEM_LOAD_RETIRED.L3_MISS``
208+
inside ``lua53.dll``/``dfhooks_dfhack.dll``. Each item above is
209+
independently revertable.
210+
211+
Lua 5.4/5.5 upgrade assessment
212+
==============================
213+
214+
The ``__ipairs`` blocker is narrower than it appears:
215+
216+
- ``__ipairs`` is registered on all generated metatables
217+
(``LuaTypes.cpp`` ``SetPairsMethod`` x4; ``LuaWrapper.cpp``
218+
``wtype_ipairs``/``complex_enum_ipairs``). It exists because some
219+
proxy ``__index`` implementations never return nil for integer
220+
keys (enum attrs — DFHack/dfhack#1860), so plain ``ipairs`` would
221+
not terminate.
222+
- Under 5.4, ``__ipairs`` is ignored entirely; ``ipairs`` indexes
223+
``t[i]`` until nil. Migration path: make every proxy ``__index``
224+
bounded for integer keys (containers already bound by item_count;
225+
enum attrs need an explicit bound — the enum count is known), then
226+
audit direct ``__ipairs`` users (one script:
227+
``scripts/test/fix/stuck-written-materials.lua``; plus
228+
``test/structures/enum_attrs.lua`` semantics).
229+
- Real benefits for our profile: **generational GC** (young-gen
230+
collections over small dense sets directly attack the
231+
cold-sweep problem), ~10-25% faster VM, ``lua_newuserdatauv``
232+
uservalues (could fold ``object_ref_header`` into a uservalue —
233+
see item 12d).
234+
- Costs: semantic audit of ~500 scripts (integer/string coercion
235+
strictness, ``math.random`` algorithm change breaks seeded
236+
determinism, ``lua_resume`` signature change — we call it
237+
directly in ``LuaTools.cpp:880``), bytecode format change
238+
(``dumper.lua`` ``string.dump`` roundtrips). Moderate project.
239+
- Recommendation: pursue items 9-12 first (larger wins per risk);
240+
revisit 5.4 as a GC-locality lever afterward. 5.5 is not yet a
241+
stable target; track it but don't plan against it.
242+
243+
PR #5959 review (performance notes)
244+
===================================
245+
246+
Net-positive for the hot path, with watch-items:
247+
248+
- **Win:** ``type_identity_for<T>::lua_read``/``lua_write`` inline
249+
``*(T*)ptr`` via ``if constexpr`` — removes the second virtual
250+
call (old ``integer_identity_base::lua_read`` → virtual
251+
``read()``) on every primitive field access. This is exactly the
252+
``read_field``/``write_field`` hot path.
253+
- **Watch:** ``mapped<T>`` containers (``std::map``/
254+
``unordered_map``) now get identities with O(n) ``item_pointer``
255+
iteration. Any struct field of map type newly exposed to Lua
256+
(check the ``library/xml`` submodule bump) becomes O(n^2) under
257+
``ipairs``-style iteration — audit which types gain this.
258+
- **Watch:** ``container_storage<C<E*,A...>>`` generalizes the
259+
"vector<T*> == vector<void*>" layout assumption to *any*
260+
random-access pointer container (deque etc.). The assumption
261+
holds for same-width pointer instantiations in practice, but the
262+
blast radius widened.
263+
- **Note:** removing ``BUILD_DFHACK_LIB`` guards lets plugins
264+
instantiate ``container_impl``/``type_identity_for`` locally,
265+
creating non-canonical identity objects — ``is_type_compatible``
266+
compares identity *pointers*, so a plugin-side local instance of
267+
an existing identity could yield false type mismatches. Document
268+
that identities must come from ``identity_traits``/canonical
269+
sources only.
270+
- ``field_error``/``get_object_internal`` newly ``DFHACK_EXPORT`` —
271+
required by header-side template instantiation; fine.
272+
- ``LuaTools.h`` now includes ``DataIdentity.h`` — heavier include
273+
in every consumer TU; compile-time only.
274+
- ``bool`` moved to ``NUMBER_IDENTITY_TRAITS`` with correct
275+
``isInteger()==false`` and bool-specific read/write — semantics
276+
preserved (``type()`` still ``IDTYPE_PRIMITIVE``, name "bool").
277+
- ``c_string`` → ``primitive_identity_base`` preserves
278+
"raw pointer string" write error and push-string read semantics.
279+
- ``push_adhoc_pointer`` retains the "GC never moves objects" hack —
280+
reinforces that any allocator work must stay non-moving.
281+
282+
Suggested order
283+
===============
284+
285+
1. Callsite attribution (who owns the Lua stalls).
286+
2. Cheap knobs: hook count, GC tuning, string-work audit.
287+
3. Field-lookup inline cache (item 11).
288+
4. Arena allocator (items 9-10).
289+
5. C++ container swaps + ``matchFocusString`` (items 14-15).
290+
6. Overlay batching per attribution (item 16).
291+
7. Lua 5.4 evaluation (generational GC) once 1-6 land.

0 commit comments

Comments
 (0)