|
| 1 | +========================= |
| 2 | +Performance plan |
| 3 | +========================= |
| 4 | + |
| 5 | +Working plan for improving DFHack's per-tick CPU budget. Based on |
| 6 | +microarchitecture analysis (Intel VTune capture ``r120ue``, Jan 2026) |
| 7 | +and code review of the Lua integration. Not user-facing documentation. |
| 8 | + |
| 9 | +.. contents:: |
| 10 | + |
| 11 | +Baseline findings |
| 12 | +================= |
| 13 | + |
| 14 | +From ``r120ue`` (uarch exploration, whole-system sampling; ~24k |
| 15 | +``CPU_CLK_UNHALTED.THREAD`` samples): |
| 16 | + |
| 17 | +- DFHack-side code is ~21% of sampled CPU: ``lua53.dll`` 3322, |
| 18 | + ``dfhooks_dfhack.dll`` 1750, plugins ~110. |
| 19 | +- DFHack owns ~11% of ``STALLS_L3_MISS`` and ~11% of DTLB |
| 20 | + ``WALK_ACTIVE`` samples. |
| 21 | +- Top lua53 stallers are data-layout problems, not dispatch: |
| 22 | + ``luaH_getshortstr`` (hash-chain walk, 17 stall samples), |
| 23 | + ``luaD_precall``, ``lua_geti``, ``luaH_next`` (``pairs`` iteration), |
| 24 | + ``lua_getmetatable``, ``mainposition``, ``luaC_checkfinalizer``. |
| 25 | +- Top CPU consumers are the C<->Lua transition layer: |
| 26 | + ``luaV_execute`` 759, ``luaD_precall`` 260, ``index2addr`` 115, |
| 27 | + ``luaV_finishget`` (metamethod resolution) 90, ``luaG_traceexec`` |
| 28 | + 82 (the always-armed ``LUA_MASKCOUNT`` interrupt hook), |
| 29 | + ``match``/``singlematch`` ~88 (Lua pattern matching per frame), |
| 30 | + plus ``lua_tolstring``/``luaS_hash`` churn. |
| 31 | +- C++ side: MSVC ``std::_Hash``/``std::_Tree`` internals, |
| 32 | + ``std::string`` construct/destroy churn, ``matchFocusString``, |
| 33 | + ``Units::isActive`` — chained-node STL containers have the same |
| 34 | + miss problem as Lua's tables. |
| 35 | + |
| 36 | +Working hypothesis (operator model): DF's own memory traffic evicts |
| 37 | +essentially all DFHack data from L3 between simulation ticks, so every |
| 38 | +tick starts cold and each pointer-chase (hash chain, GC list, |
| 39 | +metatable walk) pays full miss latency. Goal: make the per-tick cold |
| 40 | +walk *sequential* (prefetcher-trackable) rather than pointer-chasing, |
| 41 | +and reduce the number of misses by shortening dependent-load chains. |
| 42 | + |
| 43 | +Overlay processing is believed to dominate the per-tick budget; this |
| 44 | +is consistent with the data but not yet proven — see the attribution |
| 45 | +step below. |
| 46 | + |
| 47 | +Constraints and rejected directions |
| 48 | +=================================== |
| 49 | + |
| 50 | +- **No PGO/LTO (deferred).** DFHack does questionable things to DF's |
| 51 | + vtables and binary; whole-program optimization changes codegen |
| 52 | + assumptions the interpose/binpatch machinery relies on. Would need |
| 53 | + a dedicated validation effort before it is safe to enable. Revisit |
| 54 | + only if cheaper items are exhausted. |
| 55 | +- **Hugepages: optional-only.** ``VirtualAlloc(MEM_LARGE_PAGES)`` |
| 56 | + needs ``SeLockMemoryPrivilege`` and often fails for normal users; |
| 57 | + Linux needs THP/hugetlbfs. Steam Deck (Linux) is a significant |
| 58 | + user population and we have no Linux profiling. Any allocator work |
| 59 | + must function with ordinary pages; hugepages may be an opt-in |
| 60 | + bonus path (``madvise(MADV_HUGEPAGE)`` on Linux, |
| 61 | + ``MEM_LARGE_PAGES`` on Windows) with mandatory fallback. |
| 62 | +- **No LLVM-JIT Lua (Ravi etc.).** The miss profile is in |
| 63 | + ``ltable``/``lgc``/``ldo`` (data layout + transitions), not |
| 64 | + ``lvm`` dispatch. A JIT keeps identical object layout and would |
| 65 | + also have to reproduce our ``lua_lock`` threading patch, |
| 66 | + ``LUA_MASKCOUNT`` interrupt hooks, ``lstate.h`` field access, |
| 67 | + yieldable pcall, and debug-hook fidelity. Wrong tool for this |
| 68 | + problem. |
| 69 | +- **LuaJIT** has the best-in-class data layout but is Lua 5.1 |
| 70 | + semantics; incompatible with our 5.3 corpus. |
| 71 | + |
| 72 | +Action items |
| 73 | +============ |
| 74 | + |
| 75 | +Ordered by leverage-per-effort. All are architecture-independent |
| 76 | +unless noted. |
| 77 | + |
| 78 | +Measurement |
| 79 | +----------- |
| 80 | + |
| 81 | +1. Per-tick attribution. ``perf_counters.update_lua_ms`` / |
| 82 | + ``update_plugin_ms`` (``Core.cpp:1668-1682``) already isolate Lua |
| 83 | + update time; ``script-manager.lua`` prints them. Baseline a real |
| 84 | + fort first. |
| 85 | +2. Callsite-stack attribution on VTune captures: ``dd_callsite`` |
| 86 | + parent links in ``dicer.db`` allow rebuilding stacks and splitting |
| 87 | + Lua-side stalls between overlay widgets, timers, and plugin |
| 88 | + events. (Scratch query script: ``%TEMP%/vtune_agg.py``; accepts a |
| 89 | + ``dicer.db`` path for other captures.) |
| 90 | +3. Lua-level: run ``profiler.lua`` scoped to the update window, or a |
| 91 | + count hook aggregating by Proto, to identify which script |
| 92 | + functions dominate ticks. |
| 93 | +4. Object-level (if needed): we own the Lua source — log |
| 94 | + ``luaC_newobj``/``luaM_realloc_`` (addr, type, size, phase) during |
| 95 | + a session, dump live objects by walking ``allgc`` at a tick |
| 96 | + boundary, cross-reference VTune miss addresses. |
| 97 | + |
| 98 | +Cheap knobs |
| 99 | +----------- |
| 100 | + |
| 101 | +5. Raise the interrupt-hook interval. ``interrupt_init`` arms |
| 102 | + ``LUA_MASKCOUNT, 256`` unconditionally (``LuaTools.cpp:508``) — |
| 103 | + ``luaG_traceexec`` shows 82 samples. Raising the count trades |
| 104 | + runaway-script interrupt latency for less hook dispatch. Tune and |
| 105 | + verify interrupt still fires acceptably. |
| 106 | +6. GC tuning on the core state (``lua_gc`` pause/stepmul) — never |
| 107 | + configured today; only ``scripts/dwarf-op.lua`` calls |
| 108 | + ``collectgarbage``. One-line experiments; watch memory growth. |
| 109 | +7. Audit per-frame Lua string work: ``match``/``singlematch`` |
| 110 | + samples imply ``string.find``/patterns in a per-frame path; |
| 111 | + ``luaS_hash``/``lua_tolstring`` churn implies string building per |
| 112 | + frame. Fix at the call site once attribution identifies it. |
| 113 | +8. Audit ``pairs()`` iteration over hash-part tables in hot loops |
| 114 | + (``luaH_next`` stalls): array-part iteration is contiguous and |
| 115 | + prefetchable; hash-part is not. |
| 116 | + |
| 117 | +Structural: Lua heap layout |
| 118 | +--------------------------- |
| 119 | + |
| 120 | +9. Custom ``lua_Alloc`` with type/lifetime-segregated arenas. |
| 121 | + |
| 122 | + - All GC objects funnel through ``luaC_newobj(L, tt, sz)`` — one |
| 123 | + place to tag by type; other allocs via ``luaM_realloc_`` by |
| 124 | + size class. |
| 125 | + - Permanent bump region for everything allocated during |
| 126 | + ``Lua::Open`` + ``require dfhack`` + script load (protos, |
| 127 | + interned strings, metatables, fieldtables, module tables) — |
| 128 | + init order approximates tick access order; GC's ``allgc`` walk |
| 129 | + becomes quasi-sequential. Phase flag can live in |
| 130 | + ``lua_extra_state`` (already used by ``dfhack_llimits.h``). |
| 131 | + - Separate churn region for post-init garbage. |
| 132 | + - Only 3 state-creation sites: ``Core.cpp:1394``, |
| 133 | + ``LuaTools.cpp:1793``, ``LuaTools.test.cpp``. |
| 134 | + - Portable: works with plain VirtualAlloc/mmap; hugepages are a |
| 135 | + strictly optional overlay on top (see constraints). |
| 136 | + - Caveat: ``lua_Alloc`` must implement realloc semantics |
| 137 | + (alloc+copy+free is fine); Lua GC never moves objects |
| 138 | + (``push_adhoc_pointer`` relies on this already). |
| 139 | +10. Software prefetch at tick entry in ``Lua::Core::onUpdate`` |
| 140 | + (``LuaTools.cpp:2073``): prefetch ``G(State)``, registry, and |
| 141 | + each due coroutine's ``stack``/``ci`` before ``lua_resume`` — |
| 142 | + portable via ``_mm_prefetch``/``__builtin_prefetch``. Cheap; |
| 143 | + converts serial cold misses into overlapped ones. |
| 144 | + |
| 145 | +Binding layer (the userdata proxy question) |
| 146 | +------------------------------------------- |
| 147 | + |
| 148 | +11. Kill the double lookup in the field path. |
| 149 | + ``meta_struct_index`` → ``find_field`` → ``lookup_field`` |
| 150 | + (``LuaTypes.cpp:422-454``) probes a Lua field-table *and* walks |
| 151 | + the enum metatable per access. Short strings are interned, so |
| 152 | + ``TString*`` pointer equality = name equality: a small |
| 153 | + direct-mapped cache keyed ``(fieldtable, TString*)``, or a |
| 154 | + per-type dense index assigned to each field name at metatable |
| 155 | + build time, collapses this to one load. Highest-confidence fix — |
| 156 | + hits ``luaV_finishget``/``luaH_getshortstr``/``lua_getmetatable`` |
| 157 | + simultaneously. |
| 158 | +12. Reduce per-access userdata churn. ``push_object_ref`` allocates |
| 159 | + a fresh userdata (+ ``object_ref_header``: tag_ptr, tag_identity, |
| 160 | + tag_attr, field_info) for every nested struct/container access — |
| 161 | + ``unit.pos.x`` garbage-per-hop. Options, in increasing scope: |
| 162 | + |
| 163 | + a) arena allocation makes the churn cheap and local (item 9); |
| 164 | + b) memoize refs via weak-valued cache keyed ``(ptr, identity)`` |
| 165 | + — adds a hash probe, only worth it for expensive chains; |
| 166 | + c) bulk getters for hot access patterns (e.g. one call returning |
| 167 | + ``x,y,z`` for ``pos``) — semantic addition, avoids N |
| 168 | + allocations and N metamethod hops per vector; |
| 169 | + d) pack/shrink ``object_ref_header`` — review which of its four |
| 170 | + pointers are needed on the common path; union-tag fields could |
| 171 | + live in a side table populated only for unions. |
| 172 | +13. Keep the proxy model; make the ref cheaper. A full move away |
| 173 | + from userdata proxies isn't practical (scripts rely on |
| 174 | + reference identity, ``__index`` polymorphism, ``_field``), but |
| 175 | + the identity-layer work in PR #5959 already removes one virtual |
| 176 | + dispatch from every primitive field read — that direction |
| 177 | + (more ``if constexpr`` static knowledge in the access path) is |
| 178 | + the right way to slim the proxy rather than replacing it. |
| 179 | + |
| 180 | +C++-side containers |
| 181 | +------------------- |
| 182 | + |
| 183 | +14. MSVC STL ``unordered_map``/``unordered_set``/``map`` on per-tick |
| 184 | + paths are chained-node structures with the same miss signature |
| 185 | + (``_Fnv1a_append_value``, ``_Hash`` loops appear in the stall |
| 186 | + list — e.g. inside ``Units::isActive``). Swap hot-path instances |
| 187 | + for flat maps / sorted vectors / precomputed indices. |
| 188 | +15. ``matchFocusString`` allocates ``std::string`` per call per |
| 189 | + frame — make it allocation-free (``string_view``, cached split). |
| 190 | + |
| 191 | +Overlay / script layer |
| 192 | +---------------------- |
| 193 | + |
| 194 | +16. Once attribution lands (item 2/3): batch the overlay render |
| 195 | + path. ``dfhack.penarray`` bulk tile ops already exist |
| 196 | + (``LuaApi.cpp:997+``). Per-cell ``Pen`` writes through |
| 197 | + metamethods per frame should move to bulk APIs or C++ rendering. |
| 198 | + If overlay is indeed the budget owner, this is likely the |
| 199 | + largest single user-visible win. |
| 200 | + |
| 201 | +Verification methodology |
| 202 | +======================== |
| 203 | + |
| 204 | +For each change: measure ``update_lua_ms`` / ``update_plugin_ms`` |
| 205 | +before/after on the same save; VTune Hotspots (not uarch exploration |
| 206 | +— less noise) filtered to the sim thread; compare |
| 207 | +``MEMORY_ACTIVITY.STALLS_L3_MISS`` and ``MEM_LOAD_RETIRED.L3_MISS`` |
| 208 | +inside ``lua53.dll``/``dfhooks_dfhack.dll``. Each item above is |
| 209 | +independently revertable. |
| 210 | + |
| 211 | +Lua 5.4/5.5 upgrade assessment |
| 212 | +============================== |
| 213 | + |
| 214 | +The ``__ipairs`` blocker is narrower than it appears: |
| 215 | + |
| 216 | +- ``__ipairs`` is registered on all generated metatables |
| 217 | + (``LuaTypes.cpp`` ``SetPairsMethod`` x4; ``LuaWrapper.cpp`` |
| 218 | + ``wtype_ipairs``/``complex_enum_ipairs``). It exists because some |
| 219 | + proxy ``__index`` implementations never return nil for integer |
| 220 | + keys (enum attrs — DFHack/dfhack#1860), so plain ``ipairs`` would |
| 221 | + not terminate. |
| 222 | +- Under 5.4, ``__ipairs`` is ignored entirely; ``ipairs`` indexes |
| 223 | + ``t[i]`` until nil. Migration path: make every proxy ``__index`` |
| 224 | + bounded for integer keys (containers already bound by item_count; |
| 225 | + enum attrs need an explicit bound — the enum count is known), then |
| 226 | + audit direct ``__ipairs`` users (one script: |
| 227 | + ``scripts/test/fix/stuck-written-materials.lua``; plus |
| 228 | + ``test/structures/enum_attrs.lua`` semantics). |
| 229 | +- Real benefits for our profile: **generational GC** (young-gen |
| 230 | + collections over small dense sets directly attack the |
| 231 | + cold-sweep problem), ~10-25% faster VM, ``lua_newuserdatauv`` |
| 232 | + uservalues (could fold ``object_ref_header`` into a uservalue — |
| 233 | + see item 12d). |
| 234 | +- Costs: semantic audit of ~500 scripts (integer/string coercion |
| 235 | + strictness, ``math.random`` algorithm change breaks seeded |
| 236 | + determinism, ``lua_resume`` signature change — we call it |
| 237 | + directly in ``LuaTools.cpp:880``), bytecode format change |
| 238 | + (``dumper.lua`` ``string.dump`` roundtrips). Moderate project. |
| 239 | +- Recommendation: pursue items 9-12 first (larger wins per risk); |
| 240 | + revisit 5.4 as a GC-locality lever afterward. 5.5 is not yet a |
| 241 | + stable target; track it but don't plan against it. |
| 242 | + |
| 243 | +PR #5959 review (performance notes) |
| 244 | +=================================== |
| 245 | + |
| 246 | +Net-positive for the hot path, with watch-items: |
| 247 | + |
| 248 | +- **Win:** ``type_identity_for<T>::lua_read``/``lua_write`` inline |
| 249 | + ``*(T*)ptr`` via ``if constexpr`` — removes the second virtual |
| 250 | + call (old ``integer_identity_base::lua_read`` → virtual |
| 251 | + ``read()``) on every primitive field access. This is exactly the |
| 252 | + ``read_field``/``write_field`` hot path. |
| 253 | +- **Watch:** ``mapped<T>`` containers (``std::map``/ |
| 254 | + ``unordered_map``) now get identities with O(n) ``item_pointer`` |
| 255 | + iteration. Any struct field of map type newly exposed to Lua |
| 256 | + (check the ``library/xml`` submodule bump) becomes O(n^2) under |
| 257 | + ``ipairs``-style iteration — audit which types gain this. |
| 258 | +- **Watch:** ``container_storage<C<E*,A...>>`` generalizes the |
| 259 | + "vector<T*> == vector<void*>" layout assumption to *any* |
| 260 | + random-access pointer container (deque etc.). The assumption |
| 261 | + holds for same-width pointer instantiations in practice, but the |
| 262 | + blast radius widened. |
| 263 | +- **Note:** removing ``BUILD_DFHACK_LIB`` guards lets plugins |
| 264 | + instantiate ``container_impl``/``type_identity_for`` locally, |
| 265 | + creating non-canonical identity objects — ``is_type_compatible`` |
| 266 | + compares identity *pointers*, so a plugin-side local instance of |
| 267 | + an existing identity could yield false type mismatches. Document |
| 268 | + that identities must come from ``identity_traits``/canonical |
| 269 | + sources only. |
| 270 | +- ``field_error``/``get_object_internal`` newly ``DFHACK_EXPORT`` — |
| 271 | + required by header-side template instantiation; fine. |
| 272 | +- ``LuaTools.h`` now includes ``DataIdentity.h`` — heavier include |
| 273 | + in every consumer TU; compile-time only. |
| 274 | +- ``bool`` moved to ``NUMBER_IDENTITY_TRAITS`` with correct |
| 275 | + ``isInteger()==false`` and bool-specific read/write — semantics |
| 276 | + preserved (``type()`` still ``IDTYPE_PRIMITIVE``, name "bool"). |
| 277 | +- ``c_string`` → ``primitive_identity_base`` preserves |
| 278 | + "raw pointer string" write error and push-string read semantics. |
| 279 | +- ``push_adhoc_pointer`` retains the "GC never moves objects" hack — |
| 280 | + reinforces that any allocator work must stay non-moving. |
| 281 | + |
| 282 | +Suggested order |
| 283 | +=============== |
| 284 | + |
| 285 | +1. Callsite attribution (who owns the Lua stalls). |
| 286 | +2. Cheap knobs: hook count, GC tuning, string-work audit. |
| 287 | +3. Field-lookup inline cache (item 11). |
| 288 | +4. Arena allocator (items 9-10). |
| 289 | +5. C++ container swaps + ``matchFocusString`` (items 14-15). |
| 290 | +6. Overlay batching per attribution (item 16). |
| 291 | +7. Lua 5.4 evaluation (generational GC) once 1-6 land. |
0 commit comments