Issue Description
A single corrupt LZ4-compressed node object in the NuDB nodestore makes xrpld abort the process instead of treating the object as corrupt or missing. Because the object is read again on every startup during ledger acquisition, this turns into a crash loop under systemd (Restart=on-failure). In our case: 24 SIGABRT core dumps in ~70 minutes. The only recovery was to delete the entire nodestore and resync from peers.
The backend already has a graceful path for corrupt data, but LZ4 failures bypass it:
NuDBBackend::fetch handles one kind of corruption gracefully. If DecodedBlob::wasOk() is false, it returns Status::DataCorrupt, and DatabaseRotatingImp logs Corrupt NodeObject #<hash> and returns nullptr. The caller then treats the object as missing, so it can be re-acquired from peers.
NuDBFactory.cpp#L209-L233
- But
nodeobjectDecompress() is called on the line above that check (L218) with no try/catch. When the stored bytes aren't valid LZ4, lz4Decompress throws std::runtime_error("lz4_decompress: LZ4_decompress_safe").
codec.h#L45-L50
DatabaseRotatingImp::fetchNodeObject catches that exception, logs it, and then calls rethrow(). It escapes to std::terminate → abort.
DatabaseRotatingImp.cpp#L151-L158
So the same class of fault (corrupt stored bytes) is either recoverable or fatal, depending only on which decode step notices it first. The code is unchanged on develop as of 2026-09-30. NuDBFactory.cpp L297, a second nodeobjectDecompress call site, has the same pattern.
A related problem affects repair even if the exception were caught: doInsert ignores nudb::error::key_exists (L244). A correct copy re-fetched from peers therefore can't replace the corrupt value in the same backend, and the next cold read of that key would hit the bad bytes again.
Steps to Reproduce
We hit this in production rather than by design; the trigger was silent corruption of one stored value (see Environment). A deterministic repro should be:
- Unit-level: store a blob for some key whose varint length prefix is valid but whose LZ4 payload is garbage, then call
NuDBBackend::fetch on that key. Expected: Status::DataCorrupt. Actual: throws.
- Node-level: on a stopped node, flip bytes inside the LZ4 payload of one record in the NuDB
.dat file, for a node that will be read during acquisition. Start xrpld; it aborts on each start.
Expected Result
- A decompression failure is handled the same way as
!decoded.wasOk(): return Status::DataCorrupt and log the hash at fatal/error.
- The object is treated as missing and re-acquired from peers, and the server keeps running.
- The re-acquired object can replace the corrupt one on disk. For example, write it to the writable backend and prefer that copy, or keep a small list of known-corrupt keys that fetch reports as not found until rotation drops the bad backend. Otherwise the corruption comes back on the next cold read.
Actual Result
- Every crash shows the same message:
terminate called after throwing an instance of 'std::runtime_error'
what(): lz4_decompress: LZ4_decompress_safe
xrpld.service: Main process exited, code=dumped, status=6/ABRT
- The first crash came from a process that had been healthy for ~1.5 days, with no preceding restart. After that, every start aborted within ~40–90 s, during ledger acquisition.
- The hash of the corrupt object is not logged on the exception path, so the bad record can't be identified or targeted, for example with
ledger_cleaner.
ledger_cleaner can't be used as a workaround, because it reads through the same fetchNodeObject path and the server never stays up long enough.
- Recovery required stopping xrpld, deleting the whole nodestore (NuDB +
ledger.db/transaction.db/state.db, keeping wallet.db), and doing a full from-scratch state sync from peers (~1 h to full with node_size=medium).
Crashing thread, from coredumpctl info (symbolized frames):
#7 xrpl::rethrow()
#8 n/a (xrpld + 0x5728c2f) <- unsymbolized; presumably the fetch lambda in DatabaseRotatingImp
#9 xrpl::node_store::DatabaseRotatingImp::fetchNodeObject(xrpl::base_uint<256> const&, unsigned int, xrpl::node_store::FetchReport&, bool)
#10 xrpl::node_store::Database::fetchNodeObject(xrpl::base_uint<256> const&, unsigned int, xrpl::node_store::FetchType, bool)
Concurrent thread at the same moment (acquisition driving the read):
xrpl::SHAMap::descendAsync(...)
xrpl::SHAMap::gmnProcessNodes(...)
xrpl::SHAMap::getMissingNodes(int, xrpl::SHAMapSyncFilter const*)
xrpl::InboundLedger::trigger(std::shared_ptr<xrpl::Peer> const&, xrpl::InboundLedger::TriggerReason)
xrpl::InboundLedger::runData()
Environment
xrpld version 3.4.1, git d147fccf54a500fce586522f28d6044c37fd8d29, package xrpld 3.4.1-1 from packages.xrplf.org. That commit isn't public yet, so the links above point to 3.4.0; the relevant code is identical on develop.
- Ubuntu 24.04.4 LTS, kernel 6.8.0, x86_64, 12 vCPU / 47 GB, KVM VPS.
- Nodestore:
type=NuDB on ext4, online_delete=5000, advisory_delete=0. node_size was small at the time of the crash (now medium). Non-validator, serving a co-located Clio via gRPC.
- No disk I/O or filesystem errors in
dmesg, and the disk wasn't full. As a KVM guest the node can't see host ECC/MCE events, so the corruption source (memory bit-flip vs. storage bit-rot vs. software) can't be determined from the guest. The NuDB structure was intact (the varint length prefix parsed); only the LZ4 payload was bad.
Possibly related: #4883 (corruption correlated with tight online_delete intervals; this crash was in the rotating backend) and #7246 (LZ4 error handling on the overlay path; different code, same class of gap). The nodestore compress side does check for failure (lz4Compress throws on outSize == 0), so a failed compression can't silently write a bad blob.
Supporting Files
- Full
coredumpctl info backtrace available on request.
- One core dump (745 MB,
.zst) retained. We can share it privately with maintainers, but won't post it publicly because it contains process memory.
Issue Description
A single corrupt LZ4-compressed node object in the NuDB nodestore makes xrpld abort the process instead of treating the object as corrupt or missing. Because the object is read again on every startup during ledger acquisition, this turns into a crash loop under
systemd(Restart=on-failure). In our case: 24SIGABRTcore dumps in ~70 minutes. The only recovery was to delete the entire nodestore and resync from peers.The backend already has a graceful path for corrupt data, but LZ4 failures bypass it:
NuDBBackend::fetchhandles one kind of corruption gracefully. IfDecodedBlob::wasOk()is false, it returnsStatus::DataCorrupt, andDatabaseRotatingImplogsCorrupt NodeObject #<hash>and returnsnullptr. The caller then treats the object as missing, so it can be re-acquired from peers.NuDBFactory.cpp#L209-L233
nodeobjectDecompress()is called on the line above that check (L218) with no try/catch. When the stored bytes aren't valid LZ4,lz4Decompressthrowsstd::runtime_error("lz4_decompress: LZ4_decompress_safe").codec.h#L45-L50
DatabaseRotatingImp::fetchNodeObjectcatches that exception, logs it, and then callsrethrow(). It escapes tostd::terminate→abort.DatabaseRotatingImp.cpp#L151-L158
So the same class of fault (corrupt stored bytes) is either recoverable or fatal, depending only on which decode step notices it first. The code is unchanged on
developas of 2026-09-30.NuDBFactory.cppL297, a secondnodeobjectDecompresscall site, has the same pattern.A related problem affects repair even if the exception were caught:
doInsertignoresnudb::error::key_exists(L244). A correct copy re-fetched from peers therefore can't replace the corrupt value in the same backend, and the next cold read of that key would hit the bad bytes again.Steps to Reproduce
We hit this in production rather than by design; the trigger was silent corruption of one stored value (see Environment). A deterministic repro should be:
NuDBBackend::fetchon that key. Expected:Status::DataCorrupt. Actual: throws..datfile, for a node that will be read during acquisition. Start xrpld; it aborts on each start.Expected Result
!decoded.wasOk(): returnStatus::DataCorruptand log the hash atfatal/error.Actual Result
ledger_cleaner.ledger_cleanercan't be used as a workaround, because it reads through the samefetchNodeObjectpath and the server never stays up long enough.ledger.db/transaction.db/state.db, keepingwallet.db), and doing a full from-scratch state sync from peers (~1 h tofullwithnode_size=medium).Crashing thread, from
coredumpctl info(symbolized frames):Concurrent thread at the same moment (acquisition driving the read):
Environment
xrpld version 3.4.1, gitd147fccf54a500fce586522f28d6044c37fd8d29, packagexrpld 3.4.1-1from packages.xrplf.org. That commit isn't public yet, so the links above point to3.4.0; the relevant code is identical ondevelop.type=NuDBon ext4,online_delete=5000,advisory_delete=0.node_sizewassmallat the time of the crash (nowmedium). Non-validator, serving a co-located Clio via gRPC.dmesg, and the disk wasn't full. As a KVM guest the node can't see host ECC/MCE events, so the corruption source (memory bit-flip vs. storage bit-rot vs. software) can't be determined from the guest. The NuDB structure was intact (the varint length prefix parsed); only the LZ4 payload was bad.Possibly related: #4883 (corruption correlated with tight
online_deleteintervals; this crash was in the rotating backend) and #7246 (LZ4 error handling on the overlay path; different code, same class of gap). The nodestore compress side does check for failure (lz4Compressthrows onoutSize == 0), so a failed compression can't silently write a bad blob.Supporting Files
coredumpctl infobacktrace available on request..zst) retained. We can share it privately with maintainers, but won't post it publicly because it contains process memory.