Skip to content

Corrupt NuDB node object aborts xrpld (uncaught lz4_decompress exception) instead of returning DataCorrupt, causing a restart crash loop (Version: 3.4.1) #8343

Description

@nsmithau

Issue Description

A single corrupt LZ4-compressed node object in the NuDB nodestore makes xrpld abort the process instead of treating the object as corrupt or missing. Because the object is read again on every startup during ledger acquisition, this turns into a crash loop under systemd (Restart=on-failure). In our case: 24 SIGABRT core dumps in ~70 minutes. The only recovery was to delete the entire nodestore and resync from peers.

The backend already has a graceful path for corrupt data, but LZ4 failures bypass it:

  1. NuDBBackend::fetch handles one kind of corruption gracefully. If DecodedBlob::wasOk() is false, it returns Status::DataCorrupt, and DatabaseRotatingImp logs Corrupt NodeObject #<hash> and returns nullptr. The caller then treats the object as missing, so it can be re-acquired from peers.
    NuDBFactory.cpp#L209-L233
  2. But nodeobjectDecompress() is called on the line above that check (L218) with no try/catch. When the stored bytes aren't valid LZ4, lz4Decompress throws std::runtime_error("lz4_decompress: LZ4_decompress_safe").
    codec.h#L45-L50
  3. DatabaseRotatingImp::fetchNodeObject catches that exception, logs it, and then calls rethrow(). It escapes to std::terminate → abort.
    DatabaseRotatingImp.cpp#L151-L158

So the same class of fault (corrupt stored bytes) is either recoverable or fatal, depending only on which decode step notices it first. The code is unchanged on develop as of 2026-09-30. NuDBFactory.cpp L297, a second nodeobjectDecompress call site, has the same pattern.

A related problem affects repair even if the exception were caught: doInsert ignores nudb::error::key_exists (L244). A correct copy re-fetched from peers therefore can't replace the corrupt value in the same backend, and the next cold read of that key would hit the bad bytes again.

Steps to Reproduce

We hit this in production rather than by design; the trigger was silent corruption of one stored value (see Environment). A deterministic repro should be:

  1. Unit-level: store a blob for some key whose varint length prefix is valid but whose LZ4 payload is garbage, then call NuDBBackend::fetch on that key. Expected: Status::DataCorrupt. Actual: throws.
  2. Node-level: on a stopped node, flip bytes inside the LZ4 payload of one record in the NuDB .dat file, for a node that will be read during acquisition. Start xrpld; it aborts on each start.

Expected Result

  • A decompression failure is handled the same way as !decoded.wasOk(): return Status::DataCorrupt and log the hash at fatal/error.
  • The object is treated as missing and re-acquired from peers, and the server keeps running.
  • The re-acquired object can replace the corrupt one on disk. For example, write it to the writable backend and prefer that copy, or keep a small list of known-corrupt keys that fetch reports as not found until rotation drops the bad backend. Otherwise the corruption comes back on the next cold read.

Actual Result

  • Every crash shows the same message:
    terminate called after throwing an instance of 'std::runtime_error'
      what():  lz4_decompress: LZ4_decompress_safe
    xrpld.service: Main process exited, code=dumped, status=6/ABRT
    
  • The first crash came from a process that had been healthy for ~1.5 days, with no preceding restart. After that, every start aborted within ~40–90 s, during ledger acquisition.
  • The hash of the corrupt object is not logged on the exception path, so the bad record can't be identified or targeted, for example with ledger_cleaner.
  • ledger_cleaner can't be used as a workaround, because it reads through the same fetchNodeObject path and the server never stays up long enough.
  • Recovery required stopping xrpld, deleting the whole nodestore (NuDB + ledger.db/transaction.db/state.db, keeping wallet.db), and doing a full from-scratch state sync from peers (~1 h to full with node_size=medium).

Crashing thread, from coredumpctl info (symbolized frames):

#7  xrpl::rethrow()
#8  n/a (xrpld + 0x5728c2f)          <- unsymbolized; presumably the fetch lambda in DatabaseRotatingImp
#9  xrpl::node_store::DatabaseRotatingImp::fetchNodeObject(xrpl::base_uint<256> const&, unsigned int, xrpl::node_store::FetchReport&, bool)
#10 xrpl::node_store::Database::fetchNodeObject(xrpl::base_uint<256> const&, unsigned int, xrpl::node_store::FetchType, bool)

Concurrent thread at the same moment (acquisition driving the read):

xrpl::SHAMap::descendAsync(...)
xrpl::SHAMap::gmnProcessNodes(...)
xrpl::SHAMap::getMissingNodes(int, xrpl::SHAMapSyncFilter const*)
xrpl::InboundLedger::trigger(std::shared_ptr<xrpl::Peer> const&, xrpl::InboundLedger::TriggerReason)
xrpl::InboundLedger::runData()

Environment

  • xrpld version 3.4.1, git d147fccf54a500fce586522f28d6044c37fd8d29, package xrpld 3.4.1-1 from packages.xrplf.org. That commit isn't public yet, so the links above point to 3.4.0; the relevant code is identical on develop.
  • Ubuntu 24.04.4 LTS, kernel 6.8.0, x86_64, 12 vCPU / 47 GB, KVM VPS.
  • Nodestore: type=NuDB on ext4, online_delete=5000, advisory_delete=0. node_size was small at the time of the crash (now medium). Non-validator, serving a co-located Clio via gRPC.
  • No disk I/O or filesystem errors in dmesg, and the disk wasn't full. As a KVM guest the node can't see host ECC/MCE events, so the corruption source (memory bit-flip vs. storage bit-rot vs. software) can't be determined from the guest. The NuDB structure was intact (the varint length prefix parsed); only the LZ4 payload was bad.

Possibly related: #4883 (corruption correlated with tight online_delete intervals; this crash was in the rotating backend) and #7246 (LZ4 error handling on the overlay path; different code, same class of gap). The nodestore compress side does check for failure (lz4Compress throws on outSize == 0), so a failed compression can't silently write a bad blob.

Supporting Files

  • Full coredumpctl info backtrace available on request.
  • One core dump (745 MB, .zst) retained. We can share it privately with maintainers, but won't post it publicly because it contains process memory.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions