Skip to content

Evaluate inclusive metrics from borrowed file statistics #3346

Description

@unikdahal

Part of #3343.

What's the feature are you trying to implement?

InclusiveMetricsEvaluator decides whether a data file might contain rows matching a predicate. It only reads the file's record count and its per-column value, null and NaN counts and lower/upper bounds, yet its entry point requires a complete DataFile. Statistics that live outside a DataFile — for example whole-file statistics carried with a scan task for execution-time pruning — cannot be evaluated without building one.

Proposal: introduce a crate-internal view that borrows exactly the statistics the evaluator reads, evaluate through it, and keep the existing DataFile entry point as a thin conversion.

Expected behaviour:

  • No change for existing callers: evaluating a DataFile gives the same results as today.
  • A record count known to be zero still excludes the file unless empty files are included.
  • An unknown record count does not by itself exclude a file; the available bounds and counts may still prune it.
  • No public API change (pub(crate) only).

This is a small refactor that lets later #3343 work prune whole files from task-level statistics without reconstructing DataFiles.

Willingness to contribute

I can contribute to this feature independently; an implementation with tests is ready.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions