diff --git a/docs/src/index.md b/docs/src/index.md index 7a3ff0f..0ad4b62 100644 --- a/docs/src/index.md +++ b/docs/src/index.md @@ -38,15 +38,23 @@ Xnew = transform(mach, X) ## Available Transformers See [complete list](transformers/all_transformers) of transformers in this package. -In `MLJTransforms` we denote transformers that can operate on columns with `Continuous` and/or `Count` [scientific types](https://juliaai.github.io/ScientificTypes.jl/dev/) as *numerical transformers*. Meanwhile, *categorical transformers* operate on `Multiclass` and/or `OrderedFactor` [scientific types](https://juliaai.github.io/ScientificTypes.jl/dev/). Most categorical transformers in this package operate by converting categorical values into numerical values or vectors, and are therefore considered categorical encoders. We categorize categorical encoders as follows: +In `MLJTransforms` we denote transformers that can operate on columns with `Continuous` and/or `Count` [scientific types](https://juliaai.github.io/ScientificTypes.jl/dev/) as *numerical transformers*. Meanwhile, *categorical transformers* operate on `Multiclass` and/or `OrderedFactor` [scientific types](https://juliaai.github.io/ScientificTypes.jl/dev/). Most categorical transformers in this package operate by converting categorical values into numerical values or vectors, and are therefore considered categorical encoders. +Some transformers in this package can operate on both `Finite` and `Infinite` scientific +types or other special scientific types (eg, to represent time). To learn more about +scientific types see [the official +documentation](https://juliaai.github.io/ScientificTypes.jl/dev/#Type-hierarchy). -| **Category** | **Description** | -|:---------------------------:|:-------------------------------------------------------------------------------:| -| [Classical Encoders](transformers/classical.md) | Traditional categorical encoding algorithms and techniques. | -| [Neural-based Encoders](transformers/neural) | Categorical encoders based on neural networks. | -| [Contrast Encoders](transformers/contrast.md) | Categorical encoders that could be modeled via a contrast matrix. | -| [Utility Encoders](transformers/utility.md) | Categorical encoders meant to be used as preprocessors for other transformers or models.| +### Categorical encoders + +The categorical encoders in this package can be further broken down as follows: + + +| **Category** | **Description** | +|:-----------------------------------------------:|:----------------------------------------------------------------------------------------:| +| [Classical Encoders](transformers/classical.md) | Traditional categorical encoding algorithms and techniques. | +| [Neural-based Encoders](transformers/neural) | Categorical encoders based on neural networks. | +| [Contrast Encoders](transformers/contrast.md) | Categorical encoders that could be modeled via a contrast matrix. | +| [Utility Encoders](transformers/utility.md) | Categorical encoders meant to be used as preprocessors for other transformers or models. | -Some transformers in this package can even operate on both `Finite` and `Infinite` scientific types or other special scientific types (eg, to represent time). To learn more about scientific types see [the official documentation](https://juliaai.github.io/ScientificTypes.jl/dev/#Type-hierarchy). \ No newline at end of file diff --git a/docs/src/transformers/all_transformers.md b/docs/src/transformers/all_transformers.md index ae548fa..9a96c92 100644 --- a/docs/src/transformers/all_transformers.md +++ b/docs/src/transformers/all_transformers.md @@ -1,27 +1,30 @@ ### Summary Table -| Transformer | Brief Description | -|:----------:|:----------:| -| [Standardizer](@ref) | Transforming columns of numerical features by standardization | -| [UnivariateBoxCoxTransformer](@ref) | Apply BoxCox transformation given a single vector | -| [InteractionTransformer](@ref) | Transforming columns of numerical features to create new interaction features | -| [UnivariateDiscretizer](@ref) | Discretize a continuous vector into an ordered factor | -| [FillImputer](@ref) | Fill missing values of features belonging to any scientific type | -| [UnivariateTimeTypeToContinuous](@ref) | Transform a vector of time type into continuous type | -| [UnivariateFillImputer](@ref) | Fill in missing values in a single vector | -| [OneHotEncoder](@ref) | Encode categorical variables into one-hot vectors | -| [ContinuousEncoder](@ref) | Adds type casting functionality to OnehotEncoder | -| [OrdinalEncoder](@ref) | Encode categorical variables into ordered integers | -| [FrequencyEncoder](@ref) | Encode categorical variables into their normalized or unormalized frequencies | -| [TargetEncoder](@ref) | Encode categorical variables into relevant target statistics | -| [DummyEncoder](@ref ContrastEncoder) | Encodes by comparing each level to the reference level, intercept being the cell mean of the reference group | -| [SumEncoder](@ref ContrastEncoder) | Encodes by comparing each level to the reference level, intercept being the grand mean | -| [HelmertEncoder](@ref ContrastEncoder) | Encodes by comparing levels of a variable with the mean of the subsequent levels of the variable -| [ForwardDifferenceEncoder](@ref ContrastEncoder) | Encodes by comparing adjacent levels of a variable (each level minus the next level) -| [ContrastEncoder](@ref) | Allows defining a custom contrast encoder via a contrast matrix | -| [HypothesisEncoder](@ref ContrastEncoder) | Allows defining a custom contrast encoder via a hypothesis matrix | -| [EntityEmbedder](@ref) | Encode categorical variables into dense embedding vectors | -| [CardinalityReducer](@ref) | Reduce cardinality of high cardinality categorical features by grouping infrequent categories | -| [MissingnessEncoder](@ref) | Encode missing values of categorical features into new values | +| Transformer | Brief Description | +|:------------------------------------------------:|:------------------------------------------------------------------------------------------------------------:| +| [Standardizer](@ref) | Transforming columns of numerical features by standardization | +| [UnivariateBoxCoxTransformer](@ref) | Apply BoxCox transformation given a single vector | +| [PolynomialTransformer](@ref) | Transformer to add columns that are products of other columns | +| [InteractionTransformer](@ref) | Transforming columns of numerical features to create new interaction features¹ | +| [UnivariateDiscretizer](@ref) | Discretize a continuous vector into an ordered factor | +| [FillImputer](@ref) | Fill missing values of features belonging to any scientific type | +| [UnivariateTimeTypeToContinuous](@ref) | Transform a vector of time type into continuous type | +| [UnivariateFillImputer](@ref) | Fill in missing values in a single vector | +| [OneHotEncoder](@ref) | Encode categorical variables into one-hot vectors | +| [ContinuousEncoder](@ref) | Adds type casting functionality to OnehotEncoder | +| [OrdinalEncoder](@ref) | Encode categorical variables into ordered integers | +| [FrequencyEncoder](@ref) | Encode categorical variables into their normalized or unormalized frequencies | +| [TargetEncoder](@ref) | Encode categorical variables into relevant target statistics | +| [DummyEncoder](@ref ContrastEncoder) | Encodes by comparing each level to the reference level, intercept being the cell mean of the reference group | +| [SumEncoder](@ref ContrastEncoder) | Encodes by comparing each level to the reference level, intercept being the grand mean | +| [HelmertEncoder](@ref ContrastEncoder) | Encodes by comparing levels of a variable with the mean of the subsequent levels of the variable | +| [ForwardDifferenceEncoder](@ref ContrastEncoder) | Encodes by comparing adjacent levels of a variable (each level minus the next level) | +| [ContrastEncoder](@ref) | Allows defining a custom contrast encoder via a contrast matrix | +| [HypothesisEncoder](@ref ContrastEncoder) | Allows defining a custom contrast encoder via a hypothesis matrix | +| [EntityEmbedder](@ref) | Encode categorical variables into dense embedding vectors | +| [CardinalityReducer](@ref) | Reduce cardinality of high cardinality categorical features by grouping infrequent categories | +| [MissingnessEncoder](@ref) | Encode missing values of categorical features into new values | + +¹The `InteractionTransformer` is deprecated. Use `PolynomialTransformer(interactions_only=true)` instead. ### All Transformers @@ -37,6 +40,10 @@ MLJTransforms.UnivariateStandardizer MLJTransforms.UnivariateBoxCoxTransformer ``` +```@docs; canonical = false +MLJTransforms.PolynomialTransformer +``` + ```@docs; canonical = false MLJTransforms.InteractionTransformer ``` @@ -87,4 +94,4 @@ MLJTransforms.CardinalityReducer ```@docs; canonical = false MLJTransforms.MissingnessEncoder -``` \ No newline at end of file +``` diff --git a/src/MLJTransforms.jl b/src/MLJTransforms.jl index dd4a051..9c540ca 100644 --- a/src/MLJTransforms.jl +++ b/src/MLJTransforms.jl @@ -62,6 +62,7 @@ export ContrastEncoder # MLJModels transformers include("transformers/other_transformers/continuous_encoder.jl") include("transformers/other_transformers/interaction_transformer.jl") +include("transformers/other_transformers/polynomial_transformer.jl") include("transformers/other_transformers/univariate_time_type_to_continuous.jl") include("transformers/other_transformers/fill_imputer.jl") include("transformers/other_transformers/one_hot_encoder.jl") @@ -72,5 +73,5 @@ include("transformers/other_transformers/univariate_discretizer.jl") export UnivariateDiscretizer, UnivariateStandardizer, Standardizer, UnivariateBoxCoxTransformer, OneHotEncoder, ContinuousEncoder, FillImputer, UnivariateFillImputer, - UnivariateTimeTypeToContinuous, InteractionTransformer + UnivariateTimeTypeToContinuous, InteractionTransformer, PolynomialTransformer end diff --git a/src/transformers/other_transformers/interaction_transformer.jl b/src/transformers/other_transformers/interaction_transformer.jl index 20d36ca..4086bbc 100644 --- a/src/transformers/other_transformers/interaction_transformer.jl +++ b/src/transformers/other_transformers/interaction_transformer.jl @@ -1,31 +1,24 @@ +# The model implementation here is deprecated @mlj_model mutable struct InteractionTransformer <: Static order::Int = 2::(_ > 1) features::Union{Nothing, Vector{Symbol}} = nothing::(_ !== nothing ? length(_) > 1 : true) end -infinite_scitype(col) = eltype(scitype(col)) <: Infinite - -actualfeatures(features::Nothing, table) = - filter(feature -> infinite_scitype(Tables.getcolumn(table, feature)), Tables.columnnames(table)) - -function actualfeatures(features::Vector{Symbol}, table) - diff = setdiff(features, Tables.columnnames(table)) - diff != [] && throw(ArgumentError(string("Column(s) ", join([x for x in diff], ", "), " are not in the dataset."))) - - for feature in features - infinite_scitype(Tables.getcolumn(table, feature)) || throw(ArgumentError("Column $feature's scitype is not Infinite.")) - end - return Tuple(features) -end - interactions(columns, order::Int) = collect(Iterators.flatten(combinations(columns, i) for i in 2:order)) interactions(columns, variables...) = .*((Tables.getcolumn(columns, var) for var in variables)...) +const WARN_INTERACTION_DEPRECATED = """ + `InteractionTransformer(; kwargs...)` is deprecated. Instead use + `PolynomialTransformer(; interactions_only=true, kwargs...)`. The + `PolynomialTransformer` type is also provided by the MLJTransforms module. + """ + function MMI.transform(model::InteractionTransformer, _, X) + Base.depwarn(WARN_INTERACTION_DEPRECATED, :transform) features = actualfeatures(model.features, X) interactions_ = interactions(features, model.order) interaction_features = Tuple(Symbol(join(inter, "_")) for inter in interactions_) diff --git a/src/transformers/other_transformers/polynomial_transformer.jl b/src/transformers/other_transformers/polynomial_transformer.jl new file mode 100644 index 0000000..bb3c2bd --- /dev/null +++ b/src/transformers/other_transformers/polynomial_transformer.jl @@ -0,0 +1,192 @@ +# # STRUCT AND CONSTRUCTORS + +const WARN_DEGREE = "The `degree` must be at least 1. "* + "Reset `degree=2`. " + + +mutable struct PolynomialTransformer <: Static + degree::Int + features::Union{Nothing, Vector{Symbol}} + interactions_only::Bool +end + +function MMI.clean!(model::PolynomialTransformer) + message = "" + if model.degree ≤ 0 + model.degree = 2 + message *= WARN_DEGREE + end + return message +end + +function PolynomialTransformer( + ; order=2, + degree=order, + features=nothing, + interactions_only=false, + ) + model = PolynomialTransformer(degree, features, interactions_only) + message = MMI.clean!(model) + isempty(message) || @warn message + return model +end + + +# # HELPERS + +abstract type Selection end +struct WithRepetitions <: Selection end +struct WithoutRepetitions <: Selection end + +""" + premonomials(alphabet, degree, kind_of_selection::Selection) + +*Private method* to help generate monomials. A **pre-monomial** is a vector with elements + from the alphabet, with possible repetitions, but with no element predecessor coming + *after* the element itself in the alphabet. + +Note degree one "pre-monomials" are excluded. + +# Example + +```julia-repl +julia> premonomials((:x, :y, :z), 3, WithoutRepetitions()) +4-element Vector{Vector{Symbol}}: + [:x, :y] + [:x, :z] + [:y, :z] + [:x, :y, :z] + +julia> premonomials((:x, :y, :z), 3, WithRepetitions()) +16-element Vector{Vector{Symbol}}: + [:x, :x] + [:x, :y] + [:x, :z] + [:y, :y] + [:y, :z] + [:z, :z] + [:x, :x, :x] + [:x, :x, :y] + [:x, :x, :z] + [:x, :y, :y] + [:x, :y, :z] + [:x, :z, :z] + [:y, :y, :y] + [:y, :y, :z] + [:y, :z, :z] + [:z, :z, :z] +``` +""" +premonomials(alphabet, degree, ::WithoutRepetitions) = + premonomials(alphabet, degree, Combinatorics.combinations) +premonomials(alphabet, degree, ::WithRepetitions) = + premonomials(alphabet, degree, Combinatorics.with_replacement_combinations) +premonomials(alphabet, degree, fnctn) = + collect(Iterators.flatten(fnctn(alphabet, i) for i in 2:degree)) + +column_product(columns, premonomial...) = + .*((Tables.getcolumn(columns, feature) for feature in premonomial)...) + + +# # CORE IMPLEMENTATION + +function MMI.transform(model::PolynomialTransformer, _, X) + features = MLJTransforms.actualfeatures(model.features, X) + kind_of_selection = model.interactions_only ? WithoutRepetitions() : WithRepetitions() + premonomials = MLJTransforms.premonomials(features, model.degree, kind_of_selection) + new_features = Tuple(Symbol(join(premon, "_")) for premon in premonomials) + materializer = Tables.materializer(X) + columns = Tables.Columns(X) + table_addendum = + NamedTuple{new_features}( + [column_product(columns, premon...) for premon in premonomials], + ) + return merge(Tables.columntable(X), table_addendum) |> materializer +end + + +# # TRAITS + +metadata_model(PolynomialTransformer, + input_scitype = Tuple{Table}, + output_scitype = Table, + human_name = "polynomial transformer", + load_path = "MLJTransforms.PolynomialTransformer") + +# Package metadata for docstring generation +metadata_pkg(PolynomialTransformer, + package_name = "MLJTransforms", + package_uuid = "23777cdb-d90c-4eb0-a694-7c2b83d5c1d6", + package_url = "https://github.com/JuliaAI/MLJTransforms.jl", + is_pure_julia = true, + package_license = "MIT") + +""" +$(MLJModelInterface.doc_header(PolynomialTransformer)) + +This `Static` transformer generates new features comprised of monomials in existing +features that have `Continuous` or `Count` scitype, up to some specified degree. A +restricted set of features may be specified, and one may elect to generate only +interaction monomials (no feature appearing with degree higher than one). + +In MLJ or MLJBase, you can transform features `X` with the single call + + transform(machine(model), X) + +See also the example below. + + +# Hyper-parameters + +- `degree=2`: maximum degree of monomials to be generated + +- `features=nothing`: vector of features for which monomials should be generated; if + `nothing` (unspecified) then all `Continuous` and `Count` features are used. + +# Operations + +- `transform(machine(model), X)`: Generate a new table from `X` with the monomial columnn + specified by hyper-parameters. + +# Example + +``` +using MLJ + +X = ( + A = [1, 2, 3], + B = [4, 5, 6], + C = [7, 8, 9], + D = ["cat", "dog", "rat"] +) + +transformer = PolynomialTransformer(degree=2, features=[:A, :B]) +mach = machine(transformer) +julia> transform(mach, X) |> pretty +┌───────┬───────┬───────┬─────────┬───────┬───────┬───────┐ +│ A │ B │ C │ D │ A_A │ A_B │ B_B │ +│ Int64 │ Int64 │ Int64 │ String │ Int64 │ Int64 │ Int64 │ +│ Count │ Count │ Count │ Textual │ Count │ Count │ Count │ +├───────┼───────┼───────┼─────────┼───────┼───────┼───────┤ +│ 1 │ 4 │ 7 │ cat │ 1 │ 4 │ 16 │ +│ 2 │ 5 │ 8 │ dog │ 4 │ 10 │ 25 │ +│ 3 │ 6 │ 9 │ rat │ 9 │ 18 │ 36 │ +└───────┴───────┴───────┴─────────┴───────┴───────┴───────┘ + +transformer = PolynomialTransformer(degree=3, interactions_only=true) +mach = machine(transformer) +julia> transform(mach, X) |> pretty +┌───────┬───────┬───────┬─────────┬───────┬───────┬───────┬───────┐ +│ A │ B │ C │ D │ A_B │ A_C │ B_C │ A_B_C │ +│ Int64 │ Int64 │ Int64 │ String │ Int64 │ Int64 │ Int64 │ Int64 │ +│ Count │ Count │ Count │ Textual │ Count │ Count │ Count │ Count │ +├───────┼───────┼───────┼─────────┼───────┼───────┼───────┼───────┤ +│ 1 │ 4 │ 7 │ cat │ 4 │ 7 │ 28 │ 28 │ +│ 2 │ 5 │ 8 │ dog │ 10 │ 16 │ 40 │ 80 │ +│ 3 │ 6 │ 9 │ rat │ 18 │ 27 │ 54 │ 162 │ +└───────┴───────┴───────┴─────────┴───────┴───────┴───────┴───────┘ + +``` + +""" +PolynomialTransformer diff --git a/src/utils.jl b/src/utils.jl index 8ac976c..6f0c9ac 100644 --- a/src/utils.jl +++ b/src/utils.jl @@ -1 +1,24 @@ -# add utility functions here \ No newline at end of file +# add utility functions here + +has_infinite_scitype(col) = scitype(col) <:AbstractVector{<:Union{Missing,Infinite}} + +# method to extrac, `Infinite` scitype features from a table, given a subset of features. +actualfeatures(features::Nothing, table) = + filter(Tables.columnnames(table)) do feature + MLJTransforms.has_infinite_scitype(Tables.getcolumn(table, feature)) + end +function actualfeatures(features::Vector{Symbol}, table) + diff = setdiff(features, Tables.columnnames(table)) + diff != [] && + throw(ArgumentError(string( + "Column(s) ", + join([x for x in diff], ", "), + " are not in the dataset."), + ) + ) + for feature in features + MLJTransforms.has_infinite_scitype(Tables.getcolumn(table, feature)) || + throw(ArgumentError("Column $feature's scitype is not Infinite.")) + end + return Tuple(features) +end diff --git a/test/runtests.jl b/test/runtests.jl index 8925eb8..3a022cc 100644 --- a/test/runtests.jl +++ b/test/runtests.jl @@ -20,6 +20,7 @@ stable_rng = StableRNGs.StableRNG(123) using Dates: DateTime, Date, Time, Day, Hour _get(x) = CategoricalArrays.DataAPI.unwrap(x) +include("test_utils.jl") include("utils.jl") include("generic.jl") @@ -36,6 +37,7 @@ include("transformers/other_transformers/fill_imputer.jl") include("transformers/other_transformers/one_hot_encoder.jl") include("transformers/other_transformers/univariate_time_type_to_continuous.jl") include("transformers/other_transformers/interaction_transformer.jl") +include("transformers/other_transformers/polynomial_transformer.jl") include("transformers/other_transformers/continuous_encoder.jl") include("transformers/other_transformers/univariate_boxcox_transformer.jl") include("transformers/other_transformers/standardizer.jl") diff --git a/test/test_utils.jl b/test/test_utils.jl new file mode 100644 index 0000000..febcf15 --- /dev/null +++ b/test/test_utils.jl @@ -0,0 +1,117 @@ +# Function to create a dummy dataset +function create_dummy_dataset(target_type::Symbol; as_dataframe::Bool = false, return_y::Bool=true) + # Define categorical features with shorter names + A = ["g", "b", "g", "r", "r", "r", "r", "b", "b", "r"] + B = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] + C = ["f", "f", "f", "m", "f", "m", "f", "m", "f", "m"] + D = [true, false, true, false, true, false, true, false, true, false] + E = [1, 2, 3, 4, 5, 6, 6, 3, 2, 1] + F = ["s", "s", "s", "s", "l", "m", "s", "l", "m", "m"] + + # Define the target variable based on the target type + if target_type == :binary + y = [0, 1, 1, 1, 0, 1, 0, 1, 1, 1] + y = coerce(y, Multiclass) + elseif target_type == :binary_str + y = ["no", "yes", "yes", "yes", "no", "yes", "no", "yes", "yes", "yes"] + y = coerce(y, Multiclass) + elseif target_type == :multiclass + y = [0, 1, 2, 0, 1, 2, 0, 1, 2, 0] + y = coerce(y, Multiclass) + elseif target_type == :multiclass_str + y = ["c1", "c2", "c3", "c1", "c2", "c3", "c1", "c2", "c3", "c1"] + y = coerce(y, Multiclass) + elseif target_type == :regression + y = [10.5, 15.2, 13.1, 14.7, 11.9, 16.8, 17.0, 19.3, 18.1, 20.0] + y = coerce(y, Continuous) + else + error("Unsupported target type.") + end + + # Combine into a named tuple + X = (A = A, B = B, C = C, D = D, E = E, F = F) + + # Coerce A, C, D, F to multiclass and B to continuous and E to ordinal + X = coerce(X, + :A => Multiclass, + :B => Continuous, + :C => Multiclass, + :D => Multiclass, + :E => OrderedFactor, + :F => Multiclass, + ) + + as_dataframe && (X = DataFrame(X)) + + return (return_y) ? (X, y) : X; +end + + +struct Object{I<:Integer} + A::I +end +# Create dummy dataset but with high cardinality +function generate_high_cardinality_table(num_rows; obj=false, special_cat='E') + # Set the random seed for reproducibility + Random.seed!(stable_rng, 123) + + # Define the categories for the categorical features with their respective probabilities + low_card_categories = ['A', 'B', 'C', 'D', special_cat] + low_card_probs = [0.7, 0.2, 0.05, 0.04, 0.01] # Imbalanced distribution + + + + high_card_categories1 = [((obj) ? Object(i) : i) for i in 1:100] + high_card_probs1 = vcat(fill(0.01, 90), fill(0.1, 10)) # Last 10 categories more frequent + + high_card_categories2 = [string("Group", i) for i in 1:200] + high_card_probs2 = vcat(fill(0.005, 190), fill(0.05, 10)) # Last 10 categories more frequent + + # Function to generate a weighted random sample + function weighted_sample(categories, probs) + cumulative_probs = cumsum(probs) + rand_val = rand() + for (i, p) in enumerate(cumulative_probs) + if rand_val <= p + return categories[i] + end + end + end + + # Generate the categorical features with imbalanced distributions + low_card_feature = [weighted_sample(low_card_categories, low_card_probs) for _ in 1:num_rows] + high_card_feature1 = [weighted_sample(high_card_categories1, high_card_probs1) for _ in 1:num_rows] + high_card_feature2 = [weighted_sample(high_card_categories2, high_card_probs2) for _ in 1:num_rows] + + dataset = DataFrame( + LowCardFeature = low_card_feature, + HighCardFeature1 = high_card_feature1, + HighCardFeature2 = high_card_feature2 + ) + + dataset = coerce(dataset, + :LowCardFeature => Multiclass, + :HighCardFeature1 => Multiclass, + :HighCardFeature2 => Multiclass, + ) + + return dataset + +end + + +function generate_X_with_missingness(;john_name="John") + Xm = ( + A = categorical(["Ben", john_name, missing, missing, "Mary", "John", missing]), + B = [1.85, 1.67, missing, missing, 1.5, 1.67, missing], + C= categorical([7, 5, missing, missing, 10, 5, missing]), + D = [23, 23, 44, 66, 14, 23, 11], + E = categorical([missing, 'g', 'r', missing, 'r', 'g', 'p']) + ) + + return Xm +end + + +# Display the dataset +dataset = generate_high_cardinality_table(1000; obj=false) diff --git a/test/transformers/other_transformers/interaction_transformer.jl b/test/transformers/other_transformers/interaction_transformer.jl index 0561a2a..757d41d 100644 --- a/test/transformers/other_transformers/interaction_transformer.jl +++ b/test/transformers/other_transformers/interaction_transformer.jl @@ -1,17 +1,3 @@ - -@testset "Interaction Transformer functions" begin - # No column provided, A has scitype Continuous, B has scitype Count - table = (A = [1., 2., 3.], B = [4, 5, 6], C = ["x₁", "x₂", "x₃"]) - @test MLJTransforms.actualfeatures(nothing, table) == (:A, :B) - # Column provided - @test MLJTransforms.actualfeatures([:A, :B], table) == (:A, :B) - # Column provided, not in table - @test_throws ArgumentError("Column(s) D are not in the dataset.") MLJTransforms.actualfeatures([:A, :D], table) - # Non Infinite scitype column provided - @test_throws ArgumentError("Column C's scitype is not Infinite.") MLJTransforms.actualfeatures([:A, :C], table) -end - - @testset "Interaction Transformer" begin # Check constructor sanity checks: order > 1, length(features) > 1 @test_logs (:warn, "Constraint `model.order > 1` failed; using default: order=2.") InteractionTransformer(order = 1) diff --git a/test/transformers/other_transformers/polynomial_transformer.jl b/test/transformers/other_transformers/polynomial_transformer.jl new file mode 100644 index 0000000..95710e7 --- /dev/null +++ b/test/transformers/other_transformers/polynomial_transformer.jl @@ -0,0 +1,174 @@ +using MLJTransforms +import MLJBase +using Test +import DataFrames.DataFrame +using Combinatorics + +@testset "PolynomialTransformer" begin + @testset "helper functions" begin + @test collect(Combinatorics.with_replacement_combinations((:x, :y, :z), 2)) == + [[:x, :x], [:x, :y], [:x, :z], [:y, :y], [:y, :z], [:z, :z]] + @test MLJTransforms.premonomials( + (:x, :y, :z), + 3, + MLJTransforms.WithoutRepetitions(), + ) == [[:x, :y], [:x, :z], [:y, :z], [:x, :y, :z]] + @test MLJTransforms.premonomials( + (:x, :y, :z), + 3, + MLJTransforms.WithRepetitions(), + ) == [[:x, :x], [:x, :y], [:x, :z], [:y, :y], [:y, :z], [:z, :z], [:x, :x, :x], + [:x, :x, :y], [:x, :x, :z], [:x, :y, :y], [:x, :y, :z], [:x, :z, :z], + [:y, :y, :y], [:y, :y, :z], [:y, :z, :z], [:z, :z, :z]] + end + + + # # INTERACTIONS ONLY + + @testset "interactions only" begin + # Check constructor sanity checks: + @test_logs( + (:warn, MLJTransforms.WARN_DEGREE), + PolynomialTransformer(interactions_only=true, degree = 0), + ) + + X = (A = [1, 2, 3], B = [4, 5, 6], C = [7, 8, 9]) + # Default degree=2, features=nothing, i.e., all columns + Xt = MLJBase.transform(PolynomialTransformer(interactions_only=true), nothing, X) + @test Xt == ( + A = [1, 2, 3], + B = [4, 5, 6], + C = [7, 8, 9], + A_B = [4, 10, 18], + A_C = [7, 16, 27], + B_C = [28, 40, 54] + ) + # degree=3, features=nothing, ie all columns + Xt = MLJBase.transform( + PolynomialTransformer(interactions_only=true,degree=3), + nothing, + X, + ) + @test Xt == ( + A = [1, 2, 3], + B = [4, 5, 6], + C = [7, 8, 9], + A_B = [4, 10, 18], + A_C = [7, 16, 27], + B_C = [28, 40, 54], + A_B_C = [28, 80, 162] + ) + # degree=2, features=[:A, :B], ie all columns + Xt = MLJBase.transform( + PolynomialTransformer(interactions_only=true, degree=2, features=[:A, :B]), + nothing, + X, + ) + @test Xt == ( + A = [1, 2, 3], + B = [4, 5, 6], + C = [7, 8, 9], + A_B = [4, 10, 18] + ) + # degree=3, features=[:A, :B, :C], some non continuous columns + X = merge(X, (D = ["x₁", "x₂", "x₃"],)) + Xt = MLJBase.transform( + PolynomialTransformer(interactions_only=true, degree=3, features=[:A, :B, :C]), + nothing, + X, + ) + @test Xt == ( + A = [1, 2, 3], + B = [4, 5, 6], + C = [7, 8, 9], + D = ["x₁", "x₂", "x₃"], + A_B = [4, 10, 18], + A_C = [7, 16, 27], + B_C = [28, 40, 54], + A_B_C = [28, 80, 162] + ) + # degree=2, features=nothing, only continuous columns are dealt with + Xt = MLJBase.transform( + PolynomialTransformer(interactions_only=true, degree=2), + nothing, + X, + ) + @test Xt == ( + A = [1, 2, 3], + B = [4, 5, 6], + C = [7, 8, 9], + D = ["x₁", "x₂", "x₃"], + A_B = [4, 10, 18], + A_C = [7, 16, 27], + B_C = [28, 40, 54], + ) + end + + @testset "all terms" begin + # Check constructor sanity checks: + @test_logs( + (:warn, MLJTransforms.WARN_DEGREE), + PolynomialTransformer(degree = 0), + ) + + X = (A = [1, 2, 3], B = [4, 5, 6]) + # Default degree=2, features=nothing, i.e., all columns + Xt = MLJBase.transform(PolynomialTransformer(), nothing, X) + @test Xt == ( + A = [1, 2, 3], + B = [4, 5, 6], + A_A = [1, 4, 9], + A_B = [4, 10, 18], + B_B = [16, 25, 36], + ) + + # degree=3, features=nothing, ie all columns + Xt = MLJBase.transform( + PolynomialTransformer(degree=3), + nothing, + X, + ) + @test Xt == ( + A = [1, 2, 3], + B = [4, 5, 6], + A_A = [1, 4, 9], + A_B = [4, 10, 18], + B_B = [16, 25, 36], + A_A_A = [1, 8, 27], + A_A_B = [4, 20, 54], + A_B_B = [16, 50, 108], + B_B_B = [64, 125, 216], + ) + + # degree=2, some non continuous columns + X = merge(X, (D = ["x₁", "x₂", "x₃"],)) + Xt = MLJBase.transform( + PolynomialTransformer(degree=2), + nothing, + X, + ) + @test Xt == ( + A = [1, 2, 3], + B = [4, 5, 6], + D = ["x₁", "x₂", "x₃"], + A_A = [1, 4, 9], + A_B = [4, 10, 18], + B_B = [16, 25, 36], + ) + end + + @testset "non-native table types" begin + X = (A = [1, 2, 3], B = [4, 5, 6], C = [7, 8, 9]) |> DataFrame + Xt = MLJBase.transform(PolynomialTransformer(interactions_only=true), nothing, X) + @test Xt == DataFrame(( + A = [1, 2, 3], + B = [4, 5, 6], + C = [7, 8, 9], + A_B = [4, 10, 18], + A_C = [7, 16, 27], + B_C = [28, 40, 54] + )) + end +end + +true diff --git a/test/utils.jl b/test/utils.jl index fb36f43..73375d0 100644 --- a/test/utils.jl +++ b/test/utils.jl @@ -1,117 +1,26 @@ -# Function to create a dummy dataset -function create_dummy_dataset(target_type::Symbol; as_dataframe::Bool = false, return_y::Bool=true) - # Define categorical features with shorter names - A = ["g", "b", "g", "r", "r", "r", "r", "b", "b", "r"] - B = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] - C = ["f", "f", "f", "m", "f", "m", "f", "m", "f", "m"] - D = [true, false, true, false, true, false, true, false, true, false] - E = [1, 2, 3, 4, 5, 6, 6, 3, 2, 1] - F = ["s", "s", "s", "s", "l", "m", "s", "l", "m", "m"] - - # Define the target variable based on the target type - if target_type == :binary - y = [0, 1, 1, 1, 0, 1, 0, 1, 1, 1] - y = coerce(y, Multiclass) - elseif target_type == :binary_str - y = ["no", "yes", "yes", "yes", "no", "yes", "no", "yes", "yes", "yes"] - y = coerce(y, Multiclass) - elseif target_type == :multiclass - y = [0, 1, 2, 0, 1, 2, 0, 1, 2, 0] - y = coerce(y, Multiclass) - elseif target_type == :multiclass_str - y = ["c1", "c2", "c3", "c1", "c2", "c3", "c1", "c2", "c3", "c1"] - y = coerce(y, Multiclass) - elseif target_type == :regression - y = [10.5, 15.2, 13.1, 14.7, 11.9, 16.8, 17.0, 19.3, 18.1, 20.0] - y = coerce(y, Continuous) - else - error("Unsupported target type.") - end - - # Combine into a named tuple - X = (A = A, B = B, C = C, D = D, E = E, F = F) - - # Coerce A, C, D, F to multiclass and B to continuous and E to ordinal - X = coerce(X, - :A => Multiclass, - :B => Continuous, - :C => Multiclass, - :D => Multiclass, - :E => OrderedFactor, - :F => Multiclass, +@testset "actualfeatures" begin + # No column provided, A has scitype Continuous, B has scitype Count + table = (A = [1., 2., 3.], B = [4, 5, 6], C = ["x₁", "x₂", "x₃"]) + @test MLJTransforms.actualfeatures(nothing, table) == (:A, :B) + # Column provided + @test MLJTransforms.actualfeatures([:A, :B], table) == (:A, :B) + # Column provided, not in table + @test_throws( + ArgumentError("Column(s) D are not in the dataset."), + MLJTransforms.actualfeatures([:A, :D], table), ) - - as_dataframe && (X = DataFrame(X)) - - return (return_y) ? (X, y) : X; -end - - -struct Object{I<:Integer} - A::I -end -# Create dummy dataset but with high cardinality -function generate_high_cardinality_table(num_rows; obj=false, special_cat='E') - # Set the random seed for reproducibility - Random.seed!(stable_rng, 123) - - # Define the categories for the categorical features with their respective probabilities - low_card_categories = ['A', 'B', 'C', 'D', special_cat] - low_card_probs = [0.7, 0.2, 0.05, 0.04, 0.01] # Imbalanced distribution - - - - high_card_categories1 = [((obj) ? Object(i) : i) for i in 1:100] - high_card_probs1 = vcat(fill(0.01, 90), fill(0.1, 10)) # Last 10 categories more frequent - - high_card_categories2 = [string("Group", i) for i in 1:200] - high_card_probs2 = vcat(fill(0.005, 190), fill(0.05, 10)) # Last 10 categories more frequent - - # Function to generate a weighted random sample - function weighted_sample(categories, probs) - cumulative_probs = cumsum(probs) - rand_val = rand() - for (i, p) in enumerate(cumulative_probs) - if rand_val <= p - return categories[i] - end - end - end - - # Generate the categorical features with imbalanced distributions - low_card_feature = [weighted_sample(low_card_categories, low_card_probs) for _ in 1:num_rows] - high_card_feature1 = [weighted_sample(high_card_categories1, high_card_probs1) for _ in 1:num_rows] - high_card_feature2 = [weighted_sample(high_card_categories2, high_card_probs2) for _ in 1:num_rows] - - dataset = DataFrame( - LowCardFeature = low_card_feature, - HighCardFeature1 = high_card_feature1, - HighCardFeature2 = high_card_feature2 - ) - - dataset = coerce(dataset, - :LowCardFeature => Multiclass, - :HighCardFeature1 => Multiclass, - :HighCardFeature2 => Multiclass, - ) - - return dataset - -end - - -function generate_X_with_missingness(;john_name="John") - Xm = ( - A = categorical(["Ben", john_name, missing, missing, "Mary", "John", missing]), - B = [1.85, 1.67, missing, missing, 1.5, 1.67, missing], - C= categorical([7, 5, missing, missing, 10, 5, missing]), - D = [23, 23, 44, 66, 14, 23, 11], - E = categorical([missing, 'g', 'r', missing, 'r', 'g', 'p']) + # Non Infinite scitype column provided + @test_throws( + ArgumentError("Column C's scitype is not Infinite."), + MLJTransforms.actualfeatures([:A, :C], table), ) - - return Xm end +@testset "has_infinite_scitype" begin + @test MLJTransforms.has_infinite_scitype([1.6, 1.7]) + @test MLJTransforms.has_infinite_scitype([42, missing]) + @test !MLJTransforms.has_infinite_scitype(["cat", "dog"]) + @test !MLJTransforms.has_infinite_scitype(["cat", missing]) +end -# Display the dataset -dataset = generate_high_cardinality_table(1000; obj=false) \ No newline at end of file +true