--- title: "Use cuda.ml with tidymodels" description: > Fit cuda.ml models through parsnip, preprocess predictors with recipes, and distinguish portable model arguments from engine-specific controls. output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Use cuda.ml with tidymodels} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", message = FALSE, eval = identical(Sys.getenv("IN_PKGDOWN"), "true") || identical(Sys.getenv("CUDA_ML_GPU_VIGNETTES"), "true") ) ``` cuda.ml registers parsnip engines for its supervised models. This gives the models the same specification, fitting, and prediction interface as other parsnip engines while training with cuML. Use the direct cuda.ml API when you need an algorithm that has no parsnip specification or detailed control over a solver. See [Getting started](cuda-ml.html) for installation and runtime setup. ## Available model specifications Loading cuda.ml registers the engine name `"cuda.ml"` for these parsnip specifications: | parsnip specification | mode | prediction types | |:--|:--|:--| | `linear_reg()` | regression | `"numeric"` | | `logistic_reg()` | classification | `"class"`, `"prob"` | | `multinom_reg()` | classification | `"class"`, `"prob"` | | `rand_forest()` | classification | `"class"`, `"prob"` | | `rand_forest()` | regression | `"numeric"` | | `nearest_neighbor()` | classification | `"class"`, `"prob"` | | `nearest_neighbor()` | regression | `"numeric"` | | `svm_rbf()` | classification | `"class"` | | `svm_rbf()` | regression | `"numeric"` | | `svm_poly()` | classification | `"class"` | | `svm_poly()` | regression | `"numeric"` | | `svm_linear()` | classification | `"class"` | | `svm_linear()` | regression | `"numeric"` | The SVM engines do not register probability predictions. In particular, `predict(fitted_svm, new_data, type = "prob")` is not supported. Use `logistic_reg()`, `multinom_reg()`, `rand_forest()`, or `nearest_neighbor()` when a classification workflow requires probabilities. Use `logistic_reg()` for a two-level outcome and `multinom_reg()` for an outcome with more than two levels. Both specifications call `cuda_ml_logistic_reg()`; the outcome determines whether cuda.ml uses its binary or multinomial loss. ## Set the engine Set the mode and common arguments in a parsnip specification, then select the cuda.ml engine: ```{r random-forest-specification} library(cuda.ml) library(parsnip) forest_spec <- rand_forest( mode = "classification", mtry = 2, trees = 500, min_n = 5 ) |> set_engine( "cuda.ml", max_depth = 20, n_bins = 256 ) set.seed(1) forest_fit <- fit(forest_spec, class ~ ., data = modeldata::hpc_data) class_predictions <- predict(forest_fit, modeldata::hpc_data, type = "class") probabilities <- predict(forest_fit, modeldata::hpc_data, type = "prob") ``` Arguments in the model specification, such as `mtry`, are common parsnip arguments. Arguments in `set_engine()`, such as `max_depth`, are specific to cuda.ml. Put each argument in only one place. ## Preprocess with a recipe cuda.ml's supervised models require numeric predictors. Scaling is especially important for KNN and SVM models because their fits depend on distances or margins. It is also usually appropriate for penalized linear and logistic models. Fit preprocessing parameters on the training data only, then apply the same recipe to assessment or production data. This example normalizes the predictors before fitting an exact KNN classifier. A workflow keeps the preprocessing and model specifications together. `fit()` estimates the recipe from the training data before fitting the cuda.ml model, and `predict()` applies the recipe to new data before calling the model engine. ```{r recipe-example} library(cuda.ml) library(parsnip) library(recipes) library(workflows) set.seed(1) training_rows <- sample( seq_len(nrow(modeldata::two_class_dat)), floor(0.8 * nrow(modeldata::two_class_dat)) ) training_data <- modeldata::two_class_dat[training_rows, ] testing_data <- modeldata::two_class_dat[-training_rows, ] classifier_recipe <- recipe(Class ~ ., data = training_data) |> step_normalize(all_numeric_predictors()) knn_spec <- nearest_neighbor( mode = "classification", neighbors = 5, dist_power = 2 ) |> set_engine( "cuda.ml", algo = "brute", metric = "euclidean" ) knn_workflow <- workflow() |> add_recipe(classifier_recipe) |> add_model(knn_spec) knn_fit <- fit(knn_workflow, data = training_data) results <- cbind( truth = testing_data$Class, predict(knn_fit, testing_data, type = "class"), predict(knn_fit, testing_data, type = "prob") ) head(results) ``` The parsnip KNN engine defaults to `algo = "ivfflat"` and `metric = "euclidean"`. `ivfflat` performs approximate neighbor search. Set `algo = "brute"`, as above, when exact search is required. The direct `cuda_ml_knn()` interface defaults to brute-force search. The cuda.ml engine does not map parsnip's `weight_func` argument; leave it as `NULL`. Tree models do not generally need normalization. A recipe can still be useful for creating numeric indicators or applying other preprocessing learned from the training set. ## Tune common arguments cuda.ml registers standard dials parameter metadata for the following parsnip arguments: | specification | registered arguments | |:--|:--| | `linear_reg()` | `penalty`, `mixture` | | `logistic_reg()`, `multinom_reg()` | `penalty`, `mixture` | | `rand_forest()` | `mtry`, `trees`, `min_n` | | `nearest_neighbor()` | `neighbors`, `dist_power` | | `svm_rbf()` | `cost`, `margin`, `rbf_sigma` | | `svm_poly()` | `cost`, `margin`, `degree`, `scale_factor` | | `svm_linear()` | `cost`, `margin` | These arguments can use `tune::tune()` in a tuning workflow. `margin` maps to the epsilon tube and affects SVM regression only; do not tune it for an SVM classifier. Engine arguments are not registered as dials parameters. Keep them fixed in `set_engine()` unless you define an explicit dials parameter and range for the tuning workflow. Useful engine arguments include: | engine | examples of engine-specific arguments | |:--|:--| | linear regression | `fit_intercept`; route-specific options described below | | logistic and multinomial regression | `fit_intercept`, `tol`, `class_weight`, `max_iter`, `linesearch_max_iter`, `lbfgs_memory`, `penalty_normalized` | | random forest | `bootstrap`, `sample_fraction`, `max_depth`, `max_leaves`, `n_bins`, `min_samples_leaf`, `split_criterion`, `min_impurity_decrease`, `max_batch_size`, `n_streams`, `seed` | | nearest neighbor | `algo`, `metric` | | SVM | `coef0`, `tol`, `max_iter`, `nochange_steps`, `cache_size` | Parsnip case weights are not currently supported by the cuda.ml engine. For per-observation weights, use `sample_weight` with `cuda_ml_logistic_reg()` or `sample_weights` with `cuda_ml_svm()` directly. Consult the corresponding `cuda_ml_*()` reference page before setting these arguments. For example, random-forest split criteria differ between classification and regression, and approximate KNN algorithms support fewer distance metrics than brute-force KNN. ### Linear regression routing The `linear_reg()` engine selects a cuda.ml solver from `penalty` and `mixture`: | values | direct function | |:--|:--| | `penalty = NULL` or `penalty = 0` | `cuda_ml_ols()` | | positive `penalty`, `mixture = 0` | `cuda_ml_ridge()` | | positive `penalty`, `mixture = 1` or `NULL` | `cuda_ml_lasso()` | | positive `penalty`, `0 < mixture < 1` | `cuda_ml_elastic_net()` | Only pass options supported by the selected function. For example, `method` is an OLS option, while `max_iter`, `tol`, and `selection` apply to the lasso and elastic-net routes. If the solver itself is part of the decision, use the named direct functions so that the relationship between the function and its arguments remains explicit. ## Choose between parsnip and the direct API Use parsnip when you want to compare engines through common specifications, use a tidymodels tuning workflow, or consume standard parsnip prediction types. Parsnip delegates training and prediction to cuda.ml's public model functions. The direct API includes several capabilities that have no parsnip specification: | Capability | parsnip | Direct cuda.ml API | |:--|:--|:--| | Linear, logistic, random-forest, KNN, and RBF, polynomial, or linear SVM models | Supported | Supported | | Clustering and dimensionality reduction | Not available | Agglomerative clustering, DBSCAN, k-means, PCA, tSVD, UMAP, and t-SNE | | Transform and inverse-transform operations | Not available | `cuda_ml_transform()` and `cuda_ml_inverse_transform()` | | Stochastic-gradient-descent linear regression | No dedicated route | `cuda_ml_sgd()` | | Hyperbolic-tangent SVM | No matching specification | `cuda_ml_svm(kernel = "tanh")` | | External tree-ensemble inference and inspection | Not available | nvForest load, predict, model-info, leaf-ID, per-tree, import, and export functions | Most algorithm-specific arguments for supported parsnip models remain available through `set_engine()`. Use the named direct function when selecting the algorithm or solver is part of the analysis, or when passing predictors and outcomes directly is preferable to a parsnip workflow. For example, cuda.ml's direct SVM interface supports a `"tanh"` kernel, but parsnip registration is limited to the RBF, polynomial, and linear SVM specifications. ```{r direct-api} direct_fit <- cuda_ml_svm( Class ~ ., data = modeldata::two_class_dat, kernel = "tanh", cost = 2, gamma = 0.1, coef0 = 0 ) direct_predictors <- subset(modeldata::two_class_dat, select = -Class) direct_predictions <- predict(direct_fit, direct_predictors) ``` See [Save and restore models](model-persistence.html) before moving a fitted model to another R process or deployment host.